2606.16262v1 Jun 15, 2026 cs.SE

UXBench: LLM이 생성한 사용자 경험(UX) 비판의 실용성을 측정하는 방법

UXBench: Measuring the Actionability of LLM-Generated UX Critiques

Xiangliang Zhang
Xiangliang Zhang
Citations: 997
h-index: 16
Hang Hua
Hang Hua
Citations: 121
h-index: 6
Wenjie Wang
Wenjie Wang
Citations: 69
h-index: 4
Zipeng Ling
Zipeng Ling
Citations: 29
h-index: 4
Yue Huang
Yue Huang
Citations: 662
h-index: 10
Shiyi Du
Shiyi Du
Citations: 47
h-index: 2
Yuexing Hao
Yuexing Hao
Citations: 17
h-index: 3
Xiaomin Li
Xiaomin Li
Citations: 129
h-index: 6
Han Bao
Han Bao
Citations: 93
h-index: 4
Yuchen Ma
Yuchen Ma
Citations: 70
h-index: 4
Yanfang Ye
Yanfang Ye
Citations: 193
h-index: 7
Xiaonan Luo
Xiaonan Luo
Citations: 22
h-index: 2
Yu Jiang
Yu Jiang
Citations: 125
h-index: 2
Di Wang
Di Wang
Citations: 7
h-index: 1

대규모 언어 모델(LLM)은 인터페이스를 검사하고 사용성 문제를 진단하며 개선 방안을 제안하는 UX 평가기로 점점 더 많이 활용되고 있습니다. 그러나 현재까지 다양한 제품 환경에서 LLM이 생성한 비판의 신뢰성과 실용성을 측정하는 표준화된 벤치마크는 존재하지 않습니다. 본 논문에서는 LLM을 상호작용 기반 UX 평가기로 평가하기 위한 벤치마크인 UXBench를 소개합니다. UXBench는 열 가지 제품 유형에 걸쳐 로컬 환경에서 실행 가능한 웹 테스트 케이스로 구성되어 있으며, 모델이 결과를 보고하기 전에 인터랙션 데이터를 수집하도록 설계된 브라우저 탐색 기능을 포함하고 있습니다. 각 평가 모델은 일곱 가지 평가 기준(rubric)에 따른 체계적인 UX 보고서를 생성하며, 보고서의 품질은 고정된 후처리 개선 도구가 비판을 바탕으로 인터페이스를 얼마나 효과적으로 개선할 수 있는지로 측정됩니다. 본 연구에서는 자동화된 개선 성능 측정 프로토콜과 익명 인간 검증 연구를 통해 8개의 최첨단 모델을 평가했습니다. 결과는 UX 평가가 완전히 정형화되지 않았으며, 다양한 측면에서 차이를 보인다는 것을 보여줍니다. 즉, 모델들은 보고서의 실용성 측면에서 의미 있는 차이를 보이며, 각 평가 기준별로 뚜렷한 개선 패턴을 나타내고, 테스트 케이스 수준에서의 신뢰성이 다르며, 제품 유형에 따라 성능이 달라지는 경향을 보였습니다.

Original Abstract

Large language models (LLMs) are increasingly deployed as UX judges that inspect interfaces, diagnose usability problems, and propose repairs. Yet no controlled benchmark measures whether the resulting critiques are reliable and actionable across heterogeneous product surfaces. We introduce UXBench, a benchmark for evaluating LLMs as interaction-grounded UX judges. UXBench comprises local-first runnable web fixtures spanning ten product-surface families, paired with coverage-gated browser exploration that forces models to collect interaction evidence before reporting. Each judge model produces a structured UX report over seven rubric dimensions; report quality is measured by whether a fixed downstream repair agent can improve the interface based on the critique. We evaluate eight frontier models under both an automated repair-lift protocol and a blind human validation study. Results show that UX judging is neither saturated nor one dimensional: models differ meaningfully in report actionability, exhibit distinct rubric-level repair signatures, vary in fixture-level reliability, and trade leadership across surface categories

0 Citations
0 Influential
8 Altmetric
40.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!