2607.28439v1 Jul 30, 2026 cs.CL

단일 심사관을 넘어: 생성 UI 평가를 위한 소셜 페르소나 패널 시뮬레이션

Beyond a Single Judge: The Evidence-Grounded, Social-Weighted Persona Panel for Generative UI Evaluation

Zheng Wu
Zheng Wu
Citations: 99
h-index: 7
Zhuosheng Zhang
Zhuosheng Zhang
Citations: 434
h-index: 11
Cheng Yang
Cheng Yang
Citations: 0
h-index: 0
Yibo Luo
Yibo Luo
Citations: 0
h-index: 0
Pu Zhang
Pu Zhang
Citations: 0
h-index: 0

생성형 UI (GenUI)는 대규모 언어 모델이 자연어 지침에서 완전하고 렌더링 가능한 인터페이스를 직접 생성할 수 있도록 하지만, 생성된 결과물의 품질을 평가하는 것은 여전히 해결해야 할 과제입니다. 인간 평가 방식은 비용이 많이 들고 평가자 간 편차가 있으며, LLM을 심사관으로 사용하는 방식은 확장성이 뛰어나지만 단일하고 숨겨진 관점만을 반영하여 실제 사용자의 다양한 그룹이 동일한 인터페이스를 어떻게 인식하는지를 제대로 파악할 수 없습니다. 우리는 심리적으로 다양하고 증거 기반의 페르소나 패널(ESPP)을 제안합니다. ESPP는 세 단계로 구성된 GenUI 평가 방법으로, 패널 구성원들은 독립적으로 스크린샷을 평가하고, 특성 기반 및 의미론적 제한이 적용된 신뢰 구간 내에서 의견을 교환하며, 델파이 방식을 활용한 사회적 가중치를 통해 최종 판단을 내립니다. ESPP는 단순한 단일 심사관 방식보다 인간의 판단과 훨씬 더 밀접하게 일치하며, 피어슨 상관 계수(r)를 0.716에서 0.922로 향상시킵니다. 프롬프트 조합 방식을 사용한 실험에서는 이 격차의 약 3분의 1만 회복되었으며, 이는 페르소나와 증거 기반 접근 방식이 개선 효과의 주요 원인임을 보여줍니다. 이러한 정확도 향상 외에도, 각 패널 참여자의 개별 평가를 유지하면 사용자 하위 그룹은 전체 모델 순위에 대해 일치하는 경향을 보이지만 특정 평가 차원에서는 뚜렷한 의견 차이를 보이는 것을 알 수 있습니다. 이러한 구조적 불일치는 단일하고 균질적인 심사관이 체계적으로 제거할 것입니다. 관련 코드는 https://github.com/Wuzheng02/ESPP 에서 확인할 수 있습니다.

Original Abstract

Generative UI (GenUI) lets large language models synthesize a complete, renderable interface directly from a natural-language instruction, but evaluating the quality of what they generate remains an open problem. Human evaluation is costly and rater-variant, while LLM-as-a-judge is scalable but reflects only a single implicit viewpoint, unable to capture how different populations of real users actually perceive the same interface. We propose the Evidence-Grounded, Social-Weighted Persona Panel (ESPP), a three-stage GenUI evaluation method in which a panel of psychologically diverse, evidence-grounded personas independently rates a screenshot, exchanges opinions under a trait-derived, semantically-gated bounded-confidence mechanism, and is aggregated via Delphi-inspired social weighting into a single judgment. ESPP tracks human judgment substantially more closely than a naive single-pass judge, raising Pearson $r$ from $0.716$ to $0.922$, and a prompt-ensemble control recovers only about a third of this gap, isolating genuine persona and evidence grounding as the dominant source of improvement. Beyond this fidelity gain, retaining each panelist's individual rating further reveals that user subgroups agree on overall model rankings yet diverge sharply on specific rating dimensions, a structural disagreement a single homogeneous judge would systematically erase. The codes are available at https://github.com/Wuzheng02/ESPP.

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!