2606.27226v1 Jun 25, 2026 cs.AI

질문하고 판단하지 마세요: 해석 가능한 LLM 평가 및 자체 개선을 위한 이진 질문

Ask, Don't Judge: Binary Questions for Interpretable LLM Evaluation and Self-Improvement

Shi-Xiong Zhang
Shi-Xiong Zhang
Citations: 122
h-index: 5
Sambit Sahu
Sambit Sahu
Citations: 87
h-index: 5
Sangwoo Cho
Sangwoo Cho
Tencent US
Citations: 769
h-index: 13
Kushal Chawla
Kushal Chawla
Citations: 8
h-index: 2
Chenyang Zhu
Chenyang Zhu
Citations: 28
h-index: 3
Pengshan Cai
Pengshan Cai
University of Massachusetts, Amherst
Citations: 432
h-index: 7
Zefang Liu
Zefang Liu
CapitalOne
Citations: 344
h-index: 9

LLM 출력 평가 문제는 자연어 처리 분야에서 큰 걸림돌입니다. 인간 평가가 비용이 많이 들고 시간이 오래 걸리며, 어휘 기반 지표는 개방형 생성에 대한 인간 판단과 낮은 상관관계를 보이며, 전체적인 LLM 평가 모델은 종종 디버깅하기 어려운 불투명한 점수를 제공합니다. 본 논문에서는 평가 기준을 원자적 이진 질문으로 분해하고, 결과 판정을 해석 가능한 다차원 점수로 집계하는 프레임워크인 BINEVAL을 제안합니다. 주어진 작업 프롬프트에 대해 메타-프롬프트는 세분화된 평가 질문을 생성하고, LLM은 각 출력에 대해 독립적으로 이러한 질문에 답변하여 투명한 질문 수준의 피드백과 함께 보정된 전체 점수를 제공합니다. 이러한 분해는 평가를 더 쉽게 검토하고 진단할 수 있게 하며, 프롬프트 개선에 직접 활용될 수 있습니다. SummEval, Topical-Chat 및 QAGS 데이터셋에서 BINEVAL은 UniEval 및 G-Eval을 포함한 강력한 기준 모델과 동등하거나 우수한 성능을 보이며, 특히 사실 일관성 벤치마크인 QAGS에서 뛰어난 결과를 얻었습니다. 기존 LLM 평가 모델보다 인간 판단과의 상관관계가 더 높을 뿐만 아니라, BINEVAL은 인간 점수 분포와 더 잘 일치하며 이전 모델에서 흔히 나타나는 천장 효과를 피하여, 경계선과 명백하게 결함 있는 출력 간의 구분을 개선합니다. 또한, 동일한 질문 수준의 피드백이 반복적인 프롬프트 최적화를 지원한다는 것을 보여줍니다. 요약 작업 및 IFBench 생성 작업에 대한 평가 프롬프트를 자체 업데이트 및 교차 모델 업데이트 설정을 통해 개선할 수 있었습니다. 종합적으로 BINEAL은 강력한 실증 성능과 함께 실제 진단 및 최적화 가치를 제공하는 작업에 구애받지 않고, 학습이 필요 없으며 해석 가능한 평가 프레임워크입니다.

Original Abstract

Evaluating LLM outputs remains a major bottleneck in NLP: human evaluation is expensive and slow, lexical metrics correlate poorly with human judgments on open-ended generation, and holistic LLM judges often produce opaque scores that are hard to debug. We propose BINEVAL, a framework that decomposes evaluation criteria into atomic binary questions and aggregates the resulting verdicts into interpretable, multi-dimensional scores. Given a task prompt, a meta-prompt generates fine-grained evaluation questions, and an LLM answers them independently for each output, yielding transparent question-level feedback together with calibrated overall scores. This decomposition makes evaluation easier to inspect, easier to diagnose, and directly usable for prompt improvement. Across SummEval, Topical-Chat, and QAGS, BINEVAL matches or outperforms strong baselines including UniEval and G-Eval, with especially strong results on factual consistency benchmarks such as QAGS. Beyond competitive correlation with human judgments, BINEVAL better matches human score distributions and avoids the ceiling effects common in prior LLM judges, leading to better discrimination between borderline and clearly flawed outputs. We further show that the same question-level feedback supports iterative prompt optimization, improving evaluator prompts on summarization and generation prompts on IFBench under both self-update and cross-model update settings. Overall, BINEVAL provides a task-agnostic, training-free, and interpretable evaluation framework that combines strong empirical performance with practical diagnostic and optimization value.

1 Citations
0 Influential
6.5 Altmetric
33.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!