한 번 더 생각하기, 후회는 줄이기: LLM의 명확화 정책을 평가하는 후회 기반 다중 회차 벤치마크
One More Turn, Less Regret: A Regret-Based Multi-Turn Benchmark for LLMs' Clarification Policies
모호한 사용자 요청은 대화형 LLM 어시스턴트에게 명확화를 순차적인 의사 결정 문제로 만듭니다. LLM은 질문해야 할지, 무엇을 질문해야 할지, 언제 중단할지, 그리고 언제 답변해야 할지를 결정해야 합니다. 본 논문에서는 명확화를 개별 질문의 품질이 아닌 정책 행동으로 평가하는 다중 회차 벤치마크인 RegretBench를 소개합니다. RegretBench는 숨겨진 의도를 활용하여 모호성을 정의하고, 의미-상태 추적을 기반으로 자유로운 상호 작용을 지원하며, 모델이 참조 명확화 정책에 비해 얼마나 많은 가치를 잃는지 측정하는 후회 기반 목표 함수를 도입합니다. 개방형 질의응답 및 제품 추천 시나리오에서의 실험 결과는 최종 성공 여부만으로는 충분하지 않음을 보여줍니다. 유사한 정확도를 가진 모델이라도 효율성, 사용자 행동에 대한 견고성, 중단 결정 측면에서 상당한 차이를 보일 수 있습니다. RegretBench는 의도 해결, 상호 작용 비용, 비효율적인 명확화, 그리고 후회를 종합적으로 측정함으로써 모델이 유용하고 효율적으로 명확화를 수행하는지 여부를 밝혀냅니다. 우리의 결과는 효과적인 명확화가 그럴듯한 질문만으로는 충분하지 않으며, 모델은 올바른 시기에 올바른 질문을 해야 하며, 사용자의 의도된 의미가 분명해지면 중단해야 함을 보여줍니다.
Ambiguous user requests make clarification a sequential decision problem for conversational LLM assistants: they must decide whether to ask, what to ask, when to stop, and when to answer. We introduce RegretBench, a multi-turn benchmark that evaluates clarification as policy behavior rather than isolated question quality. RegretBench provides a hidden-intent formulation of ambiguity, supports free-form interaction grounded in semantic-state tracking, and introduces a regret-based objective that measures how much value a model loses relative to a reference clarification policy. Experiments on open-domain QA and product recommendation scenarios show that final success alone is insufficient, as models with similar accuracy can differ substantially in efficiency, robustness to user behaviors, and stopping decisions. By jointly measuring intent resolution, interaction cost, ineffective clarification, and regret, RegretBench reveals whether models clarify usefully and efficiently. Our results show that effective clarification requires more than plausible questions: models must ask the right question at the right time and stop once the user's intended meaning is clear.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.