2608.06329v1 Aug 06, 2026 cs.CL

벤치마크의 성능 평가: 대화형 에이전트를 위한 벤치마크 평가

Benchmarking the Benchmarks: Evaluating Benchmarks for Conversational Agents

Noam Koren
Noam Koren
Citations: 27
h-index: 3
Roy Bar-Haim
Roy Bar-Haim
Citations: 905
h-index: 5
Abigail Goldsteen
Abigail Goldsteen
Citations: 239
h-index: 7

작업 지향적인 대화형 에이전트는 일반적으로 선별되거나 자동으로 생성된 벤치마크를 사용하여 평가되지만, 벤치마크의 품질은 드물게 평가됩니다. 품질이 낮은 벤치마크는 일관성 없는 작업, 단순한 시나리오 또는 제한적인 정책 적용 범위를 포함할 수 있으며, 이는 신뢰할 수 없는 평가로 이어질 수 있습니다. 본 연구에서는 LLM(대규모 언어 모델) 심판을 사용하여 벤치마크의 일관성, 복잡성 및 정책 적용 범위를 평가하는 참조 불필요 프레임워크를 제시하며, 동시에 개선점을 파악할 수 있는 실용적인 진단 정보를 제공합니다. 본 연구는 제안된 프레임워크가 독립적인 인간 어노테이션과 일치함을 입증하고, 다양한 성능을 가진 LLM에 의해 생성된 벤치마크 및 의도적으로 품질이 저하된 벤치마크를 평가하여 검증했습니다. 다양한 분야와 심판 모델에서 제안된 지표는 벤치마크의 품질 수준을 일관되게 구분합니다. 또한, 본 연구에서는 수동으로 선별된 벤치마크에도 본 프레임워크가 적용 가능하다는 것을 보여줍니다. 본 연구에서 제시하는 프레임워크는 합성 및 수동으로 생성된 대화형 에이전트 벤치마크를 평가하는 데 유용한 접근 방식을 제공합니다.

Original Abstract

Task-oriented conversational agents are evaluated using curated or automatically generated benchmarks, yet benchmark quality is rarely assessed. Poor benchmarks may contain inconsistent tasks, simplistic scenarios, or limited policy coverage, leading to unreliable evaluations. We introduce a reference-free framework that uses LLM judges to assess benchmark consistency, complexity, and policy coverage, while providing actionable diagnostics of weaknesses. We validate the framework by demonstrating agreement with independent human annotations and by evaluating benchmarks generated by LLMs of varying capabilities, as well as benchmarks subjected to controlled quality-degrading perturbations. Across domains and judge models, the proposed metrics consistently distinguish between benchmark quality levels. We further demonstrate the framework's applicability to manually curated benchmarks. Our framework offers a practical approach for evaluating synthetic and manually curated conversational-agent benchmarks.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!