Hy-MultiTurn: 심층 다중 회화 이해를 위한 6차원 벤치마크
Hy-MultiTurn: A Six-Dimensional Benchmark for Deep Multi-Turn Dialogue Understanding
챗봇 및 에이전트와의 장기간 다중 회화 상호작용은 이제 흔한 현상이며, 올바른 응답을 위해서는 이전 세부 사항을 기억하고, 후속 수정 사항을 추적하며, 의도된 객체 또는 지칭 대상을 식별하고, 필요한 조건이 충족되지 않으면 행동을 억제해야 합니다. 기존의 다중 회화 벤치마크는 일반적으로 짧은 교환에 초점을 맞추며, 특히 중국어에서 장기간의 다중 회화 상호작용에서 이러한 능력을 충분히 평가하지 못하며, 모델이 실패하는 이유와 방법에 대한 제한적인 통찰력을 제공합니다. 이러한 한계를 해결하기 위해, 우리는 실제 챗봇 오류를 분석하여 여섯 가지 반복되는 메커니즘을 식별하고 이를 사용하여 심층 다중 회화 이해를 위한 중국어 벤치마크인 Hy-MultiTurn의 여섯 가지 제어된 평가 모드를 정의했습니다. 이 여섯 가지 모드는 제약 조건 기억, 정확한 실행, 제약 조건 합성, 객체 위치 파악, 행동 억제 및 지칭 해소를 평가합니다. 우리는 12~76턴으로 구성된 209개의 제어된 작업을 만들었으며, 대화 길이, 관련 없는 주제로 인한 혼란 및 구어적인 표현이 추가적인 어려움을 더합니다. 22개의 최첨단 모델 구성에 대한 평가는 Hy-MultiTurn이 전반적으로 어려운 것을 보여줍니다. 가장 강력한 전체 구성인 GPT-5.5조차도 응답의 41.1%에서만 모든 요구 사항을 충족하며, 어떤 모델도 여섯 가지 모드 모두에서 최고의 성능을 보이지 않습니다.
Long-running multi-turn interactions with chatbots and agents are now common, and a correct response often depends on remembering earlier details, tracking later revisions, identifying intended objects or referents, and withholding action when required conditions are unmet. Existing multi-turn benchmarks typically cover short exchanges and do not fully evaluate these capabilities in long multi-turn interactions, particularly in Chinese, while offering limited insight into how and why models fail. To address these limitations, we analyze real chatbot failures to identify six recurring mechanisms and use them to define six controlled evaluation modes in Hy-MultiTurn, a Chinese benchmark for deep multi-turn dialogue understanding. The six modes evaluate constraint memory, precise execution, constraint synthesis, object localization, action suppression, and reference resolution. Across the six modes, we construct 209 controlled tasks spanning 12-76 turns, with dialogue length, irrelevant-topic distraction, and colloquial phrasing adding further difficulty. Evaluation of 22 frontier model configurations shows that Hy-MultiTurn is broadly challenging, as even GPT-5.5, the strongest overall configuration, satisfies all requirements in only 41.1 percent of responses and no model performs best in all six modes.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.