MapSatisfyBench: 행동 기반의 잠재적 의사 결정 요소를 활용한 만족도 중심 지도 에이전트 성능 평가
MapSatisfyBench: Benchmarking Satisfaction-Aware Map Agents through Behavior-Grounded Implicit Decision Factors
최근 대규모 언어 모델 기반 에이전트가 지도 서비스에 점점 더 많이 통합되고 있습니다. 지도 서비스는 전문적인 작업 환경보다는 일상생활 시나리오에 내장되어 있기 때문에, 사용자는 종종 비공식적으로 요구사항을 표현하며, 이는 많은 암묵적인 필요사항(즉, 사용자 만족도에 중요한 잠재적 의사 결정 요인)을 포함하는 불명확한 쿼리로 이어집니다. 명확화는 이러한 문제를 완화하는 효과적인 방법이지만, 일상적인 상호 작용에서 사용자의 부담을 증가시킵니다. 따라서, 능숙한 에이전트는 이용 가능한 정보 소스에서 이러한 요인을 사전에 적극적으로 파악해야 합니다. 그러나 이 능력을 평가하는 것은 어렵습니다. 첫 번째 과제는 어떤 잠재적 의사 결정 요인이 평가에 적합한지 결정하는 것입니다. 요소가 평가 가능하려면 사용자 수용에 영향을 미치고, 에이전트가 응답하기 전에 이용 가능한 정보로부터 파악될 수 있어야 합니다. 두 번째로, 사용자의 만족도는 단일의 정답으로 신뢰성 있게 표현할 수 없으므로, 만족도와 관련된 요소를 객관적이고 측정 가능한 평가 목표로 변환하는 벤치마크가 필요합니다. 이러한 과제를 해결하기 위해, 우리는 행동 연쇄 증거로부터 완전한 사용자 요구사항을 재구성하고, 잠재적 의사 결정 요인을 식별하며, 사전 쿼리 증거로 뒷받침되는 요소만 유지하는 '복원-식별-필터' 프레임워크를 제안합니다. 이 방법론을 바탕으로, 우리는 대규모의 실제 사용자 데이터를 활용하여 MapSatisfyBench를 구축하고, 5가지 차원에서 정답을 주석 처리하여 만족도 중심 지도 에이전트의 전체 연쇄 평가를 가능하게 합니다. 실험 결과, 현재 에이전트는 명시적인 작업 완료 측면에서는 일반적으로 잘 수행되지만, 암묵적인 의사 결정 요인을 충족시키고 만족도를 높이기 위한 증거를 사전에 확보하는 데는 한계가 있는 것으로 나타났습니다. 이러한 연구 결과는 MapSatisfyBench가 지도 에이전트 평가의 초점을 작업 완료에서 만족도 중심의 공간적 의사 결정으로 전환하는 벤치마크로 자리매김할 수 있음을 보여줍니다.
Large language model agents are increasingly integrated into map services. Since map services are embedded in everyday-life scenarios rather than professional task settings, users often express their needs informally, resulting in underspecified queries with many unspoken needs, namely, implicit decision factors that are critical for user satisfaction. Although clarification is an effective way to mitigate this issue, it increases user burden in daily interaction, and a capable agent should first proactively recover such factors from available information sources. However, evaluating this ability is challenging. The first challenge is to determine which implicit decision factors are suitable for evaluation. A factor is evaluable only if it affects user acceptance and can be recovered from information available to the agent before it responds. Second, user satisfaction cannot be reliably represented by a single reference answer, requiring a benchmark that converts satisfaction-relevant factors into objective and quantifiable evaluation targets. To address these challenges, we propose a restore-identify-filter framework that reconstructs complete user needs from behavior-chain evidence, identifies implicit decision factors, and retains only those supported by pre-query evidence. Building on this methodology, we construct MapSatisfyBench from large-scale, real-world anonymized user data and annotate ground truth from five dimensions and enables full-chain evaluation of satisfaction-aware map agents. Experiments show that current agents generally perform well on explicit task completion, but remain limited in satisfying implicit decision factors and proactively acquiring the evidence needed for satisfaction-aware decisions. These findings establish MapSatisfyBench as a benchmark for shifting map-agent evaluation from task completion toward satisfaction-aware spatial decision making.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.