2606.17698v1 Jun 16, 2026 cs.AI

EComAgentBench: 분산된 숨겨진 의도를 가진 장기 과제에서 쇼핑 에이전트 성능 평가

EComAgentBench: Benchmarking Shopping Agents on Long-Horizon Tasks with Distributed Hidden Intent

Tongxin Li
Tongxin Li
Citations: 11
h-index: 2
Zeyao Du
Zeyao Du
Citations: 12
h-index: 1
Yanci Zhang
Yanci Zhang
Citations: 0
h-index: 0
Haibo Zhang
Haibo Zhang
Citations: 0
h-index: 0

LLM 기반 쇼핑 에이전트가 실제 서비스에 도입되면서, 기존의 벤치마크는 사용자의 요구사항이 어떻게 나타나는지를 제대로 반영하지 못합니다. 이러한 요구사항은 검색어에 암시적으로 표현될 수도 있고, 사용자 프로필에 기록될 수도 있으며, 또는 적절한 질문을 통해야만 드러날 수도 있습니다. 기존 벤치마크는 전체 의도를 미리 제시하고 최종 선택만을 평가하기 때문에, 장기적인 과제를 제시하거나 에이전트가 어떤 요구사항을 놓쳤는지 설명할 수 없습니다. 이러한 문제를 해결하기 위해, 실제 아마존 제품 및 리뷰를 기반으로 한 662개의 작업으로 구성된 EComAgentBench를 소개합니다. 각 작업은 사용자 요구사항을 검색어, 도구로 제한된 프로필, 그리고 미리 정의된 추가 질문에 분산하여 제시하며, 에이전트는 숨겨진 의도를 파악하고, 후보 제품의 속성을 검증하며, 리뷰 증거를 확인한 후 100번의 도구 사용 내에서 단일 제품을 선택해야 합니다. 또한, 각 작업은 유형화되고 출처가 명시된 평가 기준에 따라 채점되며, 모든 실패는 특정 요구사항과 그 출처와 연결됩니다. 시스템 구축 과정은 자동화되어 있으며 신뢰성이 높으며, 텍스트 생성 전에 모든 답변이 코드에 고정되어 있고, 모든 샘플이 검증되었습니다. 7개의 모델을 평가한 결과, 가장 뛰어난 성능을 보이는 모델조차도 전체 정확도가 57.1%에 불과하며, 가시적인 출처에서 얻은 정보보다 숨겨진 출처의 정보를 활용하는 데 있어 평가 만족도가 낮아지는 것을 확인했습니다. 전반적으로, EComAgentBench는 쇼핑 에이전트가 단일 검색을 넘어 장기적인 지원을 제공할 수 있도록 발전시키는 데 필요한 재현 가능한 기반이 될 것이라고 믿습니다.

Original Abstract

As LLM-based shopping agents enter production, existing benchmarks fail to capture how a shopper's requirements arrive: stated implicitly in the query, recorded in a profile, or revealed only when the right question is asked. Benchmarks that expose full intent upfront and grade only the final choice can neither pose this long-horizon challenge nor explain which requirement an agent missed. To address this gap, we introduce EComAgentBench, a benchmark of 662 tasks grounded in real Amazon products and reviews. Each task scatters these requirements across a visible query, a tool-gated profile, and scripted clarification; an agent must uncover hidden intent, verify candidates against attributes and review evidence, and commit to a single product within 100 tool calls. Moreover, typed, source-tagged rubrics grade every task, attributing each failure to a requirement and its source. Construction is automated yet reliable, with every answer fixed in code before any text is generated and every sample validated. Our evaluation of seven models reveals that even the strongest attains only 57.1% overall accuracy, and rubric satisfaction degrades from visible to hidden sources. Overall, we believe EComAgentBench will serve as a reproducible foundation for moving shopping agents from single-query search toward dependable assistance over long horizons.

0 Citations
0 Influential
1 Altmetric
5.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!