2607.28956v1 Jul 31, 2026 cs.AI

MerchantBench: 전자상거래 운영에서 장기적 일관성을 위한 LLM 에이전트 성능 평가

MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations

Tian Pan
Tian Pan
Citations: 21
h-index: 2
Di Weng
Di Weng
Citations: 891
h-index: 16
Zhaolu Kang
Zhaolu Kang
Citations: 30
h-index: 2
Chengfu Huo
Chengfu Huo
Citations: 264
h-index: 6
Qiming Shi
Qiming Shi
Citations: 23
h-index: 2
Yulong Tao
Yulong Tao
Citations: 0
h-index: 0
Linbo Jin
Linbo Jin
Citations: 0
h-index: 0
Yibo Dou
Yibo Dou
Citations: 0
h-index: 0
J. Zhu
J. Zhu
Citations: 75
h-index: 3
Shaokang Fu
Shaokang Fu
Citations: 0
h-index: 0
Chengyu Wang
Chengyu Wang
Citations: 0
h-index: 0
Siyue Li
Siyue Li
Citations: 0
h-index: 0
Y. Cheng
Y. Cheng
Citations: 1
h-index: 1

대규모 언어 모델(LLM) 에이전트는 자율적인 도구 사용자로서 점점 더 많이 평가되고 있지만, 대부분의 벤치마크는 즉각적인 성공 기준을 가진 제한된 작업에 초점을 맞추고 있습니다. 실제 환경에서는 장기적 일관성이 필요하며, 이는 축적된 증거에 맞춰 의사 결정을 조정하면서도 광범위한 기간 동안 목적 있는 행동을 유지하는 능력입니다. 이 능력을 평가하려면, 행동이 미래의 선택을 제한하고, 피드백이 다양한 지연 시간으로 제공되며, 비일관적인 행동이 측정 가능한 누적 효과를 발생시키는 지속적인 환경이 필요합니다. 판매자 측 전자상거래는 제품 소싱, 상품 등록 및 가격 제어, 현금 흐름 관리, 그리고 다양한 지연 시간을 가진 피드백 적응에 대한 반복적이고 상호 의존적인 결정 과정을 통해 이러한 평가에 적합한 환경을 제공합니다. 우리는 98,843개의 실제 전자상거래 제품 기록을 기반으로 구축되고 에이전트와의 상호 작용을 위한 26가지 도구를 갖춘 365일 단위의 주문 시뮬레이션인 MerchantBench를 소개합니다. MerchantBench는 즉시 관찰 가능한 공급업체 관련 이벤트와 지연된 주문 결과 데이터를 결합하여, 에이전트가 개별 주문의 전체 라이프사이클을 추적하고 이전 의사 결정을 재검토하도록 요구합니다. 우리는 48개의 실행 환경에서 2가지 에이전트 프레임워크를 사용하여 8개의 LLM을 평가했으며, 각 실행은 365일 동안 시뮬레이션되었습니다. 우리의 결과는 최신 LLM과 인간 참가자 간에 상당한 격차가 있음을 보여주며, 가장 성능이 우수한 LLM 구성은 인간 참가자가 달성한 평균 최종 순 자산의 27.3%에 불과했습니다.

Original Abstract

Large language model agents are increasingly evaluated as autonomous tool users, yet most benchmarks focus on bounded tasks with immediate success criteria. Real-world deployments often require Long-Term Coherence, the capacity to preserve purposeful behavior across extended horizons while adapting decisions to accumulated evidence. Evaluating this capacity requires a persistent environment in which actions constrain future choices, feedback arrives at heterogeneous delays, and incoherent behavior produces measurable cumulative effects. Seller-side e-commerce provides a suitable setting for this evaluation through recurrent and interdependent decisions over Product Sourcing, Listing and Pricing Control, Cash-Flow Management, and Mixed-Latency Feedback Adaptation. We introduce MerchantBench, a 365-day order-level simulation grounded in 98,843 real e-commerce product records and equipped with 26 tools for agent interaction. MerchantBench couples promptly observable Upstream Supplier Events with delayed Downstream Order Outcomes, requiring agents to follow individual order lifecycles and revisit earlier decisions. We evaluate eight LLMs under two agent frameworks in 48 runs, each spanning 365 simulated days. Our results reveal a substantial gap between even the latest LLMs and human participants, with the best LLM configuration attaining only 27.3\% of the mean final net assets achieved by human participants.

0 Citations
0 Influential
8 Altmetric
40.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!