MMShopBench: 다중 모드 및 다중 대화 쇼핑 에이전트를 위한 실제 로그 기반 벤치마크
MMShopBench: A Real-Log Benchmark for Multimodal, Multi-Turn Shopping Agents
온라인 쇼핑객들은 점점 더 많은 AI 쇼핑 어시스턴트를 활용하며, 이미지와 다중 대화를 통해 텍스트만으로는 표현하기 어려운 제품 요구사항을 표현하고 구체화합니다. 그러나 기존의 벤치마크는 대부분 텍스트 기반 또는 합성된 요청에 의존하여, 이미지와 언어를 통해 복합적으로 표현되는 실제 쇼핑 요구사항을 제대로 반영하지 못합니다. 본 연구에서는 다중 모드 및 다중 대화 쇼핑 에이전트를 위한 최초의 실제 로그 기반 벤치마크인 MMShopBench를 소개합니다. 신중하게 정제되고 수동으로 주석이 달린 쇼핑 로그로 구축된 MMShopBench는 각 요청의 구매 의도와 필수 제품 요구사항에 대한 정확한 정보를 제공합니다. 에이전트는 사용자의 이미지 및 다중 대화로부터 이러한 요구사항을 종합적으로 추론하고, 이미지 및 텍스트 검색을 통해 후보 제품을 검색하며, 각 후보 제품의 이미지와 구조적 속성을 사용하여 모든 요구사항을 충족하는지 확인해야 합니다. 우리는 대표적인 오픈 소스 및 독점 모델을 증거 기반의 다중 모드 프로토콜을 사용하여 평가하고, 오픈 소스 모델을 미세 조정하기 위한 보조 학습 데이터 세트를 구축했습니다. 재현 가능한 실험을 위해 오프라인 쇼핑 샌드박스를 구축했으며, 미세 조정은 오픈 소스 모델과 선도적인 독점 모델 간의 성능 격차를 크게 줄여주어, 우리의 학습 데이터가 효과적임을 입증합니다.
Online shoppers increasingly turn to AI shopping assistants, using images and multi-turn dialogue to express and refine product needs that are difficult to articulate in text alone. However, existing benchmarks largely rely on text-only or synthetic requests, underrepresenting complex real-world shopping requirements jointly expressed through images and language. We introduce MMShopBench, the first real-log benchmark for multimodal, multi-turn shopping agents. Built from carefully cleaned and manually annotated shopping logs, MMShopBench provides ground-truth annotations of each request's purchase intent and mandatory product requirements. Agents must infer these requirements jointly from user images and multi-turn dialogue, retrieve candidate products through image and text search, and verify that each candidate satisfies all requirements using its product images and structured attributes. We evaluate representative open-source and proprietary models using an evidence-grounded multimodal protocol and construct a companion training set for fine-tuning an open-source model. To ensure reproducible experimentation, we build an offline shopping sandbox, where fine-tuning substantially narrows the performance gap between our open-source model and leading proprietary models, demonstrating the effectiveness of our training data.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.