2607.11818v1 Jul 13, 2026 cs.CV

MM-ToolSandBox: 시각적 도구 활용 에이전트 평가를 위한 통합 프레임워크

MM-ToolSandBox: A Unified Framework for Evaluating Visual Tool-Calling Agents

Di Feng
Di Feng
Citations: 89
h-index: 5
Afshin Dehghan
Afshin Dehghan
Citations: 729
h-index: 13
Kaixin Ma
Kaixin Ma
Citations: 589
h-index: 12
Alexander Metz
Alexander Metz
Citations: 0
h-index: 0
Eshan Verma
Eshan Verma
Citations: 9
h-index: 2
Jiarui Lu
Jiarui Lu
Citations: 1,404
h-index: 17

본 논문에서는 시각 정보 기반의 도구 활용 에이전트를 평가하기 위한 벤치마크 및 평가 프레임워크인 MM-ToolSandBox를 소개합니다. 이 프레임워크는 16개의 애플리케이션 영역에 걸쳐 500개 이상의 도구를 포함하는 상태 기반 실행 환경을 제공하며, 멀티 이미지, 멀티 턴 작업을 지원합니다. 에이전트는 지속적으로 입력되는 시각 정보를 해석하여 실행 가능한 도구 호출로 변환해야 하며, 현실적인 대화 현상(목표 수정, 오류 수정, 상태 변경)을 처리해야 합니다. 자동화된 시나리오 생성 파이프라인은 정보 흐름 기반 계획 및 다단계 품질 필터링을 통해 다양한 시각적 기반 시나리오를 생성하며, 258개의 인간 검증된 표준 시나리오와 인터랙티브 UI 애플리케이션을 대상으로 하는 50개의 변형 시나리오를 제공합니다. 4B 개방형 가중치 모델부터 최첨단 독점 시스템에 이르기까지 12개의 최신 모델을 평가한 결과, 현재 모델들은 여전히 견고한 시각적 도구 활용 능력이 부족하다는 것을 보여줍니다. 최고 성능의 모델조차도 50% 미만의 성공률을 달성했습니다. 실패 분석 결과, 계획 능력뿐만 아니라 시각적 정확성이 우수한 모델의 주요 병목 현상이라는 사실이 밝혀졌습니다. 전체 실패 원인의 53%가 올바른 작업 흐름에도 불구하고 이미지에서 잘못된 정보를 추출하는 데서 발생합니다. 모델 크기에 따라 계획 능력이 중요한지, 아니면 시각적 인식 능력이 중요한지가 달라지는 경향을 보이며, 이는 다양한 능력 수준의 모델 개선을 위한 근본적으로 다른 연구 방향을 제시합니다. 이 프레임워크와 벤치마크는 https://github.com/apple/ml-mmtoolsandbox 에서 공개적으로 이용할 수 있습니다.

Original Abstract

We introduce MM-ToolSandBox, a benchmark and evaluation framework for visually grounded tool-calling agents. The framework provides a stateful execution environment spanning 500+ tools across 16 application domains, supporting multi-image, multi-turn tasks where agents must ground progressively arriving visual inputs into executable tool calls while handling realistic conversational phenomena (goal revisions, error corrections, state mutations). An automated scenario generation pipeline produces diverse, visually grounded scenarios through information-flow-guided planning and multi-stage quality filtering, yielding 258 human-verified nominal scenarios and 50 variants targeting interactive UI applications. Evaluating 12 state-of-the-art models, from 4B open-weight to frontier proprietary systems, shows that current models still lack robust visual tool-calling capability: even the best model achieves below 50% success rate. Our failure analysis further reveals that visual precision, not only planning, is a primary bottleneck for capable models: 53% of failures stem from incorrect information extraction from images despite otherwise correct task workflows. A planning-to-precision crossover emerges with scale: smaller models fail at deciding what to do, while larger models fail at perceiving what they see, suggesting fundamentally different research directions for improving models at different capability levels. The framework and the benchmark are publicly available at https://github.com/apple/ml-mmtoolsandbox

0 Citations
0 Influential
36.229550745277 Altmetric
0.0 Score
Original PDF
6

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!