2606.10803v1 Jun 09, 2026 cs.CL

API를 넘어: 물리적 도구 사용에서 거대 다중 모드 언어 모델(MLLM)의 한계 탐구

Beyond APIs: Probing the Limits of MLLMs in Physical Tool Use

Zhixin Ma
Zhixin Ma
Citations: 45
h-index: 4
Wenjie Li
Wenjie Li
Citations: 23
h-index: 3
Yongqing Li
Yongqing Li
Citations: 260
h-index: 7
Yutong Zhou
Yutong Zhou
Citations: 15
h-index: 2
C. Ngo
C. Ngo
Citations: 33
h-index: 4

다중 모드 대규모 언어 모델(MLLM)은 디지털 API 활용에 뛰어나며, 점점 더 로봇에게 실제 세계와의 상호 작용을 지시하는 '두뇌' 역할을 수행하여 인공지능 시스템에 구현됩니다. 이러한 시스템에서 핵심적인 기능은 물리적 도구의 사용이며, 이는 MLLM이 인간을 돕는 데 중요한 역할을 합니다. 그러나 MLLM의 물리적 도구 사용 능력은 아직 충분히 연구되지 않았습니다. 이 문제를 해결하기 위해, 우리는 MLLM이 실제 세계 시나리오를 이해하고, 물리적 도구를 식별하며, 그 사용 계획을 수립하는 능력을 평가하도록 설계된 최초의 물리적 도구 사용 벤치마크인 PhysTool-Bench를 소개합니다. PhysTool-Bench는 제조, 전기 작업, 농업 및 의료 등 다양한 분야에 걸쳐 실제 물리적 도구 2,678개를 포함하는 2,510개의 질의로 구성됩니다. 구체적으로, 모델은 다음 두 가지 주요 측면에서 평가됩니다. 1) 장면 내 모든 물리적 도구를 인식하고, 2) 지시 및 시각적 맥락에 따라 도구 선택 및 사용 순서를 계획합니다. 13개의 선도적인 MLLM을 대상으로 한 실험 결과, 가장 우수한 모델(Gemini-3.1-Pro)조차도 장면 내 도구의 58.7%만을 식별하고 전체 질의의 21.0%만을 완수했습니다. 분석 결과, MLLM은 현실적인 장면에서 도구를 인식하는 데 어려움을 겪으며, 계획 단계에서의 성능 저하가 더욱 두드러지는 것으로 나타났습니다. 이는 인지된 도구를 작업 의미에 연결하는 기능적 상식의 부족을 보여주며, 실질적인 인공지능 시스템 개발의 중요한 난관을 지적합니다.

Original Abstract

Multimodal Large Language Models (MLLMs) excel at utilizing digital APIs and increasingly serve as the "brain" of embodied AI, instructing robots to interact with the physical world. In such embodied settings, a central capability is the use of physical tools, which underpins MLLMs' ability to assist humans in real-world tasks. Despite the importance, MLLMs' proficiency in physical tool use remains largely unexplored. To address this gap, we introduce PhysTool-Bench, the first physical tool-use benchmark designed to evaluate MLLMs' ability to comprehend real-world scenarios, identify physical tools, and plan their use. PhysTool-Bench comprises 2,510 queries over 2,678 real-world physical tools spanning diverse domains, including manufacturing, electrical work, agriculture, and healthcare. Concretely, models are evaluated along two primary dimensions: 1) recognizing all physical tools present in the scene, and 2) planning the tool selection and use sequence based on the instruction and visual context. Across 13 leading MLLMs, even the strongest model (Gemini-3.1-Pro) identifies only 58.7% of tools in a scene and completes merely 21.0% of queries end-to-end. Our analysis reveals a two-level deficit: MLLMs struggle to perceive tools in realistic scenes, and the much larger drop at the planning stage further indicates a lack of functional commonsense for mapping perceived tools onto task semantics, pinpointing a critical bottleneck for the development of practical embodied AI.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!