Beacon: 언제, 어떻게 에이전트 기반 시각적 추론을 수행할 것인가
Beacon: Knowing When and How to Perform Agentic Visual Reasoning
에이전트 기반 시각적 추론의 근본적인 목표는 다중 모드 대규모 언어 모델(MLLM)의 복잡한 작업 성공률을 향상시키는 것이며, 단순히 정교하지만 비효율적인 추론 패러다임을 제공하는 것 이상입니다. 본 연구에서는 도구 사용의 두 가지 핵심 측면, 즉 모드 적응성(Mode Adaptiveness, MA)과 도구 효과(Tool Effect, TE)를 통해 에이전트 기반 시각적 추론을 재검토합니다. 모드 적응성은 MLLM이 도구가 실제로 필요한 상황을 인식하고 이에 따라 도구를 호출할 수 있는지 여부를 나타내며, 불필요한 계산 오버헤드를 줄이는 동시에 도구 지원이 필요한 어려운 문제에 대한 성능을 향상시킵니다. 도구 효과는 실제 도구 사용의 영향을 특징지우며, 도구는 텍스트만으로는 해결할 수 없는 문제에서 모델의 기능을 확장해야 하지만, 모델이 이미 도구 없이도 해결할 수 있는 문제에서는 추가적인 오류를 발생시키지 않아야 합니다. 우리는 이러한 두 가지 특성을 정량적으로 분석하고 경험적으로 밝혀낸 결과, 기존 에이전트 기반 시각적 추론 모델은 제한된 모드 적응성을 보이는 반면, 어려운 예제에서 도구 사용으로 얻는 이점은 모델이 이미 해결할 수 있는 쉬운 예제에서 발생하는 부정적인 영향에 의해 크게 상쇄된다는 것을 확인했습니다. 이러한 관찰을 바탕으로, 우리는 전반적인 성능, 향상된 모드 적응성 및 진정한 도구 기반 성능 향상을 달성하는 새로운 에이전트 기반 시각적 추론 모델인 Beacon을 제안합니다. Beacon의 핵심은 강화 학습 단계에서 필요성을 인지한 적응형 보상(Necessity-Aware Adaptive Reward)과 힌트를 활용한 기능 확장 메커니즘(Hint-Guided Capability Expansion mechanism)이며, 이는 각각 작업의 필요성에 따라 도구 호출을 유도하고 모델이 가장 어려운 문제에 대한 도구 사용 능력을 강화합니다. 다양한 벤치마크에서 수행된 광범위한 실험은 Beacon의 강력한 전반적인 성능과 모드 적응성 및 도구 효과 측면에서의 상당한 개선을 입증했습니다.
The fundamental goal of agentic visual reasoning is to improve the success rate of multimodal large language models (MLLMs) on complex tasks, rather than merely equipping them with a sophisticated yet inefficient reasoning paradigm. In this work, we rethink agentic visual reasoning through two key dimensions of tool use: Mode Adaptiveness (MA) and Tool Effect (TE). Mode Adaptiveness characterizes whether an MLLM can recognize when tools are truly necessary and invoke them accordingly, thereby avoiding unnecessary computational overhead while improving performance on challenging problems that require tool assistance. Tool Effect characterizes the actual impact of tool use: tools should extend the model's capabilities on problems unsolvable through text-only reasoning, while avoiding additional errors on problems that the model can already solve without tools. We conduct a comprehensive analysis to quantify these two properties and empirically reveal that existing agentic visual reasoning models exhibit limited Mode Adaptiveness, while the gains produced by tool use on hard examples are largely offset by the harm introduced on easy examples that the models can already solve. Motivated by these observations, we propose Beacon, a novel agentic visual reasoning model that achieves stronger overall performance, improved Mode Adaptiveness, and genuine tool-induced performance gains. At the core of Beacon are the Necessity-Aware Adaptive Reward and the Hint-Guided Capability Expansion mechanism in the reinforcement learning stage, which respectively encourage adaptive tool invocation based on task necessity and strengthen the model's tool-use capability on the most challenging problems. Extensive experiments across diverse benchmarks demonstrate the strong overall performance of Beacon and its substantial improvements in both Mode Adaptiveness and Tool Effect.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.