DeepVoyager-VL: 시각 정보를 활용한 장기적인 다중 모드 에이전트 검색 시스템 개발
DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents
다중 모드 대규모 언어 모델(MLLM)은 시각적 이해 및 추론 능력을 향상시켰지만, 정적인 파라미터 지식으로는 지식 집약적이거나 동적으로 변화하는 개방형 문제를 해결하는 데 한계가 있습니다. 이러한 제한을 극복하기 위해, 다중 모드 심층 검색은 개방형 환경에서의 정보 접근에 있어 중요한 방향으로 부상했으며, 단일 턴의 사실 기반 검색에서 시각적 증거에 의해 안내되는 장기적이고 다중 턴 검색으로 발전하고 있습니다. 그러나 기존 방법들은 일반적으로 시각 정보를 입력 또는 답변 단계로 제한하며, 중간 추론 과정에서의 역할을 간과하고 있으며, 장기적인 상호 작용을 위한 설계가 부족합니다. 결과적으로, 시각적 증거는 지속적인 검색을 이끄는 데 거의 활용되지 않아, 상호 작용 깊이와 추론 범위를 제약하게 됩니다. 이러한 한계점을 해결하기 위해, 우리는 시각 정보를 활용하는 장기적인 다중 모드 심층 검색 프레임워크인 DeepVoyager-VL을 제안합니다. 구체적으로, 데이터 합성 및 중간 단계의 시각적 의존성과 긴 추론 체인을 갖는 문제들을 생성하기 위한 다중 모드 이벤트 그래프를 구축했습니다. 또한, 능동적인 시각 정보 획득과 필요에 따른 이미지 로딩을 위한 에이전트 프레임워크를 설계했습니다. 마지막으로, 강화 학습 없이 합성된 데이터셋으로 모델을 미세 조정했습니다. 열 가지의 다중 모드 검색 벤치마크에서 수행한 광범위한 실험 결과는 제안하는 방법의 효과성을 입증합니다.
Multimodal large language models (MLLMs) have advanced visual understanding and reasoning, yet their static parametric knowledge limits their ability to address knowledge-intensive and dynamically evolving open-world problems. To move beyond this limitation, multimodal deep search has emerged as a key direction for open-world information access, evolving from single-turn factual retrieval toward long-horizon, multi-turn search guided by visual evidence. However, existing methods typically confine vision to the input or answer stage, overlooking its role in intermediate reasoning, and lack designs tailored to long-horizon interaction. Consequently, visual evidence rarely drives continued retrieval, constraining both interaction depth and reasoning span. To address these limitations, we propose DeepVoyager-VL, a long-horizon multimodal deep-search framework for vision-in-the-loop search. Specifically, we construct a multimodal event graph to drive data synthesis, yielding problems with intermediate visual dependencies and long reasoning chains. We then design an agent framework for active visual acquisition and on-demand image loading. Finally, we fine-tune models on the synthesized data without reinforcement learning. Extensive experiments across ten multimodal search benchmarks demonstrate the effectiveness of our method.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.