2607.05943v1 Jul 07, 2026 cs.AI

SearchEyes: 검색 세계 시뮬레이션을 통한 최첨단 다중 모드 심층 검색 인텔리전스

SearchEyes: Towards Frontier Multimodal Deep Search Intelligence via Search World Simulation

Zhengbo Jiao
Zhengbo Jiao
Citations: 24
h-index: 2
Kaituo Feng
Kaituo Feng
Citations: 830
h-index: 11
Yilei Jiang
Yilei Jiang
Citations: 215
h-index: 9
Qianshan Wei
Qianshan Wei
Citations: 44
h-index: 2
Tianyi Jiang
Tianyi Jiang
Citations: 16
h-index: 3
Yunpu Ma
Yunpu Ma
Citations: 186
h-index: 4
Xiangyu Yue
Xiangyu Yue
Citations: 180
h-index: 9
Qunzhong Wang
Qunzhong Wang
Citations: 50
h-index: 3
Yangfu Li
Yangfu Li
Citations: 13
h-index: 2
Jiapeng Li
Jiapeng Li
Citations: 15
h-index: 2
Rui Huang
Rui Huang
Citations: 0
h-index: 0
Juanxi Tian
Juanxi Tian
Citations: 117
h-index: 3
Tailai Chen
Tailai Chen
Citations: 85
h-index: 5
Chuan Xiao
Chuan Xiao
Citations: 3
h-index: 1
Shanyu Rong
Shanyu Rong
Citations: 251
h-index: 6
YiMing Cheng
YiMing Cheng
Citations: 55
h-index: 3
Yanhan Zhou
Yanhan Zhou
Citations: 0
h-index: 0
Yifan Zhang
Yifan Zhang
Citations: 0
h-index: 0

다중 호프 추론을 수행하는 다중 모드 검색 에이전트를 학습시키는 것은 근본적인 구조적 단절로 인해 여전히 어려운 과제입니다. 기존 파이프라인은 학습 데이터, 검색 환경 및 보상 신호를 독립적으로 구성하므로 합성된 구조 메타데이터가 버려지고, 환경은 재현 불가능한 외부 엔진에 의존하며, 강화학습(RL) 보상은 경로 수준에서 희소하게 유지됩니다. 본 논문에서는 유형화된 지식 그래프를 기반으로 모든 구성 요소를 통합하는 *시뮬레이션된 검색 세계*의 핵심으로 작동하는 **SearchEyes**를 제시합니다. 우리는 Wikidata5M의 시각-지식 교집합에서 제약 조건이 있는 다중 호프 경로를 샘플링하기 위한 **인지-지식 체인(Perception-Knowledge Chains, PKC)**을 제안하여 호프 수준의 엔티티 메타데이터를 유지하며, 이는 자체적으로 완전한 검색 세계를 정의하고 단계별 보상 앵커 역할을 동시에 수행합니다. 또한, 별도로 학습된 프로세스 보상 모델 없이 단계별 신용 할당을 위해 이러한 앵커를 재사용하는 **호프-앵커 정책 최적화(Hop-Anchored Policy Optimization, HaPO)**를 제안합니다. 여섯 가지 다중 모드 지식 집약적인 벤치마크에 대한 실험 결과, SearchEyes는 오픈 소스 다중 모드 검색 에이전트 중에서 최고 성능을 달성했으며, SearchEyes-27B는 가장 강력한 오픈 소스 기준 모델보다 평균 6.2 포인트 향상된 결과를 보였습니다.

Original Abstract

Training multimodal search agents to perform multi-hop reasoning remains challenging due to a fundamental structural disconnect: existing pipelines construct training data, search environments, and reward signals independently, causing synthesized structural metadata to be discarded, environments to rely on irreproducible external engines, and RL rewards to remain sparse at the trajectory level. We present \textbf{SearchEyes}, which uses a typed knowledge graph as the backbone of a \emph{simulated search world} that unifies all three components. We propose \textbf{Perception-Knowledge Chains (PKC)} to sample constrained multi-hop paths over the visual-knowledge intersection of Wikidata5M, retaining hop-level entity metadata that simultaneously defines a self-contained search world and step-level reward anchors. We further propose \textbf{Hop-Anchored Policy Optimization (HaPO)}, which reuses these anchors for step-level credit assignment without a separately trained process reward model. Experiments on six multimodal knowledge-intensive benchmarks show that SearchEyes achieves state-of-the-art performance among open-source multimodal search agents, with SearchEyes-27B improving over the strongest open-source baseline by 6.2 points on average.%

1 Citations
0 Influential
5.5 Altmetric
28.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!