MetaResearcher: 적대적 가상 환경에서의 자기 성찰 강화 학습을 통한 심층 연구 확장
MetaResearcher: Scaling Deep Research via Self-Reflective Reinforcement Learning in Adversarial Virtual Environments
심층 연구 에이전트는 자율적인 정보 수집 및 통합 능력에서 놀라운 잠재력을 보여주었지만, 훈련은 여전히 시뮬레이션 환경의 정적인 특성, 사실 검색에만 국한된 작업 설계, 그리고 결과 기반 강화 학습의 비효율성으로 인해 제약을 받습니다. 본 연구에서는 심층 연구 에이전트 훈련을 네 가지 상호 보완적인 측면에서 확장하는 새로운 프레임워크인 MetaResearcher를 제안합니다. 첫째, 우리는 시간적 역동성과 적대적인 허위 정보를 학습 환경에 주입하여 에이전트가 정보 출처의 신뢰성 평가 및 시간적 충돌 해결 능력을 개발하도록 하는 '진화하는 가상 세계'를 도입했습니다. 둘째, 우리는 가설 생성 및 모순 해결과 같이 단순한 사실 검색을 넘어 실제 연구 행동으로 이어지도록 설계된 '탐색 중심 작업'을 제시합니다. 셋째, GRPO 프레임워크 내에서 자기 성찰 메타-보상 메커니즘을 제안하여 답변 정확성, 검색 경로 효율성, 분석 깊이 및 도구 활용 다양성을 동시에 최적화하고, 기존 연구에서 관찰된 반복적인 행동 루프 문제를 직접 해결합니다. 넷째, 우리는 전문화된 탐색(Scout), 필터링(Filter) 및 합성(Synthesizer) 모델로 구성된 이질적인 다중 에이전트 군집 아키텍처를 도입하여 협력적 연구 전략을 강화 학습을 통해 학습하도록 합니다. LiteResearcher 인프라를 기반으로 구축된 MetaResearcher는 훈련에 거의 비용이 들지 않으면서 GAIA 및 Xbench-DS와 같은 벤치마크 성능과 적대적인 조건 하에서의 인식론적 견고성을 크게 향상시키는 것을 목표로 합니다. 우리는 프레임워크 설계, 훈련 방법론 및 계획된 실험 검증 결과를 제시합니다.
Deep research agents have demonstrated remarkable capabilities in autonomous information gathering and synthesis, yet their training remains constrained by the static nature of simulated environments, the limits of fact-retrieval-only task designs, and the inefficiency of outcome-based reinforcement learning. In this work, we propose MetaResearcher, a novel framework that scales deep research agent training across four synergistic dimensions. First, we introduce an Evolving Virtual World that injects temporal dynamics and adversarial misinformation into the training environment, forcing agents to develop source credibility assessment and temporal conflict resolution skills. Second, we design Discovery-Oriented Tasks -- including hypothesis generation and contradiction resolution -- that transcend simple fact retrieval and push agents toward genuine research behaviors. Third, we propose a Self-Reflective Meta-Reward mechanism within the GRPO framework that jointly optimizes for answer correctness, search path efficiency, reflection depth, and tool call diversity, directly addressing the repetitive action loop problem observed in prior work. Fourth, we introduce a Heterogeneous Multi-Agent Swarm architecture comprising specialized Scout, Filter, and Synthesizer models that learn collaborative research strategies through coordinated reinforcement learning. Built upon the LiteResearcher infrastructure, MetaResearcher requires zero marginal API cost for training while targeting substantial improvements in both benchmark performance (GAIA, Xbench-DS) and epistemic robustness under adversarial conditions. We present the complete framework design, training methodology, and planned experimental validation.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.