SearchArt: 확장 가능한 합성 및 검증 작업을 통한 장기 탐색 에이전트 훈련
SearchArt: Training Long-Horizon Search Agent with Scalable Synthetic and Verified Task
최근 대규모 언어 모델(LLM)의 발전으로 인해, 탐색 에이전트는 광범위한 탐색 및 추론 범위를 가진 복잡한 문제를 자율적으로 해결할 수 있게 되었습니다. 그러나 확장 가능한 장기 작업 부족과 중간 단계의 추론 및 도구 사용 행동을 평가하고 수정하는 어려움 때문에 효과적인 탐색 에이전트를 훈련시키는 것은 여전히 어려운 과제입니다. 본 논문에서는 검증 기반 작업 합성 및 다단계 후속 훈련 파이프라인을 통해 장기 탐색 에이전트를 훈련시키는 확장 가능한 프레임워크인 SearchArt를 소개합니다. SearchArt는 웹 문서와 자동으로 생성된 증거 그래프로부터 다양한 정보 검색 질문-답변 쌍과 해당 검색 경로를 합성하여 복잡한 검색, 연구 및 사용자 지향 작업에 대한 대규모 데이터 세트를 구축합니다. 합성된 데이터의 신뢰성을 보장하기 위해, 우리는 질문-답변 일관성, 경로 품질 및 검색된 증거의 관련성을 동시에 평가하는 검증 파이프라인을 설계했습니다. 검증된 경로는 지도 학습 미세 조정과 강화 학습 기반 정책 최적화를 포함하는 다단계 훈련 과정에 사용됩니다. SearchArt로 훈련된 탐색 에이전트는 적응적인 탐색 계획, 반복적인 증거 집계 및 광범위한 상호 작용 범위를 통한 복잡한 추론 능력을 보여줍니다. 실험 결과는 Qwen3.5-27B 파라미터만 사용하여 SearchArt가 BrowseComp-ZH에서 74.39점, BrowseComp에서 70.06점, Deepresearch-bench에서 52.55점을 기록하여 심층 검색 및 심층 연구 벤치마크에서 최첨단 비공개 에이전트의 성능과 동등하거나 능가하는 것을 보여줍니다.
Recent advances in large language models (LLMs) have enabled search agents to autonomously tackle complex tasks across extended search and reasoning horizons. However, training effective search agents remains challenging due to the lack of scalable and long-horizon tasks, and the difficulty of evaluating and correcting intermediate reasoning and tool-use behaviors. We introduce SearchArt, a scalable framework for training long-horizon search agents through verification-driven task synthesis and a multi-stage post-training pipeline. SearchArt constructs large-scale datasets for complex search-, research- and user-oriented tasks by synthesizing diverse information-seeking QA pairs and corresponding search trajectories from web documents and automatically generated evidence graphs. To ensure the reliability of the synthesized data, we design a verification pipeline that jointly evaluates QA consistency, trajectory quality, and the relevance of retrieved evidence. The verified trajectories are subsequently used in a multi-stage training process comprising supervised fine-tuning and reinforcement learning-based policy optimization. Search agents trained with SearchArt exhibit adaptive search planning, iterative evidence aggregation, and complex reasoning over extended interaction horizons. Experimental results demonstrate that, with only (Qwen3.5-) 27B parameters, SearchArt scores 74.39 on BrowseComp-ZH, 70.06 on BrowseComp, and 52.55 on Deepresearch-bench, matching or surpassing frontier closed-source agents on both deepsearch and deepresearch benchmarks.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.