2607.21461v1 Jul 23, 2026 cs.AI

AREX: 심층 연구를 위한 재귀적 자기 개선 에이전트에 대한 연구

AREX: Towards a Recursively Self-Improving Agent for Deep Research

Hongjin Qian
Hongjin Qian
Citations: 65
h-index: 4
Zheng Liu
Zheng Liu
Citations: 80
h-index: 4
Yuyang Hu
Yuyang Hu
GSAI
Citations: 203
h-index: 3
Zhicheng Dou
Zhicheng Dou
Citations: 3,020
h-index: 28
Zhang Zhang
Zhang Zhang
Citations: 35
h-index: 3
Shuqi Lu
Shuqi Lu
Citations: 437
h-index: 11
Di He
Di He
Citations: 104
h-index: 3
Kang Liu
Kang Liu
Citations: 220
h-index: 9
Qiwei Ye
Qiwei Ye
Citations: 152
h-index: 5
Zhongyuan Wang
Zhongyuan Wang
Citations: 893
h-index: 11
Lei Xiong
Lei Xiong
Citations: 25
h-index: 4
Wanli Li
Wanli Li
Citations: 341
h-index: 3
Chaofan Li
Chaofan Li
Citations: 571
h-index: 8
Kun Luo
Kun Luo
Citations: 273
h-index: 2
Ziyi Xia
Ziyi Xia
Citations: 278
h-index: 3
Bingyu Yan
Bingyu Yan
Citations: 78
h-index: 2
Jiahao Wang
Jiahao Wang
Citations: 0
h-index: 0
Hui Wang
Hui Wang
Citations: 0
h-index: 0
Hongwang Xiao
Hongwang Xiao
Citations: 0
h-index: 0
Sen Wang
Sen Wang
Citations: 12
h-index: 1
Xiyan Jiang
Xiyan Jiang
Citations: 659
h-index: 3
Yingxia Shao
Yingxia Shao
Citations: 250
h-index: 5
Chaozhuo Li
Chaozhuo Li
Citations: 113
h-index: 3

심층적인 연구는 여러 제약을 동시에 만족하는 답을 찾는 것을 요구합니다. 이러한 답을 찾는 것은 비용이 많이 들지만, 후보 해의 검증은 종종 다루기 쉬운 개별 제약 조건에 따른 확인으로 분해될 수 있습니다. 이러한 탐색-검증의 비대칭성은 연구 에이전트가 단순히 더 오래 검색하는 것 이상을 해야 함을 시사합니다. 즉, 중간 결과를 검증하고 부분적으로 검증된 상태를 사용하여 후속 개선을 안내함으로써 현재 답을 재귀적으로 개선해야 합니다. 우리는 재귀적 자기 개선(RSI)을 위한 심층 연구 에이전트 패밀리인 AREX를 소개합니다. AREX는 증거를 수집하고 임시 답변을 구성하는 내부 연구 루프와, 답변을 제약 조건별로 감사하고, 해결되지 않은 주장을 식별하고, 대상 후속 연구를 시작하는 외부 자기 개선 루프를 번갈아 가며 수행합니다. AREX는 장기적인 RSI를 유지하기 위해, 검증된 증거와 해결되지 않은 제약을 보존하면서 방대한 상호 작용 기록을 간결한 개선 상태로 압축하는 자율적인 컨텍스트 업데이트 도구를 학습합니다. 우리는 AREX를 검증된 합성 작업과 에이전트 기반 중간 훈련 및 장기 강화 학습을 통해 얻은 고품질 데이터셋으로 훈련했습니다. 장기 학습 동안 희소한 최종 보상을 완화하기 위해, 결정적인 증거가 획득되거나 잘못된 연구 방향이 수정되는 주요 단계를 강조합니다. 우리는 40억 개의 파라미터를 가진 밀집 모델과 1220억 개의 파라미터를 가진 Mixture-of-Experts 모델을 구현했습니다. BrowseComp, WideSearch, DeepSearchQA, Humanity's Last Exam (HLE) 및 기타 추론 및 도구 사용 벤치마크에서 AREX는 유사한 규모의 기본 모델보다 훨씬 뛰어난 성능을 보였으며, 훨씬 더 많은 활성화 파라미터를 사용하는 모델과도 경쟁력 있는 성능을 유지합니다.

Original Abstract

Deep research requires agents to find answers that jointly satisfy multiple constraints. Discovering such answers is costly, whereas verifying a candidate can often be decomposed into tractable constraint-wise checks. This discovery--verification asymmetry suggests that a research agent should do more than simply search longer: it should recursively improve its current answer by verifying intermediate results and using the partially verified state to guide subsequent refinement. We introduce AREX, a family of Recursively Self-Improving (RSI) deep research agents. AREX alternates between an inner research loop that gathers evidence and constructs a provisional answer, and an outer self-improvement loop that audits the answer constraint-wise, identifies unresolved claims, and launches targeted follow-up research. To sustain RSI over long horizons, AREX learns an autonomous context-update tool that compresses growing interaction history into a compact improvement state preserving verified evidence and unresolved constraints, without relying on an external model. We train AREX on verified synthetic tasks and high-quality trajectories through agentic mid-training and long-horizon reinforcement learning. To mitigate sparse final rewards during long horizon learning, we emphasize key steps where decisive evidence is acquired or erroneous research directions are corrected. We instantiate a dense 4B model and a 122B-A10B Mixture-of-Experts model. Across BrowseComp, WideSearch, DeepSearchQA, Humanity's Last Exam (HLE), and other reasoning and tool-use benchmarks, AREX substantially outperforms comparable-scale baselines and remains competitive with models using substantially more activated parameters.

0 Citations
0 Influential
14 Altmetric
70.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!