2607.06223v1 Jul 07, 2026 cs.AI

정보 이득 기반 롤아웃 정책 최적화: 다중 턴 LLM 에이전트를 위한 적응형 트리 구조 롤아웃 접근 방식

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents

Luoyi Fu
Luoyi Fu
Citations: 3,774
h-index: 32
Jiaxin Ding
Jiaxin Ding
Citations: 166
h-index: 8
Xinbing Wang
Xinbing Wang
Citations: 10
h-index: 2
Fan Xu
Fan Xu
Citations: 10
h-index: 2
Yule Xie
Yule Xie
Citations: 25
h-index: 2
Shiqing Gao
Shiqing Gao
Citations: 7
h-index: 2
Xin Ding
Xin Ding
Citations: 52
h-index: 3
Yijun Zhang
Yijun Zhang
Citations: 0
h-index: 0
Haoxiang Zhang
Haoxiang Zhang
Citations: 72
h-index: 1

강화 학습은 최종 결과를 얻기 전에 일련의 중간 결정을 내려야 하는 장기적인 탐색 작업에서 대규모 언어 모델(LLM) 에이전트의 성능을 향상시키는 유망한 패러다임으로 자리 잡았습니다. 그러나 기존 방법은 여전히 중요한 한계점을 가지고 있습니다. 즉, 롤아웃 예산이 종종 중간 상태의 유용성을 명시적으로 평가하지 않고 할당됩니다. 그 결과, 정보 가치가 낮은 상태에 상당한 계산 자원이 소모될 수 있으며, 이는 서로 크게 다른 정보를 제공하는 다양한 분기에도 불구하고 발생합니다. 본 논문에서는 정보 이득 기반 롤아웃 정책 최적화(IGRPO)라는 정책 최적화 프레임워크를 제안합니다. IGRPO는 중간 상태의 정보량을 롤아웃 수집의 핵심 원칙으로 간주합니다. 구체적으로, IGRPO는 노드 수준의 정보량에 따라 확장 예산을 할당하여 예산 의식을 갖춘 트리 구조 롤아웃을 수행합니다. 이를 통해 더 많은 정보를 제공하는 분기는 더 자주 확장되는 반면, 유망하지 않은 분기는 점진적으로 억제됩니다. 또한, 정보 이득 기반 롤아웃은 명시적인 제한된 교사 분포를 경로에 유도하며, 이는 자연스럽게 명확한 정책 최적화 목표를 제공하여 적응형 트리 구조 탐색과 체계적인 정책 학습을 단일 프레임워크 내에서 통합합니다. 일곱 가지 어려운 검색 증강 질의 응답 벤치마크에서의 실험 결과는 IGRPO가 동일한 롤아웃 예산 제약 조건 하에서 강력한 기본 모델보다 지속적으로 우수한 성능을 발휘하며, 이는 유도된 교사 분포를 활용하여 장기적인 탐색 에이전트의 정책 최적화를 안내하는 것이 효과적임을 입증합니다.

Original Abstract

Reinforcement learning has become a promising paradigm for improving large language model (LLM) agents on long-horizon search tasks, where the agent must make a sequence of intermediate decisions before receiving a final outcome. However, existing methods still face a key limitation: the rollout budget is often allocated without explicitly assessing the utility of intermediate states. As a result, substantial computation may be spent on low-value states, even though different branches can vary drastically in their informativeness. In this paper, we propose Information Gain-based Rollout Policy Optimization (IGRPO), a policy optimization framework that treats intermediate-state informativeness as the organizing principle of rollout collection. Specifically, IGRPO performs budget-aware tree-structured rollouts by allocating expansion budget according to node-level informativeness, so that more informative branches are expanded more frequently while unpromising branches are progressively suppressed. We further demonstrate that the information gain-based rollout induces an explicit limiting teacher distribution over trajectories, which naturally yields a clear policy optimization target, thereby unifying adaptive tree-structured exploration with principled policy learning under a single framework. Experiments on seven challenging search-augmented QA benchmarks demonstrate that IGRPO consistently outperforms strong baselines under the same rollout budget constraints, validating the effectiveness of leveraging the induced teacher distribution to guide policy optimization for long-horizon search agents.

2 Citations
0 Influential
16 Altmetric
82.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!