2606.22995v1 Jun 22, 2026 cs.LG

장기적인 자율 에이전트 강화 학습을 위한 그룹 그래프 정책 최적화

Group-Graph Policy Optimization for Long-Horizon Agentic Reinforcement Learning

Shaohan Huang
Shaohan Huang
Citations: 482
h-index: 7
Weiwei Deng
Weiwei Deng
Citations: 1,291
h-index: 17
Yunan Wang
Yunan Wang
Citations: 138
h-index: 8
Minghui Song
Minghui Song
Citations: 191
h-index: 5
Zihan Zhang
Zihan Zhang
Citations: 315
h-index: 8
Haizhen Huang
Haizhen Huang
Citations: 620
h-index: 12
Furu Wei
Furu Wei
Citations: 2,189
h-index: 22
Feng Sun
Feng Sun
Citations: 37
h-index: 3
Qi Zhang
Qi Zhang
Citations: 43
h-index: 2

그룹 기반 강화 학습(RL)은 자율 시나리오에서 대규모 언어 모델(LLM)의 성능을 크게 향상시켰습니다. 보다 세밀한 정책 업데이트를 달성하기 위해, 최근의 자율 RL 프레임워크는 트레이젝토리 레벨 훈련에서 스텝 레벨 훈련으로 전환되었습니다. 그러나 장기적인 자율 RL은 피드백이 수십 번의 상호 작용 단계를 거쳐 지연되기 때문에 심각한 보상 희소성 문제를 겪습니다. 기존의 스텝 레벨 프레임워크는 학습의 세밀성을 향상시키지만, 여전히 에이전트 탐색을 독립적이고 선형적인 트레이젝토리로 취급하여 미세 조정된 신용 할당을 수행합니다. 이러한 단순화된 관점은 상태 전이의 본질적인 그래프 구조를 무시하여 높은 분산의 상태 가치 추정과 단기적인, 지역적인 신용 할당으로 이어집니다. 이러한 중요한 문제점을 해결하기 위해, 우리는 다중 턴 자율 작업에 특화된 새로운 그룹 기반 RL 알고리즘인 Group-Graph Policy Optimization (G2PO)을 제안합니다. G2PO는 선형적인 상호 작용 트레이젝토리를 전역 상태 전이 그래프로 명시적으로 변환합니다. 서로 다른 트레이젝토리에서 동일한 관찰값을 집계하여, 샘플링 분산을 줄이고 트레이젝토리 의존성을 완화하는 그룹 집계 상태 가치 추정을 도입했습니다. 또한, 에이전트의 액션을 상태 노드 간의 전이로 재정의하고, 엣지 중심의 어드밴티지 추정 전략을 제안합니다. G2PO는 전체 그래프에서 시간 차이(TD) 오차를 전역적으로 표준화하여, 절대적인 작업 진행에 중요한 전이를 명시적으로 식별하고 우선순위를 부여합니다. 대표적인 장기 벤치마크인 WebShop, ALFWorld 및 AppWorld에서의 광범위한 실험 결과는 G2PO가 최첨단 프롬프트 기반 및 RL 기준 성능을 크게 능가하며, GRPO에 비해 최대 22.2%의 성공률 향상을 달성했음을 보여줍니다.

Original Abstract

Group-based Reinforcement Learning (RL) has significantly enhanced Large Language Models (LLMs) in agentic scenarios. To achieve finer-grained policy updates, recent agentic RL frameworks have shifted from trajectory-level to step-level training. However, long-horizon agentic RL suffers from severe reward sparsity and delay, as feedback is often deferred for dozens of interaction steps. While existing step-level frameworks refine training granularity, their credit assignment remains coarse-grained and still treats agent exploration as isolated, linear trajectories. This oversimplified perspective ignores the inherent graph structure of state transitions, leading to high-variance state-value estimation and myopic, localized credit assignment. To overcome these critical bottlenecks, we propose Group-Graph Policy Optimization (G2PO), a novel group-based RL algorithm tailored for multi-turn agentic tasks. G2PO explicitly transforms linear interaction trajectories into a global state-transition graph. By aggregating identical observations across different trajectories, we introduce group-aggregation state-value estimation that reduces sampling variance and trajectory-dependent bias. Furthermore, we redefine agent actions as transitions between state nodes and propose an edge-centric advantage estimation strategy. By globally standardizing Temporal Difference (TD) errors across the entire graph, G2PO explicitly identifies and prioritizes critical transitions that drive absolute task progress. Extensive experiments on representative long-horizon benchmarks-WebShop, ALFWorld, and AppWorld-demonstrate that G2PO substantially outperforms state-of-the-art prompt-based and RL baselines, achieving remarkable success rate improvements of up to 22.2% over GRPO.

0 Citations
0 Influential
11 Altmetric
55.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!