2607.28076v1 Jul 30, 2026 cs.AI

에이전트 기반 강화 학습을 위한 그룹 반사적 자기 증류

Group-Reflective Self-Distillation for Agentic Reinforcement Learning

Xiaoliang Fu
Xiaoliang Fu
Fudan University
Citations: 48
h-index: 4
Binbin Zheng
Binbin Zheng
Citations: 11
h-index: 2
Xing Ma
Xing Ma
Citations: 47
h-index: 4
Guanqun Zhao
Guanqun Zhao
Citations: 0
h-index: 0
Zijun Xie
Zijun Xie
Citations: 0
h-index: 0
Enlei Gong
Enlei Gong
Citations: 43
h-index: 2
Zeyu Chen
Zeyu Chen
Citations: 0
h-index: 0

검증 가능한 보상을 활용한 강화 학습(RLVR)은 대규모 언어 모델 에이전트를 훈련하는 데 효과적입니다. 그러나 최종 보상은 거친 수준의 경로 정보를 제공하며, 성공적인 행동, 반복되는 실수 및 우연한 선택 사항들이 동일한 결과 신호에 혼합되어 있습니다. 기존의 에이전트 기반 자기 증류 방법은 자연어 처리 능력을 활용하여 희소한 감독 신호를 풍부하게 하지만, 외부에서 검색하거나 더 강력한 모델로부터 추출된 기술들은 현재 경험과 일치하지 않거나, 정책의 능력 범위를 초과하거나, 특정 경로에 국한될 수 있습니다. 본 논문에서는 정책 자체의 검증된 시뮬레이션 데이터로부터 능력을 고려하고 결과 차이를 반영하는 지침을 도출하는 그룹 반사적 자기 증류(GRSD) 방법을 제안합니다. 각 프롬프트에 대해, 정책은 온-폴리시 그룹 내의 각 검증된 경로를 분석하며, 성공적인 시뮬레이션과 실패한 시뮬레이션을 비교하여 그룹 수준의 특권 정보 지침을 생성합니다. 이 지침을 기반으로, 자기 학습 모델은 결과 기반의 장점을 조절하면서 검증기가 결정한 학습 방향을 유지하며 턴 단위의 기여도 할당을 개선합니다. 다양한 에이전트 환경 및 모델 크기에서의 실험 결과는 GRSD가 경쟁적인 기본 방법보다 꾸준히 우수한 성능을 보이며, 새롭지 않은 작업에 대한 일반화 능력이 더 뛰어나다는 것을 보여줍니다.

Original Abstract

Reinforcement learning with verifiable rewards (RLVR) is effective for training large language model agents. However, terminal rewards provide only coarse trajectory-level supervision, leaving successful behaviors, recurring mistakes, and incidental choices entangled in the same outcome signal. Existing agentic self-distillation methods enrich sparse supervision with natural-language skills, but skills retrieved externally or extracted from a single trajectory by stronger models may mismatch current experience, exceed the policy's capability, or remain path-specific. We propose Group-Reflective Self-Distillation (GRSD), which derives capability-aligned and outcome-discriminative guidance from the policy's own verified rollouts. For each prompt, the policy reflects on each verified trajectory in an on-policy group, and a stop-gradient snapshot contrasts the resulting reflections from successful and failed rollouts to construct group-level privileged guidance. Conditioned on this guidance, a self-teacher refines turn-level credit assignment by modulating outcome-based advantages while preserving the verifier-determined learning direction. Experiments across multiple agentic environments and model scales demonstrate that GRSD consistently outperforms competitive baselines and generalizes more effectively to unseen tasks.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!