2608.13179v1 Aug 13, 2026 cs.AI

방향이 아닌 크기를 가르치십시오: 검증 기반 성능 제한을 갖는 다중 단계 LLM 에이전트의 신뢰도 기반 보상 할당

Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents

Leilei Gan
Leilei Gan
Citations: 9
h-index: 2
Chenyi Zhuang
Chenyi Zhuang
Citations: 351
h-index: 12
Siyuan Lu
Siyuan Lu
Citations: 10
h-index: 2
Linjian Mo
Linjian Mo
Citations: 0
h-index: 0
Zechuan Wang
Zechuan Wang
Citations: 50
h-index: 4
Hongxuan Zhang
Hongxuan Zhang
Citations: 55
h-index: 4

검증 가능한 보상을 사용하는 강화 학습(RLVR)은 다중 회전 도구 사용 에이전트 훈련에 대한 검증 기반의 성능 상한을 제공하지만, 경로 수준의 보상 할당 방식은 다양한 회전 단계별 결과를 하나의 보상 신호로 통합합니다. 온정책 증류는 토큰 단위의 상세한 감독 신호를 제공하지만, 이는 교사 모델에 의해 제한되거나 기울기 집중 현상을 유발할 수 있습니다. 본 논문에서는 RL의 검증 기반 성능 상한을 유지하면서, 특권적인 자체 교사를 통해 얻은 토큰 수준의 상세 정보를 통합하는 계층적 보상 할당 프레임워크인 $ extbf{CrEST}$를 소개합니다. $ extbf{CrEST}$는 두 가지 수준에서 보상을 할당합니다. 회전 단계별로 분할된 검증된 이점은 회전 간 희석 효과를 해결하며, 엔트로피 게이트를 사용한 자체 교사 모델 조절은 회전 내 토큰 기여도를 세밀하게 조정합니다. BFCL V3 및 WildToolBench에 대한 실험 결과, $ extbf{CrEST}$는 두 가지 모델 크기에서 RL 및 증류 기반 모델보다 일관되게 우수한 성능을 보이며, 특히 긴 경로 및 엄격한 세션 수준 지표에서 가장 큰 개선 효과를 나타냅니다. 본 연구는 정책 최적화 과정에서 교사의 역할이 업데이트 방향을 결정하는 것에서 업데이트 크기를 조절하는 것으로 변화될 수 있으며, 이를 통해 검증 기반의 성능 상한을 유지하면서 상세한 보상 할당을 가능하게 함을 보여줍니다.

Original Abstract

Reinforcement learning with verifiable rewards (RLVR) offers a verifier-bounded performance ceiling for training multi-turn tool-use agents, yet its trajectory-level credit assignment conflates heterogeneous per-turn outcomes into a single reward signal. On-policy distillation provides dense per-token supervision but is either teacher-bounded or prone to gradient concentration collapse. We introduce $\textbf{CrEST}$, a hierarchical credit assignment framework that retains RL's verifier-bounded ceiling while incorporating dense token-level signals from a privileged self-teacher. $\textbf{CrEST}$ resolves credit at two levels: turn-segmented verified advantages address inter-turn dilution, while entropy-gated self-teacher modulation refines intra-turn token contributions. Experiments on BFCL V3 and WildToolBench show that $\textbf{CrEST}$ consistently outperforms both RL and distillation baselines across two model scales, with the largest gains on long-trajectory and strict session-level metrics. Our work demonstrates that the teacher's role in policy optimization can be reduced from determining update directions to modulating update magnitudes, unlocking dense credit assignment without sacrificing the verifier-bounded ceiling.

0 Citations
0 Influential
6 Altmetric
30.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!