2606.26080v1 Jun 24, 2026 cs.LG

사후 학습에서 얻는 간과된 이점: LLM 에이전트의 성능 향상

Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents

Changdae Oh
Changdae Oh
Citations: 19
h-index: 2
Wendi Li
Wendi Li
Citations: 949
h-index: 12
S. Yeh
S. Yeh
Citations: 60
h-index: 2
Sharon Li
Sharon Li
Citations: 298
h-index: 6
S. Park
S. Park
Citations: 12
h-index: 1
Tanwi Mallick
Tanwi Mallick
Citations: 5
h-index: 2

프로세스 보상 모델은 LLM을 세밀하고 단계별로 평가할 수 있게 해주지만, 에이전트 환경에 적용하려면 매우 어렵습니다. 장기간 상호 작용, 되돌릴 수 없는 행동, 그리고 확률적인 환경 피드백은 인간의 주석 작업과 몬테카를로 추정 방식을 대규모로 실행하는 것을 불가능하게 만듭니다. 본 연구에서는 강화 학습(RL) 사후 학습이 이미 효과적인 단계별 평가를 위한 요소를 제공하며, 따라서 별도의 보상 모델 학습이 필요 없다는 것을 보여줍니다. 구체적으로, 일반적인 확률적 마르코프 결정 과정 하에서 암묵적인 이점을 도출했으며, 이를 '진전 이점(progress advantage)'이라고 명명했습니다. RL로 훈련된 정책과 참조 정책 간의 로그-확률 비율은 최적의 이점 함수를 정확하게 복원합니다. 이러한 공식화는 결과 신호가 주석 없이 생성되고, 특정 영역에 국한되지 않으며, 표준 RL 사후 학습 파이프라인의 부산물로 제공됩니다. 우리는 세 가지 다른 응용 분야(테스트 시간 스케일링, 불확실성 정량화 및 오류 원인 분석)에서 다섯 가지 벤치마크와 네 가지 모델 패밀리에 걸쳐 진전 이점의 효과를 검증했습니다. 모든 설정에서 진전 이점은 신뢰도 기반의 기존 방법보다 우수한 성능을 보였으며, 특정 작업에 대한 별도의 학습이 필요 없음에도 불구하고 전용으로 훈련된 보상 모델보다 뛰어난 결과를 보여주었습니다. 또한, 우리는 진전 이점의 특징에 대한 심층적인 분석을 제공하여 실제 에이전트 시스템 적용에 대한 실질적인 지침을 제시합니다.

Original Abstract

Process reward models enable fine-grained, step-level evaluation of LLMs, yet building them for agentic settings remains prohibitively difficult: long-horizon interactions, irreversible actions, and stochastic environment feedback make both human annotation and Monte Carlo estimation infeasible at scale. In this work, we show that reinforcement learning (RL) post-training already provides the ingredients for effective step-level scoring, eliminating the need for dedicated reward model training altogether. Concretely, we derive an implicit advantage under a general stochastic Markov decision process, which we term progress advantage -- log-probability ratio between the RL-trained policy and its reference policy exactly recovers the optimal advantage function. This formulation makes the resulting signal annotation-free, domain-agnostic, and available as a byproduct of the standard RL post-training pipeline. We validate the effectiveness of the progress advantage across three different applications: test-time scaling, uncertainty quantification, and failure attribution on five benchmarks and four model families. Across all settings, it consistently outperforms confidence-based baselines and, despite requiring no task-specific training, surpasses dedicated trained reward models. We complement these results with deeper analyses on characteristics of progress advantage, offering practical guidance for adoption in real-world agentic systems.

0 Citations
0 Influential
6 Altmetric
30.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!