2608.01667v1 Aug 03, 2026 cs.AI

TCPO: Turn 수준의 신뢰도 정책 최적화

TCPO: Turn-Level Credit Policy Optimization

Zhi Chen
Zhi Chen
Citations: 80
h-index: 4
Yaohua Tang
Yaohua Tang
Citations: 106
h-index: 3
Sicong Liao
Sicong Liao
Citations: 10
h-index: 2

검증기(verifier) 기반 강화 학습은 LLM 추론 능력을 향상시키는 강력한 방법론으로 자리 잡았습니다. 다중 회전(multi-turn) 설정에서 모델은 각 회전 후에 검증기 점수를 받고, 이 점수를 바탕으로 반복적으로 출력을 개선합니다. 이러한 점수는 상세한 피드백을 제공하지만, 직접적인 신뢰도(credit)를 제공하지는 않습니다. 점수는 현재 출력의 품질을 측정하는 반면, 신뢰도는 현재 회전이 전체 개선 경로에 미치는 영향을 측정해야 합니다. 본 논문에서는 검증기 기반 다중 회전 강화 학습을 위한 회전 수준의 신뢰도 할당 방법인 TCPO를 제안합니다. TCPO는 신뢰도 할당을 점수-신뢰도 변환 문제로 정의하고, 참조 기반 비교를 통해 회전 수준의 이점을 구축합니다. 과거(retrospective) 신뢰도는 현재 상태에서 가장 좋은 상태에 대한 즉각적인 진행 상황과 후퇴를 파악하며, 미래 지향적(hindsight) 지연 신뢰도는 나중에 보상을 받는 비개선 회전을 식별하고, 선택적 고정 이력 반사실 추정을 통해 동일한 이력을 가진 놀라운 회전을 개선합니다. 수학 추론, 코드 생성 및 AppWorld 에이전트 작업에 대한 실험 결과, TCPO는 다양한 모델 크기, 작업 도메인 및 검증기 유형에서 가장 강력한 기준 성능을 능가하거나 동등한 성능을 보였습니다. TCPO는 Qwen3-4B 및 DeepSeek-R1-Distill-Llama-8B 모델에서 최상위 또는 최고 수준의 Pass@8 성능을 달성하고, 성공에 필요한 회전 수를 줄이며, 다중 회전 에이전트 성능을 향상시킵니다. 이러한 결과는 검증기 기반 다중 회전 정책 최적화를 위한 핵심 요소로서 점수-신뢰도 변환의 중요성을 강조합니다.

Original Abstract

Verifier-guided reinforcement learning has become a powerful paradigm for improving LLM reasoning. In multi-turn settings, models receive a verifier score after each turn and iteratively refine their outputs. Although such scores provide dense feedback, they do not directly provide dense credit: a score measures the quality of the current output, while credit should measure how the current turn changes the refinement trajectory. We propose TCPO, a turn-level credit assignment method for verifier-guided multi-turn RL. TCPO casts credit assignment as score-to-credit conversion and constructs turn-level advantages through reference-based comparisons: retrospective credit captures immediate progress and regression relative to the best prior state; hindsight delayed credit identifies non-improving turns with later payoff; and selective fixed-history counterfactual estimation refines high-surprisal turns under the same history. Experiments on math reasoning, code generation, and AppWorld agent tasks show that TCPO improves or matches the strongest baselines across model scales, task domains, and verifier types. TCPO achieves the best or tied-best best-turn Pass@8 on Qwen3-4B and DeepSeek-R1-Distill-Llama-8B, reduces turns to success, and improves multi-turn agent performance. These results highlight score-to-credit conversion as a central ingredient for verifier-guided multi-turn policy optimization.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!