2606.12384v1 Jun 10, 2026 cs.LG

APPO: 능동적 절차 기반 정책 최적화

APPO: Agentic Procedural Policy Optimization

Yong Wang
Yong Wang
Citations: 472
h-index: 12
Xiangxiang Chu
Xiangxiang Chu
Citations: 278
h-index: 9
Shidong Yang
Shidong Yang
Citations: 42
h-index: 2
Yuxiang Ji
Yuxiang Ji
Citations: 164
h-index: 7
Ziyu Ma
Ziyu Ma
Citations: 96
h-index: 4
Pengkun Wang
Pengkun Wang
Citations: 723
h-index: 16
Xucong Wang
Xucong Wang
Citations: 38
h-index: 2
Guanhua Chen
Guanhua Chen
Citations: 53
h-index: 5

최근 강화 학습(RL) 분야의 발전은 대규모 언어 모델 에이전트의 다단계 도구 사용 능력에 상당한 개선을 가져왔습니다. 그러나 대부분의 기존 방법은 도구 호출 경계나 고정된 워크플로우와 같은 거친 휴리스틱 단위로 보상을 할당하여, 어떤 중간 결정이 결과에 영향을 미치는지 파악하기 어렵다는 문제가 있습니다. 본 연구에서는 능동적 RL을 두 가지 관점에서 분석합니다: extit{어디에서 분기할 것인가, 그리고 분기 후 어떻게 보상을 할당할 것인가}. 예비 분석 결과, 중요한 결정 지점은 도구 호출에 집중된 것이 아니라 생성된 시퀀스 전체에 걸쳐 넓게 분포되어 있으며, 토큰 엔트로피만으로는 최종 결과에 미치는 영향을 신뢰성 있게 반영하지 못한다는 것을 보여줍니다. 이러한 관찰을 바탕으로, 본 연구에서는 분기 및 보상 할당을 거친 상호 작용 단위에서 세분화된 결정 지점으로 이동시키는 extbf{능동적 절차 기반 정책 최적화 (APPO)}를 제안합니다. APPO는 토큰 불확실성과 정책에 의해 유도되는 후속 연속 과정의 가능성 증가를 결합한 분기 점수(Branching Score)를 사용하여 분기 위치를 선택함으로써, 더 표적화된 탐색을 가능하게 하고 동시에 잘못된 높은 엔트로피 위치를 제거합니다. 또한, APPO는 절차 수준의 장점 스케일링을 도입하여 분기된 실행 경로에 걸쳐 보상을 보다 효과적으로 분배합니다. 13개의 벤치마크 실험 결과, APPO는 기존 능동적 RL 방법보다 평균 약 4점의 성능 향상을 보여주었으며, 효율적인 도구 사용을 유지하고 행동 해석 가능성을 확보했습니다.

Original Abstract

Recent advances in agentic Reinforcement Learning (RL) have substantially improved the multi-turn tool-use capabilities of large language model agents. However, most existing methods assign credit over coarse heuristic units, such as tool-call boundaries or fixed workflows, making it difficult to identify which intermediate decisions influence downstream outcomes. In this work, we study agentic RL from two perspectives: \textit{where to branch and how to assign credit after branching}. Our pilot analysis shows that influential decision points are broadly distributed throughout the generated sequence rather than concentrated at tool calls, while token entropy alone does not reliably reflect their impact on final outcomes. Motivated by these observations, we propose \textbf{Agentic Procedural Policy Optimization (APPO)}, which shifts branching and credit assignment from coarse interaction units to fine-grained decision points in the sequence. APPO selects branching locations using a Branching Score that combines token uncertainty with policy-induced likelihood gains of subsequent continuations, enabling more targeted exploration while filtering out spurious high-entropy positions. It further introduces procedure-level advantage scaling to better distribute credit across branched rollouts. Experiments on 13 benchmarks show that APPO consistently improves strong agentic RL baselines by nearly 4 points, while keeping efficient tool-calls and maintaining behavior interpretability.

0 Citations
0 Influential
8 Altmetric
40.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!