2606.25852v1 Jun 24, 2026 cs.LG

LLM 에이전트 강화 학습을 위한 의미적 일관성 정책 최적화

Semantic Consistency Policy Optimization for Reinforcement Learning of LLM Agents

Junzhuo Li
Junzhuo Li
Citations: 43
h-index: 4
Xuming Hu
Xuming Hu
Citations: 25
h-index: 4
Peng Xu
Peng Xu
Citations: 10,019
h-index: 17
Sijia Chen
Sijia Chen
Citations: 131
h-index: 3

그룹 기반 강화 학습은 LLM 에이전트를 장기적인, 희소 보상 작업을 위해 추가적으로 훈련하며, 트레일리의 결과로부터 단계별 기여도를 파악합니다. 그러나 이 방법은 특정 단계의 기여도를 해당 트레일리의 최종 결과에 연결하기 때문에, 의미적으로 유사한 중간 단계들이 결국 성공 또는 실패하는 트레일리에 따라 반대 방향의 기여도를 받게 됩니다. 이러한 의미적 기여도 불일치는 유사한 행동에 상반된 기울기를 보내고, 실패한 트레일리 내의 부분적으로 올바른 진행 과정을 낭비하게 만듭니다. 이러한 문제점을 해결하기 위해, 우리는 동등한 가치를 사용하지 않고 단계별 기여도를 수정하는 방법인 의미적 일관성 정책 최적화 (SCPO)를 제안합니다. SCPO는 동일한 트레일리 그룹 내의 성공적인 트레일리로부터 단계별 기여도를 회복하여 이러한 불일치를 완화합니다. 구체적으로, SCPO는 각 실패한 단계를 성공적인 트레일리와 비교하고, 해당 성공 트레일리에서 새로운 진행이 이루어진 부분에 대해 긍정적인 단계별 기여도를 부여합니다. ALFWorld 및 WebShop 환경에서, SCPO는 강력한 그룹 기반 기준 모델과 동등하거나 그 이상의 성능을 보이며, 특히 가장 어려운 다단계 작업에서 큰 향상을 보여줍니다 (ALFWorld: 93.7+/-4.1%, WebShop: 74.8+/-2.0%, 파라미터 수: 1.5B).

Original Abstract

Group-based reinforcement learning effectively post-trains LLM agents for long-horizon, sparse-reward tasks by deriving step-level credit from trajectory outcomes. However, this ties a step's credit to its rollout's final outcome: semantically near-identical intermediate steps receive opposite credit depending on whether their trajectory eventually succeeded or failed. Such semantic credit inconsistency sends conflicting gradients to similar actions and wastes the partially-correct progress inside failed rollouts. Motivated by this, we propose Semantic Consistency Policy Optimization (SCPO), a value-free reward-shaping method that mitigates this inconsistency by recovering step-level credit from successful siblings in the same rollout group. Concretely, SCPO scores each failed step against a successful sibling and adds positive step-level credit for new progress along that sibling. On ALFWorld and WebShop, SCPO matches or exceeds strong group-based baselines, reaching 93.7+/-4.1 percent success on ALFWorld and 74.8+/-2.0 percent on WebShop at 1.5B parameters, with gains concentrated on the hardest multi-step tasks.

0 Citations
0 Influential
8.5 Altmetric
42.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!