2605.27140v1 May 26, 2026 cs.AI

StepOPSD: 단계 인지 온라인 선호도 증류를 통한 에이전트 강화 학습

StepOPSD: Step-Aware Online Preference Distillation for Agent Reinforcement Learning

Chenglin Wu
Chenglin Wu
Citations: 460
h-index: 10
Yanfei Zhang
Yanfei Zhang
Citations: 20
h-index: 3
Xuemin Lin
Xuemin Lin
Citations: 0
h-index: 0

다중 턴 에이전트를 위한 강화 학습은 보상 할당 불일치 문제를 겪습니다. 보상은 희소하고 전체 경로 수준에서 제공되지만, 성공은 종종 몇 가지 국지적인 결정에 달려 있습니다. 기존의 온라인 정책 증류(OPD)는 더 밀집된 토큰 수준의 감독 신호를 제공하지만, 일반적으로 다양한 에이전트 경로를 단일 문자열로 취급하는 대신 인과적 상호 작용 단위로 처리하지 않습니다. 본 논문에서는 StepOPSD라는 새로운 프레임워크를 제시합니다. StepOPSD는 에이전트의 각 단계를 보상 재분배의 기본 단위로 사용하는, 롤아웃 후 선호도 기반 자기 증류 방식입니다. StepOPSD는 경로를 행동 중심의 단계 세그먼트로 분해하고, 과거 정보가 풍부한 가이드 문맥 하에서 이러한 세그먼트를 재평가합니다. 또한 토큰 수준의 로그 확률 차이를 정방향 신호를 유지하는 방식으로 변환하여, GRPO 업데이트 전에 정규화된 단계별 보상 예산을 사용하여 이점을 형성합니다. ALFWorld 및 Search-QA 데이터셋에서 Qwen3-1.7B 및 Qwen2.5-3B-Instruct 모델을 사용하여 실험한 결과, StepOPSD는 국지적인 인과적 오류에 민감한 부분 집합에서 최고 또는 두 번째로 좋은 성능을 달성했습니다. 특히 ALFWorld Heat (79.1%), PickTwo (95.0%), Search-QA TriviaQA (61.6%)에서 1위를 차지했으며, HotpotQA에서는 공동 최고 성적 (40.4%)을 기록했습니다. 실험 결과는 또한 '두 개의 조절 변수' 법칙을 보여줍니다: 작은 α_clip 값은 국지적인 신뢰 영역을 안정화시키고, 최적의 전역 혼합 강도 λ_mix는 작업에 따라 달라집니다. 이러한 결과는 단계 인지 증류가 전체 경로 수준의 보상이 성공을 결정하는 국소적인 행동과 일치하지 않을 때 가장 유용하다는 것을 시사합니다.

Original Abstract

Reinforcement learning for multi-turn agents suffers from a credit-assignment mismatch: rewards are sparse and trajectory-level, while success often hinges on a few local decisions. Existing online policy distillation (OPD) provides denser token-level supervision, but typically treats heterogeneous agent trajectories as monolithic strings rather than causal interaction units. We present StepOPSD, a post-rollout preference self-distillation framework that takes the agent step as the unit of credit redistribution. StepOPSD decomposes trajectories into action-centered step segments, rescoring them under hindsight-enriched teacher contexts and converting token-level log-probability gaps into sign-preserving advantage shaping with a normalized per-step credit budget before the GRPO update. Across ALFWorld and Search-QA with Qwen3-1.7B and Qwen2.5-3B-Instruct, StepOPSD attains best or second-best results on subsets most sensitive to local causal errors, including first-place performance on ALFWorld Heat (79.1%), PickTwo (95.0%), Search-QA TriviaQA (61.6%), and tied-best performance on HotpotQA (40.4%). The results further reveal a consistent two-knob law: smaller α_clip acts as a broadly stabilizing local trust region, whereas the optimal global mixing strength λ_mix remains task-dependent. These findings suggest that step-aware distillation is most useful when trajectory-level rewards are weakly aligned with the local action that determines downstream success.

1 Citations
0 Influential
5 Altmetric
26.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!