에이전트 기반 강화 학습을 위한 경로 상대적 후회 증류
Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning
최근 에이전트 기반 강화 학습 방법들은 희소한 결과 보상을 보완하기 위해 후회를 활용합니다. 그러나 하나의 실행 과정에서 많은 후회 신호가 발생할 수 있으며, 이러한 신호를 각 단계에 어떻게 적절히 할당해야 하는지는 불분명합니다. 본 논문에서는 경로 상대적인 후회 증류 프레임워크인 TRIAL을 소개하며, 이는 통일된 단계 정렬 기반의 평가 프로토콜을 사용합니다. TRIAL은 각 의사 결정 단계에서 해당 의사 결정의 실제 결과에 대한 '결과 뷰'를 추출하고, 동일한 응답을 일반적인 상황과 후회 조건 하에서의 상황 모두에서 평가합니다. 부호가 있는 로그 확률 차이는 토큰 수준의 감독 신호의 방향과 지역적 강도를 결정하며, 단계 수준의 크기는 실제로 실행된 전체 경로에 대해 함께 정규화됩니다. 결과적으로 얻어지는 할당 계수는 유효한 토큰 가중 평균이 1이며, 이는 평균 곱셈 값을 유지하면서 밀집적인 감독 신호를 단계 간에 재분배합니다. WebShop과 ALFWorld에서 다양한 기반 모델을 사용하여 수행된 실험 결과, TRIAL은 모든 여덟 가지 조합 (기반 모델, 환경, 평가 지표)에서 GRPO보다 우수한 성능을 보였으며, 여섯 가지 방법 중 여섯 가지 상황에서 최상위 또는 동률의 성능을 달성했습니다. 특히 Qwen3-1.7B를 사용하여 WebShop에서 실행했을 때, TRIAL은 성공률을 56.4%에서 75.2%로, 작업 점수를 78.7%에서 85.7%로 향상시켰습니다. 추가적인 제어된 실험을 통해 경로 상대적인 단계 할당이 후회 증류만 사용하는 것보다 훨씬 더 큰 성능 향상을 제공한다는 것을 확인했습니다.
Recent agentic reinforcement learning methods use hindsight to complement sparse outcome rewards. However, a completed rollout can yield many such signals, leaving their appropriate allocation across turns unclear. We introduce TRIAL, a trajectory-relative hindsight distillation framework with a unified turn-aligned scoring protocol. For each decision turn, TRIAL extracts an outcome view of that decision's realized consequence and evaluates the same response under ordinary and hindsight-conditioned contexts. The signed log-probability gap determines the direction and local strength of token-level supervision, while turn-level magnitudes are normalized jointly over the realized trajectory. The resulting allocation multipliers have an eligible-token-weighted mean of one, redistributing dense supervision across turns while fixing its average multiplier. Experiments on WebShop and ALFWorld with different backbones show that TRIAL outperforms GRPO across all eight combinations of backbone, environment, and evaluation metric, while achieving the best or tied-best performance among six methods on six of them. On WebShop with Qwen3-1.7B, TRIAL improves the success rate from 56.4% to 75.2% and the task score from 78.7% to 85.7%. Controlled ablations further show that trajectory-relative turn allocation provides substantial gains beyond those of dense hindsight distillation alone.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.