2608.13026v1 Aug 13, 2026 cs.RO

시간적 GRPO: 시각-언어-행동 강화 학습에서 트래젝토리 수준의 보상을 넘어

Temporal GRPO: Beyond Trajectory-Level Credit in Vision-Language-Action Reinforcement Learning

Wenwen Qiang
Wenwen Qiang
Citations: 569
h-index: 15
Yao Zhou
Yao Zhou
Citations: 11
h-index: 2
Hang Gao
Hang Gao
Citations: 26
h-index: 3
Fengge Wu
Fengge Wu
Citations: 40
h-index: 4
Changwen Zheng
Changwen Zheng
Citations: 633
h-index: 14

결과 기반 강화 학습은 희소한 성공 피드백을 통해 시각-언어-행동(VLA) 정책을 후속 훈련하는 확장 가능한 방법을 제공합니다. 일반적인 GRPO 기반 VLA 후속 훈련에서, 하나의 롤아웃 레벨 어드밴티지는 트래젝토리 내의 모든 행동에 적용됩니다. 여러 유효 단계를 완료했지만 나중에 실패한 롤아웃은 초기 진행을 이끌어낸 행동에도 페널티를 줄 수 있습니다. 우리는 이를 '트래젝토리 레벨 신용 혼동(credit aliasing)'이라고 부릅니다. 시간적 GRPO는 감지 가능한 작업 단계를 구성하고, 각 롤아웃을 단계별 행동 간격과 일치시키며, 동일한 단계에 진입한 롤아웃만 비교하여 이 문제를 해결합니다. 결과적으로 얻은 단계별 어드밴티지는 단일 정책 업데이트에서 해당 간격에 적용됩니다. RoboTwin 2.0에서 시간적 GRPO는 작업 성공률과 샘플 효율성을 향상시키며, 다양한 작업 범위에서 일관된 성능 향상을 보입니다. LIBERO-Long에서 수행된 제어된 업데이트는 공유되는 필수 단계를 유지하고, 롤아웃 결과가 분기되는 첫 번째 단계에 개선을 집중합니다.

Original Abstract

Outcome-driven reinforcement learning offers a scalable way to post-train vision-language-action (VLA) policies from sparse task-success feedback. In common GRPO-based VLA post-training, one rollout-level advantage is applied to every action in the trajectory. A rollout that completes several valid stages but fails later can therefore penalize the actions that produced its earlier progress. We call this trajectory-level credit aliasing. Temporal GRPO addresses this problem by constructing detectable task stages, aligning each rollout with stage-specific action intervals, and comparing only rollouts that have entered the same stage. The resulting stage advantages are applied to their corresponding intervals in a single policy update. On RoboTwin 2.0, Temporal GRPO improves task success and sample efficiency, with consistent gains across task horizons. Controlled updates on LIBERO-Long preserve shared prerequisite stages and concentrate improvement at the first stage where rollout outcomes diverge.

0 Citations
0 Influential
7.5 Altmetric
37.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!