2607.27782v1 Jul 30, 2026 cs.RO

RedFlow: 플로우 매칭 VLA 정책의 액션 레벨 수정 사항을 위한 오류 재지향

RedFlow: Redirect Failure into Action-Level Corrections for Flow-matching VLA Policy

Zhengyang Yan
Zhengyang Yan
Citations: 44
h-index: 3
Fangqi Zhu
Fangqi Zhu
Citations: 106
h-index: 3
Quanxin Shou
Quanxin Shou
Citations: 51
h-index: 3
Yikun Miao
Yikun Miao
Citations: 8
h-index: 2
Xiaoyi Pang
Xiaoyi Pang
Citations: 6
h-index: 2
Song Guo
Song Guo
Citations: 91
h-index: 3
Junhao Li
Junhao Li
Citations: 0
h-index: 0
Zijun Wang
Zijun Wang
Citations: 0
h-index: 0
Zicong Hong
Zicong Hong
Citations: 2,088
h-index: 19

플로우 매칭 비전-언어-액션(VLA) 정책은 로봇 조작 분야에서 뛰어난 잠재력을 보여주지만, 배포 과정에서의 데이터 분포 변화로 인해 발생하는 복합적인 오류에 자주 직면합니다. 오프라인 강화 학습(RL)은 롤아웃 데이터를 활용하여 배포된 정책을 개선하는 실용적인 방법이지만, 기존의 방법들은 실패 데이터를 무시하거나 단순히 트래젝토리 레벨에서만 활용하여 학습 효율성이 낮고 지속적인 오류가 발생합니다. 본 논문에서는 **RedFlow**라는 정교한 오프라인 RL 프레임워크를 제안합니다. RedFlow는 실패 경험을 플로우 매칭 VLA 정책의 액션 레벨 수정 사항에 적용하는 방식으로 작동합니다. RedFlow는 두 가지 핵심 구성 요소로 이루어집니다: (1) **컨텍스트 인식 교정 매칭(Context-Aware Corrective Matching)** 메커니즘은 오류를 유발하는 액션을 식별하고, 유사한 컨텍스트에서 성공적인 대안을 검색하여 이를 수정 목표로 사용합니다. (2) **적응적 재지향 목적 함수(Adaptive Redirection Objective)**는 성공적인 액션을 강화하고, 바람직하지 않은 액션을 억제하며, 회복 가능한 실패를 수정 목표 방향으로 유도합니다. RedFlow는 성공적인 경험과 실패한 경험 모두를 밀집적인 지도 데이터로 변환하여, 저품질의 혼합된 데이터를 활용한 강력한 학습을 가능하게 합니다. LIBERO 벤치마크 및 세 가지 실제 로봇 조작 작업에 대한 실험 결과에서 RedFlow가 최첨단 오프라인 RL 모델들을 꾸준히 능가하며, 실제 성공률을 56.7%에서 74.7%로 향상시켰습니다. 또한 RedFlow는 PPO, GRPO 및 DDPO와 같은 강력한 온-정책 방법과 유사한 성능을 보이면서도 약 한 자리수의 더 적은 학습 샘플만을 사용합니다.

Original Abstract

Flow-matching Vision-Language-Action (VLA) policies have shown strong potential for robotic manipulation but often suffer from compounding errors caused by distribution shifts during deployment. While offline reinforcement learning (RL) provides a practical way to improve deployed policies using rollout data, existing methods either ignore failure data or exploit it only at the trajectory level, resulting in low learning efficiency and persistent errors. We propose **RedFlow**, a fine-grained offline RL framework that redirects failure experiences into action-level corrective supervision for flow-matching VLA policies. RedFlow consists of two key components: (1) a **Context-Aware Corrective Matching** mechanism that identifies failure-inducing actions and retrieves successful alternatives from similar contexts as corrective targets, and (2) an **Adaptive Redirection Objective** that jointly reinforces successful actions, suppresses undesirable ones, and redirects recoverable failures toward corrective targets. By converting both successful and failed experiences into dense supervision, RedFlow enables robust recovery learning from mixed-quality data. Experiments on the LIBERO benchmark and three real-world manipulation tasks show that RedFlow consistently outperforms state-of-the-art offline RL baselines, improving the real-world success rate from 56.7% to 74.7%. It also matches strong on-policy methods (PPO, GRPO, and DDPO) while requiring roughly an order of magnitude fewer training samples.

0 Citations
0 Influential
9.5 Altmetric
47.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!