2607.29235v1 Jul 31, 2026 cs.RO

FBFM: 월드-액션 모델 실행을 위한 학습 불필요한 비동기 피드백 메커니즘 - 플로우 매칭 기반

FBFM: A Training-Free Asynchronous Feedback Mechanism for Flow-Matching in World-Action Models Execution

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Kai Chen
Kai Chen
Citations: 96
h-index: 6
Peize Li
Peize Li
Citations: 0
h-index: 0
Ruimeng Zhang
Ruimeng Zhang
Citations: 0
h-index: 0
Ru Zhang
Ru Zhang
Citations: 0
h-index: 0

월드-액션 모델(WAM)은 시각적 변화를 예측하여 로봇 제어의 장기적인 성능을 향상시키지만, 장기적인 신뢰성을 확보하기 위해서는 실제 관찰 데이터를 통해 반복적으로 재정합해야 합니다. 기존 WAM들은 이러한 문제를 해결하기 위해 덩어리(chunk) 단위로 히스토리 또는 키-값(KV) 캐시를 실제 데이터로 업데이트합니다. 그러나 이러한 덩어리 기반 피드백은 시간적 해상도가 낮아 개별 타임 스텝 수준의 예측 오류를 수정하는 데 한계가 있습니다. 이러한 문제를 해결하기 위해, 저희는 학습이 필요 없는 추론 메커니즘인 피드백 플로우 매칭(FBFM)을 제안합니다. FBFM은 활성적으로 생성되는 덩어리 내에서 재정합을 수행합니다. 플로우 매칭 과정에서, FBFM은 조건부 속도장에 마스크된 유사 역행렬 수정법을 적용합니다. 즉, 이전 덩어리의 액션을 활용하여 다음 덩어리의 액션 생성을 유도하고, 실행된 이전 덩어리 이후에 관찰된 이미지를 사용하여 다음 프레임 예측을 안내합니다. 이러한 덩어리 간의 연결(피드백)은 덩어리 경계를 기다리지 않고 오류를 수정하는 비동기 루프를 생성합니다. FBFM은 학습이 필요 없으므로, 예기치 않은 이벤트에 대한 반응성을 향상시키고 장기적인 작업에서의 드리프트를 억제합니다. 저희는 FBFM을 공동 생성을 사용하는 WAM(DreamZero)과 단계별 WAM(LingBot-VA) 모두에서 평가했습니다. 선택된 LIBERO 및 RoboTwin2.0 작업에서, 유리한 조건에서 성공률을 5% 이상 향상시켰으며, 실제 로봇 관찰-예측 진단 결과에서도 훨씬 더 나은 추적 성능을 보였습니다. 저희는 FBFM이 개별적인 온라인 수정에 대한 새로운 패러다임을 제시하며, 개방형 루프 플로우 생성과 폐쇄형 루프의 실제 세계 역학 사이를 연결한다고 주장합니다.

Original Abstract

Although world-action models (WAMs) enhance long-horizon robot control by predicting visual evolution before acting, long-horizon reliability demands repeated re-grounding in real observations--not recursive rollout. Existing WAMs address this by refreshing history or KV cache with ground-truth data between chunks. However, such chunk-wise feedback operates at a coarse temporal granularity and thus fails to correct prediction errors at the individual time-step level. To address this, we propose Feedback Flow Matching (FBFM), a training-free inference mechanism that pushes re-grounding inside the actively generated chunk. During flow matching, FBFM applies a masked pseudoinverse correction to the conditional velocity field: it leverages the preceding action chunk to guide generation of the next action chunk, and uses the image observed after executing that preceding chunk to guide the next frame prediction. This cross-chunk pairing--where feedback from one chunk arrives in time to shape the next--creates an asynchronous loop that corrects errors without waiting for chunk boundaries. Being training-free, the mechanism improves responsiveness to unexpected events and suppresses drift in long-horizon tasks. We evaluate FBFM on both a joint-generation WAM (DreamZero) and a stage-wise WAM (LingBot-VA). On selected LIBERO and RoboTwin2.0 tasks, it improves success rates by over 5% in favorable settings, and real-world robot observation-prediction diagnostics show notably better tracking. We argue that FBFM offers a new paradigm for fine-grained online correction, bridging open-loop flow generation with closed-loop real-world dynamics.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!