2607.14739v1 Jul 16, 2026 cs.CV

FoMoVLA: 시각적 예측과 동작 지침을 결합하여 비전-언어-행동 모델 성능 향상

FoMoVLA: Bridging Visual Foresight and Motion Guidance for Vision-Language-Action Models

Bailin Li
Bailin Li
Citations: 620
h-index: 3
Kun Zhan
Kun Zhan
Citations: 1,574
h-index: 16
Wei Li
Wei Li
Citations: 210
h-index: 5
Peijin Jia
Peijin Jia
Citations: 87
h-index: 4
Xuefeng Jiang
Xuefeng Jiang
Citations: 30
h-index: 2
Titong Jiang
Titong Jiang
Citations: 79
h-index: 4
Sheng Sun
Sheng Sun
Citations: 48
h-index: 3
Xin Wen
Xin Wen
Citations: 124
h-index: 6
Han Hong
Han Hong
Citations: 0
h-index: 0
Zhikang Liu
Zhikang Liu
Citations: 608
h-index: 3
Yuan Ma
Yuan Ma
Citations: 46
h-index: 3
Yujian Li
Yujian Li
Citations: 0
h-index: 0

비전-언어-행동 (VLA) 모델은 시각-운동 정책 학습에서 놀라운 성과를 거두었지만, 여전히 근본적으로 반응적인 방식으로 작동하며, 현재의 관찰 정보와 언어를 사용하여 행동을 결정하지만, 세계의 역학에 대한 명시적인 예측을 수행하지 못합니다. 기존의 시각적 예측 방법은 미래의 시각 상태를 예측하지만, 명시적인 동작 지침이 부족합니다. 즉, 어디로 가야 하는지는 알려주지만 어떻게 이동해야 하는지는 알려주지 않습니다. 본 연구에서는 미래 특징 예측과 희소 포인트 추적이 상호 보완적이라고 주장합니다. 전자는 목표 상태를 제공하고, 후자는 목표에 도달하기 위한 연속적인 동작 경로를 파악합니다. 우리는 VLA 표현을 명시적인 시공간적 감독 신호로 확장하는 프레임워크인 FoMoVLA를 제안합니다. FoMoVLA는 미래 특징 예측과 희소 2D 포인트 추적을 동시에 학습하여 연속적인 행동 정책을 향상시키며, 이를 위해 미래 상태에 대한 짧은 예측 토큰을 도입하고, 희소한 시간 기반 2D 포인트 궤적을 디코딩하여 간결한 기하학적 동작을 모델링합니다. 또한, 가벼운 미래 조건부 크로스-어텐션 모듈을 사용하여 예상되는 상태와 포인트 역학 사이의 일관성 있는 추론을 가능하게 합니다. LIBERO, RoboCasa GR-1 Tabletop 및 LIBERO-Plus 데이터셋에 대한 광범위한 실험 결과는 최첨단 성능과 뛰어난 제로샷 일반화 능력을 보여줍니다. 프로젝트 페이지는 https://liauto-research.github.io/FoMoVLA 에서 확인할 수 있습니다.

Original Abstract

Vision-Language-Action (VLA) models have achieved impressive results in visuomotor policy learning, yet remain fundamentally reactive, mapping current observations and language to actions without explicit forward prediction of world dynamics. Existing visual foresight methods predict future visual states but lack explicit motion guidance: they show where to go but not how to get there. We argue that future feature prediction and sparse point tracking are naturally complementary: the former provides the goal state, while the latter captures the continuous motion path toward it. We propose FoMoVLA, a framework that augments VLA representations with explicit spatio-temporal supervision by jointly learning future feature foresight and sparse 2D point tracking, enhancing the continuous action policy. FoMoVLA introduces compact foresight tokens to decode future feature states, decodes sparse temporal 2D point trajectories to model compact geometric motion, and couples both through a lightweight future-conditioned cross-attention module that enables consistent reasoning between anticipated states and point dynamics. Extensive experiments on LIBERO, RoboCasa GR-1 Tabletop, and LIBERO-Plus demonstrate state-of-the-art performance and strong zero-shot generalization. Project page is available at https://liauto-research.github.io/FoMoVLA.

0 Citations
0 Influential
8 Altmetric
40.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!