2607.24159v1 Jul 27, 2026 cs.RO

DeVA: 물리적 지침을 활용한 비디오-액션 분리 모델을 이용한 로봇 정책 학습

DeVA: Decoupled Video-Action Model with physical guidance for robot policy learning

Unnat Jain
Unnat Jain
Citations: 19
h-index: 2
Judy Hoffman
Judy Hoffman
Citations: 231
h-index: 4
Meng Zhang
Meng Zhang
Citations: 0
h-index: 0
Sahil Khose
Sahil Khose
Citations: 64
h-index: 5
Simar Kareer
Simar Kareer
Citations: 297
h-index: 6
Yuchen Song
Yuchen Song
Citations: 0
h-index: 0

일반화 가능한 로봇 조작은 시각적 장면의 변화를 예측하면서 언어 명령을 실행할 수 있는 정책이 필요합니다. 최근의 비전-언어-액션 모델들은 대규모 사전 훈련으로부터 이점을 얻지만, 주로 정적인 사전 훈련 목표는 물리적 역학과 시간적 인과관계에 대한 제한적인 지침을 제공하며, 제어와 관련된 지식은 다운스트림 로봇 데모를 통해 학습되어야 합니다. 비디오 생성 모델은 미래 예측을 통해 풍부한 시공간적 선행 정보를 포함하므로 유망한 기반이 될 수 있습니다. 그러나 기존의 비디오-액션 모델들은 비디오 및 액션 예측을 공유 백본에서 결합하여 정책 적응을 최적화하기 어렵게 만들거나, 액션 브랜치를 안내할 때 비디오 정보를 충분히 활용하지 못합니다. 본 연구에서는 전문적인 비디오 및 액션 전문가, 다단계 특징 전송, 그리고 물리적으로 중요한 지침을 갖춘 분리된 비디오-액션 모델인 DeVA를 소개합니다. DeVA는 여러 비디오 레이어에서 추출한 표현을 액션 전문가로 전달하여 풍부한 정보 교환을 가능하게 하면서 정책 학습을 더 쉽게 만듭니다. 또한, 중간 비디오 특징과 액션 스트림을 물리적으로 중요한 지침(affordance/depth)으로 감독합니다. 시뮬레이션 벤치마크 및 실제 환경에서의 실험 결과는 제한된 데이터로 강력한 성능, 통합 아키텍처보다 빠른 수렴 속도, 그리고 물리적 지침으로부터 얻은 명확한 성능 향상을 보여줍니다.

Original Abstract

Generalizable robot manipulation requires policies that can anticipate how visual scenes evolve while executing language instructions. While recent Vision-Language-Action models benefit from large-scale pretraining, their predominantly static pretraining objectives provide limited supervision for physical dynamics and temporal causality, leaving control-relevant knowledge to be learned from downstream robot demonstrations. Video generative models offer a promising foundation by encoding rich spatiotemporal priors through future predictions. However, existing Video-Action Models either couple video and action prediction in a shared backbone, making policy adaptation harder to optimize, or under-utilize video information when guiding the action branch. In this work, we introduce DeVA, a Decoupled Video-Action model with specialized video and action experts, multi-level feature transfer, and physically salient guidance. DeVA transfers representations from multiple video layers to the action expert, enabling rich information exchange while making policy learning more tractable. It further supervises intermediate video features and the action stream with physically salient guidance (affordance/depth). Experiments on both simulation benchmarks and real-world deployment demonstrate strong performance with limited data, faster convergence than a unified architecture, and clear performance gains from physical guidance.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!