ContactFlow: 로봇의 다양한 형태에 적용 가능한 비디오 기반 동작 제어 방법
ContactFlow: A video action conditioning that transfers across embodiments
세계 모델은 로봇 계획에서 유망한 접근 방식으로, 에이전트가 실제 실행 전에 행동의 결과를 상상하고 검증할 수 있도록 합니다. 그러나 현재의 비디오 기반 세계 모델은 특히 접촉과 관련된 물리적 제약을 제대로 반영하는 데 어려움을 겪는 경우가 많습니다. 또한, 이러한 모델의 동작 제어는 종종 평행 그리퍼와 같은 특정 로봇 형태에 국한됩니다. 본 연구에서는 extit{Contact Flow}라는 새로운 방법을 제안합니다. 이는 로봇의 형태에 구애받지 않는 동작 표현 방식으로, 액터와 대상 물체 사이의 3차원 접촉 지점의 경로를 통해 조작을 인코딩합니다. Contact Flow는 액터의 특정 외형 및 운동학적 특성을 무시하여 인간 시연과 로봇 실행 모두에 대한 공유 조건부 신호를 제공합니다. 따라서, 우리는 Contact Flow를 기반으로 인간과 로봇의 객체 상호 작용 비디오 데이터셋을 사용하여 대규모 비디오 생성 모델을 학습할 수 있으며, 이를 통해 물리적으로 타당한 조작 결과를 예측하는 세계 모델을 구축할 수 있습니다. 제안된 모델은 '제안-상상-검증-실행' 파이프라인에 통합되어, 생성된 시뮬레이션 결과가 실제 실행 전에 시각-언어 모델에 의해 평가됩니다. DROID 데이터셋과 실제 탁상 조작 작업에서의 실험 결과는 Contact Flow가 인간 시연과 다양한 로봇 형태 간의 이전(transfer)을 가능하게 함을 보여줍니다.
World models offer a promising route toward robot planning by enabling agents to imagine and verify the consequences of actions before execution. However, current video-based world models often struggle to capture the physical constraints that govern manipulation, particularly contact. Further, their action conditioning is often constrained to specific embodiments such as parallel grippers. We propose \emph{Contact Flow}, an embodiment-agnostic action representation that encodes manipulation through the trajectory of 3D contact points between an actor and a target object. By discarding actor-specific appearance and kinematics, Contact Flow provides a shared conditioning signal for both human demonstrations and robotic execution. Therefore, we can train a large-scale video generative model on both human and robotic object interaction videos conditioned on Contact Flow, yielding a world model that predicts physically plausible manipulation outcomes. We integrate this model into a propose-imagine-verify-act pipeline, where generated rollouts are assessed by a vision-language model before execution. Experiments on the DROID dataset and real-world tabletop manipulation tasks demonstrate that Contact Flow enables transfer between human demonstrations and different robotic embodiments.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.