EmbodiedVAE: 효율적이고 제어 가능한 로봇 조작을 위한 분리된 비디오 VAE
EmbodiedVAE: Disentangled Video VAE for Efficient and Controllable Embodied Manipulation
최근, 잠재 확산 모델(LDMs)은 강력한 로봇 조작 환경 모델 구축에 기여하며, 임베디드 학습 분야에서 상당한 발전을 이루었습니다. 그러나 기존의 LDMs는 자연 장면을 위해 최적화된 변분 오토인코더(VAEs)에 주로 의존하며, 로봇 조작 시나리오의 고유한 특성을 고려하지 못하여, 결과적으로 효율적인 LDM 학습과 정밀한 로봇 제어를 방해하는 비효율적이고 제어 불가능한 잠재 표현을 생성합니다. 이러한 문제를 해결하기 위해, 본 논문에서는 로봇 조작 환경 모델에 적합한 압축적이고 제어 가능한 잠재 표현을 제공하는 새로운 비디오 VAE인 EmbodiedVAE를 제안합니다. 구체적으로, EmbodiedVAE는 비대칭 시공간 압축 모듈을 갖춘 이중 인코더-단일 디코더 구조를 채택하여 로봇 팔의 움직임을 배경 환경으로부터 자동으로 분리함으로써 전체적인 압축률을 높이는 동시에 명시적인 임베디드 잠재 표현을 제공하여 정밀한 동작 제어를 지원합니다. 또한, 학습된 로봇 동작 잠재의 시간적 일관성을 유지하기 위해, 최적 수송 기반의 일관성 모듈을 도입하여 동작의 충실성과 프레임 간의 일관성을 명시적으로 강화했습니다. 광범위한 실험 결과는 EmbodiedVAE가 우수한 재구성 품질과 높은 압축률을 달성할 뿐만 아니라, 로봇 조작 시나리오에서 기존의 최첨단 비디오 VAE에 비해 평균 2dB PSNR 향상을 통해 더욱 정밀한 동작 제어를 가능하게 함을 보여줍니다.
Latent diffusion models (LDMs) have recently significantly advanced embodied learning in constructing powerful embodied manipulation world models. However, despite the remarkable performance, existing LDMs predominantly rely on Variational Autoencoders (VAEs) optimized for natural scenes while failing to account for the unique characteristics of embodied manipulation scenarios, yielding latent representations that are neither compact nor controllable, thereby hindering efficient training of LDMs and precise robotic control. To solve this problem, we present EmbodiedVAE, a novel video VAE that provides compact yet controllable latent representations tailored for the robotic manipulation world models. Specifically, EmbodiedVAE adopts a dual-encoder, single-decoder architecture with an asymmetric spatio-temporal compression module, which automatically disentangles the robot arm's motion from background environment, resulting in overall compactness while providing explicit embodied latent to support fine-grained action control. To further preserve the temporal consistency of learned robotic motion latent, we introduce an optimal-transport-based consistency module that explicitly enforces motion fidelity and inter-frame coherence. Extensive experiments demonstrate that our proposed EmbodiedVAE achieves superior reconstruction quality with high compression rate, while enabling more precise action control in robotic manipulation scenarios with an average of 2dB PSNR improvement over state-of-the-art video VAEs.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.