강화된 조작 정책 학습을 위한 기하학적 정보를 활용한 모션 잠재 표현
Geometry-Aware Motion Latents for Learning Robust Manipulation Policies
로봇 조작을 위한 모션 잠재 표현 학습은 주로 시각 정보에서 모션 패턴을 추출하는 데 의존하지만, 효과적인 동작 추상화를 위해서는 3차원 기하학적 변환에 대한 이해가 필요합니다. 본 연구에서는 GeoMoLa (Geometry-Aware Motion Latents)를 소개하며, 이는 조작 과정 동안 점군(point cloud)의 변화를 예측하여 이산적인 모션 잠재 코드를 학습합니다. 시각적 관찰 결과를 재구성하는 대신, 시간 경과에 따른 공간 기하학적 변화라는 4차원 목표를 설정함으로써, 잠재 표현이 실제 물리적 움직임을 인코딩하도록 유도합니다. GeoMoLa는 기존 방법들이 필요로 하는 다중 카메라 이미지를 사용하지 않고 단일 RGB-D 입력만으로 최첨단 성능을 달성하며, 다양한 조작 벤치마크에서 뛰어난 결과를 보입니다. 분석 결과, 기하학적 예측이 성능 향상에 핵심적인 역할을 한다는 것을 확인했으며, 이는 조작이 공간 이해 능력에 의존한다는 점을 정량적으로 검증합니다. 또한, 학습된 코드는 효과적인 모션 추상화를 제공하며, 새로운 환경에 적용할 때 시각적 맥락과 관계없이 물리적으로 일관된 변환을 수행합니다. 실제 환경에서의 실험 결과도 이러한 강건성을 뒷받침하며, 복잡한 환경에서 최소한의 데모만으로 안정적인 조작이 가능함을 보여줍니다. 따라서, 로봇 제어를 위한 효과적인 모션 잠재 표현은 픽셀 수준의 패턴보다는 움직임 자체의 3차원 효과를 이해함으로써 더 잘 얻을 수 있음을 입증합니다.
Learning motion latents for robotic manipulation heavily relies on extracting motion patterns from visual sequences, yet effective action abstractions require understanding three-dimensional geometric transformations. Here, we introduce GeoMoLa (Geometry-Aware Motion Latents), which learns discrete motion latent codes by predicting how point clouds evolve during manipulation rather than reconstructing visual observations. This four-dimensional objective -- spatial geometry changing through time -- forces latent representations to encode actual physical motion rather than appearance patterns. GeoMoLa achieves state-of-the-art performance using only single-view RGB-D input, while existing methods require multi-view reconstruction, succeeding across diverse manipulation benchmarks. Our ablations reveal that geometric prediction is the key to driving performance, quantitatively validating that manipulation depends on spatial understanding. Furthermore, the learned codes exhibit effective motion abstraction: applying them to novel scenes produces physically consistent transformations regardless of visual context. Our real-world experiments also confirm this robustness capability, achieving robust manipulation with minimal demonstrations in cluttered environments where geometric reasoning determines success. Thus, we demonstrate that effective motion latents for robot control can better emerge from understanding motion through its three-dimensional effects rather than pixel-level patterns.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.