오픈 월드 모바일 조작을 위한 연결된 3D 장면 그래프
Articulated 3D Scene Graphs for Open-World Mobile Manipulation
시맨틱 정보는 3D 장면 이해 및 활용도 기반 객체 상호작용을 가능하게 합니다. 그러나 실제 환경에서 작동하는 로봇은 객체의 움직임을 예측할 수 없다는 중요한 제약 조건을 가지고 있습니다. 장기적인 모바일 조작은 시맨틱, 기하학, 운동학 사이의 격차를 해소해야 합니다. 본 연구에서는 다양한 상호작용 가능한 객체를 포함하는 연결된 장면의 시맨틱-운동학 3D 장면 그래프를 구축하기 위한 새로운 프레임워크인 MoMa-SG를 제시합니다. 여러 객체의 연결 정보를 포함하는 RGB-D 시퀀스가 주어지면, 객체 상호작용을 시간적으로 분할하고, 가려짐에 강건한 포인트 추적을 사용하여 객체 움직임을 추론합니다. 그런 다음, 포인트 궤적을 3D로 변환하고, 새로운 통합된 트위스트 추정 방식을 사용하여 회전 및 평행 운동 관절 매개변수를 단일 최적화 과정을 통해 강건하게 추정합니다. 다음으로, 추정된 연결 정보와 함께 객체를 연결하고, 식별된 열림 상태에서 부모-자식 관계를 기반으로 포함된 객체를 감지합니다. 또한, 본 연구에서는 계층적 객체 시맨틱 정보(부모-자식 관계 레이블 포함)와 62개의 실제 RGB-D 시퀀스에 대한 객체 축 주석을 결합하는 독특한 Arti4D-Semantic 데이터셋을 소개합니다. 이 데이터셋에는 600개의 객체 상호작용과 세 가지의 서로 다른 관찰 패러다임이 포함되어 있습니다. MoMa-SG의 성능을 두 개의 데이터셋에서 광범위하게 평가하고, 접근 방식의 핵심 설계 선택 사항에 대한 분석을 수행했습니다. 또한, 사족보행 로봇과 모바일 조작 로봇을 사용한 실제 실험을 통해, 제안하는 시맨틱-운동학 장면 그래프가 일상적인 가정 환경에서 연결된 객체의 강력한 조작을 가능하게 한다는 것을 입증했습니다. 코드와 데이터는 다음 웹사이트에서 제공합니다: https://momasg.cs.uni-freiburg.de.
Semantics has enabled 3D scene understanding and affordance-driven object interaction. However, robots operating in real-world environments face a critical limitation: they cannot anticipate how objects move. Long-horizon mobile manipulation requires closing the gap between semantics, geometry, and kinematics. In this work, we present MoMa-SG, a novel framework for building semantic-kinematic 3D scene graphs of articulated scenes containing a myriad of interactable objects. Given RGB-D sequences containing multiple object articulations, we temporally segment object interactions and infer object motion using occlusion-robust point tracking. We then lift point trajectories into 3D and estimate articulation models using a novel unified twist estimation formulation that robustly estimates revolute and prismatic joint parameters in a single optimization pass. Next, we associate objects with estimated articulations and detect contained objects by reasoning over parent-child relations at identified opening states. We also introduce the novel Arti4D-Semantic dataset, which uniquely combines hierarchical object semantics including parent-child relation labels with object axis annotations across 62 in-the-wild RGB-D sequences containing 600 object interactions and three distinct observation paradigms. We extensively evaluate the performance of MoMa-SG on two datasets and ablate key design choices of our approach. In addition, real-world experiments on both a quadruped and a mobile manipulator demonstrate that our semantic-kinematic scene graphs enable robust manipulation of articulated objects in everyday home environments. We provide code and data at: https://momasg.cs.uni-freiburg.de.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.