2602.16356v1 Feb 18, 2026 cs.RO

오픈 월드 모바일 조작을 위한 연결된 3D 장면 그래프

Articulated 3D Scene Graphs for Open-World Mobile Manipulation

A. Valada
A. Valada
Citations: 1,221
h-index: 17
Martin Buchner
Martin Buchner
Citations: 189
h-index: 4
Adrian Rofer
Adrian Rofer
Citations: 20
h-index: 2
Tim Engelbracht
Tim Engelbracht
Citations: 10
h-index: 2
T. Welschehold
T. Welschehold
Citations: 1,003
h-index: 17
Z. Bauer
Z. Bauer
Citations: 233
h-index: 7
Hermann Blum
Hermann Blum
Citations: 18
h-index: 2
Marc Pollefeys
Marc Pollefeys
Citations: 1,371
h-index: 19

시맨틱 정보는 3D 장면 이해 및 활용도 기반 객체 상호작용을 가능하게 합니다. 그러나 실제 환경에서 작동하는 로봇은 객체의 움직임을 예측할 수 없다는 중요한 제약 조건을 가지고 있습니다. 장기적인 모바일 조작은 시맨틱, 기하학, 운동학 사이의 격차를 해소해야 합니다. 본 연구에서는 다양한 상호작용 가능한 객체를 포함하는 연결된 장면의 시맨틱-운동학 3D 장면 그래프를 구축하기 위한 새로운 프레임워크인 MoMa-SG를 제시합니다. 여러 객체의 연결 정보를 포함하는 RGB-D 시퀀스가 주어지면, 객체 상호작용을 시간적으로 분할하고, 가려짐에 강건한 포인트 추적을 사용하여 객체 움직임을 추론합니다. 그런 다음, 포인트 궤적을 3D로 변환하고, 새로운 통합된 트위스트 추정 방식을 사용하여 회전 및 평행 운동 관절 매개변수를 단일 최적화 과정을 통해 강건하게 추정합니다. 다음으로, 추정된 연결 정보와 함께 객체를 연결하고, 식별된 열림 상태에서 부모-자식 관계를 기반으로 포함된 객체를 감지합니다. 또한, 본 연구에서는 계층적 객체 시맨틱 정보(부모-자식 관계 레이블 포함)와 62개의 실제 RGB-D 시퀀스에 대한 객체 축 주석을 결합하는 독특한 Arti4D-Semantic 데이터셋을 소개합니다. 이 데이터셋에는 600개의 객체 상호작용과 세 가지의 서로 다른 관찰 패러다임이 포함되어 있습니다. MoMa-SG의 성능을 두 개의 데이터셋에서 광범위하게 평가하고, 접근 방식의 핵심 설계 선택 사항에 대한 분석을 수행했습니다. 또한, 사족보행 로봇과 모바일 조작 로봇을 사용한 실제 실험을 통해, 제안하는 시맨틱-운동학 장면 그래프가 일상적인 가정 환경에서 연결된 객체의 강력한 조작을 가능하게 한다는 것을 입증했습니다. 코드와 데이터는 다음 웹사이트에서 제공합니다: https://momasg.cs.uni-freiburg.de.

Original Abstract

Semantics has enabled 3D scene understanding and affordance-driven object interaction. However, robots operating in real-world environments face a critical limitation: they cannot anticipate how objects move. Long-horizon mobile manipulation requires closing the gap between semantics, geometry, and kinematics. In this work, we present MoMa-SG, a novel framework for building semantic-kinematic 3D scene graphs of articulated scenes containing a myriad of interactable objects. Given RGB-D sequences containing multiple object articulations, we temporally segment object interactions and infer object motion using occlusion-robust point tracking. We then lift point trajectories into 3D and estimate articulation models using a novel unified twist estimation formulation that robustly estimates revolute and prismatic joint parameters in a single optimization pass. Next, we associate objects with estimated articulations and detect contained objects by reasoning over parent-child relations at identified opening states. We also introduce the novel Arti4D-Semantic dataset, which uniquely combines hierarchical object semantics including parent-child relation labels with object axis annotations across 62 in-the-wild RGB-D sequences containing 600 object interactions and three distinct observation paradigms. We extensively evaluate the performance of MoMa-SG on two datasets and ablate key design choices of our approach. In addition, real-world experiments on both a quadruped and a mobile manipulator demonstrate that our semantic-kinematic scene graphs enable robust manipulation of articulated objects in everyday home environments. We provide code and data at: https://momasg.cs.uni-freiburg.de.

6 Citations
0 Influential
9.5 Altmetric
53.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!