EgoTrack3D: 개인 시점을 이용한 3차원 객체 추적을 위한 모듈형 프레임워크
EgoTrack3D: A Modular Framework for Egocentric 3D Object Tracking
로봇 공학 및 자율 주행 시스템에서 개인 시점 영상을 통해 3차원 장면을 이해하는 것은 매우 중요하지만, 급격한 시점 변화와 부분적인 가려짐 현상은 체계적인 표현 구축을 어렵게 만듭니다. 기존의 3차원 추적 및 장면 그래프 구성 방법은 주로 명시적인 상호 작용을 다루거나 정적인 장면을 가정하며, 복잡한 동역학을 포착하는 데 한계를 가지고 있습니다. 본 논문에서는 개인 시점 RGB 영상으로부터 직접적으로 동적인 3차원 장면 표현을 재구성하고 유지하는 모듈형 프레임워크인 EgoTrack3D를 소개합니다. 이 프레임워크는 포인트 기반의 동작 점수 부여 메커니즘과 볼륨 기반 병합 휴리스틱을 사용하여 2차원 분할 마스크를 전역적인 3차원 좌표계로 변환하고, 객체 추적을 연결합니다. EgoTrack3D는 Aria Digital Twin (ADT) 데이터셋에서 가장 강력한 기준 모델 대비 정확 위치 비율(PCL)을 11% 향상시키는 성능을 보여주며, 정적 및 동적인 객체 모두에 대한 지속적인 3차원 추적이라는 더욱 일반적인 문제를 해결합니다. 또한, 실제 배포 환경의 제약 조건을 시뮬레이션하는 열악한 조건에서도 시스템의 견고성을 입증하기 위해, 자세한 깊이 정보를 희소한 3차원 바운딩 박스 추정으로 대체하고, 상호 작용 기반의 동적 연관을 통합하여, EgoTrack3D가 노이즈가 많은 관찰 데이터에도 불구하고 정확한 공간 표현을 유지할 수 있도록 합니다.
Understanding 3D scenes from egocentric video is fundamental for robotics and autonomous navigation, yet rapid viewpoint changes and partial occlusions make building structured representations challenging. Existing 3D tracking and scene graph construction methods primarily address explicit interactions or assume static scenes, limiting their ability to capture complex dynamics. We introduce EgoTrack3D, a modular framework that reconstructs and maintains a dynamic 3D scene representation directly from egocentric RGB video. The framework lifts 2D segmentation masks into a global 3D coordinate frame, using a point-based motion scoring mechanism alongside a voxel-based merging heuristic to associate object tracks. EgoTrack3D maintains accurate representations over time, achieving an 11% improvement in percentage of correct locations (PCL) relative to the strongest baseline on the Aria Digital Twin (ADT) dataset, while addressing the more general setting of persistent 3D tracking for both static and dynamic objects. Furthermore, to demonstrate the system's robustness under degraded conditions that simulate real-world deployment constraints, we replace dense depth maps with sparse 3D bounding box estimation and integrate interaction-guided dynamic association, enabling EgoTrack3D to maintain accurate spatial representations despite noisy observations.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.