FoundObj: 라벨 없이 3차원 객체 분할을 위한 자기 지도 기반 기초 모델을 활용한 보상 시스템
FoundObj: Self-supervised Foundation Models as Rewards for Label-free 3D Object Segmentation
본 논문에서는 학습 과정에서 어떠한 장면 수준의 인간 주석도 필요하지 않으면서 복잡한 장면의 점군 데이터에서 3차원 객체 분할이라는 어려운 문제를 해결합니다. 기존 방법들은 일반적으로 학습 과정에서의 객체에 대한 충분한 사전 지식 부족으로 인해 단순한 객체 식별에 제한되는 경우가 많습니다. 본 논문에서는 FoundObj라는 새로운 프레임워크를 제시합니다. 이는 초점 영역 기반 객체 발견 에이전트를 특징으로 하며, 혁신적인 의미론적 및 기하학적 보상 모듈의 지침에 따라 적절한 인접 초점 영역을 점진적으로 병합합니다. 이러한 모듈은 자기 지도 방식으로 학습된 2D/3D 기초 모델에서 얻은 의미론적 및 기하학적 정보를 활용하여 객체 발견 에이전트에 상호 보완적인 피드백을 제공하고, 강화 학습을 통해 다양한 클래스의 객체를 안정적으로 식별할 수 있도록 합니다. 다양한 벤치마크에서의 광범위한 실험 결과는 제안하는 방법이 기존의 기본 모델보다 일관되게 우수한 성능을 나타냄을 보여줍니다. 특히, 본 방법은 제로샷(zero-shot) 및 장기 꼬리(long-tail) 시나리오에서 강력한 일반화 능력을 보이며, 이는 확장 가능하고 라벨이 없는 3차원 객체 분할에 대한 잠재력을 강조합니다.
We address the challenging task of 3D object segmentation in complex scene point clouds without relying on any scene-level human annotations during training. Existing methods are typically constrained to identifying simple objects, primarily due to insufficient object priors in the learning process. In this paper, we present FoundObj, a novel framework featuring a superpoint-based object discovery agent that incrementally merges suitable neighboring superpoints, guided by our innovative semantic and geometric reward modules. These modules synergistically leverage semantic and geometric priors from self-supervised 2D/3D foundation models, providing complementary feedback to the object discovery agent and enabling robust identification of multi-class objects through reinforcement learning. Extensive experiments on diverse benchmarks demonstrate that our approach consistently outperforms existing baselines. Notably, our method exhibits strong generalization in zero-shot and long-tail scenarios, underscoring its potential for scalable, label-free 3D object segmentation.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.