LiteMVS: 파운데이션 지식 증류 및 전문가 집계를 활용한 효율적인 다중 시야 스테레오
LiteMVS: Efficient Multi-View Stereo with Foundation Distillation and Expert Aggregation
실시간 3차원 인식은 로봇 공학, 증강 현실, 그리고 인지 시스템 분야에서 매우 중요합니다. 기존의 다중 시야 스테레오(MVS) 방법은 주로 기하학적 대응 관계에 의존하는데, 이는 질감이 없거나 반복적인 영역에서는 제대로 작동하지 않습니다. 반면 단안 깊이 모델은 이미지 레벨의 강력한 사전 지식을 활용하지만, 견고한 다중 시야 기하학적 제약 조건이 부족합니다. 더욱 중요하게는 로봇 및 인체 상호 작용 시나리오에서 고품질 3차원 형상은 정적인 재구성에 필수적일 뿐만 아니라 시간적으로 일관된 4차원 표현을 학습하는 데 중요한 기반이 됩니다. 구조적 인식 능력이 강화되고 시공간 확장 가능성이 더 높은 시각적 표현을 얻기 위해, 본 논문에서는 평면 스위프 기하학적 추론과 강력한 단안 의미 및 구조적 사전 지식을 통합한 경량 다중 시야 깊이 추정 모델인 LiteMVS를 제안합니다. LiteMVS의 핵심 아이디어는 경량 분할 모델 및 대규모 비전 파운데이션 모델에서 얻은 고수준 단안 지식을 효율적으로 다중 시야 스테레오 프레임워크에 주입하는 것입니다. 특히, LiteMVS는 의미 설명자를 사용하여 비용 볼륨을 풍부하게 하고, 믹스처 오브 익스퍼트(MoE) 방식을 사용하여 깊이 가설 간의 적응형 기하학적 집계를 가능하게 합니다. 또한, 비전 파운데이션 모델에서 추출한 기하학적 사전 지식은 추론 비용을 증가시키지 않으면서 단안 가이드 기능을 더욱 강화합니다. 이러한 설계 덕분에 LiteMVS는 정적인 장면에서 깊이 추정 및 3차원 재구성에 대한 품질을 향상시킬 뿐만 아니라, 후속 시간 모델링 및 4차원 표현 학습에 더 신뢰할 수 있는 기하학적 기반을 제공합니다. ScanNetv2 및 7-Scenes 데이터셋에서의 실험 결과는 LiteMVS가 높은 품질의 깊이 예측과 3차원 재구성을 달성하면서도 경쟁력 있는 효율성을 유지함을 보여줍니다.
Real-time 3D perception is crucial for robotics, augmented reality, and embodied intelligence applications. Existing multi-view stereo (MVS) methods primarily rely on geometric correspondences, which often fail in textureless or repetitive regions, while monocular depth models leverage strong image-level priors but lack robust multi-view geometric constraints. More importantly, in robotics and embodied manipulation scenarios, high-quality 3D geometry is not only essential for static reconstruction, but also serves as a critical foundation for learning temporally consistent 4D representations. To obtain visual representations with stronger structural awareness and greater potential for spatiotemporal extension, we present LiteMVS, a lightweight multi-view depth estimation model that integrates plane-sweep geometric reasoning with strong monocular semantic and structural priors. The central idea of LiteMVS is to efficiently inject high-level monocular knowledge, obtained from lightweight segmentation models and large-scale vision foundation models, into a multi-view stereo framework. In particular, LiteMVS enriches the cost volume with semantic descriptors and employs a Mixture-of-Experts (MoE) formulation to enable adaptive geometric aggregation across depth hypotheses. Moreover, geometric priors distilled from vision foundation models further strengthen monocular guidance without increasing inference cost. Through this design, LiteMVS not only improves depth estimation and 3D reconstruction quality in static scenes, but also provides a more reliable geometric foundation for subsequent temporal modeling and 4D representation learning. Experiments on ScanNetv2 and 7-Scenes demonstrate that LiteMVS achieves high-quality depth prediction and 3D reconstruction while maintaining competitive efficiency.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.