2601.22054v1 Jan 29, 2026 cs.CV

MetricAnything: 다양한 노이즈 데이터 소스를 활용한 메트릭 심도 사전 학습의 확장

MetricAnything: Scaling Metric Depth Pretraining with Noisy Heterogeneous Sources

Donglin Di
Donglin Di
Citations: 1,137
h-index: 17
Baorui Ma
Baorui Ma
Citations: 15
h-index: 3
Jiahui Yang
Jiahui Yang
Citations: 9
h-index: 2
Xuancheng Zhang
Xuancheng Zhang
Citations: 28
h-index: 3
Jianxun Cui
Jianxun Cui
Citations: 37
h-index: 3
Hao Li
Hao Li
Citations: 55
h-index: 4
Yan Xie
Yan Xie
Citations: 89
h-index: 3
Wei Chen
Wei Chen
Citations: 89
h-index: 6

최근 비전 기반 모델의 발전은 확장(scaling)에 힘입어 이루어졌지만, 메트릭 심도 추정 분야에 이러한 방식을 적용하는 것은 여전히 어려운 과제입니다. 이는 이질적인 센서 노이즈, 카메라 의존적인 편향, 그리고 노이즈가 많은 교차 데이터 소스에서 발생하는 메트릭 모호성 때문입니다. 본 논문에서는 Metric Anything이라는 간단하고 확장 가능한 사전 학습 프레임워크를 소개합니다. 이 프레임워크는 수동으로 설계된 프롬프트, 카메라 특정 모델링 또는 작업별 아키텍처 없이, 다양한 노이즈 3D 데이터 소스에서 메트릭 심도를 학습합니다. 우리의 핵심 접근 방식은 '희소 메트릭 프롬프트(Sparse Metric Prompt)'입니다. 이는 심도 맵을 무작위로 마스킹하여 생성되며, 공간 추론을 센서 및 카메라의 편향으로부터 분리하는 보편적인 인터페이스 역할을 합니다. 약 2천만 개의 이미지-심도 쌍을 사용하여, 재구성된, 캡처된, 그리고 렌더링된 3D 데이터를 포함하며, 1만 개의 카메라 모델을 활용하여, 메트릭 심도 추정 분야에서 처음으로 명확한 확장 추세를 보여줍니다. 사전 학습된 모델은 심도 보완, 초해상도, 레이더-카메라 융합과 같은 프롬프트 기반 작업에서 뛰어난 성능을 보입니다. 또한, 프롬프트 없이 사용할 수 있도록 학습된 학생 모델은 단안 심도 추정, 카메라 내재 변수 복구, 단/다중 뷰 메트릭 3D 재구성, 그리고 VLA 계획에서 최고 수준의 결과를 달성합니다. 또한, Metric Anything의 사전 학습된 ViT 모델을 시각 인코더로 사용하는 것이 멀티모달 대규모 언어 모델의 공간 지능 능력을 크게 향상시킬 수 있음을 보여줍니다. 이러한 결과는 메트릭 심도 추정 분야가 현대적인 기반 모델을 구동하는 동일한 확장 법칙으로부터 이점을 얻을 수 있음을 보여주며, 확장 가능하고 효율적인 실세계 메트릭 인식을 위한 새로운 길을 제시합니다. MetricAnything은 http://metric-anything.github.io/metric-anything-io/ 에서 오픈 소스로 제공되어 연구 커뮤니티의 발전을 지원합니다.

Original Abstract

Scaling has powered recent advances in vision foundation models, yet extending this paradigm to metric depth estimation remains challenging due to heterogeneous sensor noise, camera-dependent biases, and metric ambiguity in noisy cross-source 3D data. We introduce Metric Anything, a simple and scalable pretraining framework that learns metric depth from noisy, diverse 3D sources without manually engineered prompts, camera-specific modeling, or task-specific architectures. Central to our approach is the Sparse Metric Prompt, created by randomly masking depth maps, which serves as a universal interface that decouples spatial reasoning from sensor and camera biases. Using about 20M image-depth pairs spanning reconstructed, captured, and rendered 3D data across 10000 camera models, we demonstrate-for the first time-a clear scaling trend in the metric depth track. The pretrained model excels at prompt-driven tasks such as depth completion, super-resolution and Radar-camera fusion, while its distilled prompt-free student achieves state-of-the-art results on monocular depth estimation, camera intrinsics recovery, single/multi-view metric 3D reconstruction, and VLA planning. We also show that using pretrained ViT of Metric Anything as a visual encoder significantly boosts Multimodal Large Language Model capabilities in spatial intelligence. These results show that metric depth estimation can benefit from the same scaling laws that drive modern foundation models, establishing a new path toward scalable and efficient real-world metric perception. We open-source MetricAnything at http://metric-anything.github.io/metric-anything-io/ to support community research.

7 Citations
0 Influential
8.5 Altmetric
49.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!