2607.12433v1 Jul 14, 2026 cs.CV

ARDepth: 점진적인 시각적 조건을 활용한 자기회귀 단안 깊이 추정

ARDepth: Auto-regressive Monocular Depth Estimation with Progressive Visual Conditioning

Guanbin Li
Guanbin Li
Citations: 3,856
h-index: 32
Xiao Tan
Xiao Tan
Citations: 28
h-index: 3
Zijie Wang
Zijie Wang
Citations: 10
h-index: 2
Wei Zhang
Wei Zhang
Citations: 79
h-index: 5
Weiming Zhang
Weiming Zhang
Citations: 158
h-index: 6
Weikai Chen
Weikai Chen
Citations: 657
h-index: 13
Xiaoxu Li
Xiaoxu Li
Citations: 0
h-index: 0

최근 확산 모델은 단안 깊이 추정(MDE)의 주류 패러다임으로 자리 잡았습니다. 그러나 이러한 모델은 일반적으로 반복적인 노이즈 제거 과정을 통해 깊이를 전역적으로 부드러운 필드로 복원한다고 가정하는데, 이는 장면 기하학의 분할적이고 크기 의존적인 구조를 명시적으로 반영하지 않습니다. 실제로 기하학적 구조는 공간 척도에 따라 점진적으로 나타나며, 거친 레이아웃, 표면 및 경계선은 계층적으로 구성됩니다. 이러한 관찰을 바탕으로, 우리는 깊이 추정을 구조화된 자기회귀 생성을 통해 정의하는 ARDepth를 제안합니다. ARDepth는 전역적인 정제 과정을 통해 깊이를 복원하는 대신, 공간 해상도가 증가함에 따라 점진적으로 깊이 표현을 구성합니다. 이 생성 프로세스를 지원하기 위해, 각 생성 단계에서 다중 척도 시각적 특징을 주입하는 스케일-점진적 조건부(SPC) 및 장면 수준의 의미론적 사전 지식을 제공하여 전역적인 구조 일관성을 향상시키는 의미론적 인식 가이드(SAG)를 도입했습니다. 이러한 설계는 모델이 미세한 로컬 디테일을 캡처하면서도 일관된 전역 기하학을 유지할 수 있도록 합니다. 실험 결과는 ARDepth가 강력한 성능을 달성하며, 다양한 스케일에서 구조적으로 일관된 깊이 예측 결과를 생성한다는 것을 보여주며, 이는 자기회귀 생성이 기하학적 모델링의 유망한 대체 패러다임임을 입증합니다.

Original Abstract

Diffusion models have recently become the dominant paradigm for monocular depth estimation (MDE). However, they implicitly assume that depth can be recovered as a globally smooth field through iterative denoising, which does not explicitly reflect the piecewise and scale-dependent organization of scene geometry. In practice, geometric structure emerges progressively across spatial scales, where coarse layout, surfaces, and boundaries are constructed in a hierarchical manner. Motivated by this observation, we introduce ARDepth, which formulates depth estimation as structured auto-regressive generation. Instead of recovering depth through global refinement, ARDepth progressively constructs depth representations as spatial resolution increases. To support this generative process, we introduce Scale-Progressive Conditioning (SPC) to inject multi-scale visual features at each generation stage, and Semantic-Aware Guidance (SAG) to provide scene-level semantic priors that enhance global structural consistency. Together, these designs enable the model to capture fine-grained local details while maintaining coherent global geometry. Empirical results demonstrate that our approach achieves strong performance and produces structurally consistent depth predictions across scales, validating auto-regressive generation as a promising alternative paradigm for geometric modeling.

0 Citations
0 Influential
16 Altmetric
80.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!