2608.02068v1 Aug 03, 2026 cs.CV

GIFT: 기하학적 불변성을 활용한 미세 조정 - 람베르트 법칙을 벗어난 단안 깊이 추정

GIFT: Geometry-Invariant Fine-Tuning for Non-Lambertian Monocular Depth Estimation

Xiangru Huang
Xiangru Huang
Citations: 6
h-index: 2
Bingqian Wu
Bingqian Wu
Citations: 44
h-index: 2
Xianghui Fan
Xianghui Fan
Citations: 13
h-index: 2
Dayu Li
Dayu Li
Citations: 0
h-index: 0
Zhaoyu Chen
Zhaoyu Chen
Citations: 0
h-index: 0
Xinlu Zeng
Xinlu Zeng
Citations: 2
h-index: 1
Huanran Cui
Huanran Cui
Citations: 0
h-index: 0
Guangzheng Xu
Guangzheng Xu
Citations: 11
h-index: 2
Hang Yang
Hang Yang
Citations: 17
h-index: 3

대규모 합성 학습 데이터를 기반으로 하는 단안 깊이 예측 모델은 뛰어난 일반화 성능을 보여주지만, 종종 람베르트 법칙을 벗어난 표면에서 잘못된 깊이를 추정하는 경향이 있습니다. 특히 거울에 반사된 내용이나 유리 뒤의 투과된 내용을 실제 표면으로 오인하는 경우가 발생합니다. 이러한 모델을 실세계 데이터로 조정하는 것은 기존 깊이 센서 역시 그러한 영역에서 신뢰성이 낮기 때문에 어렵습니다. 우리는 람베르트 법칙을 벗어난 표면의 외관은 주변 환경에 따라 달라지지만, 그 기하학적 구조는 변하지 않는다는 점을 발견했습니다. 이러한 관찰을 바탕으로, 측정된 깊이 레이블 없이도 적용 가능한 효율적인 파라미터 미세 조정 프레임워크인 GIFT (Geometry-Invariant Fine-Tuning)를 제안합니다. 우리는 카메라와 대상의 기하학적 구조는 고정한 채, RGB 이미지들을 다양한 외관 변화 조건에서 수집했습니다. GIFT는 이러한 관찰에서의 기하학적 불변성을 활용하여 람베르트 법칙을 벗어난 표면에서의 잘못된 깊이 추정을 줄이는 동시에 일반적인 깊이 예측 능력을 유지합니다. 또한, 우리는 람베르트 법칙을 벗어난 깊이 복구 성능, 외관 변화에 대한 강건성, 그리고 다른 영역에서의 성능 유지를 평가하는 제어된 벤치마크를 구축했습니다. 벤치마크와 독립적인 실세계 데이터셋에서 수행한 실험 결과, GIFT는 거울 및 투명 물체에 대한 깊이 예측 정확도를 향상시키면서도 기본 모델의 성능을 크게 유지하며, 단안 깊이 예측 모델을 람베르트 법칙을 벗어난 장면으로 적용하는 데 있어 실용적이고 저렴한 방법을 제공합니다.

Original Abstract

Monocular depth foundation models, benefiting from large-scale synthetic training data, have demonstrated strong generalization. However, they often hallucinate depth on non-Lambertian surfaces, estimating reflected content in mirrors or transmitted content behind glass rather than the physical surface itself. Adapting these models with real-world data is challenging because conventional depth sensors are also unreliable in such regions. We observe that while the appearance of a non-Lambertian surface varies with its reflected or transmitted environment, its underlying geometry remains unchanged. Based on this observation, we propose GIFT (Geometry-Invariant Fine-Tuning), a parameter-efficient post-training framework that requires no measured depth labels. We collect groups of RGB images under controlled appearance changes while keeping the camera and target geometry fixed. GIFT exploits geometric invariance across these observations to suppress non-Lambertian depth hallucinations while retaining general depth estimation capability. We further construct a controlled benchmark that evaluates non-Lambertian depth recovery, robustness to appearance changes, and performance retention in other regions. Experiments on our benchmark and an independent real-world dataset demonstrate that GIFT improves depth prediction for mirrors and transparent objects while largely preserving the base model's performance, providing a practical and low-cost approach for adapting monocular depth foundation models to non-Lambertian scenes.

0 Citations
0 Influential
1.5 Altmetric
7.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!