2602.09929v3 Feb 10, 2026 cs.CV

쉐이딩 시퀀스 추정을 통한 단안 노멀 추정

Monocular Normal Estimation via Shading Sequence Estimation

Zong-Han Li
Zong-Han Li
Citations: 101
h-index: 5
Xin Ma
Xin Ma
Citations: 807
h-index: 5
Minghui Hu
Minghui Hu
Citations: 23
h-index: 1
Yunqing Zhao
Yunqing Zhao
Citations: 31
h-index: 4
Yingchen Yu
Yingchen Yu
Citations: 51
h-index: 5
Chang Liu
Chang Liu
Citations: 3
h-index: 1
Xudong Jiang
Xudong Jiang
Citations: 77
h-index: 5
Song Bai
Song Bai
Citations: 31
h-index: 4
Qian Zheng
Qian Zheng
Citations: 64
h-index: 5

단안 노멀 추정은 다양한 조명 조건에서 단일 RGB 이미지로부터 객체의 노멀 맵을 추정하는 것을 목표로 합니다. 기존 방법들은 주로 딥러닝 모델을 사용하여 직접 노멀 맵을 예측하지만, 이러한 방법들은 종종 3차원 정렬 문제를 겪습니다. 추정된 노멀 맵은 시각적으로 올바워 보일 수 있지만, 복원된 표면은 종종 실제 기하학적 세부 사항과 일치하지 않습니다. 우리는 이러한 정렬 오류가 현재의 패러다임에서 비롯된다고 주장합니다. 모델은 노멀 맵에 표현된 다양한 기하학적 구조를 구별하고 재구성하는 데 어려움을 겪는데, 이는 기저의 기하학적 차이가 상대적으로 미묘한 색상 변화를 통해서만 반영되기 때문입니다. 이 문제를 해결하기 위해, 우리는 노멀 추정을 쉐이딩 시퀀스 추정으로 재정의하는 새로운 패러다임을 제안합니다. 쉐이딩 시퀀스는 다양한 기하학적 정보에 더 민감하게 반응합니다. 이 패러다임을 기반으로, 우리는 이미지-비디오 생성 모델을 활용하여 쉐이딩 시퀀스를 예측하는 방법인 RoSE를 제시합니다. 예측된 쉐이딩 시퀀스는 간단한 최소 제곱 문제를 해결하여 노멀 맵으로 변환됩니다. RoSE는 견고성을 향상시키고 복잡한 객체를 더 잘 처리하기 위해 다양한 모양, 재료 및 조명 조건을 갖춘 합성 데이터셋인 MultiShade를 사용하여 학습되었습니다. 실험 결과, RoSE는 객체 기반 단안 노멀 추정을 위한 실제 벤치마크 데이터셋에서 최첨단 성능을 달성하는 것으로 나타났습니다.

Original Abstract

Monocular normal estimation aims to estimate the normal map from a single RGB image of an object under arbitrary lights. Existing methods rely on deep models to directly predict normal maps. However, they often suffer from 3D misalignment: while the estimated normal maps may appear to have a correct appearance, the reconstructed surfaces often fail to align with the geometric details. We argue that this misalignment stems from the current paradigm: the model struggles to distinguish and reconstruct varying geometry represented in normal maps, as the differences in underlying geometry are reflected only through relatively subtle color variations. To address this issue, we propose a new paradigm that reformulates normal estimation as shading sequence estimation, where shading sequences are more sensitive to various geometric information. Building on this paradigm, we present RoSE, a method that leverages image-to-video generative models to predict shading sequences. The predicted shading sequences are then converted into normal maps by solving a simple ordinary least-squares problem. To enhance robustness and better handle complex objects, RoSE is trained on a synthetic dataset, MultiShade, with diverse shapes, materials, and light conditions. Experiments demonstrate that RoSE achieves state-of-the-art performance on real-world benchmark datasets for object-based monocular normal estimation.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!