2607.28581v1 Jul 30, 2026 cs.CV

ROAD: 상호 작용적 객관 정렬을 통한 판별적 의미론 활용 기반 3차원 형상 생성

ROAD: Reciprocal-Objective Alignment of Discriminative Semantics for 3D Shape Generation

Dingkang Liang
Dingkang Liang
Citations: 32
h-index: 3
Xin Zhou
Xin Zhou
Citations: 658
h-index: 11
Xiao Luo
Xiao Luo
Citations: 0
h-index: 0
Mingyang Du
Mingyang Du
Citations: 21
h-index: 2
Tianrui Feng
Tianrui Feng
Citations: 96
h-index: 3
Xiwu Chen
Xiwu Chen
Citations: 256
h-index: 5
Xiaofan Li
Xiaofan Li
Citations: 47
h-index: 3
Jiangning Zhang
Jiangning Zhang
Citations: 97
h-index: 6

고품질 3차원 형상 생성은 주로 모델의 용량과 데이터 규모를 확장하는 방식으로 이루어지며, 이는 엄청난 계산 비용을 초래합니다. 이러한 방식은 일반적으로 처음부터 기하학적 정보를 학습해야 하며, 판별적 3차원 기반 모델에 이미 내재된 풍부한 의미론적 및 구조적 사전 지식을 간과합니다. 본 연구에서는 이러한 판별적 모델이 가진 3차원 세계에 대한 깊은 이해를 활용하면 생성 비용을 크게 줄일 수 있다고 주장합니다. 이를 위해, 우리는 풍부한 판별적 사전 지식을 확산 트랜스포머로 이전하여 3차원 형상 생성을 위한 학습 비용을 절감하는 프레임워크인 ROAD를 제안합니다. 생성 모델과 판별적 모델 간의 고유한 의미론-구조적 이질성을 해결하기 위해, 우리는 상호 작용적 객관 정렬 전략을 도입했습니다. 이 방법은 전체적인 의미론적 일관성을 강화하는 전체적 의미론 압축(Holistic Semantic Condensing)과 미세한 기하학적 세부 사항 간의 엄격한 정렬을 위해 양방향 매칭 문제로 정의된 구조 최적 정렬(Structural Optimal Alignment)을 결합합니다. 3차원 기반 모델은 학습 과정에서 정렬을 감독하는 데만 사용되며, 추론 과정에서는 사용되지 않으므로 추가적인 추론 비용이 발생하지 않습니다. 제안하는 ROAD는 산업 표준인 Step1X-3D와 비교하여 훨씬 적은 양의 데이터(1.5%)로 경쟁력 있는 생성 성능을 달성하며, 학습 비용을 크게 절감하여 고품질 3차원 형상 생성이 수반하는 계산 부담을 효과적으로 줄입니다. 코드: https://github.com/H-EmbodVis/ROAD

Original Abstract

High-fidelity 3D generation predominantly relies on scaling model capacity and data, which incurs prohibitive computational costs. This paradigm typically requires learning geometry from scratch and overlooks the rich semantic and structural priors already encapsulated in discriminative 3D foundation models. We contend that leveraging the profound understanding of the 3D world possessed by these discriminative models can significantly reduce generative cost. To this end, we propose ROAD, a framework that reduces the training cost of 3D generation by transferring these rich discriminative priors into diffusion transformers. To address the inherent semantic-structural heterogeneity between generative and discriminative latents, we introduce a reciprocal-objective alignment strategy. This method synergizes Holistic Semantic Condensing to enforce global semantic coherence and Structural Optimal Alignment, which is formulated as a bipartite matching problem to rigorously align microscopic geometric details between disparate latent spaces. The 3D foundation model is only used for training-time supervision of alignment and is not used at inference, incurring no additional inference cost. Compared with the industrial baseline Step1X-3D, the proposed ROAD achieves highly competitive generation performance with only 1.5% of the training data and significantly reduces training costs, effectively reducing the computational overhead of high-fidelity 3D generation. Code is available at https://github.com/H-EmbodVis/ROAD.

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!