MSVS-VAE: 다중 스케일 앵커 기반 VecSet을 이용한 고품질 3D 재구성
MSVS-VAE: Multi-Scale Anchored VecSet for High-Fidelity 3D Reconstruction
고품질 3D 생성 모델링은 점차 잠재 확산(latent diffusion) 패러다임에 의존하고 있으며, 이 과정에서 기본이 되는 3D VAE의 재구성 품질이 주요 병목 현상으로 작용합니다. 기존 연구들은 주로 두 가지 방식으로 접근했습니다: 희소 볼륨 기반 표현은 강력한 재구성 품질을 제공하지만 상당한 메모리와 계산 오버헤드를 발생시키며, 세트 기반 표현은 압축적이고 연속적인 장점이 있지만 잠재 공간의 희소성 및 과도한 전역 부드러움으로 인해 일반적으로 정확도가 떨어집니다. 본 연구에서는 이러한 격차를 해소하고 압축성을 유지하는 계층적 세트 기반 VAE인 MSVS-VAE를 제안합니다. 핵심 아이디어는 계층적 포인트 셔플 업샘플링을 통해 앵커된 VecSet 잠재 공간을 점진적으로 밀집화하여 미세한 기하학적 모델링을 위한 공간 용량을 늘리는 것입니다. 밀집화된 계층 구조에서 효율적인 디코딩을 위해, 전역 크로스 어텐션 대신 AVS-Conv라는 기하학 정보를 고려하는 로컬 집계 연산자를 사용합니다. 이 연산자는 전체 잠재 세트가 아닌 로컬 주변 영역 내에서 작동합니다. 또한 다중 스케일 쿼리 디코딩을 도입하여 거친 스케일의 안정적인 전역 컨텍스트와 미세 스케일의 지역화된 기하학적 정보를 결합하여 지나치게 국소적인 수용 필드로부터 발생하는 아티팩트를 줄입니다. Objaverse, ABO 및 실제 환경 데이터셋에서의 광범위한 실험 결과는 MSVS-VAE가 기존의 세트 기반 및 볼륨 기반 VAE보다 일관되게 우수한 성능을 보이며, 기존 세트 기반 방법론에 비해 약 10배 빠른 디코딩 속도와 볼륨 기반 모델에 비해 약 10배 더 높은 압축률을 제공함을 보여줍니다.
High-fidelity 3D generative modeling increasingly relies on the latent diffusion paradigm, where the reconstruction quality of the underlying 3D VAE becomes a primary bottleneck. Existing approaches largely follow two paradigms: sparse voxel-based representations achieve strong reconstruction quality but incur significant memory and computational overhead, while set-based representations are compact and continuous yet typically lag in fidelity due to latent sparsity and excessive global smoothness. We propose MSVS-VAE, a hierarchical set-based VAE that closes this fidelity gap without sacrificing compactness. Our key idea is to progressively densify anchored VecSet latents via hierarchical point-shuffle upsampling, increasing spatial capacity for fine-grained geometry modeling. To efficiently decode from the densified hierarchy, we replace global cross-attention with AVS-Conv, a geometry-aware local aggregation operator operating within local neighborhoods rather than the exhaustive latent set. We further introduce multi-scale query decoding to fuse coarse-to-fine latent features, where coarse scales provide stable global context, and fine scales refine localized geometry, reducing artifacts from overly local receptive fields. Extensive experiments on Objaverse, ABO, and in-the-wild benchmarks demonstrate that MSVS-VAE consistently outperforms prior set-based and voxel-based VAEs, delivering approximately 10x faster decoding than prior set-based methods and approximately 10x higher compactness than voxel-based baselines.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.