글로벌 잠재 변수를 넘어선: 확장 가능한 3차원 모델링을 위한 청크 기반 희소 그리드 VAE
Beyond Global Latents: Chunk-Based Sparse Grid VAE for Scalable 3D Modeling
희소 볼륨 그리드는 상세한 3차원 재구성에 필요한 공간 구조를 유지하지만, 활성 표면 셀의 증가에 따라 해상도가 높아질수록 메모리 사용량이 급격히 증가합니다. 본 논문에서는 전체 잠재 볼륨 대신 로컬 청크를 중심으로 구성된 희소 그리드 변분 오토인코더인 ChunkVAE를 제안합니다. 로컬 학습 연산자를 통해 독립적으로 선택된 인코더 및 디코더 파티션을 사용하고, 추론 시 청크 크기가 훈련 시와 다를 수 있도록 설계했습니다. 두 가지 상호 보완적인 데이터 연산자가 이러한 유연성을 실현하는 데 기여합니다. Balanced Binary Object Partitioning은 활성 셀을 분배하면서 중복을 제한하고, S-Curve 가중치 스티칭은 전체 잠재 변수 또는 재구성을 조립할 때 신뢰할 수 없는 경계 특징을 완화합니다. 세 가지 객체 벤치마크에서 ChunkVAE는 $512^3$부터 $1536^3$까지의 해상도에서 강력한 기준 모델과 경쟁하거나 더 나은 성능을 보입니다. 작은 청크 크기는 최대 할당 메모리를 줄이고 청크별 연산 시간을 단축시켜 병렬 추론 속도를 향상시킵니다. 안정적인 스티칭된 잠재 변수와 개선된 이미지-3차원 지표는 로컬 압축이 전체 인터페이스를 유지하면서 기하학적 구조의 확장을 가능하게 한다는 것을 나타냅니다.
Sparse voxel grids preserve the spatial structure needed for detailed 3D reconstruction, but their memory still grows rapidly with resolution as active surface cells increase. We introduce ChunkVAE, a sparse grid variational autoencoder organized around local chunks rather than a global latent volume. Local learned operators permit independently chosen encoder and decoder partitions and allow inference chunk sizes to differ from training. Two complementary data operators make this flexibility practical: Balanced Binary Object Partitioning distributes active cells while limiting replicated overlap, while S-Curve weighted stitching attenuates unreliable boundary features when assembling a global latent or reconstruction. Across three object benchmarks, ChunkVAE is competitive with or better than strong baselines from $512^3$ to $1536^3$; smaller chunks lower peak allocated memory and shorten per-chunk compute, enabling faster parallel inference. Stable stitched latents and improved image to 3D metrics indicate that local compression can scale geometry while retaining the global interface required downstream.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.