2608.02016v1 Aug 03, 2026 cs.CV

글로벌 잠재 변수를 넘어선: 확장 가능한 3차원 모델링을 위한 청크 기반 희소 그리드 VAE

Beyond Global Latents: Chunk-Based Sparse Grid VAE for Scalable 3D Modeling

Chunchao Guo
Chunchao Guo
Citations: 1,097
h-index: 15
Haolin Liu
Haolin Liu
Citations: 580
h-index: 7
Qin Lin
Qin Lin
Citations: 2,336
h-index: 11
Xianghui Yang
Xianghui Yang
Citations: 787
h-index: 11
Yunfei Zhao
Yunfei Zhao
Citations: 617
h-index: 8
Zeqiang Lai
Zeqiang Lai
Citations: 703
h-index: 10
Zhihao Liang
Zhihao Liang
Citations: 47
h-index: 4
Zibo Zhao
Zibo Zhao
Citations: 740
h-index: 10
Long Quan
Long Quan
Citations: 295
h-index: 4
Kaiyi Zhang
Kaiyi Zhang
Citations: 134
h-index: 7
Bowen Zhang
Bowen Zhang
Citations: 253
h-index: 5

희소 볼륨 그리드는 상세한 3차원 재구성에 필요한 공간 구조를 유지하지만, 활성 표면 셀의 증가에 따라 해상도가 높아질수록 메모리 사용량이 급격히 증가합니다. 본 논문에서는 전체 잠재 볼륨 대신 로컬 청크를 중심으로 구성된 희소 그리드 변분 오토인코더인 ChunkVAE를 제안합니다. 로컬 학습 연산자를 통해 독립적으로 선택된 인코더 및 디코더 파티션을 사용하고, 추론 시 청크 크기가 훈련 시와 다를 수 있도록 설계했습니다. 두 가지 상호 보완적인 데이터 연산자가 이러한 유연성을 실현하는 데 기여합니다. Balanced Binary Object Partitioning은 활성 셀을 분배하면서 중복을 제한하고, S-Curve 가중치 스티칭은 전체 잠재 변수 또는 재구성을 조립할 때 신뢰할 수 없는 경계 특징을 완화합니다. 세 가지 객체 벤치마크에서 ChunkVAE는 $512^3$부터 $1536^3$까지의 해상도에서 강력한 기준 모델과 경쟁하거나 더 나은 성능을 보입니다. 작은 청크 크기는 최대 할당 메모리를 줄이고 청크별 연산 시간을 단축시켜 병렬 추론 속도를 향상시킵니다. 안정적인 스티칭된 잠재 변수와 개선된 이미지-3차원 지표는 로컬 압축이 전체 인터페이스를 유지하면서 기하학적 구조의 확장을 가능하게 한다는 것을 나타냅니다.

Original Abstract

Sparse voxel grids preserve the spatial structure needed for detailed 3D reconstruction, but their memory still grows rapidly with resolution as active surface cells increase. We introduce ChunkVAE, a sparse grid variational autoencoder organized around local chunks rather than a global latent volume. Local learned operators permit independently chosen encoder and decoder partitions and allow inference chunk sizes to differ from training. Two complementary data operators make this flexibility practical: Balanced Binary Object Partitioning distributes active cells while limiting replicated overlap, while S-Curve weighted stitching attenuates unreliable boundary features when assembling a global latent or reconstruction. Across three object benchmarks, ChunkVAE is competitive with or better than strong baselines from $512^3$ to $1536^3$; smaller chunks lower peak allocated memory and shorten per-chunk compute, enabling faster parallel inference. Stable stitched latents and improved image to 3D metrics indicate that local compression can scale geometry while retaining the global interface required downstream.

0 Citations
0 Influential
7.5 Altmetric
37.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!