2607.27634v1 Jul 30, 2026 cs.CV

4DHumanDiff: 일관성 있는 360도 동적 인간 모델을 위한 텍스트 기반 직접 4D 가우시안 스플래팅 생성

4DHumanDiff: Direct Text-to-4DGS Generation for Consistent 360-Degree Dynamic Humans

Wangmeng Zuo
Wangmeng Zuo
Citations: 109
h-index: 4
Ren-Rong Wu
Ren-Rong Wu
Citations: 210
h-index: 7
Haoran Chen
Haoran Chen
Citations: 0
h-index: 0
Yuxiang Wei
Yuxiang Wei
Harbin Institute of Technology
Citations: 1,287
h-index: 14
Xiaowei Jin
Xiaowei Jin
Citations: 1,887
h-index: 11
Hui Li
Hui Li
Citations: 184
h-index: 5

텍스트 프롬프트를 사용하여 고품질의 360도 동적 인간 자산을 생성하는 것은 어려운 과제입니다. 기존 방법은 일반적으로 단안 또는 다중 시점 영상을 먼저 합성한 다음, 이를 4D 표현으로 변환하는데, 이는 비용이 많이 들고 종종 불완전한 기하 구조나 시점에 따라 일관되지 않은 결과물을 초래합니다. 본 논문에서는 텍스트 프롬프트로부터 직접 동적 인간을 생성하는 확산 모델 프레임워크인 4DHumanDiff를 제시합니다. 4DHumanDiff는 전체적인 4D 표현 공간을 엔드투엔드로 모델링하여, 영상의 사전 생성 및 장면별 재구성을 피함으로써 시점에 일관되고 시간적으로 응집된 자산 생성이 가능합니다. 이 모델은 동작 인지 능력을 향상시키기 위해 시간적 주의 메커니즘을 포함하는 3D U-Net 기반 구조를 사용합니다. 또한, 6만 개의 고품질 데이터 쌍으로 구성된 대규모 텍스트-4D 가우시안 스플래팅 데이터셋을 구축하고, 렌더링 품질과 동작의 부드러움을 향상시키기 위해 2D 정규화 및 학습이 필요 없는 4D 보간 기법을 도입했습니다. 실험 결과, 4DHumanDiff는 1분 이내에 일관성 있는 360도 동적 인간을 생성하며, 시간적 및 다중 시점 일관성이 뛰어나고 추론 시간을 10배 이상 단축할 수 있음을 확인했습니다.

Original Abstract

Generating high-quality 360-degree dynamic human assets from text prompts is challenging. Existing methods usually synthesize monocular or multi-view videos first and then fit a 4D representation, which is expensive and often causes incomplete geometry or view-inconsistent renderings. We present 4DHumanDiff, a diffusion framework that directly generates dynamic humans represented by 4D Gaussian Splatting (4DGS) from text prompts. By modeling the structured 4D representation space end-to-end, 4DHumanDiff avoids video pre-generation and per-scene reconstruction, making it better suited for view-consistent and temporally coherent asset generation. The model uses a 3D U-Net backbone with temporal attention for motion-aware generation. We further construct a large-scale text-to-4DGS dataset with 60,000 high-quality pairs, and introduce 2D regularization and training-free 4D interpolation to improve rendering quality and motion smoothness. Experiments show that 4DHumanDiff generates consistent 360-degree dynamic humans within one minute, achieves better temporal and multi-view consistency, and reduces inference time by more than 10x.

0 Citations
0 Influential
7 Altmetric
35.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!