2607.12752v1 Jul 14, 2026 cs.CV

Hallo4D: 일관성 있는 시공간 생성 모델의 다중 모달 환각 완화 기술

Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation

Jie Cao
Jie Cao
Citations: 42
h-index: 3
Haoyang Tong
Haoyang Tong
Citations: 71
h-index: 5
Hongbo Wang
Hongbo Wang
Citations: 6
h-index: 1
Huaibo Huang
Huaibo Huang
Citations: 250
h-index: 7
Jin Liu
Jin Liu
Citations: 33
h-index: 3
Ran He
Ran He
Citations: 358
h-index: 9

최근 3차원 생성 분야에서 눈부신 발전이 있었지만, 기존 방법들은 종종 명시적인 기하학적 일관성 메커니즘 없이 2차원 확산 기반 감독 방법을 사용하며, 이로 인해 중복 구조 및 부정확한 기하학 등과 같은 공간 환각 현상이 발생합니다. 이러한 문제는 시점과 시간 변화에 걸쳐 일관성을 유지해야 하는 4차원 생성에서 더욱 심각해지며, 흔들림, 동일성 깜빡임, 구조적 드리프트 등의 문제가 발생합니다. 본 논문에서는 3차원 및 4차원 콘텐츠 생성이 시 발생하는 시공간 환각 현상을 완화하기 위한 통합적이고 모델에 독립적인 프레임워크인 extbf{Hallo4D}를 제시합니다. Hallo4D는 대규모 다중 모달 언어 모델(LMM)을 활용하여 다양한 시점과 프레임에서 생성된 결과물로부터 공간적 및 시간적 불일치를 식별하고 요약하는 생성-탐지-수정 패러다임을 도입합니다. 이러한 정보는 LMM 기반 선택기가 다중 모델 투표를 통해 후보 수정 사항을 평가하는 이미지 공간 일관성 최적화에 활용되며, 재학습이나 구조 변경 없이 작동합니다. 또한 Hallo4D는 시간적 일관성을 향상시키고 최적화 효율성을 높이기 위해 움직임 인지 키프레임 샘플링, LMM 기반 초기화, 외형 정렬 기능을 포함하고 있습니다. 더불어, 어려운 시점에서 견고성을 높이기 위해 노출 인식 최적화 및 가시성 가지치기 기술을 추가로 도입했습니다. 광범위한 실험 결과는 Hallo4D가 다양한 3차원 및 4차원 생성 환경에서 강력한 기준 모델보다 우수한 성능을 보이며, 일관성을 고려한 콘텐츠 생성을 위한 확장 가능하고 일반적인 솔루션을 제공한다는 것을 입증합니다.

Original Abstract

While recent advances in 3D generation have enabled impressive visual synthesis, existing methods often rely on 2D diffusion supervision without explicit mechanisms for geometric consistency, leading to spatial hallucinations such as duplicated structures and misaligned geometry. These issues become more severe in 4D generation, where maintaining consistency across viewpoints and temporal evolution introduces additional challenges, including jitter, identity flicker, and structural drift. We present \textbf{Hallo4D}, a unified and model-agnostic framework for mitigating spatiotemporal hallucinations in 3D and 4D content generation. Hallo4D introduces a generation-detection-correction paradigm that leverages large multimodal language models (LMMs) to identify and summarize spatial and temporal inconsistencies from multi-view and multi-frame renderings. These insights guide a consensus-driven image-space consistency optimization, where an LMM-based selector evaluates candidate corrections through multi-model voting, without requiring retraining or architectural modifications. To further improve temporal consistency and optimization efficiency, Hallo4D incorporates motion-aware keyframe sampling, LMM-guided initialization, and appearance alignment. We additionally introduce exposure-aware optimization and visibility pruning to enhance robustness under challenging viewpoints. Extensive experiments demonstrate that Hallo4D consistently outperforms strong baselines across diverse 3D and 4D generation settings, providing a scalable and generalizable solution for consistency-aware content generation.

0 Citations
0 Influential
4.5 Altmetric
22.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!