포괄적인 상호 작용 기반 충돌을 통한 다중 시점 일관성 있는 조립형 3차원 객체 생성
Inclusive Interactive Collisions for Multi-View Consistent Compositional 3D Generation
최근 텍스트-이미지 확산 모델의 발전으로 3차원 객체 생성이 크게 발전했지만, 기존 방법은 다음과 같은 두 가지 주요 문제점을 가지고 있습니다. (1) 대부분 단일 3차원 객체를 생성하며, 가우시안 원리를 사용하여 합리적인 상호 작용을 모델링하는 데 어려움이 있어 여러 객체가 포함된 조립형 3차원 자산 생성이 어렵습니다. (2) 3차원 최적화 과정에서 종종 시점 간 불일치가 발생하는데, 이는 Score Distillation Sampling이 각 시점에 대해 개별적으로 수행되기 때문에 필연적으로 발생하는 문제입니다. 위 문제점을 해결하기 위해, 저희는 여러 시점에서 일관성을 유지하며 합리적인 상호 작용을 갖는 조립형 3차원 객체를 생성하는 새로운 최적화 기반 방법인 I2C-3D를 제안합니다. 구체적으로, 가우시안 원리가 자연스럽게 합리적인 상호 작용 영역에 나타나도록 유도하는 포괄적인 상호 작용 기반 충돌(Inclusive Interactive Collisions) 전략을 제안하여, 조립된 장면 내 객체들이 물리적으로 타당하고 시각적으로 일관되게 상호 작용하도록 합니다. 또한, 다중 시점 간 일관성을 향상시키기 위해, 사전 학습된 확산 모델에서 인스턴스 토큰과 공간 토큰의 어텐션 맵을 조절하여 다중 시점 일관성 우선순위와 레이아웃 우선순위를 추출하는 Multi-View Adaptive Score Distillation Sampling 기법을 개발했습니다. 이러한 정교한 설계 덕분에 I2C-3D는 고품질의 다중 시점 일관성을 갖는 조립형 3차원 객체를 생성할 뿐만 아니라, 유연한 3차원 편집 기능을 지원하여 복잡한 장면 생성을 용이하게 합니다. 광범위한 실험 결과, I2C-3D가 기존 방법보다 생성 품질 및 다중 시점 일관성 측면에서 우수한 성능을 보였습니다.
Recent breakthroughs in 3D generation have advanced notably with the development of text-to-image diffusion model. However, existing methods remain two practical challenges: (1) They primarily generate single 3D object, but struggle to generate multi-object compositional 3D assets due to the lack of the modeling for Gaussian primitives in reasonable interactions. (2) They often suffer from cross-view inconsistency during 3D optimization, as Score Distillation Sampling inherently performs on each single view, inevitably resulting in cross-view hallucinations. To solve above issues, we propose I2C-3D, a novel optimization-based method to generate multi-view consistent compositional 3D assets with reasonable interactions. Specifically, we propose an Inclusive Interactive Collisions strategy to guide Gaussian primitives appearing in reasonable interaction regions naturally, thereby ensuring objects in the compositional scene interact in a physically plausible and visually coherent way. Additionally, to enhance multi-view consistency, Multi-View Adaptive Score Distillation Sampling is devised to distill multi-view consistency prior and layout prior from pre-trained diffusion model by modulating attention map of instance token and spatial token across viewpoints. Benefiting from above elaborate designs, I2C-3D not only generates high-fidelity multi-view consistent compositional 3D assets but also supports 3D editing flexibly, facilitating complex scene generation. Extensive experiments demonstrate our I2C-3D outperforms existing methods in generation quality and multi-view consistency.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.