2603.16085v1 Mar 17, 2026 cs.CV

Interact3D: 상호작용하는 3D 객체의 조립적 생성

Interact3D: Compositional 3D Generation of Interactive Objects

Hui Shan
Hui Shan
Citations: 2
h-index: 1
Keyang Luo
Keyang Luo
Citations: 6
h-index: 2
Ming Li
Ming Li
Citations: 274
h-index: 11
Sizhe Zheng
Sizhe Zheng
Citations: 20
h-index: 1
Yanwei Fu
Yanwei Fu
Citations: 2
h-index: 1
Zhen Chen
Zhen Chen
Citations: 71
h-index: 4
Xiangru Huang
Xiangru Huang
Citations: 2
h-index: 1

최근 3D 생성 분야의 발전은 고품질의 개별 자산 생성을 가능하게 했습니다. 그러나 단일 이미지로부터 3D 조립 객체를 생성하는 것, 특히 가려진 영역에서의 생성은 여전히 어려운 과제입니다. 기존 방법은 종종 숨겨진 영역의 기하학적 세부 정보를 손상시키고, 객체 간의 공간적 관계(OOR)를 유지하는 데 실패합니다. 본 논문에서는 물리적으로 타당한 상호작용을 하는 3D 조립 객체를 생성하도록 설계된 새로운 프레임워크인 Interact3D를 제시합니다. 저희의 접근 방식은 먼저 고급 생성 모델을 활용하여 통일된 3D 가이드 장면을 사용하여 고품질의 개별 자산을 생성합니다. 이러한 자산을 물리적으로 조립하기 위해, 견고한 두 단계로 구성된 조립 파이프라인을 도입합니다. 3D 가이드 장면을 기반으로, 주 객체는 정밀한 전역-국소 기하학적 정렬(등록)을 통해 고정되며, 이후의 기하학적 요소는 미분 가능한 Signed Distance Field (SDF) 기반 최적화를 사용하여 통합됩니다. 이 최적화는 기하학적 교차를 명시적으로 처벌합니다. 어려운 충돌을 줄이기 위해, 폐쇄 루프의 에이전트 기반 개선 전략을 추가로 사용합니다. Vision-Language Model (VLM)은 생성된 장면의 다중 뷰 렌더링을 자율적으로 분석하고, 목표로 하는 수정 프롬프트를 생성하며, 이미지 편집 모듈을 통해 반복적으로 생성 파이프라인을 자체 수정하도록 안내합니다. 광범위한 실험 결과는 Interact3D가 향상된 기하학적 충실도와 일관된 공간적 관계를 갖춘 유망한 충돌 인식 조립 결과를 성공적으로 생성한다는 것을 보여줍니다.

Original Abstract

Recent breakthroughs in 3D generation have enabled the synthesis of high-fidelity individual assets. However, generating 3D compositional objects from single images--particularly under occlusions--remains challenging. Existing methods often degrade geometric details in hidden regions and fail to preserve the underlying object-object spatial relationships (OOR). We present a novel framework Interact3D designed to generate physically plausible interacting 3D compositional objects. Our approach first leverages advanced generative priors to curate high-quality individual assets with a unified 3D guidance scene. To physically compose these assets, we then introduce a robust two-stage composition pipeline. Based on the 3D guidance scene, the primary object is anchored through precise global-to-local geometric alignment (registration), while subsequent geometries are integrated using a differentiable Signed Distance Field (SDF)-based optimization that explicitly penalizes geometry intersections. To reduce challenging collisions, we further deploy a closed-loop, agentic refinement strategy. A Vision-Language Model (VLM) autonomously analyzes multi-view renderings of the composed scene, formulates targeted corrective prompts, and guides an image editing module to iteratively self-correct the generation pipeline. Extensive experiments demonstrate that Interact3D successfully produces promising collsion-aware compositions with improved geometric fidelity and consistent spatial relationships.

0 Citations
0 Influential
5.5 Altmetric
27.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!