2607.06565v1 Jul 07, 2026 cs.CV

ELSA3D: 탄성 기반 의미 연결을 통한 통합 3차원 이해 및 생성

ELSA3D: Elastic Semantic Anchoring for Unified 3D Understanding and Generation

Yifan Shen
Yifan Shen
Citations: 88
h-index: 5
Yuanzhe Liu
Yuanzhe Liu
Citations: 55
h-index: 4
Xinzhuo Li
Xinzhuo Li
Citations: 38
h-index: 4
Ismini Lourentzou
Ismini Lourentzou
Citations: 1,446
h-index: 19
Onkar Susladkar
Onkar Susladkar
Citations: 257
h-index: 8
Tianjiao Yu
Tianjiao Yu
Citations: 91
h-index: 5
Xiaona Zhou
Xiaona Zhou
Citations: 59
h-index: 5

통합 3차원 모델은 단일 구조 내에서 3차원 자산을 생성하고 언어를 통해 이를 추론하는 것을 목표로 하지만, 이러한 텍스트-3차원 상호작용은 여전히 암묵적인 경우가 많습니다. 기존 방법들은 텍스트와 3차원 토큰을 평탄한 시퀀스로 연결하고 자기 주의 메커니즘에 의존하여, 거친 구조적 단서와 미세한 기하학적 세부 정보를 하나의 차별화되지 않은 표현으로 통합합니다. 본 연구에서는 탄성 기반 의미 연결(elastic semantic anchoring) 방식을 통해 이러한 문제를 해결하는 통합 3차원 모델인 ELSA3D를 소개합니다. ELSA3D는 스케일 인지 옥트리 토크나이저를 사용하여 기하학적 정보를 표현하며, 의미적인 단서를 선택하고 가장 관련 있는 3차원 스케일로 전달하여 해당 스케일에 맞는 기하학적 증거를 검색하고, 이 정보를 통합된 표현으로 다시 쓰는 희소(sparse)한 크로스-모달 유닛인 Anchor Tokens을 도입합니다. 가벼운 블록 단위 라우터는 계산 및 추론 과정에서 탄력성을 제공하며, 어떤 텍스트 토큰이 어느 기하학적 스케일에서 Anchor를 생성할지 선택하여, 크로스-모달 능력을 필요한 곳에 집중시킵니다. ELSA3D는 이미지-3차원 생성, 텍스트-3차원 생성 및 3차원 캡셔닝 작업에서 최첨단 성능을 달성했으며, 동일 모델의 비-탄성 버전보다 FLOPs와 추론 지연 시간을 약 절반으로 줄이면서 가장 강력한 통합 기준 모델보다 우수한 성능을 보였습니다.

Original Abstract

Unified 3D foundation models aspire to generate 3D assets and reason about them in language within a single backbone, but their text-3D interaction remains largely implicit. Existing methods concatenate text and 3D tokens into a flat sequence and rely on self-attention, collapsing coarse structural cues and fine geometric details into one undifferentiated representation. We introduce ELSA3D, a unified 3D model that addresses this with elastic semantic anchoring, structuring language and geometric reasoning jointly along matched abstraction scales. ELSA3D represents geometry with a scale-aware octree tokenizer and introduces Anchor Tokens, sparse cross-modal units that select semantic cues, route them to the most relevant 3D scale, retrieve scale-specific geometric evidence, and write the fused signal back into the unified representation, keeping interaction sparse yet precise. A lightweight per-block router makes both computation and reasoning elastic, choosing which text tokens instantiate anchors at which geometric scale so that cross-modal capacity concentrates where alignment is most needed. ELSA3D achieves state-of-the-art performance across image-to-3D generation, text-to-3D generation, and 3D captioning, outperforming the strongest unified baseline while roughly halving FLOPs and inference latency relative to the non-elastic version of the same model.

0 Citations
0 Influential
9.5 Altmetric
47.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!