2602.23153v1 Feb 26, 2026 cs.CV

효율적인 인코더 없는 푸리에 기반 3D 대규모 다중 모드 모델

Efficient Encoder-Free Fourier-based 3D Large Multimodal Model

Guofeng Mei
Guofeng Mei
Citations: 55
h-index: 4
Luigi Riz
Luigi Riz
Citations: 83
h-index: 5
Yujiao Wu
Yujiao Wu
Citations: 282
h-index: 6
Yiming Wang
Yiming Wang
Citations: 33
h-index: 3
Fabio Poiesi
Fabio Poiesi
Citations: 126
h-index: 6
Wei Lin
Wei Lin
Citations: 8
h-index: 1

3차원 데이터를 처리하는 대규모 다중 모드 모델(LMM)은 일반적으로 기하학적 특징을 추출하기 위해 크고 사전 훈련된 시각 인코더에 의존합니다. 최근 2D LMM은 효율성과 확장성을 위해 이러한 인코더를 제거하기 시작했지만, 포인트 클라우드의 무순서 및 대규모 특성으로 인해 이러한 패러다임을 3D로 확장하는 것은 여전히 어려운 과제입니다. 이는 중요한 미해결 질문을 남깁니다. 즉, 복잡한 인코더 없이 무순서 3D 데이터를 효과적이고 효율적으로 토큰화하는 LMM을 어떻게 설계할 수 있을까요? 우리는 최초의 효율적인 인코더 없는 푸리에 기반 3D 장면 LMM인 Fase3D를 제안합니다. Fase3D는 새로운 토크나이저를 사용하여 확장성 및 순열 불변성 문제를 해결하며, 이 토크나이저는 포인트 클라우드 직렬화와 고속 푸리에 변환(FFT)을 결합하여 자기 주의(self-attention)를 근사합니다. 이러한 설계는 효과적이고 계산적으로 최소한의 아키텍처를 가능하게 하며, 이는 세 가지 주요 혁신을 기반으로 합니다. 첫째, 우리는 구조화된 초점(superpoint)을 사용하여 대규모 장면을 간결하게 표현합니다. 둘째, 공간 채우기 곡선 직렬화 후 FFT를 통해 효율적인 전역 컨텍스트 모델링 및 그래프 기반 토큰 병합을 가능하게 합니다. 마지막으로, 푸리에 보강 LoRA 어댑터는 LLM에 전역 주파수 인지 상호 작용을 무시할 수 있는 비용으로 주입합니다. Fase3D는 인코더 기반 3D LMM과 비교 가능한 성능을 달성하면서 계산 및 파라미터 측면에서 훨씬 더 효율적입니다. 프로젝트 웹사이트: https://tev-fbk.github.io/Fase3D.

Original Abstract

Large Multimodal Models (LMMs) that process 3D data typically rely on heavy, pre-trained visual encoders to extract geometric features. While recent 2D LMMs have begun to eliminate such encoders for efficiency and scalability, extending this paradigm to 3D remains challenging due to the unordered and large-scale nature of point clouds. This leaves a critical unanswered question: How can we design an LMM that tokenizes unordered 3D data effectively and efficiently without a cumbersome encoder? We propose Fase3D, the first efficient encoder-free Fourier-based 3D scene LMM. Fase3D tackles the challenges of scalability and permutation invariance with a novel tokenizer that combines point cloud serialization and the Fast Fourier Transform (FFT) to approximate self-attention. This design enables an effective and computationally minimal architecture, built upon three key innovations: First, we represent large scenes compactly via structured superpoints. Second, our space-filling curve serialization followed by an FFT enables efficient global context modeling and graph-based token merging. Lastly, our Fourier-augmented LoRA adapters inject global frequency-aware interactions into the LLMs at a negligible cost. Fase3D achieves performance comparable to encoder-based 3D LMMs while being significantly more efficient in computation and parameters. Project website: https://tev-fbk.github.io/Fase3D.

2 Citations
0 Influential
3 Altmetric
17.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!