2607.25948v1 Jul 28, 2026 cs.CV

MODUS: 디코더 전용, 다양한 모달리티 간의 모든-에서-모든 모델링

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities

Zhaochong An
Zhaochong An
Citations: 63
h-index: 6
Serge Belongie
Serge Belongie
Citations: 70
h-index: 3
Amir Zadeh
Amir Zadeh
Citations: 161
h-index: 6
Chuan Li
Chuan Li
Citations: 40
h-index: 4
Zhitong Gao
Zhitong Gao
Citations: 104
h-index: 6
Mingqiao Ye
Mingqiao Ye
Citations: 1,146
h-index: 5
Jesse Allardice
Jesse Allardice
Citations: 131
h-index: 3
Afshin Dehghan
Afshin Dehghan
Citations: 729
h-index: 13
Amir Zamir
Amir Zamir
Citations: 261
h-index: 6
Roman Bachmann
Roman Bachmann
EPFL
Citations: 1,341
h-index: 11
Xian Liu
Xian Liu
Citations: 0
h-index: 0
François Fleuret
François Fleuret
Citations: 0
h-index: 0
David Mizrahi
David Mizrahi
EPFL
Citations: 683
h-index: 6
Oğuzhan Fatih Kar
Oğuzhan Fatih Kar
EPFL
Citations: 454
h-index: 8

다중 모달 데이터 처리 모델은 단일 네트워크 내에서 어떤 입력 모달리티로부터도 어떠한 출력 모달리티라도 예측하는 방식으로 작동하며, 이는 다중 모드 비전 및 비전-언어 모델뿐만 아니라 생태학 및 천문학과 같은 과학 분야에서도 점점 더 많이 활용되고 있습니다. 기존의 모든-에서-모든 모델은 주로 인코더-디코더 또는 확산 구조를 사용하여 처음부터 학습되므로 성능에 영향을 미치며, 강력하게 사전 훈련된 디코더 전용 모델을 초기 설정으로 사용하기 어렵다는 단점이 있습니다. 본 연구에서는 디코더 전용의 모든-에서-모든 다중 모달 모델링 방식을 제안합니다. 이 방식은 모든 모달리티를 동등하게 취급하며, 특정 모달리티에 특화된 헤드, 손실 함수 또는 작업 파이프라인 없이 임의의 모달리티를 입력 및 출력으로 사용하도록 설계되었습니다. 각 모달리티가 동일한 모델에서 입력과 출력 역할을 모두 수행하므로, Modus라는 이름의 결과 모델은 중간 모달리티를 통한 체인 생성 또는 다른 생성된 모달리티를 사용하여 모델 자체의 출력을 검증하는 등 다양한 응용 분야에 활용될 수 있습니다. Modus는 사전 학습 없이도 뛰어난 성능을 보이며, 다양한 벤치마크에서 전문화된 모델이나 다중 작업 모델과 경쟁할 수 있는 결과를 보여줍니다. 모든 관련 자료는 https://modus-multimodal.epfl.ch/ 에서 공개적으로 이용 가능합니다.

Original Abstract

Any-to-any models predict any modality from any combination of others within a single network, a formulation used in multimodal vision and vision-language models, and increasingly in scientific domains such as ecology and astronomy. Existing any-to-any models are typically trained from scratch using encoder-decoder or diffusion architectures, impacting their performance and preventing them from using strong pre-trained decoder-only models as a prior. In this work, we investigate decoder-only any-to-any multimodal modeling, which treats all modalities symmetrically and supports arbitrary modalities as inputs and outputs without modality-specific heads, losses, or task pipelines. Because every modality is both an input and an output of the same model, the resulting model, named Modus, can support a range of applications, such as chained generation through intermediate modalities or cross-modal self-verification by scoring the model's own outputs with another generated modality. Modus demonstrates strong out-of-the-box performance and is competitive with specialist and multitask baselines using a single model across various benchmarks. All materials are open-sourced at https://modus-multimodal.epfl.ch/.

0 Citations
0 Influential
6.5 Altmetric
32.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!