MODUS: 디코더 전용, 다양한 모달리티 간의 모든-에서-모든 모델링
MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities
다중 모달 데이터 처리 모델은 단일 네트워크 내에서 어떤 입력 모달리티로부터도 어떠한 출력 모달리티라도 예측하는 방식으로 작동하며, 이는 다중 모드 비전 및 비전-언어 모델뿐만 아니라 생태학 및 천문학과 같은 과학 분야에서도 점점 더 많이 활용되고 있습니다. 기존의 모든-에서-모든 모델은 주로 인코더-디코더 또는 확산 구조를 사용하여 처음부터 학습되므로 성능에 영향을 미치며, 강력하게 사전 훈련된 디코더 전용 모델을 초기 설정으로 사용하기 어렵다는 단점이 있습니다. 본 연구에서는 디코더 전용의 모든-에서-모든 다중 모달 모델링 방식을 제안합니다. 이 방식은 모든 모달리티를 동등하게 취급하며, 특정 모달리티에 특화된 헤드, 손실 함수 또는 작업 파이프라인 없이 임의의 모달리티를 입력 및 출력으로 사용하도록 설계되었습니다. 각 모달리티가 동일한 모델에서 입력과 출력 역할을 모두 수행하므로, Modus라는 이름의 결과 모델은 중간 모달리티를 통한 체인 생성 또는 다른 생성된 모달리티를 사용하여 모델 자체의 출력을 검증하는 등 다양한 응용 분야에 활용될 수 있습니다. Modus는 사전 학습 없이도 뛰어난 성능을 보이며, 다양한 벤치마크에서 전문화된 모델이나 다중 작업 모델과 경쟁할 수 있는 결과를 보여줍니다. 모든 관련 자료는 https://modus-multimodal.epfl.ch/ 에서 공개적으로 이용 가능합니다.
Any-to-any models predict any modality from any combination of others within a single network, a formulation used in multimodal vision and vision-language models, and increasingly in scientific domains such as ecology and astronomy. Existing any-to-any models are typically trained from scratch using encoder-decoder or diffusion architectures, impacting their performance and preventing them from using strong pre-trained decoder-only models as a prior. In this work, we investigate decoder-only any-to-any multimodal modeling, which treats all modalities symmetrically and supports arbitrary modalities as inputs and outputs without modality-specific heads, losses, or task pipelines. Because every modality is both an input and an output of the same model, the resulting model, named Modus, can support a range of applications, such as chained generation through intermediate modalities or cross-modal self-verification by scoring the model's own outputs with another generated modality. Modus demonstrates strong out-of-the-box performance and is competitive with specialist and multitask baselines using a single model across various benchmarks. All materials are open-sourced at https://modus-multimodal.epfl.ch/.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.