정보 이론 기반의 다중 모드 상호 작용 학습을 위한 분해 방법
Information-Theoretic Decomposition for Multimodal Interaction Learning
다중 모드 학습은 여러 모드에 걸쳐 중복, 고유 및 시너지 정보를 파악하는 데 의존하며, 이들은 종합적으로 다중 모드 상호 작용을 구성합니다. 중요한 과제이지만 아직 충분히 연구되지 않은 부분은 이러한 암묵적인 상호 작용이 샘플별로 동적으로 변화한다는 것입니다. 본 논문에서는 이러한 동적이고 샘플 특유의 상호 작용을 학습하는 것이 효과적인 다중 모드 학습에 왜 중요한지를 보여주는 최초의 체계적인 정보 이론 분석을 제시합니다. 또한, 우리의 분석은 기존 방법들이 이러한 다양한 유형의 상호 작용을 학습하는 데 있어 부족한 점을 드러냅니다. 즉, 모드 앙상블 방식은 시너지 효과를 제대로 포착하지 못하고, 공동 학습 패러다임은 종종 중복 정보를 충분히 활용하지 못합니다. 이는 각 샘플에 따라 다양한 유형의 상호 작용으로부터 적응적으로 학습할 수 있는 접근 방식이 필요함을 강조합니다. 이를 위해, 본 논문에서는 샘플 특유의 상호 작용을 명시적으로 모델링하고 학습하는 새로운 패러다임인 분해 기반 다중 모드 상호 작용 학습 (DMIL)을 제안합니다. 먼저, 구성 요소 상호 작용을 분리하기 위한 변분 분해 아키텍처를 설계했습니다. 둘째, 이러한 명시적인 상호 작용 요소를 활용하여 미세 조정 과정을 통해 포괄적인 상호 작용 학습을 달성하는 새로운 학습 전략을 사용합니다. 다양한 작업 및 아키텍처에서의 광범위한 실험 결과는 DMIL이 전체 샘플 특유의 상호 작용에 적응함으로써 일관되게 우수한 성능을 달성함을 보여줍니다. 우리의 프레임워크는 유연하며 광범위하게 적용 가능하며, 다중 모드 학습을 위한 상호 작용 중심적인 패러다임을 확립합니다. 코드는 https://github.com/GeWu-Lab/DMIL 에서 확인할 수 있습니다.
Multimodal learning hinges on capturing redundant, unique, and synergistic information across modalities, which collectively constitute multimodal interactions. A critical yet underexplored challenge is that these implicit interactions vary dynamically across samples. In this work, we present the first systematic, information-theoretic analysis highlighting why learning these dynamic, sample-specific interactions is critical for effective multimodal learning. Our analysis further reveals deficits in conventional paradigms at learning these distinct interaction types: modality ensemble approaches struggle to capture synergy, while joint learning paradigms often under-utilize redundant information. This highlights the need for an approach that can adaptively learn from different interaction types on a per-sample basis. To this end, we propose Decomposition-based Multimodal Interaction Learning (DMIL), a novel paradigm that explicitly models and learns from sample-specific interactions. First, we design a variational decomposition architecture to isolate the constituent interaction components. Second, we employ a new learning strategy that leverages these explicit interaction components in a fine-tuning process to achieve comprehensive interaction learning. Extensive experiments across diverse tasks and architectures demonstrate that DMIL consistently achieves superior performance by adapting to holistic sample-specific interactions. Our framework is flexible and broadly applicable, establishing an interaction-centric paradigm for multimodal learning. The code is available at https://github.com/GeWu-Lab/DMIL.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.