2607.12787v1 Jul 14, 2026 cs.AI

10억 개 이상의 파라미터를 가진 다중 모드 감정 언어 모델이 정말 필요한가?

Do We Really Need Multimodal Emotion Language Models Larger Than 1B Parameters?

Xuri Ge
Xuri Ge
Citations: 436
h-index: 12
Kaiwen Zheng
Kaiwen Zheng
Citations: 35
h-index: 4
Junchen Fu
Junchen Fu
University of Glasgow
Citations: 708
h-index: 8
Wenhao Deng
Wenhao Deng
Citations: 152
h-index: 4
Hu Han
Hu Han
Citations: 7,636
h-index: 41
Joemon M. Jose
Joemon M. Jose
Citations: 282
h-index: 9

최근 다중 모드 대규모 언어 모델(MLLM)의 발전은 비디오, 오디오 및 언어 등을 공동으로 모델링하여 다중 모드 감정 인식(MER) 성능을 크게 향상시키고 해석 가능한 설명 생성 기능을 가능하게 했습니다. 그러나 이러한 성능 향상은 종종 모델 파라미터 크기 증가(예: 최소 70억 개)를 동반하며, 이는 높은 계산 비용을 초래하고 추론 효율성을 저하시켜 로봇 및 모바일 기기와 같은 자원 제약 환경에서의 실시간 배포를 어렵게 만듭니다. 이러한 상황에서 다음과 같은 근본적인 질문이 제기됩니다: 고품질 MER을 위해서는 10억 개 이상의 파라미터를 가진 다중 모드 MER 모델이 정말로 필요한가? 본 논문에서는 더 큰 모델이 반드시 필요하다는 가정에 도전하며, 지식 증류를 통해 향상된 다중 모드 감정 이해 및 인식을 달성하는 경량화된 MER 프레임워크(Light-MER)를 제안합니다. Light-MER는 강력하고 대규모의 교사 모델로부터 지식을 가벼운 10억 개 미만의 파라미터를 가진 학생 모델로 전달하여 풍부한 다중 모드 감정 추론 및 인식을 유지하면서 배포 효율성을 크게 향상시키는 것을 목표로 합니다. 특히, 우리는 지식 전달을 강화하기 위한 두 가지 새로운 최적화 전략을 소개합니다: (1) Sliced Wasserstein Distance와 hidden-state alignment를 결합한 새로운 optimal transport loss, 그리고 (2) GRPO 기반의 새로운 다중 보상 최적화 전략으로 MER 성능과 효율성의 균형을 맞추어 학생 모델의 학습 능력을 더욱 향상시킵니다. 9개의 표준 데이터 세트에 대한 광범위한 실험 결과는 Light-MER가 최첨단 성능을 달성하면서 추론 효율성을 크게 개선한다는 것을 보여줍니다. 이는 미래 연구를 위한 작고 효율적인 다중 모드 감정 언어 모델의 강력한 잠재력을 강조합니다. 코드: https://github.com/GAIR-Lab/Light-MER

Original Abstract

Recent advances in multimodal large language models (MLLMs) have significantly improved the performance of multimodal emotion recognition (MER) and enabled interpretable description generation by jointly modeling video, audio, and language, etc. However, these performance improvements are often accompanied by an increase in model parameter size (e.g, at least 7B), which simultaneously incurs high computational costs and reduces inference efficiency, thereby hindering real-time deployment on resource-constrained platforms such as robots and mobile devices. This raises a fundamental question: do we really need the multimodal MER model larger than 1B parameters for high-quality MER? In this paper, we challenge the assumption that larger models are inherently necessary and proposes a lightweight MER framework (called Light-MER), which achieves better and faster multimodal sentiment understanding and recognition through knowledge distillation. It can transfer knowledge from a strong, large-scale teacher model to a lightweight sub-billion-parameter student model, aiming to preserve rich multimodal emotion reasoning and recognition while substantially improving deployment efficiency. Specifically, we introduce two new optimization strategies to enhance knowledge transfer: (1) a new optimal transport loss that combines Sliced Wasserstein Distance with hidden-state alignment, and (2) a new multi-reward optimization strategy based on GRPO that balances MER performance and efficiency, aimed at further enhancing the learning capabilities of student models. Extensive experiments on nine benchmark datasets demonstrate that Light-MER achieves state-of-the-art performance while significantly improving inference efficiency. This highlights the strong potential of small multimodal emotion language models for future research. Code is available at https://github.com/GAIR-Lab/Light-MER.

0 Citations
0 Influential
49.45879734614 Altmetric
0.0 Score
Original PDF
5

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!