2606.25325v1 Jun 24, 2026 cs.AI

다중 모드 감정 추론을 위한 전방향 인지 정책 최적화

Omni-Perception Policy Optimization for Multimodal Emotion Reasoning

Beier Zhu
Beier Zhu
Citations: 303
h-index: 7
Lewei Lu
Lewei Lu
Citations: 36
h-index: 3
Wenwen Tong
Wenwen Tong
Citations: 1,791
h-index: 7
Zhi Han
Zhi Han
Citations: 11
h-index: 2
Pengyan Shao
Pengyan Shao
Citations: 2
h-index: 1
Peipei Song
Peipei Song
Citations: 55
h-index: 4
Xinyi Wang
Xinyi Wang
Citations: 30
h-index: 3
Jiangnan Chen
Jiangnan Chen
Citations: 37
h-index: 2
Xun Yang
Xun Yang
Citations: 2,162
h-index: 26

본 연구에서는 현재의 감정을 다루는 통합 다중 모드 대규모 언어 모델(Omni-MLLM)들이 여전히 신뢰할 수 있는 전방향 인지를 부족하게 수행한다는 점을 발견했습니다. 이러한 모델들은 (i) 추론 과정에서 다중 모드 정보를 충분히 활용하지 못하고, (ii) 종종 다른 모드의 정보를 기반으로 사실과 다른 내용을 생성하는 경향이 있습니다(모달리티 특이적 환각). 이러한 문제점을 해결하기 위해, 본 연구에서는 다중 모드 인지를 명시적으로 최적화하는 강화 학습 프레임워크인 OPPO(Omni-Perception Policy Optimization)를 제안합니다. 첫째, Omni-Perception Reward는 정답 추론 과정을 세분화된 시각, 음향 및 감정 정보로 분해하고, 이러한 정보를 의미적으로 재현하는 경로에 대해 보상을 제공합니다. 둘째, Omni-Perception Loss는 전체 입력과 단일 모드 마스킹된 입력을 비교하며, 모달리티 특이적 증거 토큰에 대해서만 KL 페널티를 적용하여 다중 모드 간의 환각을 억제합니다. 또한, 다중 모드 활용도와 신뢰성을 정량적으로 평가하기 위한 진단 벤치마크인 MEP-Bench를 소개합니다. 실험 결과, OPPO는 MER-UniBench 및 MME-Emotion 데이터셋에서 최고 성능을 달성했으며, MEP-Bench에서 활용도 및 신뢰성 점수를 크게 향상시켰습니다. 이는 다중 모드 감정 추론에 있어 충분하고 신뢰할 수 있는 전방향 인지의 중요성을 강조합니다.

Original Abstract

We find that current emotion-oriented Omni-MLLMs still lack reliable omni-modal perception: they (i) underutilize multimodal cues in their reasoning trajectories and (ii) exhibit unfaithful behavior, often hallucinating modality-specific statements from other modalities. Building on these insights, we propose OPPO (Omni-Perception Policy Optimization), a reinforcement learning framework that explicitly optimizes multimodal perception. First, an Omni-Perception Reward decomposes ground-truth reasoning into fine-grained visual, acoustic, and emotion cues and rewards trajectories that semantically recover these cues. Second, an Omni-Perception Loss compares the policy under full and unimodally masked inputs, applying a KL penalty only to modality-specific evidence tokens to suppress cross-modal hallucination. We further introduce MEP-Bench, a diagnostic benchmark that quantifies utilization and faithfulness. Experiments show that OPPO achieves state-of-the-art performance on MER-UniBench and MME-Emotion, while substantially improving utilization and faithfulness scores on MEP-Bench, highlighting the importance of sufficient and faithful omni perception for multimodal emotion reasoning.

1 Citations
0 Influential
13 Altmetric
66.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!