2601.06870v1 Jan 11, 2026 cs.LG

DaQ-MSA: 다중 모드 감성 분석을 위한 디노이징 및 품질 향상 확산 증강

DaQ-MSA: Denoising and Qualifying Diffusion Augmentations for Multimodal Sentiment Analysis

Jiazhang Liang
Jiazhang Liang
Citations: 0
h-index: 0
Jianheng Dai
Jianheng Dai
Citations: 2
h-index: 1
Miaosen Luo
Miaosen Luo
Citations: 11
h-index: 2
Menghua Jiang
Menghua Jiang
Citations: 3
h-index: 1
Sijie Mai
Sijie Mai
Citations: 26
h-index: 3

다중 모드 대규모 언어 모델(MLLM)은 시각-언어 작업에서 뛰어난 성능을 보이지만, 다중 모드 감성 분석에서의 효과는 고품질 훈련 데이터의 부족으로 인해 제한됩니다. 이는 정확한 다중 모드 이해와 일반화 능력을 저해합니다. 이러한 문제점을 해결하기 위해, 우리는 비디오 및 오디오 모드에 대해 의미를 보존하는 확산 모델 기반 증강을 적용하여 다중 모드 훈련 데이터의 분포를 확장합니다. 그러나 데이터 양만 늘리는 것은 충분하지 않습니다. 확산 모델로 생성된 샘플은 품질 편차가 크고, 노이즈가 많은 증강은 성능을 저하시킬 수 있습니다. 따라서 우리는 다중 모드 감성 분석을 위한 디노이징 및 품질 향상 확산 증강(DaQ-MSA)을 제안합니다. DaQ-MSA는 증강된 샘플의 신뢰도를 평가하는 품질 평가 모듈을 도입하여, 적응적인 훈련 가중치를 부여합니다. 낮은 품질의 샘플의 가중치를 낮추고, 고품질 샘플의 가중치를 높임으로써, DaQ-MSA는 더욱 안정적인 학습을 가능하게 합니다. 우리의 접근 방식은 확산 모델의 생성 능력을 MLLM의 의미 이해와 통합하여, 인간의 주석이나 추가적인 감독 없이 MLLM을 훈련하기 위한 강력하고 일반화 가능한 자동 증강 전략을 제공합니다.

Original Abstract

Multimodal large language models (MLLMs) have demonstrated strong performance on vision-language tasks, yet their effectiveness on multimodal sentiment analysis remains constrained by the scarcity of high-quality training data, which limits accurate multimodal understanding and generalization. To alleviate this bottleneck, we leverage diffusion models to perform semantics-preserving augmentation on the video and audio modalities, expanding the multimodal training distribution. However, increasing data quantity alone is insufficient, as diffusion-generated samples exhibit substantial quality variation and noisy augmentations may degrade performance. We therefore propose DaQ-MSA (Denoising and Qualifying Diffusion Augmentations for Multimodal Sentiment Analysis), which introduces a quality scoring module to evaluate the reliability of augmented samples and assign adaptive training weights. By down-weighting low-quality samples and emphasizing high-fidelity ones, DaQ-MSA enables more stable learning. By integrating the generative capability of diffusion models with the semantic understanding of MLLMs, our approach provides a robust and generalizable automated augmentation strategy for training MLLMs without any human annotation or additional supervision.

0 Citations
0 Influential
1.5 Altmetric
7.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!