2601.20618v1 Jan 28, 2026 cs.CV

GDCNet: 다중 모드 풍자 탐지를 위한 생성적 불일치 비교 네트워크

GDCNet: Generative Discrepancy Comparison Network for Multimodal Sarcasm Detection

Xiang Ao
Xiang Ao
Citations: 1,383
h-index: 19
Shuguang Zhang
Shuguang Zhang
Citations: 2
h-index: 1
Junhong Lian
Junhong Lian
Institute of Computing Technology, Chinese Academy of Sciences
Citations: 16
h-index: 3
Guoxin Yu
Guoxin Yu
Citations: 29
h-index: 3
Baoxun Xu
Baoxun Xu
Citations: 393
h-index: 6

다중 모드 풍자 탐지(MSD)는 이미지-텍스트 쌍 내에서 의미적 불일치를 모델링하여 풍자를 식별하는 것을 목표로 합니다. 기존 방법은 종종 교차 모드 임베딩 불일치를 활용하여 불일치를 감지하지만, 시각적 및 텍스트 내용이 느슨하게 관련되거나 의미적으로 간접적인 경우 어려움을 겪습니다. 최근 접근 방식은 대규모 언어 모델(LLM)을 활용하여 풍자적 단서를 생성하지만, 이러한 생성의 고유한 다양성과 주관성은 종종 노이즈를 유발합니다. 이러한 제한 사항을 해결하기 위해, 우리는 생성적 불일치 비교 네트워크(GDCNet)를 제안합니다. 이 프레임워크는 다중 모드 LLM에 의해 생성된 설명적이고 사실에 기반한 이미지 캡션을 사용하여 안정적인 의미적 앵커로 활용함으로써 교차 모드 충돌을 포착합니다. 구체적으로, GDCNet은 생성된 객관적인 설명과 원래 텍스트 간의 의미적 및 감정적 불일치를 계산하고, 동시에 시각-텍스트 일관성을 측정합니다. 이러한 불일치 특징은 게이트 모듈을 통해 시각 및 텍스트 표현과 융합되어 모달리티 기여도를 적응적으로 균형 있게 조정합니다. MSD 벤치마크에 대한 광범위한 실험 결과, GDCNet은 우수한 정확도와 안정성을 보여주며, MMSD2.0 벤치마크에서 새로운 최고 성능을 달성했습니다.

Original Abstract

Multimodal sarcasm detection (MSD) aims to identify sarcasm within image-text pairs by modeling semantic incongruities across modalities. Existing methods often exploit cross-modal embedding misalignment to detect inconsistency but struggle when visual and textual content are loosely related or semantically indirect. While recent approaches leverage large language models (LLMs) to generate sarcastic cues, the inherent diversity and subjectivity of these generations often introduce noise. To address these limitations, we propose the Generative Discrepancy Comparison Network (GDCNet). This framework captures cross-modal conflicts by utilizing descriptive, factually grounded image captions generated by Multimodal LLMs (MLLMs) as stable semantic anchors. Specifically, GDCNet computes semantic and sentiment discrepancies between the generated objective description and the original text, alongside measuring visual-textual fidelity. These discrepancy features are then fused with visual and textual representations via a gated module to adaptively balance modality contributions. Extensive experiments on MSD benchmarks demonstrate GDCNet's superior accuracy and robustness, establishing a new state-of-the-art on the MMSD2.0 benchmark.

2 Citations
0 Influential
9.5 Altmetric
49.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!