MM-StanceDet: 검색 기반 다중 모드 다중 에이전트 태도 감지
MM-StanceDet: Retrieval-Augmented Multi-modal Multi-agent Stance Detection
다중 모드 태도 감지(MSD)는 공론을 이해하는 데 매우 중요하지만, 특히 상반된 정보를 포함하는 텍스트와 이미지를 효과적으로 융합하는 것은 여전히 어려운 과제입니다. 기존 방법은 종종 문맥적 이해 부족, 모달 간 해석의 모호성, 그리고 단일 단계 추론의 취약성으로 인해 어려움을 겪습니다. 이러한 문제점을 해결하기 위해, 우리는 검색 기반 다중 모드 다중 에이전트 태도 감지(MM-StanceDet)라는 새로운 다중 에이전트 프레임워크를 제안합니다. MM-StanceDet은 문맥적 이해를 위한 검색 증강, 미묘한 해석을 위한 특수화된 다중 모드 분석 에이전트, 다양한 관점을 탐색하기 위한 추론 강화 토론 단계, 그리고 견고한 판단을 위한 자기 성찰 기능을 통합합니다. 다섯 가지 데이터 세트에 대한 광범위한 실험 결과, MM-StanceDet이 최첨단 모델보다 훨씬 뛰어난 성능을 보였으며, 이는 다중 에이전트 아키텍처와 구조화된 추론 단계가 복잡한 다중 모드 태도 문제를 해결하는 데 효과적임을 입증합니다.
Multimodal Stance Detection (MSD) is crucial for understanding public discourse, yet effectively fusing text and image, especially with conflicting signals, remains challenging. Existing methods often face difficulties with contextual grounding, cross-modal interpretation ambiguity, and single-pass reasoning fragility. To address these, we propose Retrieval-Augmented Multi-modal Multi-agent Stance Detection (MM-StanceDet), a novel multi-agent framework integrating Retrieval Augmentation for contextual grounding, specialized Multimodal Analysis agents for nuanced interpretation, a Reasoning-Enhanced Debate stage for exploring perspectives, and Self-Reflection for robust adjudication. Extensive experiments on five datasets demonstrate MM-StanceDet significantly outperforms state-of-the-art baselines, validating the efficacy of its multi-agent architecture and structured reasoning stages in addressing complex multimodal stance challenges.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.