ReMMD: 현실적인 다국어, 다중 이미지 기반 에이전트 검증을 통한 다중 모드 허위 정보 탐지
ReMMD: Realistic Multilingual Multi-Image Agentic Verification for Multimodal Misinformation Detection
다국어 서술, 여러 이미지, 다양한 출처 및 미묘한 텍스트-이미지 프레임 오류를 결합한 바이럴 게시물이 증가하면서 다중 모드 허위 정보 탐지의 중요성이 점점 커지고 있습니다. 기존의 벤치마크와 방법론은 이러한 환경에 제대로 부합하지 못하며, 일반적으로 짧은 캡션, 단일 이미지, 이진 레이블 또는 하나의 조작 출처만을 고려합니다. 또한, 현실적인 증거 검색 하에서 에이전트 기반 검증은 비용이 많이 듭니다. 본 논문에서는 다중 모드 허위 정보 탐지를 위한 현실적인 다국어, 다중 이미지 기반 에이전트 검증 프레임워크인 ReMMD를 제시합니다. ReMMD는 500개의 샘플, 2,756개의 이미지, 다섯 가지 단일 언어, 두 가지 교차 언어 설정, 세 단계의 텍스트 길이, 다중 이미지 게시물, 다섯 단계의 진실성 레이블, 여덟 단계의 왜곡 레이블, 증거 출처 및 근거를 포함하는 실제 환경 기반의 다중 모드 허위 정보 탐지 벤치마크인 ReMMDBench를 포함합니다. 또한, ReMMD는 게시물을 원자적 요소로 분해하고 재사용 가능한 증거 세트를 구축하며 구조화된 L1/L2/L3 출력을 예측하는 지속 메모리 기반 검증기인 ReMMD-Agent를 포함합니다. 독점 시스템, 오픈 LVLM, MMD-Agent 및 T2-Agent를 통해 ReMMD-Agent는 GPT-5.2를 사용하여 41.80%의 정확도와 39.12%의 매크로 F1 점수를 달성하며, 가장 우수한 다섯 단계 진실성 성능을 보입니다. 또한, MMD-Agent에 비해 17.5%, T2-Agent에 비해 79.9% 더 낮은 비용으로 운영됩니다. 본 프로젝트는 https://dang-ai.github.io/ReMMD 에서 확인할 수 있습니다.
Multimodal misinformation detection is increasingly important because viral posts now combine long multilingual narratives, several images, mixed provenance, and subtle cross-modal framing errors. Existing benchmarks and methods remain poorly matched to this setting: they usually isolate short captions, single images, binary labels, or one manipulation source, while agentic verification remains costly under realistic evidence search. We present ReMMD, a realistic multilingual multi-image agentic verification framework for multimodal misinformation detection. ReMMD includes ReMMDBench, a real-world multimodal misinformation detection benchmark with 500 samples, 2,756 images, five monolingual evaluations, two cross-lingual settings, three text-length tiers, multi-image posts, five-way veracity labels, eight distortion labels, evidence provenance, and rationales. It also includes ReMMD-Agent, a persistent-memory verifier that decomposes posts into atomic points, builds a reusable evidence set, and predicts structured veracity verdicts, fine-grained distortion diagnoses, and explanatory rationales. Across proprietary systems, open LVLMs, MMD-Agent, and T$^2$-Agent, ReMMD-Agent obtains the best five-way veracity performance, with 41.80% accuracy and 39.12% macro-F1 using GPT-5.2, while reducing cost by 17.5% relative to MMD-Agent and 79.9% relative to T$^2$-Agent. The project is available at https://dang-ai.github.io/ReMMD.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.