2607.18080v1 Jul 20, 2026 cs.CV

희소 증거도 충분하다: 다중 모드 동영상 허위 정보 탐지를 위한 능동적인 증거 탐색

Sparse Evidence Can Suffice: Agentic Evidence Seeking for Multimodal Video Misinformation Detection

Yongxiu Xu
Yongxiu Xu
Citations: 130
h-index: 4
Yubin Wang
Yubin Wang
Citations: 10
h-index: 2
Gaopeng Gou
Gaopeng Gou
Citations: 2,149
h-index: 22
Xinkui Lin
Xinkui Lin
Citations: 8
h-index: 2
Dong Xie
Dong Xie
Citations: 0
h-index: 0
Yuqi Qian
Yuqi Qian
Citations: 0
h-index: 0
Hongbo Xu
Hongbo Xu
Citations: 9
h-index: 2
Haochen Zhao
Haochen Zhao
Citations: 5
h-index: 1
Jiarui Lu
Jiarui Lu
Citations: 12
h-index: 2

다중 모드 동영상 허위 정보 탐지는 일반적으로 전체 동영상과 관련된 콘텐츠를 한 번에 처리하고 판단하는 통합적인 동영상 이해 작업으로 구성됩니다. 그러나 실제 허위 정보는 종종 희소하고 조합된 증거 구조를 가지고 있습니다. 즉, 신뢰할 수 있는 결정은 몇 가지 연관된 단서에만 의존하며, 대부분의 동영상 콘텐츠는 제한적인 추가 정보를 제공합니다. 따라서 포괄적인 다중 모드 추론은 상당한 중복을 초래하고 결정적인 증거를 가릴 수 있습니다. 이에 따라 우리는 증거 획득과 검증을 분리하는 방법을 제안합니다. 즉, 먼저 의사 결정에 중요한 희소한 단서를 식별한 다음, 획득된 증거를 기반으로 진실성을 판단합니다. 이에 따라 다중 모드 동영상 허위 정보 탐지를 위한 희소 상호 작용 증거 검증 프레임워크인 SIEVE를 제안합니다. 증거 탐색 에이전트는 사용 가능한 다중 모드 증거를 능동적으로 탐색하고, 작은 크기의 증거 패키지를 구성하며, 이 패키지는 검증기가 진실성을 판단하는 데 사용됩니다. 에이전트는 지도 학습을 통해 획득된 증거 탐색 경로로 훈련되며, 정보적인 증거 획득을 촉진하고 불필요하거나 잘못된 상호 작용을 방지하는 증거 기반 강화 학습 목표를 사용합니다. 여러 동영상 허위 정보 평가 데이터 세트에 대한 실험 결과, SIEVE는 평가된 기본 모델보다 일관되게 우수한 성능을 보이며, 작은 크기의 증거 패키지를 사용하여 신뢰할 수 있는 검증을 지원합니다. 또한, 결과적으로 생성되는 획득 과정은 명시적이고 검토 가능한 증거 경로를 제공하여 다중 모드 허위 정보 탐지의 투명성과 근거성을 향상시킵니다.

Original Abstract

Multimodal video misinformation detection is commonly formulated as a holistic video-understanding task, where the entire video and its associated content are processed and judged in a single pass. However, real-world misinformation often exhibits a sparse and compositional evidence structure: a reliable decision may depend on only a few coupled clues, while most video content contributes limited additional information. Exhaustive multimodal reasoning may therefore introduce substantial redundancy and obscure decisive evidence. This motivates decoupling evidence acquisition from verification: first identifying sparse, decision-relevant clues and then judging veracity based on the acquired evidence. Accordingly, we propose SIEVE, a framework for Sparse Interactive Evidence Verification via Extraction in multimodal video misinformation detection. An evidence-seeking agent actively explores the available multimodal evidence and constructs a compact evidence package, which is then used by a verifier to determine veracity. The agent is trained with supervised evidence-seeking trajectories and an evidence-aware reinforcement learning objective that promotes informative evidence acquisition while discouraging unnecessary or invalid interactions. Experiments on multiple video misinformation benchmarks show that SIEVE consistently outperforms the evaluated baselines and supports reliable verification using compact evidence packages. Moreover, the resulting acquisition process provides an explicit and inspectable evidence trail, improving the transparency and groundedness of multimodal misinformation detection.

0 Citations
0 Influential
11 Altmetric
55.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!