HyperClaim: 비디오 허위 정보 탐지를 위한 세밀한 교차 모드 하이퍼 그래프 추론
HyperClaim: Fine-Grained Cross-Modal Hypergraph Reasoning for Video Misinformation Detection
비디오 허위 정보 탐지는 종종 전역적인 다중 모드 융합 또는 자유 형식의 다중 모드 추론을 통해 수행됩니다. 이러한 접근 방식 모두 질의 구절, 문맥적 텍스트 및 짧은 시간 간격 내의 프레임 간의 상호 작용에서 발생하는 국소화된 진위 여부 단서를 충분히 반영하지 못할 수 있습니다. 이러한 상호 작용은 본질적으로 고차원적이므로 쌍방향 그래프 표현 방식으로는 다중 모드 간의 복잡한 의존성을 포착하기 어렵습니다. 반면 하이퍼그래프는 이러한 관계를 나타내는 데 적합합니다. 우리는 표본 수준의 진위 여부 분류를 위한 차별적인 시간 기반 하이퍼그래프 프레임워크인 HyperClaim을 제안합니다. HyperClaim은 제목 또는 벤치마크에서 제공된 짝된 텍스트를 유사한 질의로 사용하여 질의 토큰, 증거 토큰 및 표본으로 추출된 프레임 간의 희소 이질적 하이퍼그래프를 구성합니다. 또한 신뢰도 기반 필터링과 소스 할당을 적용하여 텍스트-프레임 쌍 및 단거리 시간 정보 단위로 구성하고, 잔여 텍스트-비디오 보정을 통한 적응형 소프트 인시던스 추론을 수행하며, 불일치에 민감한 방식으로 텍스트, 시각 및 하이퍼엣지 상태를 집계합니다. HyperClaim은 생성된 근거 또는 외부 도구 호출에 의존하지 않고도 전역적인 융합 방식으로는 평탄화되는 세밀한 교차 모드 및 시간 구조를 유지합니다. FactGuard 시간 프로토콜에서 HyperClaim은 FakeSV, FakeTT 및 FakeVV 데이터셋에서 각각 83.7%, 82.0% 및 87.3%의 정확도를 달성하여 강력한 차별적 모델 및 추론 중심의 기준 모델보다 우수한 성능을 보였습니다. 학습된 인시던스 및 어텐션 가중치는 토큰 및 프레임 수준의 구조를 더욱 명확하게 보여줍니다.
Video misinformation detection is often approached through global multimodal fusion or free-form multimodal reasoning. Both paradigms can under-represent localized authenticity cues that arise from coupled interactions among query phrases, contextual text, and short temporal spans of frames. Because such interactions are inherently higher-order, pairwise graph formulations are insufficient to capture multi-way cross-modal dependencies, whereas hypergraphs offer a suitable representation for these relations. We propose HyperClaim, a discriminative temporal hypergraph framework for sample-level authenticity classification. Using the title or benchmark-provided paired text as a claim-like query, HyperClaim constructs a sparse heterogeneous hypergraph over query tokens, evidence tokens, and sampled frames; applies confidence-aware filtering and source budgeting to form compact text-frame and short-range temporal evidence units; performs adaptive soft-incidence reasoning with residual text-video calibration; and aggregates textual, visual, and hyperedge states through a discrepancy-aware readout. Without relying on generated rationales or external tool calls, HyperClaim preserves fine-grained cross-modal and temporal structure that global fusion tends to flatten. Under the FactGuard temporal protocol, it achieves 83.7%, 82.0%, and 87.3% accuracy on FakeSV, FakeTT, and FakeVV, respectively, outperforming strong discriminative and reasoning-centric baselines. Learned incidence and attention weights further reveal token- and frame-level structure.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.