2601.11178v1 Jan 16, 2026 cs.AI

TANDEM: 멀티모달 혐오 표현을 위한 시간 인식 신경망 탐지

TANDEM: Temporal-Aware Neural Detection for Multimodal Hate Speech

Girish A. Koushik
Girish A. Koushik
Citations: 39
h-index: 3
Helen Treharne
Helen Treharne
Citations: 18
h-index: 2
Diptesh Kanojia
Diptesh Kanojia
IITB-Monash Research Academy
Citations: 1,292
h-index: 18

소셜 미디어 플랫폼은 오디오, 시각, 텍스트 단서들의 복잡한 상호작용을 통해 유해한 서사가 형성되는 긴 형식의 멀티모달 콘텐츠가 점차 주류를 이루고 있습니다. 자동화된 시스템은 혐오 표현을 높은 정확도로 식별할 수 있지만, 효과적인 인간 개입(human-in-the-loop) 중재에 필요한 정확한 타임스탬프 및 대상 식별과 같은 세밀하고 해석 가능한 증거를 제공하지 못하는 "블랙박스"로 작동하는 경우가 많습니다. 본 연구에서는 시청각 혐오 탐지를 단순한 이진 분류 작업에서 구조화된 추론 문제로 전환하는 통합 프레임워크인 TANDEM을 소개합니다. 우리의 접근 방식은 시각-언어 모델과 오디오-언어 모델이 자기 제약적(self-constrained) 교차 모달 문맥을 통해 서로를 최적화하는 새로운 탠덤 강화 학습 전략을 채택하여, 조밀한 프레임 단위의 지도 학습 없이도 긴 시간적 시퀀스에 대한 추론을 안정화합니다. 세 가지 벤치마크 데이터셋에 대한 실험 결과, TANDEM은 정확한 시간적 근거(grounding)를 유지하면서도 제로샷 및 문맥 증강 베이스라인을 크게 능가하였으며, 특히 HateMM 데이터셋의 대상 식별에서 최신 기술 대비 30% 향상된 0.73의 F1 점수를 달성했습니다. 또한 우리는 이진 탐지는 견고한 반면, 다중 클래스 환경에서 모욕적인 콘텐츠와 혐오스러운 콘텐츠를 구별하는 것은 내재된 레이블 모호성과 데이터 불균형으로 인해 여전히 어려운 과제임을 확인했습니다. 더 넓은 관점에서 본 연구 결과는 복잡한 멀티모달 환경에서도 구조화되고 해석 가능한 정렬이 달성 가능함을 시사하며, 투명하고 실질적인 차세대 온라인 안전 중재 도구를 위한 청사진을 제시합니다.

Original Abstract

Social media platforms are increasingly dominated by long-form multimodal content, where harmful narratives are constructed through a complex interplay of audio, visual, and textual cues. While automated systems can flag hate speech with high accuracy, they often function as "black boxes" that fail to provide the granular, interpretable evidence, such as precise timestamps and target identities, required for effective human-in-the-loop moderation. In this work, we introduce TANDEM, a unified framework that transforms audio-visual hate detection from a binary classification task into a structured reasoning problem. Our approach employs a novel tandem reinforcement learning strategy where vision-language and audio-language models optimize each other through self-constrained cross-modal context, stabilizing reasoning over extended temporal sequences without requiring dense frame-level supervision. Experiments across three benchmark datasets demonstrate that TANDEM significantly outperforms zero-shot and context-augmented baselines, achieving 0.73 F1 in target identification on HateMM (a 30% improvement over state-of-the-art) while maintaining precise temporal grounding. We further observe that while binary detection is robust, differentiating between offensive and hateful content remains challenging in multi-class settings due to inherent label ambiguity and dataset imbalance. More broadly, our findings suggest that structured, interpretable alignment is achievable even in complex multimodal settings, offering a blueprint for the next generation of transparent and actionable online safety moderation tools.

1 Citations
0 Influential
9 Altmetric
46.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!