2603.23916v1 Mar 25, 2026 cs.CV

DecepGPT: 다문화 데이터셋과 강력한 다중 모드 학습을 활용한 스키마 기반 기만 탐지

DecepGPT: Schema-Driven Deception Detection with Multicultural Datasets and Robust Multimodal Learning

Jiajia Huang
Jiajia Huang
Citations: 1
h-index: 1
Dongliang Zhu
Dongliang Zhu
Citations: 19
h-index: 3
Zitong Yu
Zitong Yu
Citations: 23
h-index: 3
Huimin Ma
Huimin Ma
Citations: 254
h-index: 10
Jiayu Zhang
Jiayu Zhang
Citations: 3
h-index: 1
Chunmei Zhu
Chunmei Zhu
Citations: 29
h-index: 2
Xiaochun Cao
Xiaochun Cao
Citations: 136
h-index: 4

다중 모드 기만 탐지는 법의학 및 보안 분야에서 시청각 단서를 분석하여 기만적인 행동을 식별하는 것을 목표로 합니다. 이러한 중요한 상황에서 조사관은 시청각 단서와 최종 결정 간의 연결을 입증할 수 있는 검증 가능한 증거가 필요하며, 다양한 분야 및 문화적 맥락에서의 신뢰할 수 있는 일반화가 중요합니다. 그러나 기존 벤치마크는 중간 추론 단서 없이 이진 레이블만 제공합니다. 또한, 데이터셋은 규모가 작고 시나리오 범위가 제한적이어서 단관 학습(shortcut learning)을 유발합니다. 우리는 세 가지 주요 기여를 통해 이러한 문제를 해결합니다. 첫째, 기존 벤치마크에 구조화된 단서 수준 설명과 추론 체인을 추가하여 추론 데이터셋을 구축함으로써 모델의 출력에 대한 감사 가능 보고서를 생성합니다. 둘째, 네 국가에서 통일된 ``To Tell The Truth'' 텔레비전 형식을 기반으로 한 다문화 데이터셋인 T4-Deception을 공개합니다. 1695개의 샘플로 구성된 이 데이터셋은 가장 큰 비-실험실 기만 탐지 데이터셋입니다. 셋째, 소규모 데이터 환경에서의 강력한 학습을 위한 두 가지 모듈을 제안합니다. Stabilized Individuality-Commonality Synergy (SICS)는 학습 가능한 전역 사전 지식을 샘플 적응 잔차와 결합하여 다중 모드 표현을 개선하고, 양극성 인지 조정(polarity-aware adjustment)을 통해 표현을 양방향으로 재조정합니다. Distilled Modality Consistency (DMC)는 지식 증류를 통해 모달리티별 예측을 융합된 다중 모드 예측과 일치시켜 단일 모달리티 단관 학습을 방지합니다. 세 가지 기존 벤치마크와 새로운 데이터셋에 대한 실험 결과, 제안된 방법은 도메인 내 및 도메인 간 시나리오 모두에서 최첨단 성능을 달성하며, 다양한 문화적 맥락에서 우수한 일반화 성능을 보입니다. 데이터셋과 코드는 공개될 예정입니다.

Original Abstract

Multimodal deception detection aims to identify deceptive behavior by analyzing audiovisual cues for forensics and security. In these high-stakes settings, investigators need verifiable evidence connecting audiovisual cues to final decisions, along with reliable generalization across domains and cultural contexts. However, existing benchmarks provide only binary labels without intermediate reasoning cues. Datasets are also small with limited scenario coverage, leading to shortcut learning. We address these issues through three contributions. First, we construct reasoning datasets by augmenting existing benchmarks with structured cue-level descriptions and reasoning chains, enabling model output auditable reports. Second, we release T4-Deception, a multicultural dataset based on the unified ``To Tell The Truth'' television format implemented across four countries. With 1695 samples, it is the largest non-laboratory deception detection dataset. Third, we propose two modules for robust learning under small-data conditions. Stabilized Individuality-Commonality Synergy (SICS) refines multimodal representations by synergizing learnable global priors with sample-adaptive residuals, followed by a polarity-aware adjustment that bi-directionally recalibrates representations. Distilled Modality Consistency (DMC) aligns modality-specific predictions with the fused multimodal predictions via knowledge distillation to prevent unimodal shortcut learning. Experiments on three established benchmarks and our novel dataset demonstrate that our method achieves state-of-the-art performance in both in-domain and cross-domain scenarios, while exhibiting superior transferability across diverse cultural contexts. The datasets and codes will be released.

0 Citations
0 Influential
5 Altmetric
25.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!