훈련 없이 다중 모달 대규모 언어 모델을 활용한 음성 딥페이크 탐지를 위한 설명 기반의 설명 생성
XAI-Grounded Explanation Generation for Speech Deepfake Detection with Training-Free Multimodal Large Language Models
음성 딥페이크 탐지(SDD) 시스템은 신뢰할 수 있는 의사 결정을 위해 투명하고 이해하기 쉬운 설명을 제공해야 합니다. 기존의 설명 방식은 크게 두 가지 범주로 나뉩니다. 전통적인 설명 가능한 인공지능(XAI) 방법, 예를 들어 그래디언트 기반 기여도 분석은 모델의 결정과 밀접하게 연관된 저수준 신호 정보를 제공하며, 자연어 설명을 통해 인간이 이해하기 어려울 수 있습니다. 반면, 대규모 언어 모델(LLM)을 활용한 설명 생성 방식은 종종 휴리스틱 증거 및 작업별 감독 정보 부족으로 인해 일반적이고 근거 없는 설명을 생성하는 경향이 있으며, 이는 SDD를 위한 제한적인 설명 데이터셋에서 비롯됩니다. 따라서 본 연구에서는 XAI 기반 증거와 다중 모달 LLM을 통합하여 보다 구체적이고 신뢰할 수 있는 설명을 생성하는 훈련 없이 사용할 수 있는 설명 프레임워크를 제안합니다. PartialSpoof 데이터셋을 사용하여 근거 기반의 설명 데이터셋을 구축하고, XAI 방법을 사용한 모델이 내부 정확도를 45% 이상 향상시켰으며, 이는 인간 평가 및 신뢰성 검증을 통해 확인되었습니다.
Speech deepfake detection (SDD) systems require trustworthy explanations for reliable decision-making. Existing explanation ways mainly fall into two categories. Traditional explainable AI (XAI), such as gradient-based attribution, produces low-level attribution signals tightly coupled with model decisions, and harder to be understood by human than natural language explanations. Meanwhile, large language model (LLM)-based explanation generation often produces generic and ungrounded descriptions due to the lack of heuristic evidence and task-specific supervision, stemming from limited grounded explanation datasets for SDD. We therefore propose a training-free explanation framework that integrates XAI evidence with multimodal LLMs to generate grounded and specific explanations. Using the PartialSpoof dataset, we construct a grounded explanation dataset and show that methods with XAI increase inside accuracy by over 45\%, verified through human evaluation and faithfulness checks.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.