화자 인식에서의 설명 가능한 인공지능: 어텐션 맵 시각화 및 평가
Explainable AI in Speaker Recognition -- Attention Map Visualisation and Evaluation
인공지능(AI) 시스템, 특히 신경망으로 구현된 시스템의 의사 결정 과정을 설명하고 이해하는 것은 설명 가능한 인공지능(XAI) 분야에 속합니다. 인간의 주의 메커니즘과 유사하게, 신경망은 의사 결정을 내리는 동안 정보를 선택적으로 처리하는 자체적인 주의 메커니즘을 가지고 있다고 가정됩니다. 본 연구에서는 XAI의 한 영역인 신경망의 주의 메커니즘 분석 및 시각화를 제안합니다. 우리의 실험은 특정 발화로부터 화자 신원을 식별하도록 훈련된 화자 인식 신경망에 대해 수행되었습니다. 이전 연구에서는 신경망의 주의 메커니즘을 분석하고 시각화하기 위해 클래스 활성화 맵(CAM) 기반 방법이 광범위하게 사용되었습니다. 이러한 각 방법은 네트워크 입력에 대한 어텐션 맵을 생성하며, 화자 인식 네트워크가 의사 결정을 내릴 때 어떤 입력 영역이 선택적으로 처리되는지를 강조합니다. 그러나 이러한 방법에서 생성된 어텐션 맵의 평가는 아직 충분히 연구되지 않았습니다. 본 연구에서는 기존의 어텐션 맵 평가 알고리즘을 체계적으로 검토하고, 주요 개념을 확립하며, 그 한계를 파악합니다. 이 기존 평가 알고리즘을 바탕으로, 식별된 한계를 해결하기 위한 새로운 버전인 '수정된 랜덤 입력 샘플링 기반 설명-평가(Modified RISE-eval)' 알고리즘을 제안합니다. Modified RISE-eval을 사용하여 특정 화자 인식 네트워크에 적용된 두 가지 대표적인 CAM 기반 방법인 GradCAM과 LayerCAM에서 생성된 어텐션 맵을 평가했습니다. 평가 결과는 GradCAM과 LayerCAM이 화자 인식 작업의 다양한 실험 조건 하에서 각각 뚜렷한 장점을 가지고 있음을 보여줍니다.
Explaining and understanding the decision-making process of artificial intelligence (AI) systems, particularly those implemented by neural networks, falls within the field of explainable AI (XAI). Analogous to the human attention mechanism, neural networks are assumed to possess their own attention mechanisms that selectively process information during decision-making. This work proposes to study one XAI topic: analysing and visualising the attention mechanisms of neural networks. Our experiments are performed on speaker recognition neural networks that are trained to identify speaker identity from a given utterance. Previous studies have widely used class activation map (CAM)-based methods to analyse and visualise the attention mechanisms of neural networks. Each of these methods produces an attention map for each network input, highlighting which input regions are selectively processed when the speaker recognition network makes decisions. However, the evaluation of attention maps produced by these methods remains largely underexplored. This work systematically reviews an existing attention map evaluation algorithm, establishing key concepts and identifying its shortcomings. On the basis of this existing evaluation algorithm, a new version is then proposed to address the identified shortcomings, called the Modified Randomised Input Sampling for Explanation - Evaluation algorithm (Modified RISE-eval). Using Modified RISE-eval, we evaluate the attention maps produced by two representative CAM-based methods, GradCAM and LayerCAM, applied to a certain speaker recognition network. The evaluation results demonstrate that GradCAM and LayerCAM each exhibit distinct advantages when applied under different experimental conditions in the speaker recognition task.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.