개념 활성화 벡터를 이용한 L2 구두 평가 시스템의 편향 분석
Bias Analysis of L2 Speaking Assessment Systems Using Concept Activation Vectors
자동화된 구두 평가 시스템이 제2 언어(L2) 학습자의 구두 시험을 채점하는 데 점점 더 많이 사용되고 있으며, 이러한 시스템의 점수가 구두 능력에 의존하고 모국어(L1) 또는 나이와 같은 관련 없는 화자 속성에 의존하지 않는다는 것을 입증하는 것이 중요합니다. 트랜스포머 기반 기초 모델은 이러한 L2 구두 평가기의 정확성을 향상시켰지만, 이들의 블랙박스 특성으로 인해 공정성과 해석 가능성 분석이 더욱 어려워졌습니다. 본 연구는 특징 기반 평가기에서 원치 않는 속성('개념')에 대한 편향을 감지하기 위해 사용된 개념 활성화 벡터(CAV)에 기반한 기존 연구를 바탕으로, 텍스트 기반 BERT 평가기와 Whisper 기반 음성 및 텍스트 다중 모드 평가 시스템의 두 가지 신경망 구두 평가 시스템에 CAV 기반 분석을 확장합니다. CAV는 모델의 활성화 공간에서 인간이 해석할 수 있는 개념을 방향으로 표현하여, 모델의 내부 표현에 개념이 인코딩되어 있는지 여부와 예측된 점수에 영향을 미치는지 여부를 구별할 수 있게 합니다. 후자는 기울기 기반 감도 메트릭을 사용하여 정량화됩니다. CAV는 선형 분리성에 의존하는데, 이는 복잡한 신경망 임베딩 공간에서 덜 가능하므로, 희소 오토인코더(SAE)가 희소 잠재 공간에서 CAV를 학습하고 이를 활성화 공간으로 매핑함으로써 더 깨끗한 개념 방향을 제공하는지 조사합니다. 분석 결과, 개념 재현성은 탐색되는 표현 및 아키텍처에 따라 크게 달라지며, 개념 자체에는 덜 의존적인 것으로 나타났습니다. 개념에 대한 민감성 또한 아키텍처에 따라 달라집니다. SAE는 개념을 더 선형적으로 재현 가능하게 만들지만, 특히 저차원 레이어에서 원래 활성화 공간의 민감도를 감소시킵니다. 이러한 결과는 구두 평가 시스템의 편향을 감사할 때 개념 재현성과 개념 영향력을 구별해야 할 필요성을 강조합니다.
Automatic speaking assessment systems are increasingly deployed in high-stakes settings to mark second language (L2) learners' speaking tests, making it critical to show that their scores depend on speaking proficiency rather than irrelevant speaker attributes such as first language (L1) or age. Transformer-based foundation models have improved the accuracy of these L2 speaking graders, but their black-box representations make fairness and interpretability analysis more difficult. Building on prior work that used Concept Activation Vectors (CAVs) to detect bias towards unwanted attributes (`concepts') in feature-based graders, we extend CAV-based analysis to two neural speaking assessment systems: a text-based BERT grader and a speech-and-text multimodal grader based on Whisper. CAVs represent human-interpretable concepts as directions in a model's activation space, allowing us to distinguish between whether a concept is encoded in a model's internal representations and whether it influences the predicted score, the latter quantified using a gradient-based sensitivity metric. Since CAVs rely on linear separability, which is less likely in complex neural embedding spaces, we also investigate whether sparse autoencoders (SAEs) provide cleaner concept directions by learning CAVs in a sparse latent space and mapping them back to activation space. Our analysis shows that concept recoverability depends strongly on the representation and architecture being probed, rather than on the concept alone. Sensitivity to concepts is also architecture-dependent. SAEs make concepts more linearly recoverable, but attenuate the original activation-space sensitivity, especially in low-dimensional layers. These findings highlight the need to distinguish concept recoverability from concept influence when auditing bias in speaking assessment systems.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.