부재의 소리: 오디오-언어 임베딩 모델은 부정(否定) 표현에 어려움을 겪는다
The Sound of Absence: Audio-Language Embedding Models Struggle with Negation
CLAP과 같은 오디오-언어 임베딩 모델은 주로 현재 발생하는 음향 이벤트 매칭 성능 평가에 사용되지만, 부정 표현에 대한 평가는 드물다. 본 연구에서는 이러한 긍정(肯定) 중심의 평가 방식이 중요한 한계점을 숨기고 있음을 보여준다. 즉, 해당 모델들은 부정된 음향 개념을 제대로 인코딩하지 못하며, 긍정과 부정 문구 설명을 거의 동일한 표현으로 매핑한다. 이러한 문제점을 밝히기 위해, 본 연구에서는 기존 데이터셋을 두 가지의 부정 인식 태스크, 즉 Retrieval-Neg 및 Multiple-Choice Negation (MCQ-Neg)으로 변환하는 프레임워크인 NegEval-Audio를 제안하여 모델들이 현재 존재하거나 부재 상태인 음향 이벤트 간의 구분을 제대로 수행하는지 확인한다. AudioCaps와 Clotho 데이터셋에 대한 실험 결과, 부정 표현이 포함될 경우 성능이 현저하게 저하되었으며, 특히 MCQ 유형의 정확도는 우연히 맞힐 확률보다 훨씬 낮았다. 이러한 문제는 최근 등장한 멀티모달 LLM 기반 임베딩 모델에서도 지속적으로 나타났다. 훈련 과정 없이 부정에 대한 지침(steering)을 적용하는 방법이 MCQ-Neg 성능 향상에 일부 도움이 되었지만, Retrieval-Neg에는 큰 효과를 보이지 않았다. 이는 긍정 편향이 표현 공간의 근본적인 결함이며, 명시적인 부정 인식 학습 목표가 필요함을 시사한다.
Audio-language embedding models such as CLAP are widely evaluated on matching present sound events, but rarely on negation. We show this affirmation-only evaluation hides a key limitation: these models fail to encode negated sound concepts, mapping affirmative and negated captions to nearly identical representations. To expose this blind spot, we introduce NegEval-Audio, a framework that converts existing datasets into two negation-aware tasks, Retrieval-Neg and Multiple-Choice Negation (MCQ-Neg), to probe whether models distinguish present from absent events. On AudioCaps and Clotho, performance degrades sharply under negation, with negation-type MCQ accuracy falling far below chance, and the failure persists even for a recent multimodal LLM-based embedding model. While a training-free steering method improves MCQ-Neg, it yields marginal gains for Retrieval-Neg. This indicates that affirmation bias is a fundamental flaw in the representation geometry, necessitating explicit negation-aware training objectives.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.