대규모 오디오 언어 모델의 오디오 제약 우회 공격: 분류, 공격-방어 분석 및 비용 고려 평가
Audio Jailbreaks in Large Audio-Language Models: Taxonomy, Attack-Defense Analysis, and Cost-Aware Evaluation
대규모 오디오 언어 모델(LALMs)은 제약 우회 위험을 토큰 수준 프롬프트에서 음성 인식부터 추론에 이르는 전체 파이프라인으로 확장하며, 여기서 부적절한 행동은 의미, 음향 스타일, 신호 왜곡 또는 내부 표현을 통해 유발될 수 있습니다. 기존 연구는 다양한 위협 모델 및 평가 프로토콜 하에서 이러한 위험을 다루고 있으며, 이는 공격의 실용성이나 방어의 효용성을 비교하기 어렵게 만듭니다. 본 논문에서는 LALM 제약 우회 공격 및 방어에 대한 통일된 분류 체계와 통제된 경험적 평가를 제공합니다. 우리는 기존 연구를 의미 기반, 음향 기반, 신호 기반, 임베딩 레이어 공격으로; 가드 기반, 학습 불필요형, 학습 기반 방어로; 그리고 교차 모달, 오디오 전용, 인터랙티브 벤치마크로 분류했습니다. 그런 다음, 열 가지 오픈 소스 LALM에서 대표적인 공격 및 방어를 평가하여 공격 성공률뿐만 아니라 정상적인 거부 반응과 지연 시간도 측정합니다. 우리의 결과는 Acoustic Best-of-N이 오디오 공간의 심각한 취약점을 드러내고, Narrative Framing은 효과적인 저지연 의미 기반 위협이며, 현재 방어 방법은 견고성과 유용성 간의 균형을 맞추기 어렵다는 것을 보여줍니다. 이러한 발견은 성공률만을 기준으로 하는 LALM 안전성 벤치마크에 더하여 비용 및 유용성을 고려한 평가가 필요함을 뒷받침합니다.
Large Audio Language Models (LALMs) expand jailbreak risks from token-level prompting to the full speech perception-to-reasoning pipeline, where unsafe behavior can be induced through semantics, acoustic style, signal artifacts, or internal representations. Existing work studies these risks under heterogeneous threat models and evaluation protocols, making it difficult to compare attack practicality or defense utility. This paper provides a unified taxonomy and a controlled empirical evaluation of LALM jailbreak attacks and defenses. We organize prior work into semantic, acoustic, signal, and embedding-layer attacks; guard-based, training-free, and training-based defenses; and cross-modal, audio-native, and interactive benchmarks. We then evaluate representative attacks and defenses across ten open-source LALMs, measuring not only attack success rate but also benign refusal and latency. Our results show that Acoustic Best-of-N reveals strong worst-case audio-space vulnerabilities, Narrative Framing is an effective low-latency semantic threat, and current defenses trade robustness against benign usability. These findings support cost- and utility-aware evaluation as a necessary complement to success-rate-only LALM safety benchmarks.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.