V-DEAL: 비디오 안전성 오류를 이해 거부 결합 실패로 진단하는 방법
V-DEAL: Diagnosing Video Safety De-Calibration as an Understanding-Refusal Coupling Failure
비디오 대규모 언어 모델(Video Large Language Models)이 실제 응용 분야에 점점 더 많이 사용됨에 따라, 이들의 안전성을 확보하는 것이 매우 중요해졌습니다. 흥미롭게도, 우리는 유해한 비디오와 함께 무해한 질문을 사용할 때 동일한 비디오를 사용하여 명시적으로 유해한 질문과 함께 사용하는 경우보다 공격 성공률이 더 높다는 것을 발견했습니다. 이러한 취약점의 근본적인 메커니즘을 이해하기 위해, 모델의 행동, 이해 능력 및 내부 표현 전반에 걸쳐 이 실패를 공동으로 분석하는 세 단계로 구성된 진단 프레임워크인 V-DEAL을 제시합니다. V-DEAL은 인지 오류를 점진적으로 배제하고 모델의 내부 거부 경향을 정량화함으로써, 관찰된 취약점의 근본적인 메커니즘을 분석하는 새로운 진단적 관점을 제공합니다. 우리는 세 개의 공개 벤치마크에서 여섯 가지 비디오 LLM을 테스트했으며, 모델이 유해한 비디오 콘텐츠를 81% 이상의 정확도로 올바르게 인식한다는 것을 확인했지만, 유해한 비디오와 무해한 질문을 함께 사용하는 조건에서 평균 공격 성공률은 여전히 48.33%에 달했습니다. 추가적인 은닉 상태 분석 결과, 시각적 이해는 텍스트 기반 이해보다 약한 거부 경향을 활성화한다는 것을 보여줍니다. 또한, 우리는 프롬프트 주입 방지 방법을 도입하여 공격 성공률을 평균 48.24%만큼 줄였으며, 이는 기존의 파인튜닝 기반 방법과 유사한 성능을 제공합니다. 이 방법은 비디오 LLM에서 이러한 안전 관련 위험을 해결하는 효과적이고 실용적인 수단입니다.
As Video Large Language Models are increasingly deployed in real-world applications, ensuring their safety alignment has become critical. Counterintuitively, we find that harmful videos paired with benign queries achieve higher attack success rates than the same videos paired with explicitly harmful queries. To understand the underlying mechanism of this vulnerability, we present V-DEAL, a three-level diagnostic framework that jointly analyzes this failure across model behaviour, understanding, and internal representations. By progressively ruling out perception failure and quantifying the model's internal refusal tendency, V-DEAL provides a new diagnostic perspective for analyzing the underlying mechanism of the observed vulnerability. We tested six Video LLMs on three public benchmarks and observed that models correctly recognize harmful video content with over 81\% accuracy, yet the average attack success rate still reaches 48.33\% under the condition pairing harmful videos with benign queries. Hidden-state analysis further shows that visual understanding activates a weaker refusal tendency than textual understanding. Furthermore, we introduce a prompt injection intervention method that reduces attack success rates by an average of 48.24 percentage points and achieves performance comparable to prior fine-tuning-based methods, providing an effective and practical means to address such safety risks in Video LLMs.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.