RedDiffuser: 강화 학습 기반 확산 모델을 활용한 시각-언어 모델의 다중 모드 안전성 오류 감사
RedDiffuser: Auditing Multimodal Safety Failures in Vision-Language Models via Reinforced Diffusion
최근 대규모 시각-언어 모델(VLM)이 다양한 환경에 광범위하게 사용되면서, 다중 모드 입력을 통해 안정적인 안전성을 확보하는 것이 중요해졌습니다. 그러나 기존 평가 방법은 주로 명시적인 악의적 질문에 초점을 맞추고 있으며, 보다 현실적이고 간과된 위험인 '유해한 맥락 노출' 하에서의 안전성 유지를 제대로 고려하지 못하고 있습니다. 특히 다중 모드 시스템에서 시각적 입력은 모델의 동작을 크게 변화시키므로 텍스트만으로 감사를 수행하는 것은 충분하지 않습니다. 본 연구에서는 유해한 맥락 노출 환경에서 다중 모드 안전성을 감사하며, 부분적으로 유해한 텍스트가 포함된 시각적 맥락과 결합되었을 때 VLM이 안전한 동작을 유지하는지 여부를 조사합니다. 체계적인 감사를 위해, 우리는 블랙박스 안전 테스트를 위한 의미적으로 일관성 있는 시각적 입력을 생성하는 강화 학습 기반 프레임워크인 RedDiffuser (RedDiff)를 제안합니다. RedDiffuser는 탐욕적인 프롬프트 검색과 강화 최적화를 결합하여 잠재적인 안전성 오류를 드러내는 고위험 다중 모드 입력을 찾아냅니다. 공개 소스 및 상용 VLM에 대한 광범위한 실험 결과, 이러한 맥락 의존적인 오류가 널리 존재한다는 것을 확인했습니다. LLaVA에서 RedDiffuser는 원래 데이터셋에서 최대 10.69%, 보류 데이터셋에서 8.91%까지 안전하지 않은 응답률을 증가시켰으며, Gemini 및 LLaMA-Vision으로 높은 수준의 성능을 유지했습니다. 이러한 취약점은 외부 안전 장치 하에서도 지속적으로 나타나므로, 현재 시스템 수준의 안전 메커니즘이 실제 다중 모드 위험에 충분히 대응할 수 없음을 시사합니다. 본 연구는 기존 안전 평가에서 간과된 중요한 부분을 밝혀냈으며, 현대 VLM 시스템의 숨겨진 취약점을 진단하기 위한 맥락 인식 다중 모드 감사 방법을 제시합니다.
Large Vision-Language Models (VLMs) are increasingly deployed in open-ended environments, where ensuring reliable safety under multimodal inputs is critical. However, existing evaluations remain largely instruction-centric, focusing on explicit malicious queries while overlooking a more realistic and underexplored risk: whether safety alignment remains robust under harmful contextual exposure. This limitation is particularly important for multimodal systems, where visual inputs can substantially steer model behavior and render text-only auditing insufficient. In this work, we study multimodal safety auditing under harmful contextual exposure, asking whether VLMs can maintain safe behavior when partial toxic text is paired with visual context. To enable systematic auditing, we propose RedDiffuser (RedDiff), a reinforcement-based framework that leverages diffusion models to generate semantically coherent visual inputs for black-box safety testing. By combining greedy prompt search with reinforcement optimization, RedDiffuser uncovers high-risk multimodal inputs that expose latent safety failures. Extensive experiments on both open-source and commercial VLMs show that such context-conditioned failures are widespread. On LLaVA, RedDiffuser increases unsafe response rates by up to 10.69% on the original set and 8.91% on a hold-out set, with strong transferability to Gemini and LLaMA-Vision. These vulnerabilities persist even under external safety guardrails, suggesting that current system-level safety mechanisms remain insufficient for realistic multimodal risks. Our findings reveal a critical blind spot in existing safety evaluations and establish context-aware multimodal auditing as an essential paradigm for diagnosing hidden vulnerabilities in modern VLM systems.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.