UniSAFE: 통합 다중 모드 모델의 안전성 평가를 위한 종합적인 벤치마크
UniSAFE: A Comprehensive Benchmark for Safety Evaluation of Unified Multimodal Models
통합 다중 모드 모델(Unified Multimodal Models, UMMs)은 강력한 교차 모달성 기능을 제공하지만, 단일 작업 모델에서는 관찰되지 않는 새로운 안전성 위험을 초래합니다. 하지만 현재의 안전성 벤치마크는 여전히 작업 및 모달리티별로 분산되어 있어 복잡한 시스템 수준의 취약점을 종합적으로 평가하는 데 한계가 있습니다. 이러한 문제를 해결하기 위해, 우리는 7가지 I/O 모달리티 조합에 걸쳐 기존 작업과 새로운 다중 모드 컨텍스트 이미지 생성 환경을 포괄하는 시스템 수준의 UMM 안전성 평가를 위한 최초의 종합 벤치마크인 UniSAFE를 소개합니다. UniSAFE는 공통된 위험 시나리오를 작업별 I/O 구성에 투영하는 공유 대상 설계를 사용하여 안전성 실패에 대한 제어된 교차 작업 비교를 가능하게 합니다. 6,802개의 선별된 데이터 인스턴스로 구성된 UniSAFE를 사용하여 15개의 최첨단 UMM(특허 및 오픈 소스)을 평가했습니다. 우리의 결과는 현재의 UMM에서 발견되는 중요한 취약점을 보여주며, 특히 다중 이미지 구성 및 다중 턴 설정에서 안전성 위반이 심각하며, 텍스트 출력 작업보다 이미지 출력 작업이 더 취약하다는 것을 알 수 있습니다. 이러한 결과는 UMM의 시스템 수준 안전성 정렬을 강화할 필요성을 강조합니다. 저희의 코드와 데이터는 https://github.com/segyulee/UniSAFE 에서 공개적으로 이용하실 수 있습니다.
Unified Multimodal Models (UMMs) offer powerful cross-modality capabilities but introduce new safety risks not observed in single-task models. Despite their emergence, existing safety benchmarks remain fragmented across tasks and modalities, limiting the comprehensive evaluation of complex system-level vulnerabilities. To address this gap, we introduce UniSAFE, the first comprehensive benchmark for system-level safety evaluation of UMMs across 7 I/O modality combinations, spanning conventional tasks and novel multimodal-context image generation settings. UniSAFE is built with a shared-target design that projects common risk scenarios across task-specific I/O configurations, enabling controlled cross-task comparisons of safety failures. Comprising 6,802 curated instances, we use UniSAFE to evaluate 15 state-of-the-art UMMs, both proprietary and open-source. Our results reveal critical vulnerabilities across current UMMs, including elevated safety violations in multi-image composition and multi-turn settings, with image-output tasks consistently more vulnerable than text-output tasks. These findings highlight the need for stronger system-level safety alignment for UMMs. Our code and data are publicly available at https://github.com/segyulee/UniSAFE
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.