2603.17476v1 Mar 18, 2026 cs.CV

UniSAFE: 통합 다중 모드 모델의 안전성 평가를 위한 종합적인 벤치마크

UniSAFE: A Comprehensive Benchmark for Safety Evaluation of Unified Multimodal Models

Hojung Jung
Hojung Jung
Citations: 55
h-index: 5
Se-Young Yun
Se-Young Yun
Citations: 28
h-index: 1
Segyu Lee
Segyu Lee
Citations: 2
h-index: 1
Sang-Sub Jang
Sang-Sub Jang
Citations: 88
h-index: 6
Boryeong Cho
Boryeong Cho
Citations: 8
h-index: 1
S. An
S. An
Citations: 0
h-index: 0
Juhyeong Kim
Juhyeong Kim
Citations: 2
h-index: 1
Jaehyun Kwak
Jaehyun Kwak
Citations: 12
h-index: 2
Yongjin Yang
Yongjin Yang
Citations: 190
h-index: 5
Young-Bin Park
Young-Bin Park
Citations: 33
h-index: 3
Wonjune Chang
Wonjune Chang
Citations: 8
h-index: 1

통합 다중 모드 모델(Unified Multimodal Models, UMMs)은 강력한 교차 모달성 기능을 제공하지만, 단일 작업 모델에서는 관찰되지 않는 새로운 안전성 위험을 초래합니다. 하지만 현재의 안전성 벤치마크는 여전히 작업 및 모달리티별로 분산되어 있어 복잡한 시스템 수준의 취약점을 종합적으로 평가하는 데 한계가 있습니다. 이러한 문제를 해결하기 위해, 우리는 7가지 I/O 모달리티 조합에 걸쳐 기존 작업과 새로운 다중 모드 컨텍스트 이미지 생성 환경을 포괄하는 시스템 수준의 UMM 안전성 평가를 위한 최초의 종합 벤치마크인 UniSAFE를 소개합니다. UniSAFE는 공통된 위험 시나리오를 작업별 I/O 구성에 투영하는 공유 대상 설계를 사용하여 안전성 실패에 대한 제어된 교차 작업 비교를 가능하게 합니다. 6,802개의 선별된 데이터 인스턴스로 구성된 UniSAFE를 사용하여 15개의 최첨단 UMM(특허 및 오픈 소스)을 평가했습니다. 우리의 결과는 현재의 UMM에서 발견되는 중요한 취약점을 보여주며, 특히 다중 이미지 구성 및 다중 턴 설정에서 안전성 위반이 심각하며, 텍스트 출력 작업보다 이미지 출력 작업이 더 취약하다는 것을 알 수 있습니다. 이러한 결과는 UMM의 시스템 수준 안전성 정렬을 강화할 필요성을 강조합니다. 저희의 코드와 데이터는 https://github.com/segyulee/UniSAFE 에서 공개적으로 이용하실 수 있습니다.

Original Abstract

Unified Multimodal Models (UMMs) offer powerful cross-modality capabilities but introduce new safety risks not observed in single-task models. Despite their emergence, existing safety benchmarks remain fragmented across tasks and modalities, limiting the comprehensive evaluation of complex system-level vulnerabilities. To address this gap, we introduce UniSAFE, the first comprehensive benchmark for system-level safety evaluation of UMMs across 7 I/O modality combinations, spanning conventional tasks and novel multimodal-context image generation settings. UniSAFE is built with a shared-target design that projects common risk scenarios across task-specific I/O configurations, enabling controlled cross-task comparisons of safety failures. Comprising 6,802 curated instances, we use UniSAFE to evaluate 15 state-of-the-art UMMs, both proprietary and open-source. Our results reveal critical vulnerabilities across current UMMs, including elevated safety violations in multi-image composition and multi-turn settings, with image-output tasks consistently more vulnerable than text-output tasks. These findings highlight the need for stronger system-level safety alignment for UMMs. Our code and data are publicly available at https://github.com/segyulee/UniSAFE

0 Citations
0 Influential
23 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!