통합은 비용을 수반하는가? Uni-SafeBench: 통합형 다중 모드 대규모 모델을 위한 안전성 벤치마크
Does Unification Come at a Cost? Uni-SafeBench: A Safety Benchmark for Unified Multimodal Large Models
통합형 다중 모드 대규모 모델(UMLM)은 단일 아키텍처 내에서 이해 및 생성 능력을 통합합니다. 이러한 아키텍처 통합은 다중 모드 특징의 심층적인 융합을 통해 모델 성능을 향상시키지만, 동시에 중요한 안전 문제를 야기합니다. 기존의 안전성 벤치마크는 주로 개별적인 이해 또는 생성 작업에 초점을 맞추고 있어, 통합 프레임워크 하에서 다양한 작업을 처리할 때 UMLM의 전반적인 안전성을 평가하는 데 한계가 있습니다. 이러한 문제를 해결하기 위해, 우리는 7가지 작업 유형에 걸쳐 6가지 주요 안전 범주를 포함하는 포괄적인 벤치마크인 Uni-SafeBench를 소개합니다. 엄격한 평가를 위해, 우리는 맥락적 안전성과 고유한 안전성을 효과적으로 분리하는 프레임워크인 Uni-Judger를 개발했습니다. Uni-SafeBench를 통한 종합적인 평가 결과, 통합 과정은 모델 능력을 향상시키지만, 기본 LLM의 고유한 안전성을 크게 저하시킨다는 것을 발견했습니다. 또한, 공개 소스 UMLM은 생성 또는 이해 작업에 특화된 다중 모드 대규모 모델보다 안전 성능이 현저히 낮습니다. 우리는 이러한 위험을 체계적으로 드러내고 더 안전한 AGI 개발을 촉진하기 위해 모든 리소스를 공개합니다.
Unified Multimodal Large Models (UMLMs) integrate understanding and generation capabilities within a single architecture. While this architectural unification, driven by the deep fusion of multimodal features, enhances model performance, it also introduces important yet underexplored safety challenges. Existing safety benchmarks predominantly focus on isolated understanding or generation tasks, failing to evaluate the holistic safety of UMLMs when handling diverse tasks under a unified framework. To address this, we introduce Uni-SafeBench, a comprehensive benchmark featuring a taxonomy of six major safety categories across seven task types. To ensure rigorous assessment, we develop Uni-Judger, a framework that effectively decouples contextual safety from intrinsic safety. Based on comprehensive evaluations across Uni-SafeBench, we uncover that while the unification process enhances model capabilities, it significantly degrades the inherent safety of the underlying LLM. Furthermore, open-source UMLMs exhibit much lower safety performance than multimodal large models specialized for either generation or understanding tasks. We open-source all resources to systematically expose these risks and foster safer AGI development.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.