2605.25534v1 May 25, 2026 cs.AI

StructBreak: 구조적 인지 과부하로 인한 안전 실패 사례 연구 – 멀티모달 대규모 언어 모델(MLLM)을 중심으로

StructBreak: Structural Cognitive Overload-Induced Safety Failures in MLLMs

Zhiyi Yin
Zhiyi Yin
Citations: 73
h-index: 6
Yang Luo
Yang Luo
Citations: 240
h-index: 7
S. Li
S. Li
Citations: 809
h-index: 13
Xinran Liu
Xinran Liu
Citations: 202
h-index: 8
Tiantian Ji
Tiantian Ji
Citations: 0
h-index: 0
Lingyun Peng
Lingyun Peng
Citations: 63
h-index: 1

멀티모달 대규모 언어 모델(MLLM)은 구조적 추론에 뛰어난 성능을 보이지만, 구조적 일관성 측면에서 심각한 논리적 취약점을 가지고 있습니다. 우리는 이러한 현상을 '구조적 인지 과부하(SCO)'라고 명명하며, 이는 깊이 있는 추론과 안전 정렬 간의 충돌로 인해 발생하는 부산물입니다. 기존 연구는 주로 텍스트 및 픽셀 수준의 교란에 초점을 맞추었으며, SCO에 대한 연구는 상대적으로 부족했습니다. 이에 따라, 우리는 SCO를 정량화하기 위한 자동화된 엔드-투-엔드 프레임워크인 StructBreak를 제안합니다. StructBreak를 활용하여, 우리는 새로운 고차원 인지 과부하 공격 패러다임을 발견했으며, 이는 실제 블랙박스 환경에서 작동하며 모델 내부 정보에 대한 접근이 필요하지 않습니다. 따라서, 이 프레임워크를 사용하여 10가지 다양한 위협 시나리오를 포괄하는 종합적인 벤치마크를 구축했습니다. 6개의 선도적인 MLLM에 대한 실험적 평가 결과, SCO는 쉽게 유해한 콘텐츠 생성을 유발하며, 평균적으로 92%의 ASR(Gemini 2.5에서는 최대 97%)을 보입니다. SCO의 작동 메커니즘을 규명하기 위해, 우리는 어텐션 다이내믹스, 잠재 공간 토폴로지 및 기하학적 분석 등 모델 수준의 해석을 추가적으로 수행했습니다. 우리의 연구 결과는 StructBreak가 안전 필터를 우회하는 새로운 구조적 채널 역할을 한다는 것을 보여줍니다. 또한, 기존 안전 메커니즘의 한계는 현재의 정렬 패러다임이 복잡한 멀티모달 추론 시대에 충분하지 않음을 강조합니다.

Original Abstract

Multimodal Large Language Models (MLLMs) excel at structural reasoning yet suffer from a sharp logical brittleness in structural consistency. We term this phenomenon Structural Cognitive Overload (SCO), a byproduct of the contention between deep reasoning and safety alignment. However, prior work has predominantly targeted typographic and pixel-level perturbations, leaving the study of SCO largely unexplored. To this end, we propose StructBreak, an automated end-to-end framework designed to quantify SCO. By leveraging StructBreak, we uncover a novel higher-order cognitive overload attack paradigm; notably, this attack operates under a practical black-box setting, requiring no internal model access. Consequently, we utilize this framework to establish a comprehensive benchmark spanning ten diverse threat scenarios. Empirical evaluations on six leading MLLMs reveal that SCO readily triggers toxic generation, yielding a 92% average ASR (up to 97% on Gemini 2.5). To elucidate the mechanism of SCO, we further conduct model-level interpretations spanning attention dynamics, latent space topology, and geometric analysis. Our findings reveal that StructBreak acts as a novel structural channel to circumvent safety filters. Furthermore, the limited efficacy of inherent safety mechanisms underscores that current alignment paradigms are insufficient for the era of complex multimodal reasoning.

0 Citations
0 Influential
6.5 Altmetric
32.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!