2605.27932v1 May 27, 2026 cs.CV

이미지 기반 추론과 안전: 다중 모드 탈옥 공격에 대한 강건성을 결정하는 요인은 무엇인가?

When Think-with-Image Meets Safety: What Determines Multimodal Jailbreak Robustness?

Fangzhou Wu
Fangzhou Wu
Citations: 337
h-index: 7
Binghan Lu
Binghan Lu
Citations: 1
h-index: 1
Bing Hu
Bing Hu
Citations: 17
h-index: 2
Yuan Tian
Yuan Tian
Citations: 31
h-index: 3
Xiaomin Li
Xiaomin Li
Citations: 129
h-index: 6
N. Gong
N. Gong
Citations: 15,417
h-index: 64

이미지 기반 추론은 대규모 시각-언어 모델을 위한 새로운 추론 패러다임으로 부상하고 있지만, 그 안전성에 대한 이해는 아직 부족합니다. 기존 시스템들은 직접 응답 생성, 텍스트 전용 이전 단계, 시각적 상태 조작, 그리고 명시적인 외부 이미지 도구 호출 등 다양한 설계 방식을 포함합니다. 본 논문에서는 이러한 평가된 패러다미터 중 어떤 것이 다중 모드 탈옥 공격에 대한 강건성을 향상시키는지, 그리고 그 이유는 무엇인지 질문합니다. 여러 시각-언어 모델을 대상으로 실험한 결과, 명시적인 이미지 도구 상호 작용은 가장 낮은 공격 성공률을 보였으며, 평균적으로 평가된 모델에서 약 30%의 탈옥 성공률 감소를 나타냈습니다. 이 결과는 초기에는 놀랍게 느껴질 수 있습니다. 왜냐하면 반환된 이미지 도구의 출력이 수동으로 변경되거나 자체적으로 안전하지 않아 보이는 경우에도 공격 성공률(ASR)은 여전히 낮지만, 텍스트 전용 이전 단계에서는 거의 직접 답변 수준의 ASR을 보이기 때문입니다. 이러한 결과는 낮은 ASR이 단순히 반환된 이미지의 의미론적이거나 텍스트 기반의 이미지 도구 추적에 의해 설명될 수 없음을 시사합니다. 이 패턴을 설명하기 위해, 우리는 이미지 도구 호출을 안전과 관련된 방향으로 숨겨진 표현의 잔류적인 변화로 모델링하는 이미지 도구 안전 벡터 프레임워크를 제안합니다. 표현 수준 분석 및 활성화 개입 연구는 이러한 설명을 뒷받침합니다. 전반적으로, 우리의 결과는 명시적인 이미지 도구 상호 작용이 탈옥 공격에 대한 강건성을 향상시키는 유망한 설계 방식임을 시사하며, 동시에 파이프라인별 안전성 평가의 필요성을 강조합니다.

Original Abstract

Think-with-image reasoning is emerging as a new inference paradigm for large vision-language models, but its safety implications remain poorly understood. Existing systems already span multiple process designs, including direct response generation, text-only prior turn, visual-state manipulation, and explicit external image-tool invocation. In this paper, we ask which of these evaluated paradigms improves multimodal jailbreak robustness, and why. Across multiple vision-language models, explicit image-tool interaction yields the lowest attack success rates in our experiments, reducing jailbreak success by around 30% relative on average across the evaluated models. This finding is initially surprising: ASR remains low even when the returned image-tool output is manually overridden or itself unsafe-looking, but returns near direct-answering levels under text-only prior turn controls. These results indicate that the lower ASR is not explained by benign returned-image semantics or by the textual image-tool trace alone. To explain the pattern, we introduce an image-tool safety vector framework that models image-tool invocation as a residual shift in hidden representations toward a safety-relevant direction. Representation-level analyses and activation interventions support this account. Overall, our results suggest that explicit image-tool interaction is a promising design pattern for improving jailbreak robustness, while also motivating pipeline-specific safety evaluation.

0 Citations
0 Influential
30 Altmetric
150.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!