2603.19092v1 Mar 19, 2026 cs.CV

SAVeS: 시맨틱 힌트를 활용하여 비전-언어 모델의 안전성 판단을 제어하는 방법

SAVeS: Steering Safety Judgments in Vision-Language Models via Semantic Cues

Carlos Hinojosa
Carlos Hinojosa
Citations: 71
h-index: 4
Bernard Ghanem
Bernard Ghanem
Citations: 33
h-index: 3
Clemens Grange
Clemens Grange
Citations: 1
h-index: 1

비전-언어 모델(VLMs)은 실제 환경 및 로봇 시스템에서 점점 더 많이 활용되고 있으며, 이때 안전 결정은 시각적 맥락에 크게 의존합니다. 그러나 이러한 판단에 어떤 시각적 정보가 영향을 미치는지는 아직 명확하지 않습니다. 본 연구에서는 VLMs의 다중 모달 안전 행동이 간단한 시맨틱 힌트를 통해 제어될 수 있는지 조사합니다. 우리는 장면 내용의 변화 없이 제어된 텍스트, 시각, 인지적 개입을 적용하는 시맨틱 스티어링 프레임워크를 제안합니다. 이러한 효과를 평가하기 위해, 우리는 시맨틱 힌트 하에서의 상황 안전성을 측정하는 벤치마크인 SAVeS를 제안하며, 행동 거부, 근거 기반 안전 추론, 오탐을 분리하여 평가하는 프로토콜을 함께 제시합니다. 여러 VLMs 및 최첨단 벤치마크를 대상으로 한 실험 결과, 안전 결정은 시맨틱 힌트에 매우 민감하게 반응하며, 이는 학습된 시각-언어 연관성에 대한 의존성을 나타내며, 실제 시각적 이해에 기반한 것이 아님을 시사합니다. 또한, 자동화된 스티어링 파이프라인이 이러한 메커니즘을 활용할 수 있음을 보여주며, 이는 다중 모달 안전 시스템의 잠재적인 취약점을 강조합니다.

Original Abstract

Vision-language models (VLMs) are increasingly deployed in real-world and embodied settings where safety decisions depend on visual context. However, it remains unclear which visual evidence drives these judgments. We study whether multimodal safety behavior in VLMs can be steered by simple semantic cues. We introduce a semantic steering framework that applies controlled textual, visual, and cognitive interventions without changing the underlying scene content. To evaluate these effects, we propose SAVeS, a benchmark for situational safety under semantic cues, together with an evaluation protocol that separates behavioral refusal, grounded safety reasoning, and false refusals. Experiments across multiple VLMs and an additional state-of-the-art benchmark show that safety decisions are highly sensitive to semantic cues, indicating reliance on learned visual-linguistic associations rather than grounded visual understanding. We further demonstrate that automated steering pipelines can exploit these mechanisms, highlighting a potential vulnerability in multimodal safety systems.

2 Citations
0 Influential
2 Altmetric
12.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!