2607.15218v1 Jul 16, 2026 cs.AI

단어는 안전하지만 행동은 위험할 때: 숨겨진 상태 위험 공간에서 텍스트 안전성 너머의 물리적 위험에 대한 탐구

When Words Are Safe But Actions Kill: Probing Physical Danger Beyond Text Safety in Hidden-State Risk Space

Qi Li
Qi Li
Citations: 95
h-index: 5
Ke Xu
Ke Xu
Citations: 87
h-index: 5
Zihang Zhan
Zihang Zhan
Citations: 3
h-index: 1
Chuanpu Fu
Chuanpu Fu
Citations: 874
h-index: 11
Weimeng Wang
Weimeng Wang
Citations: 0
h-index: 0
Ziqiang Wang
Ziqiang Wang
Citations: 36
h-index: 5

대규모 언어 모델(LLM)이 점점 더 많은 역할을 수행하며, 특히 구체적인 환경에서의 에이전트 제어에 사용되면서, 언어적으로는 무해해 보이는 명령조차도 실제 세계에서는 위험한 결과를 초래할 수 있습니다. 본 연구에서는 이러한 물리적 환경에서의 위험이 일반적인 텍스트 수준의 유해 콘텐츠 위험과 동일한 문제인지 조사합니다. 숨겨진 상태 방향 분석 및 랜덤 분할 기반의 검증을 통해 Qwen2.5-3B/7B/14B/32B, Phi-3.5 및 SmolLM2 모델에서 콘텐츠 위험(CD)과 물리적 위험(PD)이 서로 다른 신호로 나타나는 것을 확인했습니다. 이러한 CD/PD 분리 현상을 바탕으로, 전체 숨겨진 상태에 대한 단일 레이어의 L2 정규화된 로지스틱 탐색 모델인 PRISM을 제안합니다. PRISM은 SafeAgentBench에서 86.2~87.7%의 정확도를 달성했으며, 오탐율(FPR)은 11.7~13.7%입니다. 반면, 동일 규모의 LLM 평가 모델은 안전한 작업에 대해 24.7~39.0%의 FPR을 보였습니다. 또한 직접적인 유해 키워드가 없는 1,000개의 물리적 위험 쌍으로 구성된 비교 벤치마크인 PhysicalSafetyBench-1K (PSB-1K)를 도입하여 방법론이 명시적인 유해 표현 대신 실제 물리적 위험을 감지하는지 테스트했습니다. PSB-1K에서 PRISM은 99.6%의 정확도와 0.7%의 FPR을 달성했으며, Qwen2.5-3B 모델은 안전한 작업 중 67.8%를 거부했습니다. 또한 PRISM은 SafeText 및 EARBench에서도 동일한 성능을 보여주며, 숨겨진 상태 탐색이 텍스트 필터링을 넘어 실제 물리적 안전성을 평가하는 데 유용한 방법임을 입증합니다.

Original Abstract

Large language models (LLMs) increasingly serve as high-level planners for embodied agents, where linguistically benign instructions can become unsafe once grounded in the physical world. We study whether this physically grounded danger is the same safety problem as ordinary text-level content danger. Through hidden-state direction analysis and random-split null tests, we show that content danger (CD) and physical danger (PD) form separable signals in LLM representations across Qwen2.5-3B/7B/14B/32B, Phi-3.5 and SmolLM2. Building on the CD/PD separability, we propose PRISM, a single-layer L2-regularized logistic probe over full hidden states. PRISM achieves 86.2--87.7\% accuracy on SafeAgentBench with 11.7--13.7\% FPR, while same-scale LLM judges over-block safe tasks at 24.7--39.0\% FPR. We further introduce PhysicalSafetyBench-1K (PSB-1K), a contrastive benchmark of 1{,}000 physical-risk pairs without direct harm keywords, to test whether methods detect physically grounded danger rather than explicit unsafe wording. On PSB-1K, PRISM reaches 99.6\% accuracy and 0.7\% FPR, whereas a Qwen2.5-3B judge rejects 67.8\% of safe tasks. PRISM also replicates on SafeText and EARBench, supporting hidden-state probing as a representation-level method for physical safety beyond text moderation.

0 Citations
0 Influential
5.5 Altmetric
27.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!