단어는 안전하지만 행동은 위험할 때: 숨겨진 상태 위험 공간에서 텍스트 안전성 너머의 물리적 위험에 대한 탐구
When Words Are Safe But Actions Kill: Probing Physical Danger Beyond Text Safety in Hidden-State Risk Space
대규모 언어 모델(LLM)이 점점 더 많은 역할을 수행하며, 특히 구체적인 환경에서의 에이전트 제어에 사용되면서, 언어적으로는 무해해 보이는 명령조차도 실제 세계에서는 위험한 결과를 초래할 수 있습니다. 본 연구에서는 이러한 물리적 환경에서의 위험이 일반적인 텍스트 수준의 유해 콘텐츠 위험과 동일한 문제인지 조사합니다. 숨겨진 상태 방향 분석 및 랜덤 분할 기반의 검증을 통해 Qwen2.5-3B/7B/14B/32B, Phi-3.5 및 SmolLM2 모델에서 콘텐츠 위험(CD)과 물리적 위험(PD)이 서로 다른 신호로 나타나는 것을 확인했습니다. 이러한 CD/PD 분리 현상을 바탕으로, 전체 숨겨진 상태에 대한 단일 레이어의 L2 정규화된 로지스틱 탐색 모델인 PRISM을 제안합니다. PRISM은 SafeAgentBench에서 86.2~87.7%의 정확도를 달성했으며, 오탐율(FPR)은 11.7~13.7%입니다. 반면, 동일 규모의 LLM 평가 모델은 안전한 작업에 대해 24.7~39.0%의 FPR을 보였습니다. 또한 직접적인 유해 키워드가 없는 1,000개의 물리적 위험 쌍으로 구성된 비교 벤치마크인 PhysicalSafetyBench-1K (PSB-1K)를 도입하여 방법론이 명시적인 유해 표현 대신 실제 물리적 위험을 감지하는지 테스트했습니다. PSB-1K에서 PRISM은 99.6%의 정확도와 0.7%의 FPR을 달성했으며, Qwen2.5-3B 모델은 안전한 작업 중 67.8%를 거부했습니다. 또한 PRISM은 SafeText 및 EARBench에서도 동일한 성능을 보여주며, 숨겨진 상태 탐색이 텍스트 필터링을 넘어 실제 물리적 안전성을 평가하는 데 유용한 방법임을 입증합니다.
Large language models (LLMs) increasingly serve as high-level planners for embodied agents, where linguistically benign instructions can become unsafe once grounded in the physical world. We study whether this physically grounded danger is the same safety problem as ordinary text-level content danger. Through hidden-state direction analysis and random-split null tests, we show that content danger (CD) and physical danger (PD) form separable signals in LLM representations across Qwen2.5-3B/7B/14B/32B, Phi-3.5 and SmolLM2. Building on the CD/PD separability, we propose PRISM, a single-layer L2-regularized logistic probe over full hidden states. PRISM achieves 86.2--87.7\% accuracy on SafeAgentBench with 11.7--13.7\% FPR, while same-scale LLM judges over-block safe tasks at 24.7--39.0\% FPR. We further introduce PhysicalSafetyBench-1K (PSB-1K), a contrastive benchmark of 1{,}000 physical-risk pairs without direct harm keywords, to test whether methods detect physically grounded danger rather than explicit unsafe wording. On PSB-1K, PRISM reaches 99.6\% accuracy and 0.7\% FPR, whereas a Qwen2.5-3B judge rejects 67.8\% of safe tasks. PRISM also replicates on SafeText and EARBench, supporting hidden-state probing as a representation-level method for physical safety beyond text moderation.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.