2602.19396v1 Feb 23, 2026 cs.AI

글 속에 숨기: 활성화 분리(Activation Disentanglement)를 통한 은닉된 탈옥 공격 탐지

Hiding in Plain Text: Detecting Concealed Jailbreaks via Activation Disentanglement

Amirhossein Farzam
Amirhossein Farzam
Citations: 14
h-index: 2
Majid Behabahani
Majid Behabahani
Citations: 0
h-index: 0
Mani Malek
Mani Malek
Citations: 1,274
h-index: 6
Yuriy Nevmyvaka
Yuriy Nevmyvaka
Citations: 3
h-index: 1
Guillermo Sapiro
Guillermo Sapiro
Citations: 76
h-index: 3

대규모 언어 모델(LLM)은 여전히 유창하고 의미적으로 일관된 탈옥 프롬프트에 취약하며, 따라서 표준적인 휴리스틱으로는 탐지하기 어렵습니다. 특히, 공격자가 요청의 형식을 조작하여 모델이 요청을 준수하도록 유도함으로써 악의적인 의도를 숨기려고 할 때 심각한 문제가 발생합니다. 이러한 공격은 유연한 표현 방식을 통해 악의적인 의도를 유지하므로, 구조적 특징이나 목표별 특징에 의존하는 방어 체계는 실패할 수 있습니다. 이러한 문제에 착안하여, 우리는 추론 과정에서 LLM의 활성화를 통해 의미적 요인 쌍을 분리하는 자기 지도(self-supervised) 프레임워크를 소개합니다. 우리는 이 프레임워크를 목표와 표현 방식에 적용하고, 제어된 목표 및 표현 방식의 변형을 포함하는 프롬프트 모음인 GoalFrameBench를 구축했습니다. 이 데이터셋을 사용하여, Representation Disentanglement on Activations (ReDAct) 모듈을 학습시켜 동결된 LLM에서 분리된 표현을 추출합니다. 그런 다음, 표현 방식의 표현을 기반으로 작동하는 이상 탐지기인 FrameShield를 제안합니다. FrameShield는 최소한의 계산 오버헤드로 다양한 LLM 계열에서 모델에 독립적인 탐지를 향상시킵니다. ReDAct에 대한 이론적 보장과 광범위한 실증적 검증 결과는 ReDAct의 분리가 FrameShield의 성능을 효과적으로 향상시킨다는 것을 보여줍니다. 마지막으로, 우리는 분리 과정을 해석 가능성 분석 도구로 활용하여 목표 및 표현 방식 신호에 대한 뚜렷한 특징을 밝히고, 의미적 분리 과정을 LLM 안전 및 메커니즘 해석 가능성을 위한 핵심 요소로 제시합니다.

Original Abstract

Large language models (LLMs) remain vulnerable to jailbreak prompts that are fluent and semantically coherent, and therefore difficult to detect with standard heuristics. A particularly challenging failure mode occurs when an attacker tries to hide the malicious goal of their request by manipulating its framing to induce compliance. Because these attacks maintain malicious intent through a flexible presentation, defenses that rely on structural artifacts or goal-specific signatures can fail. Motivated by this, we introduce a self-supervised framework for disentangling semantic factor pairs in LLM activations at inference. We instantiate the framework for goal and framing and construct GoalFrameBench, a corpus of prompts with controlled goal and framing variations, which we use to train Representation Disentanglement on Activations (ReDAct) module to extract disentangled representations in a frozen LLM. We then propose FrameShield, an anomaly detector operating on the framing representations, which improves model-agnostic detection across multiple LLM families with minimal computational overhead. Theoretical guarantees for ReDAct and extensive empirical validations show that its disentanglement effectively powers FrameShield. Finally, we use disentanglement as an interpretability probe, revealing distinct profiles for goal and framing signals and positioning semantic disentanglement as a building block for both LLM safety and mechanistic interpretability.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!