RouteGuard: LLM 에이전트의 스킬 포이즈닝 공격에 대한 내부 신호 탐지
RouteGuard: Internal-Signal Detection of Skill Poisoning in LLM Agents
에이전트 스킬은 LLM 에이전트에 대한 새로운 형태의 간접 주입 공격을 야기하며, 기존의 간접 프롬프트 주입과 달리 공격자는 악성 명령어를 정당한 명령어 소스로 기능하는 복잡하고 행동 지향적인 스킬 내부에 숨길 수 있습니다. 본 연구에서는 스킬 포이즈닝 공격을 사전에 탐지하고, 성공적인 스킬 포이즈닝이 신뢰할 수 있는 컨텍스트에서 악성 스킬 영역으로 응답 시간 주의가 이동하는 '주의력 탈취'라는 구조적인 내부 효과를 유발하여 유해한 동작을 초래한다는 것을 확인했습니다. 이러한 메커니즘에 기반하여, 본 연구에서는 응답 컨텍스트에 따른 주의력과 숨겨진 상태 정렬을 신뢰도 게이팅을 통한 후처리 융합으로 결합하는 고정 백본 탐지기인 RouteGuard를 제안합니다. 실제 및 합성 오픈 소스 스킬 벤치마크를 통해, RouteGuard는 가장 강력하거나 견고한 탐지기로 꾸준히 우수한 성능을 보였습니다. 특히, 중요한 Skill-Inject 채널에서 0.8834의 F1 값을 달성했으며, 어휘 검사를 통해 놓친 설명 공격의 90.51%를 복구하여, 스킬 포이즈닝 방어에는 텍스트 필터링만으로는 부족하며 내부 신호 탐지가 필요하다는 것을 보여줍니다.
Agent skills introduce a new and more severe form of indirect injection for LLM agents: unlike traditional indirect prompt injection, attackers can hide malicious instructions inside a dense, action-oriented skill that already functions as a legitimate instruction source. We study pre-execution skill-poison detection and show that successful skill poisoning induces a structured internal effect, attention hijacking, in which response-time attention shifts from trusted context to malicious skill spans and drives harmful behavior. Motivated by this mechanism, we propose RouteGuard, a frozen-backbone detector that combines response-conditioned attention and hidden-state alignment through reliability-gated late fusion. Across both real and synthetic open-source skill benchmarks, RouteGuard is consistently the strongest or most robust detector; on the critical Skill-Inject channel slice, it reaches 0.8834 F1 and recovers 90.51% of description attacks missed by lexical screening, showing that defending against skill poisoning requires internal-signal detection rather than text-only filtering
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.