2605.28201v1 May 27, 2026 cs.AI

식재, 유지, 트리거: 대규모 언어 모델 에이전트에 대한 잠복 공격

Plant, Persist, Trigger: Sleeper Attack on Large Language Model Agents

Fuli Feng
Fuli Feng
Citations: 1,688
h-index: 21
Dongrui Liu
Dongrui Liu
Citations: 49
h-index: 3
Wenjie Wang
Wenjie Wang
Citations: 31
h-index: 3
Fengbin Zhu
Fengbin Zhu
National University of Singapore
Citations: 1,375
h-index: 14
Yongxiang Li
Yongxiang Li
Citations: 156
h-index: 8
Moxin Li
Moxin Li
Citations: 770
h-index: 14
Z. Ma
Z. Ma
Citations: 19
h-index: 2

대규모 언어 모델(LLM) 에이전트는 외부 환경으로부터의 보안 위협에 여전히 취약합니다. 공격자는 도구에서 반환된 데이터, 웹페이지 또는 MCP 컨텍스트와 같은 외부 정보에 악성 콘텐츠를 주입하여 안전하지 않은 행동이나 부정확한 출력과 같은 유해한 에이전트 동작을 유발할 수 있습니다. 기존 연구는 주로 단일 상호 작용 공격에 초점을 맞추는데, 이는 에이전트가 악성 콘텐츠를 관찰하고 즉시 사용자 요청 내에서 유해한 행동을 보이는 경우입니다. 그러나 본 연구에서는 악성 콘텐츠가 동일한 에이전트에 의해 제공되는 여러 상호 작용 동안 지속될 수 있으며, 이로 인해 이러한 위협을 탐지하고 완화하기가 더 어렵다는 것을 보여줍니다. 구체적으로, 악성 콘텐츠는 에이전트 상태에 유지되어 여러 상호 작용 동안 잠재 상태로 남아 있다가 나중에 정상적인 사용자 쿼리에 의해 활성화될 수 있습니다. 우리는 이 유형의 보안 위협을 '잠복 공격(Sleeper Attack)'으로 정의합니다. 이를 평가하기 위해, 실제 유해 결과 6가지, 공격 전략 3가지, 그리고 에이전트 상태 대상 3가지(세션 컨텍스트, 메모리 및 재사용 가능한 기술)를 포괄하는 1,896개의 인스턴스를 포함하는 벤치마크를 구축했습니다. 7개의 강력한 오픈 소스 및 상용 LLM에 대한 실험 결과, 최첨단 LLM 에이전트가 단일 상호 작용 기준에서 낮은 공격 성공률을 달성하더라도 잠복 공격에 여전히 취약하다는 것을 확인했습니다. 본 연구의 코드와 데이터는 다음 주소에서 확인할 수 있습니다: https://anonymous.4open.science/r/skdvnfu23ihr9wdscnksf1asdffsaef.

Original Abstract

Large Language Model (LLM) agents remain vulnerable to safety threats from the external environment, where attackers inject adversarial content into external observations such as tool-returned data, webpages, or MCP context, causing harmful agentic behaviors such as unsafe actions or incorrect outputs. Existing studies typically focus on single-interaction attacks, where the agent observes adversarial content and immediately exhibits harmful behavior within one user request. However, we show that adversarial content can also persist across interactions served by the same agent, making such threats harder to detect and mitigate. Specifically, adversarial content may persist in the agent state, remain dormant across interactions, and later be activated by a benign user query. We formalize this type of safety threat as Sleeper Attack. To evaluate it, we construct a benchmark with 1,896 instances covering six real-world harmful outcomes, three attack strategies, and three agent state targets: session context, memory, and reusable skills. Experiments on seven strong open-source and closed-source LLMs show that state-of-the-art LLM agents remain vulnerable to Sleeper Attack, even when they achieve low attack success rates under a single-interaction baseline. Our code and data are available at https://anonymous.4open.science/r/skdvnfu23ihr9wdscnksf1asdffsaef.

0 Citations
0 Influential
10.5 Altmetric
52.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!