식재, 유지, 트리거: 대규모 언어 모델 에이전트에 대한 잠복 공격
Plant, Persist, Trigger: Sleeper Attack on Large Language Model Agents
대규모 언어 모델(LLM) 에이전트는 외부 환경으로부터의 보안 위협에 여전히 취약합니다. 공격자는 도구에서 반환된 데이터, 웹페이지 또는 MCP 컨텍스트와 같은 외부 정보에 악성 콘텐츠를 주입하여 안전하지 않은 행동이나 부정확한 출력과 같은 유해한 에이전트 동작을 유발할 수 있습니다. 기존 연구는 주로 단일 상호 작용 공격에 초점을 맞추는데, 이는 에이전트가 악성 콘텐츠를 관찰하고 즉시 사용자 요청 내에서 유해한 행동을 보이는 경우입니다. 그러나 본 연구에서는 악성 콘텐츠가 동일한 에이전트에 의해 제공되는 여러 상호 작용 동안 지속될 수 있으며, 이로 인해 이러한 위협을 탐지하고 완화하기가 더 어렵다는 것을 보여줍니다. 구체적으로, 악성 콘텐츠는 에이전트 상태에 유지되어 여러 상호 작용 동안 잠재 상태로 남아 있다가 나중에 정상적인 사용자 쿼리에 의해 활성화될 수 있습니다. 우리는 이 유형의 보안 위협을 '잠복 공격(Sleeper Attack)'으로 정의합니다. 이를 평가하기 위해, 실제 유해 결과 6가지, 공격 전략 3가지, 그리고 에이전트 상태 대상 3가지(세션 컨텍스트, 메모리 및 재사용 가능한 기술)를 포괄하는 1,896개의 인스턴스를 포함하는 벤치마크를 구축했습니다. 7개의 강력한 오픈 소스 및 상용 LLM에 대한 실험 결과, 최첨단 LLM 에이전트가 단일 상호 작용 기준에서 낮은 공격 성공률을 달성하더라도 잠복 공격에 여전히 취약하다는 것을 확인했습니다. 본 연구의 코드와 데이터는 다음 주소에서 확인할 수 있습니다: https://anonymous.4open.science/r/skdvnfu23ihr9wdscnksf1asdffsaef.
Large Language Model (LLM) agents remain vulnerable to safety threats from the external environment, where attackers inject adversarial content into external observations such as tool-returned data, webpages, or MCP context, causing harmful agentic behaviors such as unsafe actions or incorrect outputs. Existing studies typically focus on single-interaction attacks, where the agent observes adversarial content and immediately exhibits harmful behavior within one user request. However, we show that adversarial content can also persist across interactions served by the same agent, making such threats harder to detect and mitigate. Specifically, adversarial content may persist in the agent state, remain dormant across interactions, and later be activated by a benign user query. We formalize this type of safety threat as Sleeper Attack. To evaluate it, we construct a benchmark with 1,896 instances covering six real-world harmful outcomes, three attack strategies, and three agent state targets: session context, memory, and reusable skills. Experiments on seven strong open-source and closed-source LLMs show that state-of-the-art LLM agents remain vulnerable to Sleeper Attack, even when they achieve low attack success rates under a single-interaction baseline. Our code and data are available at https://anonymous.4open.science/r/skdvnfu23ihr9wdscnksf1asdffsaef.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.