2606.10525v1 Jun 09, 2026 cs.CR

에이전트 환경에서의 자동 프롬프트 주입 공격 평가

Assessing Automated Prompt Injection Attacks in Agentic Environments

F. Tramèr
F. Tramèr
Citations: 1,509
h-index: 16
David K. Hofer
David K. Hofer
Citations: 22
h-index: 2
Edoardo Debenedetti
Edoardo Debenedetti
ETH Zürich
Citations: 2,958
h-index: 13

간접적인 프롬프트 주입은 신뢰할 수 없는 외부 데이터와 상호 작용하는 LLM 에이전트에 심각한 위협을 가하지만, 탈옥(jailbreaking)에 효과적임이 입증된 자동화된 공격 방법은 실제적인 에이전트 환경에서 아직 충분히 연구되지 않았습니다. 본 논문에서는 AgentDojo 프레임워크 내에서 LLM 에이전트를 대상으로 한 자동 프롬프트 주입 공격에 대한 종합적인 실증적 평가를 수행합니다. 화이트박스(GCG) 및 블랙박스(TAP) 방법을 모두 에이전트 환경에 적용하여 평가했습니다. 80개의 작업 쌍을 사용하여 네 가지 영역과 여러 모델에서 평가한 결과, 블랙박스 최적화 방법이 그래디언트 기반 방법에 비해 현저히 우수한 성능을 보였으며, 이는 GCG가 합리적인 계산 자원 내에서 불안정한 최적화 특성을 갖기 때문이라고 판단됩니다. 또한 TAP의 효과는 공격에 사용되는 모델에 따라 달라지며, 일반적인 능력과 안전성 튜닝 모두 공격 성공 여부에 영향을 미칩니다. 더 강력한 모델은 더욱 효과적인 주입을 생성하는 반면, 안전성에 맞춰 조정된 공격자는 적대적 프롬프트를 생성하기를 거부할 수 있습니다. 작업에 국한되지 않는 공격은 새로운 작업 및 데이터 분포 영역으로 효과적으로 전이되지만, 작은 오픈 소스 모델에서 최적화된 공격은 GPT-5와 같은 최첨단 모델로는 잘 전이되지 않습니다. 이러한 결과는 자동 프롬프트 주입이 모델에 따라 달라지는 신뢰할 수 있는 위협이지만, 모델에 독립적인 악용에는 여전히 상당한 장벽이 존재함을 보여줍니다.

Original Abstract

Indirect prompt injection poses a critical threat to LLM agents that interact with untrusted external data, yet automated attack methods--proven effective for jailbreaking--remain underexplored in realistic agentic settings. We present a comprehensive empirical evaluation of automated prompt injection attacks against LLM agents, adapting both white-box (GCG) and black-box (TAP) methods to the agentic setting within the AgentDojo framework. We evaluate across 80 task pairs spanning four domains and multiple models, and find that black-box optimization substantially outperforms gradient-based methods, a gap we attribute to GCG's optimization instability under reasonable compute budgets. We also find that TAP's effectiveness depends on the attacker model, as both general capability and safety tuning affect attack success--stronger models produce more effective injections, while safety-tuned attackers can refuse to generate adversarial prompts. Task-universal attacks transfer effectively to unseen tasks and out-of-distribution domains, but attacks optimized on smaller open-source models do not transfer to frontier models like GPT-5. These findings highlight automated prompt injection as a credible but model-dependent threat, with significant barriers remaining for model-agnostic exploitation.

0 Citations
0 Influential
8 Altmetric
40.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!