누가 그 대가를 치르는가? 실제 웹 에이전트 시스템을 위한 이해관계자 중심의 프롬프트 주입 공격 성능 평가
Who Pays the Price? Stakeholder-Centric Prompt Injection Benchmarking for Real-world Web Agents
대규모 언어 모델(LLM)에 의해 구동되는 웹 에이전트는 점점 더 많은 실제 환경에서 사용되고 있으며, 이들은 신뢰할 수 없는 웹 콘텐츠를 통해 작동하며 직접적인 결과를 초래하는 작업을 수행합니다. 이는 프롬프트 주입 공격에 취약하게 만듭니다. 이러한 공격은 겉보기에는 무해한 콘텐츠 내에 적대적인 지침을 숨겨 에이전트의 동작을 조작합니다. 기존 보안 성능 평가는 주로 기술적 타당성에 초점을 맞춘 extit{공격 중심} 관점을 채택하여, 결과적으로 발생하는 피해의 미묘한 분포를 간과합니다. 그러나 실제로는 프롬프트 주입 위험은 피해자에게 따라 달라집니다. 단일 공격으로 인해 다양한 이해관계자에 대해 비대칭적인 결과를 초래할 수 있으며, 동일한 공격 패턴이 누구를 대상으로 하는지에 따라 효과가 크게 다를 수 있습니다. 이러한 특성을 파악하기 위해, 우리는 실제 웹 에이전트 시스템에서 피해를 체계적으로 분류하고 귀속시키는 extit{이해관계자 중심} 성능 평가 도구인 extbf{\sysname}을 소개합니다. 이 도구는 영향을 받는 개체(예: 사용자, 판매자, 플랫폼)를 구별하고, 공격을 구체적인 목표로 분해하며, 결과 및 프로세스 수준의 상호 보완적인 지표를 사용하여 각 사례를 평가합니다. 우리의 결과는 상당하고 다양한 취약점을 드러냅니다. 현재 에이전트가 안정적으로 방어할 수 있는 단일 공격 목표도 없습니다. 또한 실패 모드는 사용자의 위임된 작업에 영향을 주지 않고 성공하는 \emph{은밀한 기생}부터, 공격의 성공 없이 작업을 중단시키는 \emph{목표 불일치 중단}, 그리고 적대적인 목표와 작업 무결성이 동시에 침해되는 \emph{복합 실패}까지 다양한 질적 모드로 나타납니다. 이러한 패턴은 기존 평가 방법으로는 파악하기 어려우며, 이는 실제 환경에 LLM 기반 에이전트를 배포할 때 이해관계자를 고려한 평가의 필요성을 강조합니다. 성능 평가는 다음 위치에서 사용할 수 있습니다: https://github.com/StakeBench/SBC.
Web agents driven by large language models (LLMs) are increasingly deployed in real-world environments, where they operate over untrusted web content and execute actions with direct consequences. This makes them vulnerable to prompt-injection attacks, in which seemingly benign content embeds adversarial instructions that manipulate agent behaviour. Existing security benchmarks adopt an \textit{attack-centric} perspective, focusing on the technical feasibility of injections while overlooking the nuanced distribution of resulting harms. In practice, however, prompt-injection risk is victim-dependent: a single exploit can produce asymmetric consequences for different stakeholders, and the same attack pattern may exhibit substantially different effectiveness depending on whom it targets. To capture these properties, we introduce \textbf{\sysname}, a \textit{stakeholder-centric} benchmark to systematically categorize and attribute harm in real-world web agent systems. It distinguishes between affected entities (e.g., user, seller, platform), decomposes the attacks into concrete objectives, and evaluates each case with complementary outcome- and process-level metrics. Our results reveal substantial and heterogeneous vulnerabilities: not a single attack objective is reliably resisted by current agents, and failures distribute across qualitatively distinct modes ranging from \emph{stealthy parasitism} (attack succeeds without disrupting the user's delegated task) to \emph{misaligned disruption} (task disrupted without attack success) and \emph{compounded failure} (both adversarial objective and task integrity simultaneously violated). These patterns are missed by conventional evaluation, highlighting the need for stakeholder-aware assessment of LLM-based agents in real-world deployments. Benchmark is available at https://github.com/StakeBench/SBC.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.