적응형 공격: LLM 에이전트 보안을 위한 다단계, 다중 LLM 벤치마크
Adaptive Adversaries: A Multi-Turn, Multi-LLM Benchmark for LLM Agent Security
LLM 기반 에이전트는 외부 콘텐츠를 처리하면서 프롬프트 주입 및 다단계 조작에 노출됩니다. 대부분의 안전성 벤치마크는 평가 전에 수집된 고정된 공격 패턴 또는 단일/다단계 상호 작용을 대상으로 방어자를 평가합니다. 본 논문에서는 메모리가 없는 LLM 방어자에 대한 extit{적응형 다중 라운드 공격}을 위한 21가지 시나리오로 구성된 벤치마크를 제시합니다. 자율적인 LLM 공격자가 이전 방어자 응답을 관찰하고 라운드마다 전략을 변경하며, 각 방어자 응답은 독립적인 상호 작용으로 평가됩니다. 21가지 시나리오, 공격자, 방어자 및 구조화된 출력 점수를 고정하고 첫 번째 공격자 턴의 점수만 고려하면 공격 성공률(ASR)은 0~1%에 불과합니다. 반면, 15라운드의 적응형 공격을 허용하면 ASR이 5.4~14.0%로 증가합니다. 최첨단 LLM 공격자 3개를 결합하면 가장 뛰어난 단일 공격자가 생성하는 고유한 성공 공격보다 1.4~2.2배 많은 공격을 발견할 수 있으며, 생성된 공격은 기존 벤치마크의 공격과 낮은 코사인 유사성(0.02~0.14)을 보입니다. Claude Opus 4.6과 GPT-5.4는 전체적으로 동일한 ASR(각각 5.4%; 95% 신뢰 구간 중첩)을 보이지만, 그 취약점은 크게 다릅니다. 특정 시나리오에서 Opus는 60%의 ASR(95% 신뢰 구간 36~80%)을 달성하는 반면, GPT-5.4와 Gemini는 각각 7%(신뢰 구간 1~30%)로 유지되며, 이 차이는 더 많은 반복 실험에서도 유지됩니다. 21가지 시나리오 중 13가지에서 최소한 하나의 방어자 쌍을 구별할 수 있지만, 시나리오에 따라 순위가 달라집니다(Kendall's W = 0.19). 본 논문에서는 벤치마크(21가지 평가 시나리오, 10가지 공개 개발 시나리오, 오케스트레이터, 기본 테스트 코드, 다중 공격자 CLI)와 함께, 3x3 최첨단 LLM 매트릭스에서 생성된 945개의 트랜스크립트, 공격 재현 데이터셋 및 오픈 경쟁의 최종 점수 라운드에서 추출한 18,422개의 gpt-oss-20b 전투 기록을 공개합니다.
LLM-based agents process external content, exposing them to prompt injection and multi-turn manipulation. Most safety benchmarks evaluate defenders against fixed attack pools collected before evaluation, single-turn or multi-turn. We present a 21-scenario benchmark for \emph{adaptive multi-round attacks against memoryless LLM defenders}: an autonomous LLM attacker observes prior defender responses and pivots across rounds, while each defender response is evaluated as a fresh interaction. Holding the 21 scenarios, attackers, defenders, and structured-output scoring fixed, restricting scoring to the first attacker turn yields $0$-$1\%$ attack success rate (ASR); allowing 15 rounds of adaptive attack yields $5.4$-$14.0\%$. Pooling three frontier attacker LLMs uncovers $1.4$-$2.2\times$ as many unique successful attacks as the best single attacker, and the generated attacks have low cosine similarity ($0.02$-$0.14$) to attacks in existing benchmarks. Claude Opus 4.6 and GPT-5.4 are tied in aggregate ($5.4\%$ each; overlapping $95\%$ CIs), but their weaknesses differ sharply: on one scenario Opus reaches $60\%$ ASR ($95\%$ CI $36$--$80\%$) while GPT-5.4 and Gemini each stay at $7\%$ (CI $1$-$30\%$; the gap is preserved in a higher-$N$ replication). $13$ of $21$ scenarios distinguish at least one defender pair, yet rankings disagree across scenarios (Kendall's $W = 0.19$). We release the benchmark -- 21 evaluation scenarios, 10 public development scenarios, the orchestrator, baseline harnesses, and a multi-attacker CLI -- plus 945 transcripts from the 3$\times$3 frontier matrix, an attack-replay dataset, and 18{,}422 gpt-oss-20b battles from an open competition's final scoring rounds.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.