2607.18063v1 Jul 20, 2026 cs.CR

적응형 공격: LLM 에이전트 보안을 위한 다단계, 다중 LLM 벤치마크

Adaptive Adversaries: A Multi-Turn, Multi-LLM Benchmark for LLM Agent Security

Devina Jain
Devina Jain
Citations: 39
h-index: 2
Chuan Li
Chuan Li
Citations: 18
h-index: 1
David Hartmann
David Hartmann
Citations: 4
h-index: 1

LLM 기반 에이전트는 외부 콘텐츠를 처리하면서 프롬프트 주입 및 다단계 조작에 노출됩니다. 대부분의 안전성 벤치마크는 평가 전에 수집된 고정된 공격 패턴 또는 단일/다단계 상호 작용을 대상으로 방어자를 평가합니다. 본 논문에서는 메모리가 없는 LLM 방어자에 대한 extit{적응형 다중 라운드 공격}을 위한 21가지 시나리오로 구성된 벤치마크를 제시합니다. 자율적인 LLM 공격자가 이전 방어자 응답을 관찰하고 라운드마다 전략을 변경하며, 각 방어자 응답은 독립적인 상호 작용으로 평가됩니다. 21가지 시나리오, 공격자, 방어자 및 구조화된 출력 점수를 고정하고 첫 번째 공격자 턴의 점수만 고려하면 공격 성공률(ASR)은 0~1%에 불과합니다. 반면, 15라운드의 적응형 공격을 허용하면 ASR이 5.4~14.0%로 증가합니다. 최첨단 LLM 공격자 3개를 결합하면 가장 뛰어난 단일 공격자가 생성하는 고유한 성공 공격보다 1.4~2.2배 많은 공격을 발견할 수 있으며, 생성된 공격은 기존 벤치마크의 공격과 낮은 코사인 유사성(0.02~0.14)을 보입니다. Claude Opus 4.6과 GPT-5.4는 전체적으로 동일한 ASR(각각 5.4%; 95% 신뢰 구간 중첩)을 보이지만, 그 취약점은 크게 다릅니다. 특정 시나리오에서 Opus는 60%의 ASR(95% 신뢰 구간 36~80%)을 달성하는 반면, GPT-5.4와 Gemini는 각각 7%(신뢰 구간 1~30%)로 유지되며, 이 차이는 더 많은 반복 실험에서도 유지됩니다. 21가지 시나리오 중 13가지에서 최소한 하나의 방어자 쌍을 구별할 수 있지만, 시나리오에 따라 순위가 달라집니다(Kendall's W = 0.19). 본 논문에서는 벤치마크(21가지 평가 시나리오, 10가지 공개 개발 시나리오, 오케스트레이터, 기본 테스트 코드, 다중 공격자 CLI)와 함께, 3x3 최첨단 LLM 매트릭스에서 생성된 945개의 트랜스크립트, 공격 재현 데이터셋 및 오픈 경쟁의 최종 점수 라운드에서 추출한 18,422개의 gpt-oss-20b 전투 기록을 공개합니다.

Original Abstract

LLM-based agents process external content, exposing them to prompt injection and multi-turn manipulation. Most safety benchmarks evaluate defenders against fixed attack pools collected before evaluation, single-turn or multi-turn. We present a 21-scenario benchmark for \emph{adaptive multi-round attacks against memoryless LLM defenders}: an autonomous LLM attacker observes prior defender responses and pivots across rounds, while each defender response is evaluated as a fresh interaction. Holding the 21 scenarios, attackers, defenders, and structured-output scoring fixed, restricting scoring to the first attacker turn yields $0$-$1\%$ attack success rate (ASR); allowing 15 rounds of adaptive attack yields $5.4$-$14.0\%$. Pooling three frontier attacker LLMs uncovers $1.4$-$2.2\times$ as many unique successful attacks as the best single attacker, and the generated attacks have low cosine similarity ($0.02$-$0.14$) to attacks in existing benchmarks. Claude Opus 4.6 and GPT-5.4 are tied in aggregate ($5.4\%$ each; overlapping $95\%$ CIs), but their weaknesses differ sharply: on one scenario Opus reaches $60\%$ ASR ($95\%$ CI $36$--$80\%$) while GPT-5.4 and Gemini each stay at $7\%$ (CI $1$-$30\%$; the gap is preserved in a higher-$N$ replication). $13$ of $21$ scenarios distinguish at least one defender pair, yet rankings disagree across scenarios (Kendall's $W = 0.19$). We release the benchmark -- 21 evaluation scenarios, 10 public development scenarios, the orchestrator, baseline harnesses, and a multi-attacker CLI -- plus 945 transcripts from the 3$\times$3 frontier matrix, an attack-replay dataset, and 18{,}422 gpt-oss-20b battles from an open competition's final scoring rounds.

0 Citations
0 Influential
1 Altmetric
5.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!