2607.11698v1 Jul 13, 2026 cs.CR

에이전트 해킹 에이전트: 프로덕션 에이전트 레드 팀링을 위한 자동 연구

Agent Hacks Agent: Autoresearch for Production-Agent Red-Teaming

Cong Wang
Cong Wang
Citations: 126
h-index: 4
Xutao Mao
Xutao Mao
Citations: 83
h-index: 2
Xiang Zheng
Xiang Zheng
City University of Hong Kong
Citations: 495
h-index: 6

Claude Code 및 Codex와 같은 프로덕션 LLM 에이전트는 신뢰할 수 없는 콘텐츠, 파일, 명령 및 작업 공간 상태에서 작동하므로 안전 관련 오류가 직접적으로 발생 가능합니다. 따라서 레드 팀 활동은 진화하는 모델과 도구에 발맞춰 이루어져야 합니다. 기존 접근 방식은 주로 공격 성공률을 최적화하고 벤치마크, 페이로드 또는 공격 프로그램과 같은 아티팩트를 보존하는데 초점을 맞추고 있습니다. 이러한 아티팩트는 공격이 성공한 위치를 기록하지만, 안전하지 않은 에이전트 행동의 원인이 되는 조건을 기록하지는 않습니다. 본 연구에서는 하나의 에이전트 기반 연구 환경을 사용하여 다른 프로덕션 LLM 에이전트에 대한 재사용 가능한 취약점 지식을 발견하는 자동화된 레드 팀 활동을 수행합니다. 우리는 AHA라는 검증 가능한 발견 루프를 제시합니다. 이 루프는 취약점 가설을 제안하고, 반증자를 구성하며, 유효한 공격을 실행하고, 안전한 격리 환경에서 실행 결과를 분석하고, 확인된 내용을 취약점 개념 그래프(VCG)에 통합합니다. 각 개념은 공격 대상 인터페이스와 안전하지 않은 작동 경로를 연결하며, 주张, 조건, 반证자, 전이 예측 및 증거를 포함합니다. Claude Code 및 Codex를 사용하여 직접 및 간접 공격을 포함하는 세 가지 시나리오에서 발견된 개념들은 모델과 에이전트 전체에 걸쳐 재사용 가능한 취약점 코어를 드러냅니다. 고정된 VCG는 추가 검색 없이 가장 강력한 기존 발견 방법보다 동일한 단일 샷 프로토콜에서 14.2% 더 높은 성능을 보이며, 다양한 시나리오와 공격 채널로 전이됩니다. 결과적으로 생성된 VCG는 프로덕션 안전 팀이 취약점을 검사하고, 패치를 확인하며, 재사용 가능한 안전 지식을 축적하는 데 사용될 수 있는 감사 가능한 자료를 제공합니다. 저희 코드는 https://github.com/henrymao2004/Auto-research-red-teaming-in-sleep 에서 확인할 수 있습니다.

Original Abstract

Production LLM agents such as Claude Code and Codex operate over untrusted content, files, commands, and workspace state, making safety failures directly actionable. Red-teaming must therefore keep pace with evolving models and tools. Existing approaches mainly optimize attack success and preserve artifacts such as benchmarks, payloads, or attack programs, which record where attacks succeed but not the enabling conditions behind unsafe agent behavior. We study automated red-teaming for production LLM agents using one agentic research environment to discover reusable vulnerability knowledge about another. We present AHA, a falsifiable discovery loop that proposes a vulnerability hypothesis, constructs a falsifier, instantiates a valid attack, executes it in a sandboxed harness, reflects on the trajectory, and promotes confirmed findings into a Vulnerability Concept Graph (VCG). Each concept links an attacker-facing surface to an unsafe trajectory through a claim, enabling condition, falsifier, transfer prediction, and supporting evidence. Across Claude Code and Codex on three scenarios covering direct and indirect attacks, the discovered concepts reveal a reusable vulnerability core across models and agents. A frozen VCG requires no further search and outperforms the strongest frozen discovery baseline by 14.2 percentage points under the same single-shot protocol, while transferring across scenarios and attack channels. The resulting VCG provides an auditable artifact for production safety teams to inspect vulnerabilities, validate patches, and accumulate reusable safety knowledge. Our code is available at https://github.com/henrymao2004/Auto-research-red-teaming-in-sleep.

0 Citations
0 Influential
23 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!