2607.26998v1 Jul 29, 2026 cs.CR

AgentSnare: 자율적인 침투 에이전트의 지연, 회피 및 무력화를 위한 학습

AgentSnare: Learning to Delay, Divert, and Defuse Autonomous Penetration Agents

Tianhang Zheng
Tianhang Zheng
Citations: 45
h-index: 4
Mengnan Zhao
Mengnan Zhao
Citations: 52
h-index: 4
Ruoyu Wang
Ruoyu Wang
Citations: 0
h-index: 0
Heng Zhao
Heng Zhao
Citations: 3
h-index: 1
Renjie Wu
Renjie Wu
Citations: 0
h-index: 0
Zhixuan Chu
Zhixuan Chu
Citations: 0
h-index: 0
Wanyu Lin
Wanyu Lin
Citations: 48
h-index: 4

대규모 언어 모델(LLM) 기반 에이전트는 관찰-행동 루프를 통해 침투 테스트를 자동화하며, 도구에서 반환된 결과를 바탕으로 행동을 선택합니다. 이러한 의존성은 방어자가 공격 에이전트의 의사 결정 과정을 오도할 수 있는 속임수 정보를 주입할 수 있도록 합니다. 그러나 기존의 방어 기법은 공격 전에 환경에 미리 배치된 정적이고 고립된 요소에 크게 의존합니다. 고급 에이전트는 이러한 요소를 점진적으로 인식하고 우회하여 결국 실제 목표에 대한 공격 시도를 재개할 수 있습니다. 이 문제를 해결하기 위해, 우리는 AgentSnare라는 경로 적응형 속임수 시스템을 소개합니다. AgentSnare는 동적으로 가짜 환경을 펼쳐 침투 에이전트를 지속적으로 실제 목표에서 벗어나도록 유도합니다. 구체적으로, AgentSnare는 에이전트의 상호 작용 기록 및 가짜 환경 상태에 따라 후보 요소를 생성하는 요소 구성 정책 모델을 사용합니다. AgentSnare는 이러한 후보를 검증하고 사실적으로 일관된 가짜 환경에 유효한 요소를 점진적으로 통합하여 공격 도구 호출을 흡수하고, 가짜 환경 내에서의 침투 경로를 변경하며, 가짜 증거에 기반한 완료 보고서를 유도함으로써 공격을 지연 및 무력화합니다. 15개의 CVE-Bench 웹 애플리케이션과 세 가지 공격 모델을 대상으로 실험한 결과, AgentSnare는 가짜 환경에서 에이전트의 도구 호출의 46.8%를 흡수하고, 침투 후 행동의 55.9%를 해당 환경에 유지하며, 완료 시도의 90.0%가 가짜 증거에 기반합니다. 모든 45개의 공격자-CVE 쌍에서, 'pass@3' 기준으로 실제 목표가 성공적으로 악용된 사례는 없었습니다.

Original Abstract

Large language model (LLM) agents automate penetration testing through an observation-action loop, selecting actions based on observations returned by tools. This dependence allows defenders to inject deceptive observations that can mislead the agent's decision-making process. However, existing defenses rely heavily on static, isolated artifacts planted in the environment prior to an attack. Advanced agents can progressively recognize and bypass these artifacts, ultimately refocusing their exploitation attempts on the real target. To address this issue, we introduce AgentSnare, a trajectory-adaptive deception system that dynamically unfolds a decoy environment to continually steer the penetration agent away from the real target. Specifically, AgentSnare employs an artifact-construction policy model that constructs candidate artifacts conditioned on the agent's interaction history and decoy state. AgentSnare then validates these candidates and incrementally incorporates valid artifacts into a factually consistent decoy environment, thereby delaying the attack by absorbing its tool calls, diverting its post-entry trajectory within the decoy, and defusing it by inducing completion reports grounded in decoy evidence. Across 15 CVE-Bench web applications and three attacker models, AgentSnare absorbs 46.8% of the agent's tool calls in the decoy and retains 55.9% of post-entry actions there, while 90.0% of completion attempts are grounded in decoy evidence; across all 45 attacker-CVE pairs, no real target is successfully exploited at pass@3.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!