2606.13079v1 Jun 11, 2026 cs.CR

대규모 언어 모델 기반 AI 시스템에서 자율적인 침투 능력의 등장

The Emergence of Autonomous Penetration Capabilities in Large Language Model-Powered AI Systems

Geng Hong
Geng Hong
Citations: 364
h-index: 8
Xu Pan
Xu Pan
Citations: 127
h-index: 6
Jiarun Dai
Jiarun Dai
Citations: 265
h-index: 8
Min Yang
Min Yang
Citations: 128
h-index: 6
Jiaqi Luo
Jiaqi Luo
Citations: 26
h-index: 1
Zhile Chen
Zhile Chen
Citations: 0
h-index: 0
Yawen Duan
Yawen Duan
Citations: 398
h-index: 4
Brian Tse
Brian Tse
Citations: 409
h-index: 4
Jia Xu
Jia Xu
Citations: 62
h-index: 3
Weibing Wang
Weibing Wang
Citations: 2
h-index: 1
Yuan Zhang
Yuan Zhang
Citations: 47
h-index: 5

최근 들어, 상당한 실제 피해를 야기할 수 있는 사이버 공격을 자율적으로 수행하는 것은 최첨단 AI 시스템이 넘어서는 안 되는 중요한 경계선으로 여겨지고 있습니다. 이러한 광범위한 경계선 내에서, 자율적인 침투는 핵심적인 기능 및 하위 작업입니다. 이는 LLM 기반 AI 시스템이 인간의 개입 없이 대상 서버에 대한 적대적 작업을 독립적으로 수행하고, 취약점을 식별하고 악용하며, 권한 없는 접근 또는 제어를 획득하는 능력을 의미합니다. 많은 연구가 AI 시스템의 자율적인 침투 능력에 대한 평가를 시도해 왔습니다. 그러나 기존의 평가는 종종 불투명한 방법론을 사용하거나, 비현실적이거나 지나치게 단순화된 침투 테스트 시나리오에 의존하며, LLM에 과도한 사전 지식과 특정 작업 지침을 제공하는 경우가 많아, 현대 AI 시스템이 보다 광범위하고 고위험 사이버 공격 시나리오 내에서 이 핵심 기능을 얼마나 자율적으로 수행할 수 있는지를 정확하게 반영하지 못합니다. 이러한 한계를 해결하기 위해, 우리는 대상 서버와 에이전트 프레임워크라는 두 가지 구성 요소로 이루어진 새로운 자율적인 침투 평가 프레임워크를 구축했습니다. 구체적으로, 대상 서버 측면에서는 알려진 취약점이 없는 보안 서비스의 개수에 따라 두 가지 수준의 환경을 설계했습니다: Tier 1 (하나의 보안 서비스) 및 Tier 2 (세 개의 보안 서비스), 총 300개의 대상 서버를 구성했습니다. 한편, 에이전트 프레임워크는 특정 대상에 대한 사전 지식이 없는 일반적인 사이버 보안 도구 세트를 갖춘 범용 에이전트 아키텍처를 채택합니다. 우리는 19개의 공개 모델 및 독점 모델을 평가한 결과, 현재 모델의 침투 성공률은 10.7%에서 69.3%에 이르는 것을 확인했습니다. 또한, 전체 모델 성능 향상과 함께 자율적인 침투 능력도 계속 개선되는 경향이 있음을 관찰했습니다.

Original Abstract

Nowadays, the autonomous execution of cyberattacks capable of causing substantial real-world harm is widely regarded as one of the critical red lines that frontier AI systems must not cross. Within this broader red-line scenario, autonomous penetration represents a core enabling capability and subtask: the ability of LLM-powered AI systems to independently conduct adversarial operations against a target server without human intervention, identify and exploit vulnerabilities, and obtain unauthorized access or control. A growing body of work has sought to assess the autonomous penetration capabilities of AI systems. However, existing evaluations often employ opaque methodologies, rely on unrealistic or overly simplified penetration-testing scenarios, or provide LLMs with excessive prior knowledge and task-specific guidance, and cannot accurately capture the extent to which modern AI systems can autonomously perform this core capability within broader high-impact cyberattack scenarios. To address these limitations, we construct a new autonomous penetration evaluation framework consisting of two components: target servers and agent scaffolding. Specifically, on the target-server side, we design two levels of target environments based on the number of secure services without known vulnerabilities deployed alongside a vulnerable service: Tier~1 (one secure service) and Tier~2 (three secure services), resulting in a total of 300 target servers. Meanwhile, the agent scaffolding adopts a general-purpose agent architecture equipped with a set of general-purpose cybersecurity tools, without any target-specific prior knowledge. We evaluate 19 open-weight and proprietary LLMs, and find that current models achieve penetration success rates ranging from 10.7% to 69.3%. Moreover, we observe that autonomous penetration capability continues to improve alongside advances in overall model capability.

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!