해커인가, 환상인가? LLM 기반 자동 침투 테스트에 대한 종합 분석
Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing
대규모 언어 모델(LLM)의 빠른 발전은 자동 침투 테스트(AutoPT)에 새로운 기회를 제공했으며, 엔드 투 엔드 자율 공격을 목표로 하는 수많은 프레임워크가 등장했습니다. 그러나 관련 연구가 활발하게 진행되고 있음에도 불구하고, 기존 연구는 일반적으로 체계적인 아키텍처 분석과 통일된 벤치마크 하에서의 대규모 실증 비교가 부족합니다. 따라서 본 논문에서는 LLM 기반 AutoPT 프레임워크의 아키텍처 설계 및 종합적인 실증 평가에 초점을 맞춘 최초의 시스템 지식(SoK)을 제시합니다. 시스템 지식 수준에서, 우리는 에이전트 아키텍처, 에이전트 계획, 에이전트 메모리, 에이전트 실행, 외부 지식, 벤치마크의 여섯 가지 차원에서 기존 프레임워크 설계를 종합적으로 검토합니다. 실증 수준에서, 우리는 통일된 벤치마크를 사용하여 13개의 대표적인 오픈 소스 AutoPT 프레임워크와 2개의 기준 프레임워크에 대한 대규모 실험을 수행했습니다. 이 실험에는 총 100억 개 이상의 토큰이 사용되었으며, 1,500개 이상의 실행 로그가 생성되었습니다. 이러한 로그는 사이버 보안 분야의 전문 지식을 갖춘 15명 이상의 연구자 패널이 4개월 동안 수동으로 검토하고 분석했습니다. 본 연구는 이 빠르게 발전하는 분야의 최신 동향을 조사하고, 기존 LLM 기반 AutoPT 프레임워크에 대한 이해를 돕는 체계적인 분류 체계와 대규모 실증 벤치마크를 제공하며, 향후 연구를 위한 유망한 방향을 제시합니다.
The rapid advancement of Large Language Models (LLMs) has created new opportunities for Automated Penetration Testing (AutoPT), spawning numerous frameworks aimed at achieving end-to-end autonomous attacks. However, despite the proliferation of related studies, existing research generally lacks systematic architectural analysis and large-scale empirical comparisons under a unified benchmark. Therefore, this paper presents the first Systematization of Knowledge (SoK) focusing on the architectural design and comprehensive empirical evaluation of current LLM-based AutoPT frameworks. At systematization level, we comprehensively review existing framework designs across six dimensions: agent architecture, agent plan, agent memory, agent execution, external knowledge, and benchmarks. At empirical level, we conduct large-scale experiments on 13 representative open-source AutoPT frameworks and 2 baseline frameworks utilizing a unified benchmark. The experiments consumed over 10 billion tokens in total and generated more than 1,500 execution logs, which were manually reviewed and analyzed over four months by a panel of more than 15 researchers with expertise in cybersecurity. By investigating the latest progress in this rapidly developing field, we provide researchers with a structured taxonomy to understand existing LLM-based AutoPT frameworks and a large-scale empirical benchmark, along with promising directions for future research.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.