2602.14345v1 Feb 15, 2026 cs.CR

AXE: 능동적 익스플로잇 엔진을 활용한 제로데이 취약점 보고서 검증

AXE: An Agentic eXploit Engine for Confirming Zero-Day Vulnerability Reports

Amirali Sajadi
Amirali Sajadi
Citations: 43
h-index: 3
T. Nguyen
T. Nguyen
Citations: 92
h-index: 5
Kostadin Damevski
Kostadin Damevski
Citations: 1,608
h-index: 20
Preetha Chatterjee
Preetha Chatterjee
Citations: 491
h-index: 13

취약점 탐지 도구는 소프트웨어 프로젝트에서 널리 사용되지만, 종종 유지 관리자에게 오탐 및 실질적인 조치가 필요 없는 보고서로 인해 부담을 줍니다. 자동화된 익스플로잇 시스템은 이러한 보고서를 검증하는 데 도움이 될 수 있지만, 기존 접근 방식은 일반적으로 탐지 파이프라인과 분리되어 작동하며, 취약점 유형 및 소스 코드 위치와 같은 쉽게 얻을 수 있는 메타데이터를 활용하지 못합니다. 본 논문에서는 보고된 보안 취약점이 CWE 분류 및 취약한 코드 위치와 같은 최소한의 취약점 메타데이터를 활용하여 현실적인 그레이박스 익스플로잇 환경에서 어떻게 평가될 수 있는지 조사합니다. 우리는 다중 에이전트 프레임워크인 Agentic eXploit Engine (AXE)을 소개합니다. AXE는 웹 애플리케이션 익스플로잇을 위한 프레임워크로, 가벼운 탐지 메타데이터를 분리된 계획, 코드 탐색 및 동적 실행 피드백을 통해 구체적인 익스플로잇에 매핑합니다. CVE-Bench 데이터 세트에서 평가한 결과, AXE는 30%의 익스플로잇 성공률을 달성했으며, 이는 최첨단 블랙박스 기준보다 3배 향상된 수치입니다. 단일 에이전트 구성에서도 그레이박스 메타데이터는 1.75배의 성능 향상을 제공합니다. 체계적인 오류 분석 결과, 대부분의 실패 사례는 취약점 의미의 오해 및 충족되지 않은 실행 전제 조건과 같은 특정 추론 오류에서 비롯됩니다. 성공적인 익스플로잇의 경우, AXE는 실행 가능하고 재현 가능한 개념 증명 아티팩트를 생성하여 웹 취약점 분류 및 해결 프로세스를 간소화하는 데 유용함을 입증합니다. 또한, AXE의 일반화 가능성을 CVE-Bench에 포함되지 않은 최근의 실제 취약점 사례 연구를 통해 평가했습니다.

Original Abstract

Vulnerability detection tools are widely adopted in software projects, yet they often overwhelm maintainers with false positives and non-actionable reports. Automated exploitation systems can help validate these reports; however, existing approaches typically operate in isolation from detection pipelines, failing to leverage readily available metadata such as vulnerability type and source-code location. In this paper, we investigate how reported security vulnerabilities can be assessed in a realistic grey-box exploitation setting that leverages minimal vulnerability metadata, specifically a CWE classification and a vulnerable code location. We introduce Agentic eXploit Engine (AXE), a multi-agent framework for Web application exploitation that maps lightweight detection metadata to concrete exploits through decoupled planning, code exploration, and dynamic execution feedback. Evaluated on the CVE-Bench dataset, AXE achieves a 30% exploitation success rate, a 3x improvement over state-of-the-art black-box baselines. Even in a single-agent configuration, grey-box metadata yields a 1.75x performance gain. Systematic error analysis shows that most failed attempts arise from specific reasoning gaps, including misinterpreted vulnerability semantics and unmet execution preconditions. For successful exploits, AXE produces actionable, reproducible proof-of-concept artifacts, demonstrating its utility in streamlining Web vulnerability triage and remediation. We further evaluate AXE's generalizability through a case study on a recent real-world vulnerability not included in CVE-Bench.

0 Citations
0 Influential
10 Altmetric
50.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!