2607.27030v1 Jul 29, 2026 cs.CR

HoF-벤치: 최첨단 모델 없이 실제 AI가 발견한 CVE 재발견

HoF-Bench: Rediscovering Real AI-Discovered CVEs Without Frontier Models

P. Simecek
P. Simecek
Citations: 0
h-index: 0
Elnaz Babayeva
Elnaz Babayeva
Citations: 6
h-index: 1
Jiří Balhar
Jiří Balhar
Citations: 1
h-index: 1
Michal Bida
Michal Bida
Citations: 0
h-index: 0
Michal Buran
Michal Buran
Citations: 0
h-index: 0
Vaclav Cadek
Vaclav Cadek
Citations: 19
h-index: 1
Luigino Camastra
Luigino Camastra
Citations: 0
h-index: 0
Tomás Dulka
Tomás Dulka
Citations: 2
h-index: 1
Michal Janocko
Michal Janocko
Citations: 0
h-index: 0
Tomas Klohna
Tomas Klohna
Citations: 0
h-index: 0
Pavel Kohout
Pavel Kohout
Citations: 0
h-index: 0
O. Kokeš
O. Kokeš
Citations: 152
h-index: 3
Adam Krivka
Adam Krivka
Citations: 0
h-index: 0
Jakub Kubik
Jakub Kubik
Citations: 0
h-index: 0
Patrik Mada
Patrik Mada
Citations: 0
h-index: 0
Igor Morgenstern
Igor Morgenstern
Citations: 0
h-index: 0
Marek Pavelka
Marek Pavelka
Citations: 0
h-index: 0
Joshua Rogers
Joshua Rogers
Citations: 0
h-index: 0
Petr Stastny
Petr Stastny
Citations: 0
h-index: 0
Jan Tattermusch
Jan Tattermusch
Citations: 0
h-index: 0
Dmitrijs Trizna
Dmitrijs Trizna
Citations: 96
h-index: 4
Martin Votruba
Martin Votruba
Citations: 0
h-index: 0
Guido Vranken
Guido Vranken
Citations: 0
h-index: 0
Jakub Zikl
Jakub Zikl
Citations: 0
h-index: 0
Evelina Gabasova
Evelina Gabasova
Citations: 340
h-index: 5
Stanislav Fort
Stanislav Fort
Citations: 3,279
h-index: 4

LLM 기반 분석 도구들은 성숙한 오픈 소스 프로젝트에서 실제 취약점을 찾아내기 시작했습니다. AISLE의 분석 도구는 OpenSSL, curl, GnuTLS를 포함하여 78개의 프로젝트에 걸쳐 280개 이상의 CVE를 발견하는 데 기여했습니다. 우리는 AISLE의 공개 '명예의 전당(Hall of Fame)'에서 이름을 따온 HoF-벤치라는 벤치마크를 소개합니다. 이 벤치마크는 취약한 커밋으로 고정된 8개의 저장소에 있는 95개의 공개적으로 발견된 AI 기반 CVE로 구성되어 있습니다. 분석 도구는 소스 코드와 대상 파일을 입력으로 받지만, CVE 식별자, 설명, 수정 사항 또는 예상되는 작동 방식은 제공하지 않습니다. '맹점' 처리된 최첨단 모델을 사용한 심사관은 동일한 코드 경로, 근본 원인, 공격 조건 및 영향을 정확하게 파악하는 결과만 인정합니다. 의도적으로 최소화된 LLM 기반 분석 도구는 엄격한 프로토콜 하에서 95개의 CVE 중 최대 65개(68%)를 재발견했습니다. 이 연구의 모든 단계에서 최첨단 모델은 취약점을 탐지하지 못했습니다. 사용된 10개의 감지 엔진은 다섯 개의 공개 모델(총 매개변수 210억 ~ 2840억 개, 활성 매개변수 30억 ~ 130억 개)과 다섯 개의 독점적인 소형 또는 '플래시' 티어 모델입니다. 모든 모델은 고정된 환경에서 실행되며, 네 번 반복하고, 선택적으로 생성된 컨텍스트 단계를 사용하며, 재현 가능한 다단계 심사 단계를 거칩니다(총 7,600개의 모델-CVE 쌍에 대한 실행 기록). 난이도는 언어에 따라 크게 달라지며, 모든 모델에서 탐지되지 않은 CVE는 주로 C 기반 인프라 코드와 관련되어 있습니다. HoF-벤치는 취약점 스캐너의 성능, 반복 실행에서의 신뢰성 및 생성되는 후보 물량 비교를 위한 간결한 테스트 환경을 제공합니다. 데이터셋은 https://huggingface.co/datasets/aisleinc/HoF-Bench 에서 이용 가능합니다.

Original Abstract

LLM-based analyzers have begun finding real vulnerabilities in mature open-source projects: AISLE's analyzer is credited with more than 280 CVEs across 78 projects, including OpenSSL, curl, and GnuTLS. We introduce HoF-Bench (named after AISLE's public Hall of Fame), a benchmark built from 95 of these public AI-discovered CVEs across eight repositories pinned at vulnerable commits. Analyzers receive source and target-file scope but not CVE identifiers, descriptions, fixes, or expected mechanisms; a detector-blinded frontier-model judge credits only findings that identify the same code path, root cause, attack condition, and impact. A deliberately minimal LLM-based analyzer rediscovers up to 65 of the 95 CVEs (68%) under this strict protocol. No frontier model performs detection anywhere in the study. The ten detector backbones are five open-weight models (21B--284B total parameters, 3--13B active) and five proprietary small or "flash"-tier models. All of them run in the fixed scaffold with four repeated passes, an optional generated-context stage, and a replayable multi-round triage stage (7,600 model--CVE pass records). Difficulty is strongly structured by language; the CVEs missed by every model concentrate in C infrastructure code. HoF-Bench provides a compact test bed for comparing vulnerability scanners, their reliability across repeated runs, and the candidate volume they create. The dataset is available at https://huggingface.co/datasets/aisleinc/HoF-Bench.

0 Citations
0 Influential
22.5 Altmetric
112.5 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!