2605.29963v1 May 28, 2026 cs.CR

Honeyval: LLM 기반 HTTP 허니팟에 대한 종합적인 평가 프레임워크

Honeyval: A Comprehensive Evaluation Framework for LLM-powered HTTP Honeypots

Martin T. Vechev
Martin T. Vechev
Citations: 17,053
h-index: 67
Ilia Shumailov
Ilia Shumailov
Google DeepMind
Citations: 11,014
h-index: 28
Mark Vero
Mark Vero
Citations: 634
h-index: 12
Fabian Kaczmarczyck
Fabian Kaczmarczyck
Citations: 51
h-index: 3
Ivo Petrov
Ivo Petrov
Citations: 449
h-index: 7
Jamie Hayes
Jamie Hayes
Citations: 3,911
h-index: 13
Niels Heinen
Niels Heinen
Citations: 75
h-index: 1
Tianqi Fan
Tianqi Fan
Citations: 155
h-index: 2
Luca Invernizzi
Luca Invernizzi
Citations: 33
h-index: 3

허니팟은 실제 시스템 구성 요소를 모방하여 사이버 공격으로부터 방어하는 데 사용되는 미끼 시스템입니다. 최근에는 LLM이 허니팟의 시뮬레이션 백본으로 점점 더 많이 활용되고 있으며, 이를 통해 방어자들은 낮은 시스템 보안 위험을 가진 고상호작용 허니팟을 구축할 수 있습니다. 그러나 LLM 기반 허니팟 개발에는 통일된 평가 프레임워크가 부족합니다. 대부분의 평가는 고정된 명령에 대한 응답 유사성을 측정하거나, 수동 테스트 또는 실제 배포를 통해 이루어집니다. 이러한 방법은 개발 시 확장성이 떨어지거나, 평가 간 재현성이 낮고, 실제 공격을 대표하지 않거나, 다양한 공격자와 허니팟 구성에 적응하기 어렵습니다. 본 연구에서는 이러한 격차를 해소하고 LLM 기반 HTTP 허니팟을 위한 종합적인 평가 프레임워크인 Honeyval을 제안합니다. 우리는 기존 평가의 한계를 극복하기 위해 16개의 백엔드 애플리케이션을 기반으로 허니팟을 구축하고, AI 공격 에이전트를 공격자로 사용하며, 에이전트 및 허니팟의 기능을 다양한 사용자 정의 설정에서 모니터링하기 위한 두 가지 제어 작업을 활용하고, 공격자를 위한 명확하고 검증 가능한 악용 목표를 정의합니다. Honeyval을 사용하여 최근의 비용 효율적인 LLM을 HTTP 허니팟으로 평가하는 광범위한 실험을 수행했습니다. 우리의 실험 결과는 LLM 기반 허니팟이 규칙 기반 기본 허니팟보다 훨씬 더 긴 상호 작용 시간을 제공하며, 최첨단 모델에 의해 탐지될 가능성이 훨씬 적다는 것을 보여줍니다. 또한 평균적으로 에이전트 공격자에 비해 운영 비용의 이점을 유지합니다. 또한 다양한 반격 허니팟 구성을 실험하여, 더 긴 상호 작용 시간이 증가된 탐지 위험과 같은 고유한 트레이드오프를 관찰했습니다.

Original Abstract

Honeypots are decoy systems mimicking real system components designed to defend against cyber attacks. Recently, LLMs increasingly serve as simulation backbones for honeypots. They enable defenders to construct high-interaction honeypots with low system security risks. However, LLM-powered honeypot development lacks a unified evaluation framework. Most evaluations consist of measuring response similarity on fixed commands, manual testing, or real-world deployment. These methods are often not scalable for development, reproducible across evaluations, representative of practical attacks, or adaptable to various attacker and honeypot configurations. In this work, we bridge this gap and propose Honeyval, a comprehensive evaluation framework for LLM-powered HTTP honeypots. We address the limitations of prior evaluations by grounding the honeypots in 16 backend applications, using AI hacking agents as attackers, employing two control tasks to monitor agent and honeypot capabilities across customizations, and defining clear and verifiable exploit goals for the attacker. Using Honeyval, we conduct an extensive evaluation of recent cost-efficient LLMs as HTTP honeypots. Our experiments highlight the promise of LLM-powered honeypots; they lead to substantially longer interactions with the attacker than rule-based baseline honeypots and are far less frequently detected even by frontier models, all while, on average, preserving a running cost advantage against agentic attackers. Further, we experiment with different counter-offensive honeypots configurations, and observe unique trade-offs, such as longer interactions at the cost of increased detection.

0 Citations
0 Influential
30 Altmetric
150.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!