2607.26791v1 Jul 29, 2026 cs.CR

SecRespond: 실제 환경에서의 침해 사고 대응 AI 에이전트 성능 평가

SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response

Pengjun Xie
Pengjun Xie
Citations: 1,161
h-index: 16
Boli Chen
Boli Chen
Alibaba Group
Citations: 600
h-index: 10
Ruixue Ding
Ruixue Ding
Citations: 520
h-index: 10
Lehan Wang
Lehan Wang
Citations: 85
h-index: 6
Jinwei Huang
Jinwei Huang
Citations: 0
h-index: 0
Zhendong Liu
Zhendong Liu
Citations: 0
h-index: 0
Shuo Wang
Shuo Wang
Citations: 0
h-index: 0
Tao Lei
Tao Lei
Citations: 0
h-index: 0
Ouyang Xin
Ouyang Xin
Citations: 0
h-index: 0
Xiaomeng Li
Xiaomeng Li
Citations: 0
h-index: 0

대규모 언어 모델(LLM) 기반 에이전트는 호스트 아티팩트와 명령줄 인터페이스(CLI)에 접근하여 실제 보안 운영에서 점점 더 많이 사용되고 있으며, 따라서 이들의 보안 기능을 철저히 평가하는 것이 중요합니다. 그러나 기존의 사이버 보안 벤치마크는 주로 공격 발생 전에 에이전트를 깨끗하고 이상적인 환경에 배치하는 사전 침해 상태를 대상으로 합니다. 이러한 점은 사후 침해 상황에서의 연구가 부족하다는 것을 의미합니다. 이러한 격차를 해소하기 위해, 우리는 LLM 에이전트의 사후 침해 사고 대응 워크플로우를 평가하는 최초의 벤치마크인 SecRespond을 소개합니다. 이 벤치마크에서는 손상된 호스트의 포렌식 디스크 스냅샷과 함께 호스트 보안 제품에서 보고하는 경고, 취약점 스캔 및 기본 상태 검사를 제공하고, 에이전트는 침해 사실에 대한 포렌식 보고서, 기본 위험 및 취약점 위험을 생성하고, 이에 대한 복구 계획을 제시해야 합니다. 우리는 4가지 유형의 공격 경로, 21가지 ATT&CK 기술 및 5개의 운영 체제를 포함하는 10개의 서로 다른 손상된 클라우드 호스트로 구성된 사이버 범위를 활용하여 이 작업을 구현했습니다. OpenCode 에이전트 하니스를 사용하여 23개의 최첨단 LLM을 평가했습니다. 실험 결과, 현재 에이전트는 경고가 나타내는 문제를 안정적으로 파악할 수 있지만, 은밀한 침해를 탐지하고 포괄적이고 검증된 복구 계획을 생성하는 데 어려움을 겪으며, 어떤 모델도 단일 범위에서 완전한 탐지 및 복구를 달성하지 못했습니다. 이는 실제 사고 대응을 위한 에이전트 개발에 있어 근본적인 병목 현상을 드러냅니다. 이 벤치마크는 https://github.com/Alibaba-NLP/qqr/tree/main/data/secrespond 에서 공개적으로 사용할 수 있습니다.

Original Abstract

Large Language Model (LLM) agents are increasingly adopted in real-world security operations with access to host artifacts and command-line interfaces (CLIs), making it critical to thoroughly assess their security capabilities. However, existing cybersecurity benchmarks focus on pre-compromise settings where agents are placed in a clean and idealized environment before an attack occurs. This leaves the post-compromise setting underexplored. To address this gap, we introduce SecRespond, the first benchmark for evaluating LLM agents on the post-compromise incident-response workflow. Given a forensic disk snapshot of a compromised host together with the alerts, vulnerability scans, and baseline checks reported by a host security product, agents are required to produce forensic reports on intrusions, baseline risks, and vulnerability risks, together with a remediation plan. We instantiate this task across 10 cyber ranges, each constructed from a distinct compromised cloud host, spanning 4 entry-point types, 21 ATT&CK techniques, and 5 operating systems. We evaluate 23 frontier LLMs on the OpenCode agent harness. Experimental results show that although current agents can reliably uncover the problems exposed by alerts, they struggle to proactively investigate the disk for silent intrusions and to produce comprehensive, verified remediation plans, with no model achieving complete detection and remediation on any single range. This reveals a fundamental bottleneck in building agents for real-world incident response. The benchmark is publicly available at https://github.com/Alibaba-NLP/qqr/tree/main/data/secrespond.

0 Citations
0 Influential
47.992109794992 Altmetric
0.0 Score
Original PDF
269

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!