DiagChain: 증거 기반 공격 체인 재구성 시 LLM 에이전트 평가를 위한 진단 벤치마크
DiagChain: A Diagnostic Benchmark for Evaluating LLM Agents on Evidence-Grounded Attack Chain Reconstruction
대규모 언어 모델(LLM) 에이전트는 다양한 텔레메트리 데이터를 검색하고 해석하여 공격자의 행동 순서를 추론함으로써 공격 체인 재구성에 유망한 접근 방식을 제공합니다. 그러나 기존 벤치마크는 주로 최종 결과 또는 전체 정확도를 평가하며, 중간 추론 단계에서 오류가 발생하는 방식과 전파되는 과정에 대한 제한적인 통찰력을 제공합니다. 본 논문에서는 LLM 에이전트의 단계별 평가를 가능하게 하는 증거 기반 공격 체인 재구성을 위한 진단 벤치마크인 DiagChain을 제시합니다. DiagChain은 다양한 운영 체제, 증거 노이즈 수준 및 체인 길이를 포괄하는 69개의 시나리오 세트인 MAIN-69를 포함합니다. 또한, 본 논문에서는 검색된 증거와 재구성된 체인의 진화적인 구조적 표현을 결합하는 Evidence-Centric Retrieval-Augmented Generation (ECRAG) 방식을 소개합니다. 재구성 과정의 다양한 단계를 평가하고 체계적인 오류 분석을 지원하기 위해 다섯 가지 보완적인 지표를 제시합니다. 6개의 LLM을 사용하여 수행한 평가 결과, DiagChain은 가장 강력한 구성조차도 MAIN-69에 포함된 849개의 기준 단계 중 39.6%에서만 성공하는 것으로 나타났습니다. 분석 결과, 작은 모델은 검색된 증거를 출력에 통합하는 기본적인 작업에서 어려움을 겪는 반면, 더 큰 모델은 증거의 올바른 순서를 결정하는 것이 주요 병목 현상이 되는 후속 단계로 진행할 수 있는 것으로 밝혀졌습니다. 이러한 결과는 최종 정확도 외에도 진단 평가의 중요성을 입증하고, 증거 기반 사이버 보안 에이전트를 개선하기 위한 실질적인 통찰력을 제공합니다.
Large Language Model (LLM) agents offer a promising approach to attack chain reconstruction by retrieving and interpreting heterogeneous telemetry to infer ordered attacker actions. However, existing benchmarks mainly evaluate final outputs or aggregate accuracy, providing limited insight into how errors arise and propagate across intermediate reasoning stages. We present DiagChain, a diagnostic benchmark for evidence-grounded attack chain reconstruction that enables stage-wise evaluation of LLM agents. DiagChain includes MAIN-69, a suite of 69 scenarios spanning multiple operating systems, evidence noise levels, and chain lengths. It further introduces Evidence-Centric Retrieval-Augmented Generation (ECRAG), which couples evidence retrieval with an evolving structured representation of the reconstructed chain. Five complementary metrics are introduced to assess distinct stages of the reconstruction process and support systematic failure diagnosis. Based on evaluations using 6 LLMs, DiagChain reveals that even the strongest configuration succeeds on only 39.6% of the 849 reference steps in MAIN-69. Our analysis further shows that smaller models struggle with the more basic task of incorporating retrieved evidence into their outputs, whereas larger models can proceed to later steps, where correctly ordering that evidence becomes the main bottleneck. These results validate the importance of diagnostic evaluation beyond end-to-end accuracy and provide actionable insights for improving evidence-grounded cybersecurity agents.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.