2607.15263v1 Jul 16, 2026 cs.CR

성공률을 넘어: 공격 및 방어 보안 에이전트의 비용 효율적인 평가

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents

Paul Kassianik
Paul Kassianik
Citations: 813
h-index: 6
Blaine Nelson
Blaine Nelson
Citations: 738
h-index: 5
Yaron Singer
Yaron Singer
Citations: 752
h-index: 5

보안 에이전트 평가는 일반적으로 풍부한 추론 자원을 활용하여 최대 공격 능력을 측정하며, 취약점 발견, 익스플로잇 개발, 침투 테스트 및 CTF 성공 여부를 강조합니다. 이러한 측정은 유용하지만 불완전합니다. 실제 보안 환경에서는 모든 추론 단계, 도구 호출, 텔레메트리 쿼리 및 풍부화 요청이 예산을 소모합니다. 본 연구에서는 공격적인 Cybench 과제와 방어적인 Splunk BOTS v1 조사 과제를 통해 비용-성공 관점에서 언어 모델 기반 보안 에이전트를 평가했습니다. 단순히 최상의 성공률을 보고하는 대신, 고정된 비용 수준에서 모델을 비교하고 추론 비용과 도구 사용량을 기준으로 성능을 분석했습니다. 연구 결과는 레드팀 및 블루팀 작업에 대해 서로 다른 확장 패턴을 보여줍니다. 공격적인 CTF 성능은 추가적인 테스트 시간 컴퓨팅 자원을 통해 향상되며, 공개 가중치 모델은 기존의 독점 시스템에 근접하면서도 비용 경쟁력을 유지할 수 있습니다. 반면, 방어적인 SOC 조사 작업은 동일한 방식으로 확장되지 않습니다. 성공 여부는 단순히 추론 예산뿐만 아니라 체계적인 도구 사용, 텔레메트리 탐색 및 선택적 풍부화에 더 큰 영향을 받습니다. 본 연구는 보안 에이전트 벤치마크가 과제 성공 외에도 경제적 효율성과 실제 적용 가능성을 측정해야 한다고 주장합니다. 비용을 고려한 SOC 환경 평가를 통해 현재 실질적으로 유용한 모델과 방어 에이전트가 개선되어야 할 부분을 명확하게 파악할 수 있습니다. 연구 결과는 다음 웹사이트에서 확인할 수 있습니다: https://evals.frontier.security

Original Abstract

Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphasizing vulnerability discovery, exploit development, penetration testing, and CTF completion. Such measurements are useful but incomplete: in operational security, every reasoning step, tool call, telemetry query, and enrichment request consumes budget. We evaluate language-model security agents through this cost-success lens on offensive Cybench challenges and defensive Splunk BOTS v1 investigation challenges. Instead of reporting only best-case success, we compare models at fixed cost levels and decompose performance by inference spend and tool spend. Our results show distinct scalingregimes for red- and blue-team tasks. Offensive CTF performance improves with additional test-time compute, and scaled open-weight models can approach frontier proprietary systems while remaining cost-competitive. Defensive SOC investigation does not scale in the same way: success depends more heavily on disciplined tool use, telemetry navigation, and selective enrichment than on raw reasoning budget alone. We argue that security-agent benchmarks should measure economic efficiency and operational fit alongside task success. Cost-aware, SOC-native evaluations provide a clearer picture of which models are practically useful today and where defensive agents still need to improve. We present an interactive website with our results https://evals.frontier.security.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!