성공률을 넘어: 공격 및 방어 보안 에이전트의 비용 효율적인 평가
Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents
보안 에이전트 평가는 일반적으로 풍부한 추론 자원을 활용하여 최대 공격 능력을 측정하며, 취약점 발견, 익스플로잇 개발, 침투 테스트 및 CTF 성공 여부를 강조합니다. 이러한 측정은 유용하지만 불완전합니다. 실제 보안 환경에서는 모든 추론 단계, 도구 호출, 텔레메트리 쿼리 및 풍부화 요청이 예산을 소모합니다. 본 연구에서는 공격적인 Cybench 과제와 방어적인 Splunk BOTS v1 조사 과제를 통해 비용-성공 관점에서 언어 모델 기반 보안 에이전트를 평가했습니다. 단순히 최상의 성공률을 보고하는 대신, 고정된 비용 수준에서 모델을 비교하고 추론 비용과 도구 사용량을 기준으로 성능을 분석했습니다. 연구 결과는 레드팀 및 블루팀 작업에 대해 서로 다른 확장 패턴을 보여줍니다. 공격적인 CTF 성능은 추가적인 테스트 시간 컴퓨팅 자원을 통해 향상되며, 공개 가중치 모델은 기존의 독점 시스템에 근접하면서도 비용 경쟁력을 유지할 수 있습니다. 반면, 방어적인 SOC 조사 작업은 동일한 방식으로 확장되지 않습니다. 성공 여부는 단순히 추론 예산뿐만 아니라 체계적인 도구 사용, 텔레메트리 탐색 및 선택적 풍부화에 더 큰 영향을 받습니다. 본 연구는 보안 에이전트 벤치마크가 과제 성공 외에도 경제적 효율성과 실제 적용 가능성을 측정해야 한다고 주장합니다. 비용을 고려한 SOC 환경 평가를 통해 현재 실질적으로 유용한 모델과 방어 에이전트가 개선되어야 할 부분을 명확하게 파악할 수 있습니다. 연구 결과는 다음 웹사이트에서 확인할 수 있습니다: https://evals.frontier.security
Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphasizing vulnerability discovery, exploit development, penetration testing, and CTF completion. Such measurements are useful but incomplete: in operational security, every reasoning step, tool call, telemetry query, and enrichment request consumes budget. We evaluate language-model security agents through this cost-success lens on offensive Cybench challenges and defensive Splunk BOTS v1 investigation challenges. Instead of reporting only best-case success, we compare models at fixed cost levels and decompose performance by inference spend and tool spend. Our results show distinct scalingregimes for red- and blue-team tasks. Offensive CTF performance improves with additional test-time compute, and scaled open-weight models can approach frontier proprietary systems while remaining cost-competitive. Defensive SOC investigation does not scale in the same way: success depends more heavily on disciplined tool use, telemetry navigation, and selective enrichment than on raw reasoning budget alone. We argue that security-agent benchmarks should measure economic efficiency and operational fit alongside task success. Cost-aware, SOC-native evaluations provide a clearer picture of which models are practically useful today and where defensive agents still need to improve. We present an interactive website with our results https://evals.frontier.security.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.