2608.04317v1 Aug 05, 2026 cs.CR

Trident: 심층 강화 학습 기반 사이버 방어 시스템을 무너뜨리는 방법 (에이전트 기반)

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic)

Mahdi Imani
Mahdi Imani
Citations: 107
h-index: 6
Ryozo Masukawa
Ryozo Masukawa
Citations: 83
h-index: 5
Sanggeon Yun
Sanggeon Yun
Citations: 237
h-index: 9
Hyunwoo Oh
Hyunwoo Oh
Citations: 22
h-index: 3
Mohsen Imani
Mohsen Imani
Citations: 12
h-index: 2
Sungheon Jeong
Sungheon Jeong
Citations: 87
h-index: 6
Nathaniel D. Bastian
Nathaniel D. Bastian
Citations: 2,360
h-index: 24
Ian Bryant
Ian Bryant
Citations: 10
h-index: 2
Armita Kazeminajafabadi
Armita Kazeminajafabadi
Citations: 87
h-index: 4

심층 강화 학습(DRL)을 기반으로 하는 자율적인 사이버 방어 시스템은 많은 연구 관심을 받고 있지만, 현재까지는 정적이고 휴리스틱한 공격 에이전트에 대한 평가만 이루어져 왔으며, 적응형 위협에 대한 견고성은 거의 연구되지 않았습니다. 한편, 검증 가능한 보상을 이용한 강화 학습(RLVR)의 최근 발전은 LLM 추론 능력을 향상시켰지만, 적절한 벤치마크 환경 및 상호 작용 데이터셋의 부재로 인해 사이버 보안 분야에 통합되기는 어렵습니다. 이러한 격차를 해소하기 위해, 우리는 세 가지 구성 요소로 이루어진 에이전트 기반 LLM 레드 팀 프레임워크인 Trident를 소개합니다. 첫째, CybORG CAGE 4 및 CyberWheel을 포괄하는 동적 벤치마크 환경과 격리된 샌드박스 서버가 있습니다. 둘째, RLVR 학습을 위한 13,000개 이상의 고품질 레드-블루 상호 작용 트레일 데이터셋이 있습니다. 셋째, ``코드-정책(Code-as-Policy)`` RLVR 에이전트 아키텍처인 Trident Agentic가 있습니다. 후자는 학습 가능한 플래너를 통해 레드 에이전트 훈련을 컨텍스추얼 밴딧 문제로 재구성하며, 여기에서 학습 가능한 플래너는 압축된 실행 로그로부터 완전한 공격 전략을 생성하고, 고정된 코더는 이를 실행 가능한 Python 정책으로 변환하여 실제 DRL 방어 시스템에 배포합니다. 실험 결과, 기존 방어 시스템의 근본적인 취약성이 드러났습니다. 단일 7B 플래너를 사용하여 Trident는 평균적으로 정적 레드 에이전트 기준보다 블루 에이전트의 방어 성능을 522% 감소시켰으며, 동시에 속임수 회피 및 적응형 상태 우선순위 지정과 같은 새로운 행동 패턴을 발견했습니다. 이러한 행동은 기존의 정적 휴리스틱으로는 전혀 파악할 수 없습니다.

Original Abstract

Autonomous cyber defense systems based on Deep Reinforcement Learning (DRL) have attracted significant research attention, yet remain evaluated almost exclusively against static, heuristic red agents, leaving their robustness against adaptive threats critically understudied. Meanwhile, recent advances in Reinforcement Learning with Verifiable Rewards (RLVR) have improved LLM reasoning, but their integration into cybersecurity remains elusive due to the absence of suitable benchmark environments and interaction datasets. To bridge this gap, we introduce Trident, an agentic LLM red teaming framework comprising three components: a dynamic benchmark with isolated sandbox servers spanning CybORG CAGE 4 and CyberWheel, a dataset comprises over 13,000 high-fidelity red-blue interaction trajectories for RLVR, and a ``Code-as-Policy'' RLVR agentic architecture Trident Agentic). The latter reformulates red agent training as a contextual bandit via a tripartite Log Summarizer--Planner--Coder design, where a trainable Planner generates complete attack strategies from compressed execution logs, which a frozen Coder translates into executable Python policies deployed against live DRL defenders. Empirical evaluations reveal a fundamental brittleness in existing defenses: with a single trainable 7B planner, Trident reduces blue agent defensive performance by an average of 522% compared to static red agent baselines while autonomously discovering emergent behaviors such as decoy avoidance and adaptive state prioritization that static heuristics entirely fail to uncover.

0 Citations
0 Influential
12 Altmetric
60.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!