2606.09701v1 Jun 08, 2026 cs.CL

공격 및 방어 학습: GRPO를 이용한 언어 모델의 적응적 레드 팀 운영

Learning to Attack and Defend: Adaptive Red Teaming of Language Models via GRPO

Blake Bullwinkel
Blake Bullwinkel
Citations: 152
h-index: 7
Amanda J. Minnich
Amanda J. Minnich
Citations: 126
h-index: 7
M. Russinovich
M. Russinovich
Citations: 4,538
h-index: 23
Eugenia Kim
Eugenia Kim
Citations: 34
h-index: 2

AI 레드 팀 운영은 끊임없이 변화하는 공격자와 방어자에 대응해야 합니다. 강화학습은 새로운 공격 기법을 발견하는 데 유망한 접근 방식이며, 공동 훈련 방법은 보다 강력한 방어 시스템을 구축하는 데 도움이 될 수 있습니다. 최근 연구에서는 PPO 및 DPO를 사용하여 공격자-방어자 공동 훈련의 효과가 입증되었지만, GRPO는 이러한 환경에서 불안정하다는 보고가 있었습니다. 본 논문에서는 밀집 다중 채널 보상과 분리된 이점 정규화를 사용하는 공동 공격자-방어자 최적화를 위해 GRPO를 실현 가능한 방식으로 만드는 공동 훈련 프레임워크인 AdvGRPO를 소개합니다. 훈련은 단일 회전 공격에서 폐루프 다중 회전 공격으로 이어지는 커리큘럼을 통해 진행되며, 그 후 공동 훈련을 시작하여 공격자 및 방어 모델을 번갈아 가며 업데이트합니다. 우리는 제안하는 방법이 매우 효과적이고 전이 가능한 공격을 생성할 수 있으며, 공동 훈련된 방어 시스템이 안전성 평가 지표에서 기존 방식보다 우수한 성능을 보인다는 것을 보여줍니다.

Original Abstract

AI red teaming must continually adapt to evolving attackers and defenders. Reinforcement learning offers a promising approach to discovering novel attacks, and co-training methods can produce more robust defenders in tandem. Recent works have demonstrated the efficacy of attacker-defender co-training by applying PPO and DPO, but report that GRPO is unstable in this setting. We introduce AdvGRPO, a co-training framework that makes GRPO viable for joint attacker-defender optimization using dense multi-channel rewards and decoupled advantage normalization. Training progresses through a curriculum from single-turn to closed-loop multi-turn attacks before bootstrapping co-training, where attacker and defender models are updated in alternation. We show that our method can produce highly effective and transferable attacks and that co-trained defenders outperform baselines on safety benchmarks.

0 Citations
0 Influential
11.5 Altmetric
57.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!