공격 및 방어 학습: GRPO를 이용한 언어 모델의 적응적 레드 팀 운영
Learning to Attack and Defend: Adaptive Red Teaming of Language Models via GRPO
AI 레드 팀 운영은 끊임없이 변화하는 공격자와 방어자에 대응해야 합니다. 강화학습은 새로운 공격 기법을 발견하는 데 유망한 접근 방식이며, 공동 훈련 방법은 보다 강력한 방어 시스템을 구축하는 데 도움이 될 수 있습니다. 최근 연구에서는 PPO 및 DPO를 사용하여 공격자-방어자 공동 훈련의 효과가 입증되었지만, GRPO는 이러한 환경에서 불안정하다는 보고가 있었습니다. 본 논문에서는 밀집 다중 채널 보상과 분리된 이점 정규화를 사용하는 공동 공격자-방어자 최적화를 위해 GRPO를 실현 가능한 방식으로 만드는 공동 훈련 프레임워크인 AdvGRPO를 소개합니다. 훈련은 단일 회전 공격에서 폐루프 다중 회전 공격으로 이어지는 커리큘럼을 통해 진행되며, 그 후 공동 훈련을 시작하여 공격자 및 방어 모델을 번갈아 가며 업데이트합니다. 우리는 제안하는 방법이 매우 효과적이고 전이 가능한 공격을 생성할 수 있으며, 공동 훈련된 방어 시스템이 안전성 평가 지표에서 기존 방식보다 우수한 성능을 보인다는 것을 보여줍니다.
AI red teaming must continually adapt to evolving attackers and defenders. Reinforcement learning offers a promising approach to discovering novel attacks, and co-training methods can produce more robust defenders in tandem. Recent works have demonstrated the efficacy of attacker-defender co-training by applying PPO and DPO, but report that GRPO is unstable in this setting. We introduce AdvGRPO, a co-training framework that makes GRPO viable for joint attacker-defender optimization using dense multi-channel rewards and decoupled advantage normalization. Training progresses through a curriculum from single-turn to closed-loop multi-turn attacks before bootstrapping co-training, where attacker and defender models are updated in alternation. We show that our method can produce highly effective and transferable attacks and that co-trained defenders outperform baselines on safety benchmarks.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.