GPT-Red: 대규모 자가 학습을 통한 자동적 공격 시뮬레이션
GPT-Red: Automated Red Teaming via Self-Play at Scale
본 논문에서는 **GPT-Red**라는 자동화된 공격 시뮬레이션 에이전트를 소개합니다. GPT-Red는 최첨단 LLM에 대한 새로운 프롬프트 주입 공격을 탐지하도록 훈련되었습니다. 이 모델의 목표는 우리 생산 시스템의 안정성을 평가하고 개선하는 것입니다. 이를 위해, 우리는 GPT-Red를 사용하여 현재까지 가장 강력한 프롬프트 주입 방어 능력을 갖춘 모델인 GPT-5.6을 적대적으로 학습시켰습니다. GPT-Red를 구축하기 위해, 우리는 모델이 동시에 훈련되는 다양한 수비 에이전트 그룹을 공격하도록 하는 확장 가능한 자가 학습 알고리즘을 설계했습니다. 우리는 일부의 가장 큰 강화 학습 후처리 훈련과 동일한 규모의 컴퓨팅 리소스를 사용하여 실제 공격 시뮬레이션 환경에서 모델을 훈련시켰으며, 이는 사상 최대 규모의 LLM 안전 훈련 사례입니다. GPT-Red는 공격 시뮬레이션 분야에서 뛰어난 성능을 보입니다. 이 모델은 이전 버전의 모델(GPT-5.5까지)을 안정적으로 공격할 수 있으며, 인간 전문가보다 더 성공적인 공격을 발견하고, 또한 새로운 환경, 수비 모델 및 시스템으로 일반화됩니다. 앞으로 우리는 각 GPT 모델의 안정성이 향상됨에 따라, 이것이 더욱 강력한 공격 시뮬레이션 에이전트에 대한 더 나은 학습 신호를 제공하여 자체 개선 효과를 창출할 것이라고 예상합니다.
We introduce \textbf{GPT-Red}, an automated red-teaming agent that is trained to discover novel prompt injection attacks against frontier LLMs. The goal of this model is to evaluate and improve the robustness of our production systems. To this end, we use it to adversarially train GPT-5.6, our most robust model to prompt injections to date. To create GPT-Red, we design a scalable self-play algorithm where the model is tasked with attacking a diverse population of simultaneously-trained defender agents. We train the model on realistic red-teaming environments using compute on the same scale as some of our largest RL post-training runs, making it the single-largest LLM safety training run ever documented. GPT-Red excels at red-teaming: it reliably breaks our past models up to GPT-5.5, it finds more successful attacks than human red-teamers, and it generalizes to held-out environments, defender models, and harnesses. In the future, we expect that as we improve the robustness of each new GPT model, it will in turn will provide better learning signal for \textit{even stronger} red-teamer agents, thus unlocking a self-improvement flywheel.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.