LLM 에이전트의 안전성, 다중 라운드 레드 팀 공격, 탈옥 방지 벤치마크, 적대적 강건성, 안전 관련 시스템
LLM agent safety, multi-turn red-teaming, jailbreak benchmarks, adversarial robustness, safety-critical systems
최근 대규모 언어 모델(LLM) 에이전트는 안전 관련 시스템의 감독 구성 요소로 점점 더 많이 제안되고 있지만, 지속적인 적응형 공격에 대한 그들의 강건성은 아직 제대로 규명되지 않았습니다. 본 논문에서는 LLM 에이전트가 안전 관련 시스템 운영자로 작동하는 것을 시뮬레이션하여 다중 라운드 레드 팀 공격을 평가하기 위한 벤치마크인 NRT-Bench를 제시합니다. 여기에는 여섯 가지 핵심 안전 기능(CSF)에 의해 관리되는 플랜트를 운영하는 다섯 명의 역할 기반 운영자 팀이 포함되며, 각 운영자는 구성 가능한 LLM으로 지원됩니다. 공격자는 네 개의 채널을 통해 제한된 다중 라운드 세션에서 메시지를 주입하며, 매 턴마다 피드백을 제공합니다. 여기서 '위험'은 LLM이 판단한 텍스트가 아닌 객관적인 신호이며, 플랜트의 어떤 CSF라도 손실되는 즉시 해당 실행은 종료됩니다. 고정된 공격 및 페어링 리플레이 프로토콜 하에서 네 가지 최첨단 운영자 모델을 평가한 결과, 적응형 다중 라운드 공격이 운영자 팀의 안전 한계를 지속적으로 벗어나게 한다는 것을 확인했습니다. 네 가지 모델 모두에서 8.7%에서 12.1%의 공격 세션이 플랜트가 핵심 안전 기능을 상실하는 것으로 끝났습니다. 네 가지 모델은 전체 성공률 측면에서 거의 동일한 수준의 강건성을 보이는 것처럼 보이지만, 실패 사례는 거의 겹치지 않습니다. 즉, 149개의 세션 중 어느 것도 모든 네 가지 모델을 동시에 무력화하지 못했지만, 세 번째 모델은 적어도 하나의 모델을 무력화했습니다. 따라서 각 모델의 취약점은 서로 분리되어 있으며, 계층적으로 나타나지 않습니다. 추가된 방어 메커니즘의 효과는 모델에 따라 크게 달라집니다. 즉, 특정 가드레일 스택이나 안전 자문 에이전트가 한 모델의 공격 성공률을 낮출 수 있지만, 다른 모델에서는 오히려 상승시킬 수 있습니다. 본 논문에서는 LLM 에이전트의 재현 가능한 안전 평가를 위해 시뮬레이션 환경, 공격 데이터 세트 및 리플레이 도구를 공개합니다.
Large language model (LLM) agents are increasingly proposed as supervisory components for safety-critical systems, yet their robustness under sustained, adaptive adversarial pressure remains poorly characterized. We present NRT-Bench, a benchmark for multi-turn red-teaming of LLM agents acting as operators of a safety-critical system, instantiated in a simulated nuclear power plant control room. A five-role operator team, each backed by a configurable LLM, runs a plant governed by six critical safety functions (CSFs), while adversaries inject messages over four channels in bounded multi-turn sessions with per-turn feedback. Harm is an objective signal rather than LLM-judged text: a run terminates the moment any CSF is lost, attributed to the causing message. Evaluating four frontier operator models under a fixed-attack paired-replay protocol, we find that adaptive multi-turn attacks reliably push the operator team past a safety limit: across the four models, between 8.7% and 12.1% of attack sessions end with the plant losing a critical safety function. Although the four models look almost equally robust by this aggregate rate, their failures barely overlap: of $149$ sessions, none defeat all four models while a third defeat at least one, so vulnerabilities are nearly disjoint across models rather than nested. The effect of added defences is strongly model-dependent: the same guardrail stack or safety-advisor agent that lowers attack success for one model can raise it for another. We release the simulation venue, attack dataset, and replay tooling for reproducible safety evaluation of LLM agents.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.