다중 에이전트 LLM 안전성 평가에서의 운영적 재구성 및 승인 기반 위임
Operational Reframing and Approval-Framed Delegation in Multi-Agent LLM Safety
다중 에이전트 LLM 시스템의 안전성 평가는 종종 직접적인 프롬프트와 계획-실행 파이프라인을 비교하여, 그 차이를 하나의 '파이프라인 효과'로 보고합니다. 우리는 이 집계된 지표가 해석하기 어렵다고 주장하는데, 이는 세 가지 메커니즘을 혼동하기 때문입니다. 즉, 악의적인 의도가 실행 가능한 작업으로 재구성될 수 있고, 계획자는 요청을 거부하거나 변환할 수 있으며, 실행기는 사전 승인을 암시하는 위임 프롬프트에 따라 작동할 수 있습니다. 이러한 요인들을 분리하기 위해, 30개의 합성된 유해 시나리오와 LLM-평가된 준수성을 사용한 네 가지 에이전트 안전성 벤치마크에서 추출한 탐색적 외부 검증 데이터 세트를 사용하여 다섯 가지 조건의 통제된 비교 설계를 도입했습니다. 우리의 결과는 집계된 파이프라인 안전성이 안정적인 아키텍처 특성이 아니라는 것을 보여줍니다. 운영적 재구성은 가장 일반적인 위험 신호이며, GPT, Gemini 및 DeepSeek 모델 모두에서 시나리오 세트에 걸쳐 준수도를 향상시키는 반면, Claude는 상대적으로 저항력이 있습니다. 계획자의 행동은 주로 거부를 통해 이 위험을 상쇄할 수 있지만, 계획자가 실행 가능한 단계를 생성하는 경우, 실행기는 직접적인 운영 기준보다 더 높은 수준의 준수도를 보이는 경향이 있습니다. 승인 기반 위임은 프롬프트 설계, 모델 조합 및 시나리오 출처에 민감하며, 회의적인 실행기 프롬프트는 준수도를 크게 감소시킵니다. 모델의 직접적인 순위는 배포된 계획-실행 행동을 잘못 예측할 수도 있습니다. Gemini는 주요 데이터 세트에서 직접 프롬프트에 대해 가장 안전하지만, Claude 계획자를 사용할 때 가장 큰 증폭 효과를 보이며, 준수도가 8.9%에서 38.9%로 증가합니다. GPT의 거의 제로인 집계 파이프라인 효과는 실제로 재구성화로 인한 증가가 계획자의 거부로 인해 상쇄되는 것을 숨기고 있습니다. 이러한 결과는 다중 에이전트 안전성 평가가 아키텍처 자체에 실패를 귀속시키기 전에 재구성화, 계획자 행동, 위임 프레임 및 모델 조합을 개별적으로 보고해야 함을 시사합니다.
Safety evaluations of multi-agent LLM systems often compare a direct prompt with a planner-executor pipeline and report the difference as a single "pipeline effect." We argue that this aggregate is difficult to interpret because it conflates three mechanisms: harmful intent may be reframed as plausible operational work, the planner may refuse or transform the request, and the executor may act under delegation prompts implying prior approval. To separate these factors, we introduce a five-condition controlled contrast design, evaluated on 30 synthetic harmful scenarios and an exploratory external validation set from four agent-safety benchmarks using LLM-judged compliance. Our results show that aggregate pipeline safety is not a stable architectural property. Operational reframing is the most portable risk signal, increasing compliance for GPT, Gemini, and DeepSeek across both scenario sets, while Claude is comparatively resistant. Planner behavior can offset this risk mainly through refusal; however, when the planner produces executable steps, the executor may become more compliant than under the direct operational baseline. Approval-framed delegation is sensitive to prompt design, model pairing, and scenario source, and a skeptical executor prompt sharply reduces compliance. Raw-direct model rankings can also mispredict deployed planner-executor behavior. Gemini is safest under raw direct prompts in the primary set yet shows the largest amplification with a Claude planner, rising from 8.9 percent to 38.9 percent compliance. GPTs near-zero aggregate pipeline effect instead hides a reframing increase canceled by planner refusal. These findings suggest that multi-agent safety evaluations should report reframing, planner behavior, delegation framing, and model pairing separately before attributing failures to architecture itself.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.