2607.28685v1 Jul 30, 2026 cs.AI

안전인가, 단순한 능력인가? 에이전트 안전성 평가 기준의 타당성 검토

Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Dingyan Shang
Dingyan Shang
Citations: 0
h-index: 0
Xiao Han
Xiao Han
Citations: 3
h-index: 1
Yuan Tang
Yuan Tang
Citations: 23
h-index: 2

에이전트 안전성 평가 기준은 다양한 행동을 측정하며, 그 점수는 종종 에이전트의 안전성을 나타내는 지표로 혼용됩니다. 본 연구에서는 R-Judge, InjecAgent, AgentHarm, AgentDojo 네 가지 평가 기준을 검증 가능한 측정 도구로 간주하고, 각 평가 기준의 공식 구현 방식과 개발자가 제공하는 점수 계산 방식을 사용하여 최대 22개의 모델을 평가했습니다. MMLU와 GPQA는 저희가 정의한 단일 프로토콜에 따라 능력 지표로서 함께 측정했습니다. 여기서 가장 중요한 문제는 평가 지표 자체입니다. $F_1$ 값을 사용하는 이진 추론 기반 평가에서, '항상 긍정'인 정책은 항상 $F_1 = 2π/(1+π)$ 값을 얻습니다. R-Judge에서는 이 값이 0.690으로 나타나며, 이는 실제 차이를 보이는 21개 모델 중 5개를 능가하는 값입니다. 세 가지 광범위한 커버리지를 가진 평가 기준은 동일한 18개 모델을 서로 다른 순서로 평가하며, 이러한 불일치의 근본 원인은 작은 규모의 데이터셋에서 발생하는 현상입니다. R-Judge의 특이성과 AgentHarm의 안전성은 각각 $n=7$에서 -0.64, $n=18$에서 +0.02의 상관관계를 보이며, 무작위로 추출한 크기 7의 부분 집합 중 약 절반에서 이 값 주변으로 |ρ| ≥ 0.5를 보이는 경우도 있습니다. 검증된 타당성은 어떤 결과를 선택하느냐에 따라 달라집니다. 능력은 작업 성공을 예측하지만(ρ = +0.60), 안전성과의 부정적인 상관관계가 나타납니다(ρ = -0.44, n=21). 쌍으로 묶인 20개 모델의 경우, 이러한 차이는 Δ = -1.00 (95% CI [-1.48, -0.49], p<0.001)로 나타나며, 이는 데이터 하나를 제외하거나 조직별 클러스터링을 통해 재귀적으로 분석해도 유지됩니다. 모델 수를 41개로 늘린 경우, 안전성과의 부정적인 상관관계는 -0.16 (95% CI [-0.54, +0.22])으로 약화되고, 탈옥(jailbreak)에 대한 강도는 +0.34으로 증가하지만, 이러한 변화는 통계적으로 유의미하지 않습니다. AgentHarm는 가장 강력한 연관성을 보이는 평가 기준이며, 능력 지표를 조정한 후 세 가지 템플릿을 사용한 탈옥 안전성과 0.72의 상관관계를 보입니다. 하지만 두 평가 기준 모두 유해한 행동에 대한 준수 여부를 측정하므로, 이는 일반적인 안전성을 나타내는 증거라기보다는 수렴 타당성(convergent validity)을 보여주는 증거입니다. 안전성에 대한 주장을 할 때에는 평가 기준의 이름, 측정 지표, 목표 행동, 모델 패널 정보를 명확히 제시해야 합니다.

Original Abstract

Agent-safety benchmarks measure different behaviors, and their scores get quoted interchangeably as an agent's safety. We treat four of them (R-Judge, InjecAgent, AgentHarm, AgentDojo) as measurements to be validated, running each under its official implementation and author-provided scorer on up to 22 models, with MMLU and GPQA measured by us under one protocol as a capability composite. The metric is the first problem. On any binary trace-judgment benchmark scored by $F_1$, an ``always positive'' policy attains $F_1 = 2π/(1+π)$; on R-Judge that is $0.690$, above five of the 21 models that actually discriminate. The three broad-coverage benchmarks then rank the same 18 models differently, and the trade-off behind that disagreement is a small-panel artifact: R-Judge specificity against AgentHarm safety correlates $-0.64$ at $n{=}7$ and $+0.02$ at $n{=}18$, and a quarter of random size-7 subsets reach $|ρ| \geq 0.5$ around that near-zero value. Held-out validity turns on which outcome you pick. Capability predicts task success ($ρ{=}{+}0.60$) but correlates negatively with misalignment safety ($ρ{=}{-}0.44$, $n{=}21$). On their paired $n{=}20$ panel, the corresponding contrast is $Δ{=}{-}1.00$ (95% CI $[-1.48, -0.49]$, $p<0.001$), and it survives leave-one-organization-out and organization-clustered bootstrap analyses. On an expanded 41-model panel, the misalignment correlation weakens to $-0.16$ (95% CI $[-0.54, +0.22]$) and jailbreak strengthens to $+0.34$, though neither change is significant. \mbox{AgentHarm} shows the strongest held-out association, $ρ{=}{+}0.72$ with three-template jailbreak safety after controlling capability. But both instruments score harmful compliance, so this is evidence of convergent validity rather than general safety. Naming the benchmark, metric, target behavior, and model panel is the minimum a safety claim needs.

0 Citations
0 Influential
1 Altmetric
5.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!