AI 보안 리더보드: 방법론, 결과 및 최소 기준
AI Security Leaderboard: Methodology, Results and Minimal Standard
최첨단 AI 모델 개발자들은 파괴적인 오용을 방지하기 위해 다층적인 안전장치에 점점 더 의존하고 있지만, 이러한 안전장치가 실제로 얼마나 효과적이며, 개발자들 간에 일관성을 보이는지에 대한 공개적인 증거는 거의 존재하지 않습니다. 본 연구에서는 FAR.AI 최소 안전 기준(Version 1.0)을 소개합니다. 이는 67가지의 쉽게 접근 가능한 정적 탈옥 기법 분류, 이러한 기법들을 매우 광범위한 공격 공간으로 구성하는 방법, 그리고 대표 모델들을 대상으로 일부 샘플을 활용한 벤치마크를 포함합니다. 우리는 Claude Fable 5, GPT-5.6 Sol, Gemini 3.1 Pro, 및 Grok 4.5 모델을 두 가지 상호 보완적인 데이터 세트(총 360개의 공격 목표)를 사용하여 평가했습니다. 이 데이터 세트는 화학, 생물학, 방사선/핵 및 폭발 (CBRNE) 위협과 사이버 공격을 포괄합니다. 우리는 세 단계의 필터링을 통해 해당 도메인의 75% 이상의 목표에 대해 운영적으로 적합한 응답을 유도하는 보편적인 탈옥 기법(단일 프롬프트 템플릿)을 식별했습니다. 또한, 공격자가 실제로 지출하는 비용을 직접 모델링하는 '탈옥 비용' 지표를 도입했으며, 보편적인 탈옥 기법이 발견되지 않은 경우 하한값을 설정했습니다. 분석 결과, 모델의 안정성은 매우 불균등하며, 이러한 모델을 탈옥시키는 데 드는 비용은 최대 100배까지 차이가 납니다. 우리의 기술 풀에 대한 무작위 검색을 통해 Grok 4.5 모델에서 63개의 보편적인 탈옥 기법이 발견되었고, Gemini 3.1 Pro 모델에서는 18개가 발견되었습니다. 각각의 탈옥 기법을 찾는 데 평균적으로 약 $58 및 $278의 비용이 소요되었습니다. 전문가의 지도를 받아 조합한 경우 이 수치는 각각 385개와 231개로 증가했습니다. Claude Fable 5 및 GPT-5.6 Sol 모델에서는 어떤 전략을 사용하더라도 보편적인 탈옥 기법이 발견되지 않았습니다. 최소 안전 기준을 충족하려면 이미 공개적으로 설명되고 다른 곳에서 실제로 사용되는 방어 메커니즘만 필요하기 때문에, 이러한 격차는 현재 기술로 해결 가능할 것으로 판단됩니다. 우리는 추론, 활성화 및 입력/출력 모니터링을 결합한 다층적인 방어를 권장합니다. 분석 결과는 leaderboard.far.ai 에서 확인할 수 있습니다.
Frontier AI model developers increasingly rely on layered safeguards to prevent catastrophic misuse, but little public evidence exists on how much protection these safeguards provide, or how consistently across developers. We introduce the FAR.AI Minimal Standard for Safeguards, Version 1.0: a taxonomy of 67 readily accessible static jailbreak techniques, a method for composing them into a very large attack space, and a benchmark of flagship models against a sample of it. We evaluate Claude Fable 5, GPT-5.6 Sol, Gemini 3.1 Pro, and Grok 4.5 on two complementary datasets totalling 360 attacker goals spanning chemical, biological, radiological/nuclear and explosive (CBRNE) threats and offensive cyber, using a three-stage funnel to identify universal jailbreaks: single prompt templates that elicit operationally compliant responses on over 75% of a domain's goals. We also introduce a cost-to-jailbreak metric that models attacker spend directly, with right-censored lower bounds where no universal jailbreak was found. Robustness is highly uneven: the cost to break these models varies over a hundredfold. Random search over our technique pool found 63 universal jailbreaks against Grok 4.5 and 18 against Gemini 3.1 Pro, at an average cost of roughly $58 and $278 per jailbreak found; expert-guided composition raised these to 385 and 231. Neither Claude Fable 5 nor GPT-5.6 Sol yielded any universal jailbreak under either strategy. Because meeting the Minimal Standard requires only defenses already publicly described and deployed in production elsewhere, these gaps appear closable with current techniques. We recommend defense-in-depth combining reasoning, activation, and input/output monitoring. Results are maintained at leaderboard.far.ai.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.