모델의 행동 지오메트리를 활용한 탈주 공격 취약성 예측 및 완화
Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models
생성 시스템의 탈주 공격에 대한 취약성을 평가하고 완화하는 것은 안전한 배포를 위해 매우 중요합니다. 그러나 배포 가능한 시스템이 많기 때문에, 모든 구성 요소에 대해 완전한 평가와 최적화를 수행하는 것은 비현실적입니다. 본 논문에서는 이전에 평가되고 보호된 모델을 활용하여 모델 집단의 행동 지오메트리를 형식화함으로써, 효율적인 취약성 예측과 모델 집단 전체에 걸친 효과적인 방어 전송을 지원합니다. 저희는 이 프레임워크를 24개 공급업체의 79개 모델과 단일 기본 모델의 100가지 시스템 구성에 적용했습니다. 행동 지오메트리를 사용하는 간단한 방법은 완전 평가와 비교하여 약 98% 더 적은 테스트(probe)로 취약성 감지에 대해 AUPRC 점수 0.94를 달성합니다. 최적화된 방어를 전송할 모델을 선택하는 데 행동 지오메트리를 사용하면 동일 공급업체 할당보다 성능이 우수합니다 (+2%, p = 0.03), 추가적인 테스트 비용은 발생하지 않으며, 세 개의 모델만으로 전체 집단을 커버할 수 있습니다. 결과는 하이퍼파라미터 선택 및 평가자에 대한 강건성을 보입니다.
Evaluating and mitigating a generative system's susceptibility to jailbreak attacks is critical to its safe deployment. Given the number of deployable systems, full per-configuration evaluation and optimization is impractical. In this paper, we formalize the behavioral geometry of a population of models that, by leveraging previously evaluated and defended models, supports both efficient susceptibility prediction and effective defense transfer across a population. We apply the framework to 79 models spanning 24 providers and to 100 system configurations of a single base model. Simple methods that use the behavioral geometry reach an AUPRC of $0.94$ for susceptibility detection with $\approx98\%$ fewer probes relative to a full evaluation. Using the behavioral geometry to select which model to transfer an optimized defense from outperforms same-provider assignment ($+2\%$, $p = 0.03$) at no additional probe cost, with a set of three models sufficient to cover the population. Results are robust to hyperparameter selection and judge.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.