인공지능 안전을 위한 반응 항목 이론
Item Response Theory for AI Safety
언어 모델은 안전성 측면에서 다양한 특성을 보이며, 이러한 차이는 안전성 평가 지표를 통해 측정됩니다. 그러나 기존의 지표 점수는 중복성, 높은 상관관계 및 모델의 의도적인 성능 저하 가능성 때문에 신뢰성과 해석력이 떨어지는 문제가 있습니다. 본 연구에서는 반응 항목 이론(IRT)을 활용하여, 각 항목이 가진 통계적 특성을 추론하고 이를 바탕으로 모델의 잠재적인 안전성 속성을 측정합니다. 우리는 192개의 언어 모델에 대한 8가지 안전성 평가 지표를 분석하는 가장 큰 규모의 LLM 안전성 평가 심리측정 연구를 수행했으며, 다음과 같은 세 가지 결과를 얻었습니다. 첫째, 거부 엄격성, 진실성, 맥락적 해악이라는 세 가지 해석 가능한 요인이 모델 간의 변동성을 대부분 설명한다는 것을 확인했습니다. 둘째, 통계적으로 선택된 항목들은 무작위로 선택된 항목들보다 더 낮은 오차율로 전체 지표 점수를 재현하며, 몇몇 개별 지표의 경우 약 10개의 적응형으로 선택된 항목만으로도 충분한 성능을 얻을 수 있어 평가 비용을 97~99% 절감할 수 있습니다. 셋째, IRT는 개별 모델에 대한 감사를 지원하며, 이를 통해 단순한 성능 저하 및 API 뒤의 모델 변경 사항을 탐지할 수 있습니다. 종합적으로 볼 때, IRT는 안전성 지표를 이해하고 축소하며 감사하는 데 유용한 도구이며, 선도적인 연구 기관과 평가자들은 이러한 도구를 적극적으로 활용해야 합니다.
Language models differ in how safely they behave and these differences are measured by safety benchmarks. But aggregated benchmark scores are hard to trust and interpret, because benchmarks duplicate one another, correlate heavily, and models may sandbag when they detect evaluation. To address these issues, we draw on Item Response Theory (IRT), a statistical toolkit for measuring these latents from performance on items with inferred psychometric properties. We fit IRT models to eight safety benchmarks across 192 language models, the largest psychometric analysis of LLM safety evaluations to date, and contribute three results. First, we find that three interpretable factors of refusal strictness, truthfulness, and contextual harm explain most of the variance between models across benchmarks. Second, psychometrically selected items recover full benchmark scores with lower error than random subsets of the same size, and roughly ten adaptively chosen items suffice for several individual benchmarks, cutting evaluation cost by 97-99%. Third, IRT supports audits of individual models, showing that it can be used to detect naive sandbagging and changes of model behind APIs. Overall, we show IRT is a ready-made toolkit for reading, reducing, and auditing safety benchmarks, which we recommend frontier labs and evaluators adopt.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.