2604.22167v1 Apr 24, 2026 cs.LG

언어 모델 출력 분포에서의 꼬리 위험 추정

Estimating Tail Risks in Language Model Output Distributions

Raghav Singhal
Raghav Singhal
Citations: 344
h-index: 2
Rico Angell
Rico Angell
Citations: 58
h-index: 1
Zachary Horvitz
Zachary Horvitz
Citations: 512
h-index: 8
Zhou Yu
Zhou Yu
Citations: 209
h-index: 3
R. Ranganath
R. Ranganath
Citations: 310
h-index: 7
Kathleen McKeown
Kathleen McKeown
Citations: 333
h-index: 6
He He
He He
Citations: 179
h-index: 5

언어 모델은 점점 더 발전하고 있으며, 빠르게 대규모로 배포되고 있습니다. 그 결과, 이러한 모델의 안전성은 매우 중요한 문제가 되었습니다. 다행히도, 정렬 기술의 발전으로 유해한 모델 출력의 가능성이 크게 줄었습니다. 그러나 모델이 하루에 수십억 번 쿼리되는 경우, 극히 드문 최악의 상황도 발생할 수 있습니다. 현재 안전성 평가에서는 유해한 출력을 생성하는 입력 분포를 파악하는 데 중점을 둡니다. 이러한 평가는 모델의 확률적 특성과 꼬리 출력 행동을 고려하지 않습니다. 이러한 꼬리 위험을 측정하기 위해, 우리는 임의의 입력 쿼리에 대해 유해한 출력의 확률을 효율적으로 추정하는 방법을 제안합니다. 대상 모델에서 무차별적으로 샘플링하는 대신, 유해한 출력이 드물게 나타날 수 있는 경우, 우리는 대상 모델의 안전하지 않은 버전을 만들어 중요 샘플링을 수행합니다. 이러한 안전하지 않은 버전은 유해한 출력을 더 가능하게 만들어 샘플 효율적인 추정을 가능하게 합니다. 오용 및 불일치 정도를 측정하는 벤치마크에서, 이러한 추정치는 10~20배 적은 샘플을 사용하여 수행된 무차별 몬테카를로 추정치와 일치합니다. 예를 들어, 500개의 샘플만으로 10^-4 정도의 유해한 출력 확률을 추정할 수 있습니다. 또한, 이러한 유해성 추정치는 모델 입력의 작은 변화에 대한 모델의 민감성을 파악하고 배포 위험을 예측하는 데 도움이 될 수 있습니다. 우리의 연구는 정확한 희귀 이벤트 추정이 안전성 평가에 필수적이며 실현 가능하다는 것을 보여줍니다. 코드는 https://github.com/rangell/LMTailRisk 에서 확인할 수 있습니다.

Original Abstract

Language models are increasingly capable and are being rapidly deployed on a population-level scale. As a result, the safety of these models is increasingly high-stakes. Fortunately, advances in alignment have significantly reduced the likelihood of harmful model outputs. However, when models are queried billions of times in a day, even rare worst-case behaviors will occur. Current safety evaluations focus on capturing the distribution of inputs that yield harmful outputs. These evaluations disregard the probabilistic nature of models and their tail output behavior. To measure this tail risk, we propose a method to efficiently estimate the probability of harmful outputs for any input query. Instead of naive brute-force sampling from the target model, where harmful outputs could be rare, we operationalize importance sampling by creating unsafe versions of the target model. These unsafe versions enable sample-efficient estimation by making harmful outputs more probable. On benchmarks measuring misuse and misalignment, these estimates match brute-force Monte Carlo estimates using 10-20x fewer samples. For example, we can estimate probability of harmful outputs on the order of 10^-4 with just 500 samples. Additionally, we find that these harmfulness estimates can reveal the sensitivity of models to perturbations in model input and predict deployment risks. Our work demonstrates that accurate rare-event estimation is both critical and feasible for safety evaluations. Code is available at https://github.com/rangell/LMTailRisk

1 Citations
0 Influential
30.931471805599 Altmetric
155.7 Score
Original PDF
3

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!