2607.02510v1 Jul 02, 2026 cs.AI

LLM의 온라인 안전 모니터링

Online Safety Monitoring for LLMs

Eric Nalisnick
Eric Nalisnick
Citations: 91
h-index: 5
Alexander Timans
Alexander Timans
Citations: 65
h-index: 4
Mona Schirmer
Mona Schirmer
Citations: 9
h-index: 2
Metod Jazbec
Metod Jazbec
Citations: 141
h-index: 6
C. Naesseth
C. Naesseth
Citations: 1,488
h-index: 17
Maja Waldron
Maja Waldron
Citations: 0
h-index: 0

정렬 학습을 거치더라도 LLM은 배포 시점에 여전히 위험한 결과를 생성할 가능성이 있습니다. 따라서, LLM이 생성하는 결과물을 실시간으로 모니터링하고, 안전성을 더 이상 보장할 수 없을 때 경고를 발생시키는 것은 매우 중요합니다. 본 연구에서는 외부 모델로부터 얻은 검증 신호를 임계값 기반으로 처리하여 경고 결정을 내리는 간단한 실시간 모니터링 시스템을 제안합니다. 이때, 임계값은 위험 관리 원칙에 따라 조정됩니다. 수학적 추론 및 적대적 공격 데이터셋에 대한 실험 결과, 본 연구에서 제안하는 간단한 설계가 순차 가설 검정에 기반한 더욱 복잡한 모니터링 시스템과 경쟁력이 있음을 확인했습니다.

Original Abstract

Despite alignment training, LLMs remain prone to generating unsafe outputs at deployment time. Monitoring outputs online and raising an alarm when safety can no longer be assumed is therefore critical. We study a simple real-time monitor that turns a verifier signal from an external model into an alarm decision by thresholding, with the threshold calibrated via risk control. In experiments on mathematical reasoning and red teaming datasets, we show that this simple design is competitive with more advanced monitors based on sequential hypothesis testing.

0 Citations
0 Influential
8.5 Altmetric
42.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!