LLM의 온라인 안전 모니터링
Online Safety Monitoring for LLMs
정렬 학습을 거치더라도 LLM은 배포 시점에 여전히 위험한 결과를 생성할 가능성이 있습니다. 따라서, LLM이 생성하는 결과물을 실시간으로 모니터링하고, 안전성을 더 이상 보장할 수 없을 때 경고를 발생시키는 것은 매우 중요합니다. 본 연구에서는 외부 모델로부터 얻은 검증 신호를 임계값 기반으로 처리하여 경고 결정을 내리는 간단한 실시간 모니터링 시스템을 제안합니다. 이때, 임계값은 위험 관리 원칙에 따라 조정됩니다. 수학적 추론 및 적대적 공격 데이터셋에 대한 실험 결과, 본 연구에서 제안하는 간단한 설계가 순차 가설 검정에 기반한 더욱 복잡한 모니터링 시스템과 경쟁력이 있음을 확인했습니다.
Despite alignment training, LLMs remain prone to generating unsafe outputs at deployment time. Monitoring outputs online and raising an alarm when safety can no longer be assumed is therefore critical. We study a simple real-time monitor that turns a verifier signal from an external model into an alarm decision by thresholding, with the threshold calibrated via risk control. In experiments on mathematical reasoning and red teaming datasets, we show that this simple design is competitive with more advanced monitors based on sequential hypothesis testing.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.