선호 학습에서의 거짓 탐지 시스템을 활용한 효율적인 감시: 확장성 연구
Scaling Trends for Lie Detector Oversight in Preference Learning
LLM(대규모 언어 모델)에서 발생하는 기만적인 행동은 감시 및 예방에 상당한 비용이 소요됩니다. 이러한 문제를 해결하기 위해 '확장 가능한 거짓 탐지를 통한 감시(SOLiD)' (Cundy & Gleave, 2025)와 같은 접근 방식이 제안되었으며, 이는 거짓 탐지기를 사용하여 검토를 위한 응답을 식별하고, 고비용의 레이블러가 이를 검토하도록 합니다. 본 논문에서는 SOLiD를 더 큰 모델로 확장하고, 보다 다양하고 현실적인 선호 학습 환경에서 성능을 평가합니다. 연구 결과, 거짓 탐지가 되지 않는 비율이 10억 파라미터 모델에서 34%에서 405억 파라미터 모델에서 14%로 감소했으며, 이는 거짓 탐지기의 정확도가 99%일 때 나타나는 현상입니다. 또한, 통계적으로 유의미한 기만 증가 없이도, 비용이 많이 드는 인간 레이블러를 미세 조정 단계에서 완전히 제거할 수 있었습니다. 그러나 SOLiD는 거짓 탐지기 학습 데이터와 선호 학습 데이터 간의 분포 변화에 민감하며, 이로 인해 거짓 탐지기의 오탐율이 비현실적인 수준으로 증가할 수 있습니다.
Deceptive behavior in LLMs is costly to monitor and prevent, motivating approaches such as Scalable Oversight via Lie Detectors (SOLiD) (Cundy & Gleave, 2025), which uses lie detectors to identify responses for review by high-cost labelers. In this paper, we scale SOLiD to larger models and evaluate it in more diverse and realistic preference-learning settings. We find favorable scaling: undetected deception drops from 34% for 1B-parameter models to 14% for 405B-parameter models at a detector true positive rate of 99%, and expensive human labelers can be removed entirely from the fine-tuning phase without a statistically significant increase in deception. However, SOLiD is sensitive to distribution shift between detector training and preference-training data, which can drive detector false positive rates to impractical levels.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.