언어 모델의 잘못된 사고 과정을 탐구하다
Probing the Misaligned Thinking Process of Language Models
대규모 언어 모델은 전략적 기만, 고의적인 방해, 자기 보호 등 다양한 형태의 부합하지 않는(misaligned) 행동을 보입니다. 이러한 모델이 점점 더 중요한 상황에서 사용됨에 따라, 안전하고 책임감 있는 사용을 보장하기 위해 이러한 행동을 안정적으로 탐지하는 것이 중요합니다. 본 연구에서는 불일치(misalignment)를 세분화된 인지 과정, 즉 불일치 지표로 분해하고, 선형 프로브(linear probe)를 통해 모델의 내부 활성화 상태에서 이러한 지표의 존재 여부를 감지하여 불일치를 모니터링하는 방법을 제안합니다. 우리는 다양한 형태의 불일치 행동을 포괄하는 18개의 지표 분류 체계를 개발했으며, 이와 함께 자동화된 메타-계획 기반 파이프라인을 통해 다중 회전 학습 대화를 생성합니다. 일반화 성능을 엄격하게 평가하기 위해, 자동화된 행동 유도, 기존의 불일치 벤치마크, 그리고 자연스러운 정상적인 대화들을 결합한 외부 데이터셋을 구축했습니다. 5가지 유형의 불일치 행동에 대해, 개발된 프로브는 강력한 LLM 평가 모델과 함께 외부 데이터셋에서 0.935의 AUROC 값을 보여주었으며, 동시에 정상적인 트래픽에서는 낮은 위양성률(false positive rate)을 유지했습니다. 또한, 프로브와 모델 내부의 불일치 지표 표현 방식에 대한 심층적인 분석을 수행하여 이해도를 높였습니다.
Large language models exhibit a growing range of misaligned behaviors such as strategic deception, sandbagging, and self-preservation. As they are increasingly deployed in high-stakes settings, it is critical to reliably detect such behaviors to ensure safe and responsible use. In this work, we propose to monitor misalignment by decomposing it into fine-grained cognitive processes -- misalignment indicators -- and detecting their presence in a model's internal activations via linear probes. We develop a taxonomy of 18 indicators spanning different misaligned behaviors, paired with an automated, meta-plan-guided pipeline that generates multi-turn training conversations. To rigorously evaluate generalization, we construct an out-of-distribution suite combining automated behavioral elicitation, established misalignment benchmarks, and natural benign conversations. Across 5 misaligned behaviors, our probes match a strong LLM judge with 0.935 AUROC on out-of-distribution benchmarks while keeping a low false positive rate on benign traffic. We further perform in-depth analysis to understand the probes and the model's internal representations of misalignment indicators.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.