2602.10437v2 Feb 11, 2026 cs.LG

제어 강화 학습: 희소 오토인코더 특징을 활용한 LLM의 토큰 단위 해석 가능한 제어

Control Reinforcement Learning: Interpretable Token-Level Steering of LLMs via Sparse Autoencoder Features

Seonglae Cho
Seonglae Cho
Citations: 18
h-index: 3
Zekun Wu
Zekun Wu
Citations: 169
h-index: 7
Adriano S. Koshiyama
Adriano S. Koshiyama
Citations: 218
h-index: 8

희소 오토인코더(SAE)는 언어 모델의 활성화를 해석 가능한 특징으로 분해하지만, 기존 방법은 어떤 특징이 활성화되는지 보여줄 뿐, 어떤 특징을 증폭시켰을 때 모델의 출력이 어떻게 변하는지는 알려주지 않습니다. 본 논문에서는 제어 강화 학습(CRL)을 소개합니다. CRL은 정책을 학습하여 각 토큰에서 제어를 위해 SAE 특징을 선택하며, 해석 가능한 개입 로그를 생성합니다. 학습된 정책은 증폭되었을 때 모델의 출력을 변화시키는 특징을 식별합니다. 적응적 특징 마스킹은 다양한 특징 발견을 장려하는 동시에 단일 특징의 해석 가능성을 유지합니다. 본 프레임워크는 새로운 분석 기능을 제공합니다. 분기점 추적은 특징 선택이 출력의 정확성을 결정하는 토큰을 찾아내고, 크리틱 경로 분석은 정책의 한계를 가치 추정 오류와 분리하며, 계층별 비교는 초기 계층에서 구문 특징, 후기 계층에서 의미 특징을 드러냅니다. Gemma 2 2B 모델을 사용하여 MMLU, BBQ, GSM8K, HarmBench 및 XSTest 데이터셋에서 CRL은 성능 향상을 달성하는 동시에 토큰 단위의 개입 로그를 제공합니다. 이러한 결과는 학습된 특징 제어를 정적 특징 분석을 보완하는 메커니즘적 해석 도구로 확립합니다.

Original Abstract

Sparse autoencoders (SAEs) decompose language model activations into interpretable features, but existing methods reveal only which features activate, not which change model outputs when amplified. We introduce Control Reinforcement Learning (CRL), which trains a policy to select SAE features for steering at each token, producing interpretable intervention logs: the learned policy identifies features that change model outputs when amplified. Adaptive Feature Masking encourages diverse feature discovery while preserving singlefeature interpretability. The framework yields new analysis capabilities: branch point tracking locates tokens where feature choice determines output correctness; critic trajectory analysis separates policy limitations from value estimation errors; layer-wise comparison reveals syntactic features in early layers and semantic features in later layers. On Gemma 2 2B across MMLU, BBQ, GSM8K, HarmBench, and XSTest, CRL achieves improvements while providing per-token intervention logs. These results establish learned feature steering as a mechanistic interpretability tool that complements static feature analysis with dynamic intervention probes

1 Citations
0 Influential
4 Altmetric
21.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!