희소 오토인코더를 활용한 수백만 개의 해석 가능한 특징 발견
Discovering Millions of Interpretable Features with Sparse Autoencoders
희소 오토인코더(SAE)는 중첩된 언어 모델 표현을 희소하고 해석 가능한 특징으로 분해하는 강력한 도구로 부상했습니다. 그러나 SAE 훈련은 계산 비용이 많이 들며, 현재 공개적으로 사용 가능한 SAE 모델은 제한적입니다. 본 연구에서는 Qwen3 지시형 모델 패밀리(Qwen3-1.7B, Qwen3-4B, 및 Qwen3-8B)를 기반으로 훈련된 종합적인 SAE 모음인 **Qwen3-Instruct SAE**를 소개합니다. Qwen3-1.7B 및 Qwen3-4B의 경우, 잔류 스트림, MLP 출력, 그리고 어텐션 출력을 포함한 세 가지 주요 활성화 지점에서 계층별 SAE를 훈련했습니다. Qwen3-8B의 경우, 잔류 스트림 레이어의 일부에 대해 SAE를 훈련했습니다. 우리는 이 SAE들을 활성화 레벨 재구성 메트릭과 모델 레벨 복구 메트릭을 모두 사용하여 체계적으로 평가했으며, 이를 통해 계층 및 구성 요소 간의 뚜렷한 희소성-충실도 균형을 확인했습니다. 마지막으로, Qwen3-Instruct SAE를 거부 유도 사례 연구를 통해 활용 가능성을 보여주었으며, 선택된 SAE 특징이 지시형 Qwen3 모델을 거부 동작으로 유도할 수 있음을 입증했습니다. 본 자료는 희소 표현, 특징 레벨 메커니즘 및 지시형 언어 모델의 행동 개입 연구에 실질적인 리소스를 제공합니다.
Sparse autoencoders (SAEs) have emerged as a powerful tool for decomposing superposed language model representations into sparse and interpretable features. However, training SAEs is computationally expensive, and available open-source SAE models remain limited. In this work, we introduce \textbf{Qwen3-Instruct SAE}, a comprehensive suite of SAEs trained on the Qwen3 instruction-tuned model family, covering Qwen3-1.7B, Qwen3-4B, and Qwen3-8B. For Qwen3-1.7B and Qwen3-4B, we train layer-wise SAEs at three key activation sites: residual streams, MLP outputs, and attention outputs. For Qwen3-8B, we train SAEs on a subset of residual stream layers. We systematically evaluate these SAEs using both activation-level reconstruction metrics and model-level recovery metrics, revealing distinct sparsity--fidelity trade-offs across layers and components. Finally, we demonstrate the utility of Qwen3-Instruct SAE through a refusal-steering case study, showing that selected SAE features can causally steer instruction-tuned Qwen3 models toward refusal behavior. Our release provides a practical resource for studying sparse representations, feature-level mechanisms, and behavioral interventions in instruction-tuned language models
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.