2606.26620v1 Jun 25, 2026 cs.LG

희소 오토인코더를 활용한 수백만 개의 해석 가능한 특징 발견

Discovering Millions of Interpretable Features with Sparse Autoencoders

Xuancheng Ren
Xuancheng Ren
Citations: 18,128
h-index: 8
Bing Zhao
Bing Zhao
Citations: 103
h-index: 5
Wenbo Li
Wenbo Li
Citations: 87
h-index: 4
Hu Wei
Hu Wei
Citations: 41
h-index: 4
Wei Qiao
Wei Qiao
Citations: 53
h-index: 3
Lin Qu
Lin Qu
Citations: 164
h-index: 7
Wei Wang
Wei Wang
Citations: 11
h-index: 2
Xinyang He
Xinyang He
Citations: 5
h-index: 1

희소 오토인코더(SAE)는 중첩된 언어 모델 표현을 희소하고 해석 가능한 특징으로 분해하는 강력한 도구로 부상했습니다. 그러나 SAE 훈련은 계산 비용이 많이 들며, 현재 공개적으로 사용 가능한 SAE 모델은 제한적입니다. 본 연구에서는 Qwen3 지시형 모델 패밀리(Qwen3-1.7B, Qwen3-4B, 및 Qwen3-8B)를 기반으로 훈련된 종합적인 SAE 모음인 **Qwen3-Instruct SAE**를 소개합니다. Qwen3-1.7B 및 Qwen3-4B의 경우, 잔류 스트림, MLP 출력, 그리고 어텐션 출력을 포함한 세 가지 주요 활성화 지점에서 계층별 SAE를 훈련했습니다. Qwen3-8B의 경우, 잔류 스트림 레이어의 일부에 대해 SAE를 훈련했습니다. 우리는 이 SAE들을 활성화 레벨 재구성 메트릭과 모델 레벨 복구 메트릭을 모두 사용하여 체계적으로 평가했으며, 이를 통해 계층 및 구성 요소 간의 뚜렷한 희소성-충실도 균형을 확인했습니다. 마지막으로, Qwen3-Instruct SAE를 거부 유도 사례 연구를 통해 활용 가능성을 보여주었으며, 선택된 SAE 특징이 지시형 Qwen3 모델을 거부 동작으로 유도할 수 있음을 입증했습니다. 본 자료는 희소 표현, 특징 레벨 메커니즘 및 지시형 언어 모델의 행동 개입 연구에 실질적인 리소스를 제공합니다.

Original Abstract

Sparse autoencoders (SAEs) have emerged as a powerful tool for decomposing superposed language model representations into sparse and interpretable features. However, training SAEs is computationally expensive, and available open-source SAE models remain limited. In this work, we introduce \textbf{Qwen3-Instruct SAE}, a comprehensive suite of SAEs trained on the Qwen3 instruction-tuned model family, covering Qwen3-1.7B, Qwen3-4B, and Qwen3-8B. For Qwen3-1.7B and Qwen3-4B, we train layer-wise SAEs at three key activation sites: residual streams, MLP outputs, and attention outputs. For Qwen3-8B, we train SAEs on a subset of residual stream layers. We systematically evaluate these SAEs using both activation-level reconstruction metrics and model-level recovery metrics, revealing distinct sparsity--fidelity trade-offs across layers and components. Finally, we demonstrate the utility of Qwen3-Instruct SAE through a refusal-steering case study, showing that selected SAE features can causally steer instruction-tuned Qwen3 models toward refusal behavior. Our release provides a practical resource for studying sparse representations, feature-level mechanisms, and behavioral interventions in instruction-tuned language models

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!