2605.29358v1 May 28, 2026 cs.AI

단일 의미성을 확장하다: Claude 3 Sonnet에서 해석 가능한 특징 추출

Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet

T. Henighan
T. Henighan
Citations: 84,091
h-index: 22
Andy Jones
Andy Jones
Citations: 13,852
h-index: 13
Chris Olah
Chris Olah
Citations: 18,178
h-index: 16
Trenton Bricken
Trenton Bricken
Citations: 1,088
h-index: 8
Hoagy Cunningham
Hoagy Cunningham
Citations: 2,035
h-index: 5
C. McDougall
C. McDougall
Citations: 735
h-index: 4
Craig Citro
Craig Citro
Citations: 574
h-index: 2
Adam Pearce
Adam Pearce
Citations: 1,647
h-index: 7
Joshua Batson
Joshua Batson
Citations: 1,093
h-index: 6
Jack Lindsey
Jack Lindsey
Citations: 784
h-index: 4
E. Ameisen
E. Ameisen
Citations: 659
h-index: 5
Adly Templeton
Adly Templeton
Citations: 560
h-index: 3
Tom Conerly
Tom Conerly
Citations: 11,689
h-index: 8
Jonathan Marcus
Jonathan Marcus
Citations: 597
h-index: 2
N. Turner
N. Turner
Citations: 2,303
h-index: 21
M. MacDiarmid
M. MacDiarmid
Citations: 2,147
h-index: 10
A. Tamkin
A. Tamkin
Citations: 555
h-index: 2
Esin Durmus
Esin Durmus
Citations: 560
h-index: 3
Tristan Hume
Tristan Hume
Citations: 13,090
h-index: 13
Francesco Mosconi
Francesco Mosconi
Citations: 979
h-index: 3
C. D. Freeman
C. D. Freeman
Citations: 630
h-index: 4
T. Sumers
T. Sumers
Citations: 1,712
h-index: 17
E. Rees
E. Rees
Citations: 711
h-index: 5
Adam S. Jermyn
Adam S. Jermyn
Citations: 1,037
h-index: 3
Shan Carter
Shan Carter
Citations: 679
h-index: 3
Brian Chen
Brian Chen
Citations: 614
h-index: 4

본 연구에서는 희소 오토인코더가 생산 규모의 언어 모델인 Claude 3 Sonnet으로부터 해석 가능한 특징을 추출할 수 있음을 보여주며, 이를 통해 사전 학습 방법이 소규모 트랜스포머를 넘어 얼마나 확장될 수 있는지에 대한 미해결 질문에 답하고자 합니다. 우리는 모델의 중간 레이어 잔차 스트림에서 최대 34백만 개의 특징을 사용하여 희소 오토인코더를 학습했으며, 하이퍼파라미터 선택은 스케일링 법칙을 기반으로 이루어졌습니다. 결과적으로 추출된 특징들은 다국어 및 다중 모달(텍스트 데이터로만 학습되었음에도 이미지에 적용 가능)하며, 구체적인 예시와 개념에 대한 추상적 논의 모두에 반응합니다. 또한 이러한 특징들을 활용하여 모델의 동작을 해석과 일관성 있게 제어할 수 있습니다. 우리는 유명한 개체 및 장소뿐만 아니라 풍자 또는 코드 오류와 같은 더 추상적인 개념에 해당하는 특징들도 발견했습니다. 또한, 언어 모델이 잠재적으로 해를 끼칠 수 있는 방식과 관련된 특징들 (예: 기만, 권력 추구, 아첨, 편견)을 식별했으며, 이러한 특징들이 조작될 때 모델의 출력에 인과적으로 영향을 미치는 것을 확인했습니다. 추가적으로, 특징의 해석 가능성, 기하학적 구조 및 계산 기능에 대한 분석도 수행했습니다. 그러나 여전히 중요한 한계점이 존재합니다: 추출된 특징들의 집합은 불완전하며, 우리의 특징들이 실제로 모델의 연산을 정확하게 반영하는지 평가할 수 있는 엄격한 방법이 부족합니다.

Original Abstract

We demonstrate that sparse autoencoders can extract interpretable features from Claude 3 Sonnet, a production-scale language model, addressing the open question of whether dictionary learning methods scale beyond small transformers. We trained sparse autoencoders with up to 34 million features on the model's middle layer residual stream, using scaling laws to guide hyperparameter selection. The resulting features are multilingual and multimodal (generalizing to images despite text-only training), respond to both concrete instances and abstract discussions of concepts, and can be used to steer model behavior in ways consistent with their interpretations. We find features corresponding to famous entities and locations, as well as more abstract concepts like sarcasm or errors in code. We also identify features relevant to ways in which language models might cause harm--including features representing deception, power-seeking, sycophancy, and bias--and show that these causally influence model outputs when manipulated. Additionally, we conduct analyses of feature interpretability, geometry, and computational function. However, significant limitations remain: our suite of features is incomplete, and we lack rigorous methods for evaluating whether our features faithfully capture model computations.

698 Citations
55 Influential
11 Altmetric
863.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!