2602.04718v1 Feb 04, 2026 cs.LG

직교성 정규화를 통한 개입 가능하고 해석 가능한 특징 식별

Superposition Without Interference? Towards Isolated Interventions via Almost Orthogonal Features in Language Models

Moritz Miller
Moritz Miller
Citations: 5
h-index: 2
Florent Draye
Florent Draye
Citations: 16
h-index: 3
Bernhard Schölkopf
Bernhard Schölkopf
Citations: 999
h-index: 18

최근 고정된 희소 오토인코더를 활용한 언어 모델 미세 조정의 발전과 함께, 우리는 디코더 행렬을 거의 직교하는 특징으로 분리했습니다. 이는 특징 간의 간섭과 중첩을 줄여주면서, 대상 데이터셋에 대한 성능은 거의 변하지 않습니다. 우리의 직교성 페널티는 식별 가능한 특징을 생성하여 분해의 고유성을 보장합니다. 또한, 임베딩된 특징 설명 간의 거리가 더 엄격한 직교성 페널티에 따라 증가한다는 것을 발견했는데, 이는 해석 가능성을 위한 바람직한 특성입니다. $ extit{독립적인 인과 메커니즘}$ 원칙을 활용하여, 우리는 직교성이 인과적 개입에 적합한 모듈식 표현을 촉진한다고 주장합니다. 우리는 실험적으로 이러한 점점 더 직교화된 특징들이 개별적인 개입을 가능하게 한다는 것을 보여줍니다. 우리의 코드는 $ exttt{https://github.com/mrtzmllr/sae-icm}$에서 확인할 수 있습니다.

Original Abstract

A central premise in mechanistic interpretability is that meaningful concepts in language models are represented by linear features in activation space. For such features to support reliable interventions, manipulating one feature should not substantially alter the effects of others. In practice, however, feature entanglement leads to interference such that localized interventions can have unintended downstream effects. Motivated by the _Independent Causal Mechanisms_ principle, we propose to constrain internal features to be almost orthogonal. We argue that this promotes modular representations amenable to causal intervention. We formalize this problem by characterizing the gap between an idealized isolated intervention and its realized effect on model outputs in terms of feature interference. We upper-bound the propagation of feature interference in terms of the self-coherence of the feature dictionary, and relate this discrepancy to an explicit orthogonality regularization on the dictionary itself. Empirically, we show that this regularization enables more isolated interventions on mathematical reasoning concepts while preserving model performance. Our code is available under https://github.com/mrtzmllr/sae-icm.

2 Citations
0 Influential
55.493061443341 Altmetric
11.0 Score
Original PDF
2

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!