희소 오토인코더는 개념과 함수를 모두 인코딩한다: 특징 효과의 하위 구조 기하학
Sparse Autoencoders Encode Both Concepts and Functions: The Downstream Geometry of Feature Effects
희소 오토인코더(SAE)가 해석 도구로 광범위하게 사용되지만, SAE 특징과 모델 동작 간의 일관성 없는 연결로 인해 활용이 제한됩니다. 명확한 활성화 설명을 가진 특징이라도 약하거나 예상치 못한 인과적 영향을 미칠 수 있으며, 특정 프롬프트에 따라 제어 방향이 달라지거나 의도된 방향과 반대로 작용할 수도 있습니다. 또한, 활성화 기반 특징 선택은 원하는 출력 변화를 유발하는 특징을 놓칠 수 있습니다. 기존 연구에서는 모델 내에서 계산되는 특징의 기하학적 구조를 연구했습니다. 본 연구에서는 특징 개입으로 인해 발생하는 모델 로짓(logit) 변화의 기하학적 구조를 분석합니다. 우리는 Feature-Effect Geometry Analysis (FEGA)라는 비지도 학습 프레임워크를 제안하며, 이를 통해 다양한 맥락에서 동일한 활성 SAE 특징을 제거하고, 그 결과로 나타나는 로짓 변화들의 분포를 분석합니다. 다양한 SAE 변형 모델에서 일관된 1차원 효과는 드물게 나타나며, 대부분의 특징이 재사용 가능한 방향으로 작동하지 않습니다. 이러한 변동성을 설명하기 위해, 우리는 정적인 정보(예: 사실적 속성)와 관련된 '값'과 같은 특징과, 맥락에 따라 달라지는 연산과 관련된 '포인터'와 같은 특징을 구분합니다. '값'과 같은 특징은 구조화되고 낮은 차원의 효과를 더 자주 나타내지만, 이러한 효과는 일반적으로 여러 방향으로 걸쳐 있습니다. 반면에, '포인터'와 같은 특징은 주로 흩어진 효과를 나타냅니다. 우리의 결과는 특징이 해석 가능하고 인과적으로 관련성이 있더라도 안정적인 제어 방향을 제공하지 않을 수 있다는 것을 보여줍니다.
The wide-scale use of sparse autoencoders (SAEs) as interpretability tools is limited by inconsistent links between SAE features and model behavior. Features with clear activation descriptions may have weak or unexpected causal effects; steering can vary across prompts or oppose the intended direction; and activation-based feature selection can miss features that produce the desired output change. Prior work has studied feature geometry inside the model, where features are computed. We instead study the geometry of changes in model logits caused by feature interventions. We introduce Feature-Effect Geometry Analysis (FEGA), an unsupervised framework that removes the same active SAE feature across contexts and analyzes the resulting cloud of logit changes. Across SAE variants, consistent one-dimensional effects are rare: few features behave like reusable directions. To interpret this variation, we distinguish value-like features, tied to static information such as factual attributes, from pointer-like features, associated with context-dependent operations. Value-like features more often exhibit structured, low-dimensional effects, although these effects typically span several directions. Pointer-like features, by contrast, predominantly exhibit diffuse effects. Our results show that a feature can be interpretable and causally relevant without providing a stable direction for steering.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.