트랜스포머 표현에서 방향과 크기 분리: L2 정규화 기반 교란 분석을 통한 이중 분리 현상
Disentangling Direction and Magnitude in Transformer Representations: A Double Dissociation Through L2-Matched Perturbation Analysis
트랜스포머의 숨겨진 상태는 고차원 벡터로 정보를 인코딩하지만, 표현 공간에서의 방향(방향성)과 크기(벡터의 크기)가 서로 다른 기능적 역할을 수행하는지 여부는 명확하지 않습니다. Pythia 모델 패밀리를 연구한 결과, 놀라운 교차 분리 현상을 발견했습니다. 각도 기반 교란은 언어 모델링 손실에 최대 42.9배 더 큰 영향을 미치는 반면, 크기 기반 교란은 구문 처리(주어-동사 일치 정확도가 20.4% 감소한 반면, 1.6%만 감소)에 비례적으로 더 큰 영향을 미칩니다. 이러한 발견은 L2 정규화 기반 교란 분석이라는 방법론을 통해 가능했으며, 이 방법론은 각도 및 크기 기반 교란이 동일한 유클리드 거리 변위를 갖도록 보장합니다. 인과적 개입 분석 결과, 각도 기반 손실은 주로 어텐션 경로를 통해 발생하며, 어텐션 복구를 통해 손실의 28.4%를 회복할 수 있습니다. 반면, 크기 기반 손실은 부분적으로 LayerNorm 경로를 통해 발생하며, LayerNorm 복구를 통해 29.9%의 회복이 가능합니다. 이러한 패턴은 Pythia 아키텍처 패밀리 내의 다양한 규모에서 반복적으로 나타납니다. 이러한 결과는 방향과 크기가 LayerNorm 기반 아키텍처에서 부분적으로 서로 다른 계산적 역할을 수행한다는 증거를 제공합니다. 방향은 주로 어텐션 라우팅에 영향을 미치는 반면, 크기는 미세한 구문 판단을 위한 처리 강도를 조절합니다. RMSNorm 기반 아키텍처에서는 다른 패턴이 나타나는데, 이는 분리 현상이 아키텍처 선택에 따라 달라짐을 시사합니다. 본 연구 결과는 선형 표현 가설을 구체화하고 모델 편집 및 해석 가능성 연구에 중요한 함의를 갖습니다.
Transformer hidden states encode information as high-dimensional vectors, yet whether direction (orientation in representational space) and magnitude (vector norm) serve distinct functional roles remains unclear. Studying Pythia-family models, we discover a striking cross-over dissociation: angular perturbations cause up to 42.9 more damage to language modeling loss, while magnitude perturbations cause disproportionately more damage to syntactic processing (20.4% vs.1.6% accuracy drop on subject-verb agreement).This finding is enabled by L2-matched perturbation analysis, a methodology ensuring that an gular and magnitude perturbations achieve identical Euclidean displacements. Causal intervention reveals that angular damage flows substantially through the attention pathways (28.4% loss recovery via attention repair), while magnitude damage flows partly through the LayerNorm pathways(29.9% recovery via LayerNorm repair). These patterns replicate across scales within the Pythia architecture family. These findings provide evidence that direction and magnitude support partially distinct computational roles in LayerNorm based architectures. The direction preferentially affects attentional routing, while magnitude modulates processing intensity for fine-grained syntactic judgments. We find different patterns in RMSNorm-based architectures, suggesting that the dissociation depends on architectural choices. Our results refine the linear representation hypothesis and have implications for model editing and interpretability research
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.