확장 가능한 회로 학습을 통한 대규모 언어 모델 해석
Scalable Circuit Learning for Interpreting Large Language Models
메커니즘적 해석 분야에서 중요한 연구 방향은 LLM 구성 요소에 대한 희소 회로를 학습하여 모델의 동작 방식과 연관성을 파악하는 것입니다. 그러나 원시 뉴런은 다의적 의미를 가지므로, 학습된 회로를 해석하기 어렵습니다. 희소 오토인코더(SAE) 특징은 이러한 문제를 완화하지만, 높은 차원성은 기존의 개입 기반 회로 학습 방법을 계산적으로 비효율적으로 만듭니다. 본 논문에서는 희소 선형 회귀에 기반한 확장 가능한 회로 학습 방법인 CircuitLasso를 제안합니다. CircuitLasso는 벤치마크 데이터에서 최첨단 개입 기반 방법과 동등한 구조적 정확도를 갖는 회로를 복구하며, 계산 비용은 훨씬 적게 소요됩니다. 해석 가능성을 위해 CircuitLasso는 SAE 특징 간의 관계를 효율적으로 파악하여 인간이 이해할 수 있는 의미론적 특징이 모델을 통해 어떻게 전파되어 예측에 영향을 미치는지 보여줍니다. 마지막으로, 학습된 회로에서 얻은 통찰력을 활용하여 일반화 성능과 비용 측면에서 상당한 개선을 달성하는 도메인 일반화 작업에서 우리의 방법의 유용성을 검증합니다.
A prominent research direction in mechanistic interpretability is learning sparse circuits over LLM components to reveal how they jointly produce model behavior. However, raw neurons are polysemantic, making learned circuits hard to interpret. Sparse autoencoder (SAE) features alleviate this, but their high dimensionality makes existing intervention-based circuit learning methods computationally prohibitive. We propose CircuitLasso, a scalable circuit-learning approach based on sparse linear regression. CircuitLasso recovers circuits whose structural accuracy matches that of state-of-the-art intervention-based methods on the benchmark data, at a fraction of the computational cost. For interpretability, CircuitLasso efficiently uncovers relationships among SAE features, showing how human-interpretable semantic features propagate through the model and influence its predictions. Finally, we validate the utility of our learned circuits by leveraging their insights to achieve comparable performance at substantially lower cost on a domain-generalization task.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.