2608.03913v1 Aug 04, 2026 cs.LG

효율적인 회로 추출을 위한 희소 가중치 분해

Sparse Weight Decomposition for Efficient Circuit Extraction

Bryan Dai
Bryan Dai
Citations: 250
h-index: 4
Chuanhao Yan
Chuanhao Yan
Citations: 27
h-index: 3
Yawen Duan
Yawen Duan
Citations: 0
h-index: 0
Zhenfei Yin
Zhenfei Yin
Citations: 0
h-index: 0
Jie Fu
Jie Fu
Citations: 29
h-index: 3
Xuhan Huang
Xuhan Huang
Citations: 12
h-index: 2
Hang Zhao
Hang Zhao
Citations: 18
h-index: 2

밀집된 사전 학습 트랜스포머 모델은 해석 가능한 단위가 자연스럽게 드러나지 않아 회로 추출에 어려움이 있습니다. 기존 방법들은 보조적인 희소 표현을 학습하거나 희소 모델을 훈련하여 이러한 단위를 얻지만, 이 과정에서 상당한 추가 계산량이 발생하고 분석 대상 표현과 원래 사전 학습 모델 간의 정확도 차이가 발생할 수 있습니다. 본 연구에서는 사전 학습된 선형 투영을 각 가중치 행렬을 두 개의 희소 인자로 분해하여 재파라미터화하는 '희소 가중치 분해 (Sparse Weight Decomposition, SWD)' 방법을 제안합니다. 이 방법은 별도의 대체 네트워크를 훈련하지 않고도 파라메트릭 표현을 통해 기존의 희소 특징 학습 방법에 사용되는 동일한 점수 부여, 선택 및 제거 기반 회로 추출 워크플로우를 지원합니다. 단일 행렬 교체 시, SWD는 Transcoder 및 기타 강력한 기준 모델이 달성한 성능과 유사한 정확도를 유지하면서, 해당 기준 모델들이 대체 네트워크를 훈련하는 데 사용하는 데이터 양의 1% 미만을 사용합니다. 동일한 수준의 정확도를 달성하며, GPT-2, Qwen2.5, 및 Qwen3.5-27B 모델에서 SWD는 더 적은 활성 읽기/쓰기 연결과 선택된 단위 수를 사용하여 회로의 충분성과 필요성을 만족시킵니다. 또한, 본 연구에서는 미세 조정 후 모든 어텐션 및 MLP 가중치 행렬을 전체 모델 교체 방식으로 대체할 때도 SWD가 효과적임을 보여줍니다. 마지막으로, SWD는 제로 데이터 변형을 제공하여 메커니즘 해석 가능성 분석 (예: 단계별 분석)의 활용 범위를 넓힐 수 있습니다.

Original Abstract

Dense pretrained transformers do not naturally expose interpretable units for circuit extraction. Existing approaches obtain such units by learning auxiliary sparse representations or training sparse models, incurring substantial additional computation while potentially introducing a fidelity gap between the representation being analyzed and the original pretrained model. We propose Sparse Weight Decomposition (SWD), which reparameterizes pretrained linear projections by factorizing each weight matrix into two sparse factors whose shared intermediate coordinates serve as individually addressable circuit units. Without training a separate replacement network, this parametric representation supports the same scoring, selection, and ablation circuit extraction workflow used for methods that learn sparse features. Across single-matrix replacements, SWD matches the held-out fidelity achieved by Transcoder and other strong baselines while using less than 1% of the data that those baselines use to train their replacements. For matched replacement fidelity, SWD reaches the same circuit sufficiency and necessity targets with fewer active read/write edges and selected units across tasks on GPT-2, Qwen2.5, and Qwen3.5-27B. We further show that SWD remains effective for full-model replacement of all attention and MLP weight matrices after fine-tuning the nonzero factor values. Finally, SWD also features a zero-data variant, allowing broader use of mechanistic interpretability analysis (e.g., per-step analysis).

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!