2606.16939v1 Jun 15, 2026 cs.LG

Scalable Circuit Learning for Interpreting Large Language Models

K. Ramamurthy
K. Ramamurthy
Citations: 7,161
h-index: 33
Amit Dhurandhar
Amit Dhurandhar
Citations: 257
h-index: 6
Dennis Wei
Dennis Wei
Citations: 50
h-index: 3
Naiyu Yin
Naiyu Yin
Citations: 13
h-index: 2
Tian Gao
Tian Gao
Citations: 4
h-index: 1
Yue Yu
Yue Yu
Citations: 15
h-index: 2

A prominent research direction in mechanistic interpretability is learning sparse circuits over LLM components to reveal how they jointly produce model behavior. However, raw neurons are polysemantic, making learned circuits hard to interpret. Sparse autoencoder (SAE) features alleviate this, but their high dimensionality makes existing intervention-based circuit learning methods computationally prohibitive. We propose CircuitLasso, a scalable circuit-learning approach based on sparse linear regression. CircuitLasso recovers circuits whose structural accuracy matches that of state-of-the-art intervention-based methods on the benchmark data, at a fraction of the computational cost. For interpretability, CircuitLasso efficiently uncovers relationships among SAE features, showing how human-interpretable semantic features propagate through the model and influence its predictions. Finally, we validate the utility of our learned circuits by leveraging their insights to achieve comparable performance at substantially lower cost on a domain-generalization task.

0 Citations
0 Influential
16.5 Altmetric
82.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!