2606.16920v1 Jun 15, 2026 cs.LG

LLM 회로 발견에서의 분산 현상 해명

Demystifying Variance in Circuit Discovery of LLMs

F. Wu
F. Wu
Citations: 20
h-index: 2
V. Cevher
V. Cevher
Citations: 16,461
h-index: 62
Francesco Tonin
Francesco Tonin
Citations: 68
h-index: 4

회로 발견은 메커니즘 해석의 핵심 기술로서, 특정 작업을 수행하는 데 중요한 모델 구성 요소를 파악하는 데 사용됩니다. 현재 최고 성능을 보이는 방법(EAP-IG)은 (신뢰성/비신뢰성) 지표에서 좋은 결과를 보이지만, 상당한 변동성을 나타냅니다. 이러한 변동성은 다음과 같습니다. 첫째, 동일 분포에서 추출된 새로운 데이터 배치로 테스트할 때 회로가 변경되는 재샘플링 분산; 둘째, 프롬프트를 다시 표현했을 때 발견된 회로가 이동하는 재구성 분산; 셋째, 모집단 수준에서는 낮은 비신뢰성을 보이는 회로가 개별 샘플에서 큰 비신뢰성 변동을 나타내는 샘플 단위 분산입니다. 본 논문은 이러한 분산의 근본 원인을 연구합니다. 우리는 EAP-IG를 개선하고 이론적 보장을 제공하는 새로운 회로 발견 방법인 CEAP가 재샘플링 분산을 크게 줄일 수 있음을 보여줍니다. 또한, 프롬프트 템플릿이 다르면 모델에서 활성화되는 회로가 달라지기 때문에 재구성 분산이 발생한다고 설명합니다. 이는 다양한 템플릿으로 표현될 수 있는 작업에 대한 모델의 동작을 설명하고 제어하는 포괄적인 회로를 찾는 것이 어려울 수 있음을 시사하며, 이는 LLM을 제어하기 어려울 수 있다는 점을 암시합니다. 또한, 더 작고 해석 가능하도록 설계된 희소성이 이러한 문제를 해결하지 못한다는 것을 보여줍니다. 샘플 단위 분산에 관해서는, 이는 대부분 무해하다고 주장합니다. 극히 낮은 비신뢰성 점수는 종종 비신뢰성을 정의하는 방식에서 비롯되며, 측정된 회로의 결함 때문이 아닐 수 있습니다. 우리는 비신뢰성의 크기가 때때로 관찰되는 매우 낮은 점수에 기여하는 신경 메커니즘인 선택적 기여도 스케일링에 의해 영향을 받는다는 것을 보여줍니다.

Original Abstract

Circuit discovery is a key technique in mechanistic interpretability to pinpoint the model components that are crucial for performing a given task. Although the current state-of-the-art method (EAP-IG) performs well on the metric of (un)faithfulness, it suffers from substantial variability. This includes resampling variance, where the circuit changes when we probe with a new batch of data from the same distribution; rephrasing variance, where the discovered circuit shifts when the prompts are rephrased; and sample-wise variance, where a circuit with low population unfaithfulness exhibits large fluctuations in unfaithfulness across individual samples. This paper studies the roots of these variances. We demonstrate that CEAP, our new circuit discovery method that improves upon EAP-IG with a theoretical guarantee, can substantially lessen resampling variance. We further show that rephrasing variance arises because prompts with different templates tend to activate different circuits in the model. This leads us to argue that it may be challenging to find a comprehensive circuit that explains and controls the model's behavior on a task, which can be expressed in countless templates, suggesting that LLMs may be inherently hard to steer. We show that sparsity, which has been claimed to form more compact and interpretable task circuits, fails to solve this problem. Regarding sample-wise variance, we argue that it is largely benign: extremely poor unfaithfulness scores often stem from how unfaithfulness is defined, rather than from defects in the measured circuits. We show that the magnitude of unfaithfulness is affected by selective contribution scaling, a neural mechanism that accounts for the extremely poor scores sometimes observed.

0 Citations
0 Influential
30 Altmetric
150.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!