2607.18921v1 Jul 21, 2026 cs.LG

회로 추출 결과는 어떤 요소가 추출되었는지 및 어떻게 비교되는지에 따라 달라진다.

Circuit Claims Depend on What Is Extracted and How It Is Compared

Jie Fu
Jie Fu
Citations: 120
h-index: 4
Yang Sheng
Yang Sheng
Citations: 34
h-index: 3

회로 추출은 특정 동작을 유지하는 모델 구성 요소의 작은 집합을 식별하며, 추출된 회로는 종종 해당 동작의 메커니즘으로 해석된다. 우리는 이러한 해석이 모호하다는 점에 주목한다. 즉, 특정 동작을 유지한다고 해서 하나의 회로만이 결정되는 것은 아니기 때문이다. 왜냐하면 회로 추출 주장의 타당성은 어떤 회로가 보고되었는지 및 두 회로가 어떻게 비교되는지에 따라 달라지기 때문이다. 우리는 이를 인공적인 Lean 전술 예측 벤치마크를 통해 구체적으로 설명한다. 이 벤치마크에서는 고정된 추론 규칙과 무작위 표면 형식을 사용하여, 추출된 회로 간의 차이를 작업 자체보다는 선택에 의한 것임을 보여준다. 동일한 트랜스포머 모델에서 밀집형 및 희소 가중치 체크포인트를 사용하고, 원자적(단일 규칙) 및 합성적(다중 규칙) 추론을 평가하면서, 보고되는 추출 객체를 다양하게 변화시켰다 (예: 예측을 유지하는 간결한 회로, 주변의 읽기/쓰기/라우팅 구조를 포함하는 더 넓은 그래프, 또는 ablation 후 손실 임계값을 충족하는 가장 작은 부분 그래프). 또한 각 어텐션 헤드의 쿼리와 키가 함께 표현되는지 분리되어 표현되는지를 변경했다. 정확한 구성 요소 간의 연결 일치율은 낮으며 이러한 선택에 민감하게 반응하며, 때로는 무작위 기준선까지 떨어지기도 한다. 그러나 두 가지 더 일반적인 지표는 안정적으로 유지된다: 선택된 어텐션 헤드의 집합과 강화 학습(RL) 초기화에 사용되는 체크포인트에 따라 회로 크기 순위를 나타내는 조건. 합성적 추론에서 RL을 사용하여 얻은 가장 큰 정확도 향상은 원자 회로보다 더 많은 구조를 포함할 때 나타났다. 따라서 회로 수준의 주장은 어떤 회로가 보고되었는지, 회로 추출에 사용된 가지치기 임계값 및 회로 비교 수준을 명시했을 때만 명확하게 정의될 수 있다. 우리는 이러한 요구 사항을 회로 추출 연구를 위한 보고 방법으로 구체화한다.

Original Abstract

Circuit extraction identifies a small set of model components whose presence preserves a target behavior under ablation, and the resulting circuit is often read as the mechanism behind that behavior. We argue that this reading is under-determined: preserving behavior does not single out one circuit, because the claim it supports depends on which circuit is reported and how two circuits are compared. We make this concrete in a synthetic Lean tactic-prediction benchmark -- predicting the next step of a proof -- where fixed proof rules with randomized surface form let differences between extracted circuits be attributed to these choices rather than to the task. Across dense and weight-sparse checkpoints (most weights constrained to zero) of the same transformer, evaluated on atomic (single-rule) and compositional (multi-rule) proofs, we vary which extracted object is reported (a compact prediction-preserving circuit, a broader graph that also keeps surrounding read, write, and routing structure, or the smallest subgraph meeting a post-ablation loss threshold), and whether each attention head's query and key are represented jointly or separately. Exact component-to-component edge overlap is low and sensitive to these choices, at times dropping to a random baseline, while two coarser summaries stay stable: the set of selected attention heads, and the circuit-size ranking of conditions that differ in which supervised checkpoint initializes reinforcement learning (RL). The largest accuracy gains from RL on compositional proofs come with the most structure beyond the atomic circuits. A circuit-level claim is therefore well defined only once one states which circuit is reported, the pruning threshold used to extract it, and the level at which circuits are compared. We distill these requirements into a reporting practice for circuit-extraction studies.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!