2606.12818v1 Jun 11, 2026 cs.CL

언어 모델 내 고정 효과 유발 경로 분석

Localizing Anchoring Pathways in Language Models

Sarah Wiegreffe
Sarah Wiegreffe
University of Maryland
Citations: 6,780
h-index: 16
Hillary N. Owusu
Hillary N. Owusu
Citations: 0
h-index: 0
Naomi H. Feldman
Naomi H. Feldman
Citations: 5
h-index: 1

프롬프트 내 무관한 숫자는 언어 모델의 판단에 영향을 미쳐 수리 추론에서 고정 효과를 발생시킵니다. 본 연구에서는 공유된 답변 옵션을 갖는 통제된 객관식 설정을 통해, 언어 모델 내부에서 어떤 부분이 이러한 고정 효과에 민감하게 반응하는지 분석합니다. 정답 옵션과 고정 숫자에 해당하는 옵션 간의 로짓 차이를 비교하는 지표를 정의하고, 이 지표가 실제 행동적 고정 현상을 잘 반영하는지 검증합니다. 7B-8B Qwen 및 Llama 모델의 기본 모델과 Instruction-tuned 모델에 대해 회로 추적 기법을 적용한 결과, edge 레벨 방법이 node 레벨 방법에 비해 더 정확하게 이 신호를 복원함을 확인했습니다. 낮은 고정과 높은 고정 사이의 회로는 모델 내에서 강하게 연결되어 있어, 고정 방향에 관계없이 공유된 경로 구조를 가지고 있음을 시사합니다. 그러나 기본 모델과 Instruction-tuned 모델 간의 희소한 연결성은 덜 안정적이며, 이는 추가 학습 과정이 어떤 경로가 가장 중요한지에 영향을 미친다는 것을 나타냅니다. 전반적으로 본 연구 결과는 언어 모델 내부에서 고정과 관련된 의사 결정 신호가 어떻게 전달되는지에 대한 메커니즘적인 설명을 제공합니다.

Original Abstract

Irrelevant numbers in a prompt can shift language model judgments, producing anchoring effects in numerical reasoning. We study where this anchor-sensitive signal is carried inside language models using a controlled multiple-choice setup with shared answer options. We define a logit-difference metric comparing the correct answer option with the answer option corresponding to the anchor, and validate that it tracks behavioral anchoring. Using attribution-based circuit localization on 7B--8B Qwen and Llama base and instruction-tuned models, we find that edge-level methods recover this signal more faithfully than node-level methods. Low- and high-anchor circuits transfer strongly within a model, suggesting shared pathway structure across anchor direction. However, sparse transfer across base and instruction-tuned variants is less reliable, indicating that post-training changes which pathways matter most. Overall, our results provide a mechanistic account of how anchoring-related decision signals are carried inside language models.

0 Citations
0 Influential
8 Altmetric
40.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!