2603.18474v1 Mar 19, 2026 cs.CL

WASD: LLM의 행동을 설명하고 제어하기 위한 충분 조건으로서의 핵심 뉴런 위치 파악

WASD: Locating Critical Neurons as Sufficient Conditions for Explaining and Controlling LLM Behavior

Haonan Yu
Haonan Yu
Citations: 15
h-index: 2
Junhao Liu
Junhao Liu
Citations: 13
h-index: 2
Zhenyu Yan
Zhenyu Yan
Citations: 4
h-index: 1
Xin Zhang
Xin Zhang
Citations: 7
h-index: 1
Haoran Lin
Haoran Lin
Citations: 12
h-index: 2

대규모 언어 모델(LLM)의 정밀한 행동 제어는 복잡한 응용 분야에서 매우 중요합니다. 그러나 기존 방법들은 종종 높은 훈련 비용이 발생하거나, 자연어 제어 가능성이 부족하거나, 의미적 일관성을 저해하는 단점이 있습니다. 이러한 격차를 해소하기 위해, 우리는 모델의 행동을 토큰 생성에 대한 충분한 신경 조건으로 식별하여 설명하는 새로운 프레임워크인 WASD(unWeaving Actionable Sufficient Directives)를 제안합니다. 우리의 방법은 후보 조건을 뉴런 활성화 예측 조건으로 표현하고, 입력 변동 하에서 현재 출력을 보장하는 최소 집합을 반복적으로 검색합니다. SST-2 및 CounterFact 데이터셋에 Gemma-2-2B 모델을 사용하여 실험한 결과, 우리의 방법이 기존의 설명 기법인 어트리뷰션 그래프보다 더 안정적이고 정확하며 간결한 설명을 생성한다는 것을 확인했습니다. 또한, 다국어 출력 생성을 제어하는 사례 연구를 통해 WASD가 모델의 행동을 제어하는 데 실제로 효과적임을 검증했습니다.

Original Abstract

Precise behavioral control of large language models (LLMs) is critical for complex applications. However, existing methods often incur high training costs, lack natural language controllability, or compromise semantic coherence. To bridge this gap, we propose WASD (unWeaving Actionable Sufficient Directives), a novel framework that explains model behavior by identifying sufficient neural conditions for token generation. Our method represents candidate conditions as neuron-activation predicates and iteratively searches for a minimal set that guarantees the current output under input perturbations. Experiments on SST-2 and CounterFact with the Gemma-2-2B model demonstrate that our approach produces explanations that are more stable, accurate, and concise than conventional attribution graphs. Moreover, through a case study on controlling cross-lingual output generation, we validated the practical effectiveness of WASD in controlling model behavior.

0 Citations
0 Influential
1 Altmetric
5.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!