작업 수준의 제어 가능한 LLM을 위한 우수한 뉴런과 불량 뉴런 식별
Identifying Good and Bad Neurons for Task-Level Controllable LLMs
대규모 언어 모델(LLM)은 객관식 질문 답변 벤치마크에서 놀라운 성능을 보여주었지만, 그 방대한 규모의 뉴런에 내재된 복잡한 메커니즘은 여전히 불투명하며, LLM을 이해하고 제어하는 데 상당한 어려움을 야기합니다. 최근 연구들은 특정 능력에 대한 책임 뉴런을 식별하는 데 진전을 이루었지만, 이러한 능력 특화 방식은 여러 능력을 복합적으로 사용하는 작업 중심 시나리오에서는 비현실적입니다. 또한, 이러한 접근 방식은 작업 완료와 긍정적인 상관관계를 보이는 지원 뉴런에만 초점을 맞추고, 억제 역할을 하는 뉴런과 같은 다른 역할을 하는 뉴런을 간과하며, LLM의 우연적인 행동(즉, 진정한 이해가 아닌 우연히 정답을 맞추는 경우)으로 인해 뉴런의 중요도를 잘못 판단하는 문제를 야기합니다. 이러한 문제점을 해결하기 위해, 우리는 기능적 반대(functional antagonism)의 생물학적 원리를 LLM 뉴런 식별에 적용하는 새로운 작업 수준 LLM 이해 프레임워크인 NeuronLLM을 제안합니다. 핵심적인 아이디어는 작업 성능이 두 가지 반대 역할을 하는 뉴런에 의해 공동으로 결정된다는 것입니다. 즉, 작업 완료를 촉진하는 '우수한 뉴런'과 이를 억제하는 '불량 뉴런'이 있습니다. NeuronLLM은 우수한 뉴런과 불량 뉴런을 대비 학습하여 뉴런을 종합적으로 모델링하며, LLM의 우연적인 행동을 완화하기 위해 확장된 질문 세트를 활용합니다. 다양한 크기와 계열의 LLM에 대한 포괄적인 실험 결과, NeuronLLM이 기존 방법보다 네 가지 NLP 작업에서 우수한 성능을 보이며, LLM의 기능적 조직에 대한 새로운 통찰력을 제공합니다.
Large Language Models have demonstrated remarkable capabilities on multiple-choice question answering benchmarks, but the complex mechanisms underlying their large-scale neurons remain opaque, posing significant challenges for understanding and steering LLMs. While recent studies made progress on identifying responsible neurons for certain abilities, these ability-specific methods are infeasible for task-focused scenarios requiring coordinated use of multiple abilities. Moreover, these approaches focus only on supportive neurons that correlate positively with task completion, while neglecting neurons with other roles-such as inhibitive roles-and misled neuron attribution due to fortuitous behaviors in LLMs (i.e., correctly answer the questions by chance rather than genuine understanding). To address these challenges, we propose NeuronLLM, a novel task-level LLM understanding framework that adopts the biological principle of functional antagonism for LLM neuron identification. The key insight is that task performance is jointly determined by neurons with two opposing roles: good neurons that facilitate task completion and bad neurons that inhibit it. NeuronLLM achieves a holistic modeling of neurons via contrastive learning of good and bad neurons, while leveraging augmented question sets to mitigate the fortuitous behaviors in LLMs. Comprehensive experiments on LLMs of different sizes and families show the superiority of NeuronLLM over existing methods in four NLP tasks, providing new insights into LLM functional organization.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.