에이전트를 활용한 언어 모델의 자동화된 해석 가능성 및 특징 발견
Automated Interpretability and Feature Discovery in Language Models with Agents
본 논문에서는 대규모 언어 모델의 메커니즘적 해석 가능성을 위한 자율적인 다중 에이전트 프레임워크를 소개합니다. 이 시스템은 두 가지 주요 기능을 자동화합니다: (1) 설명 개선, 여기서 에이전트는 경쟁적인 가설을 제안하고, 표적 프롬프트 제어 및 다중 지표 평가를 통해 반복적으로 이를 검증합니다; (2) 특징 발견, 여기서 에이전트는 프롬프트 세트를 생성하고, 활성화 공간에서 k-최근접 이웃 그래프를 구성하며, 통계적 분리 가능성 및 의미적 일관성 기준을 사용하여 후보 특징을 추출합니다. Gemma-2 모델 패밀리와 가중치 희소 트랜스포머의 MLP 뉴런에 대한 실험 결과, 저희의 에이전트는 단일 해석 방식보다 성능이 우수하며, 언어별 및 안전 관련 특징을 발견하고, 감사 가능한 설명 추적을 제공합니다. 이는 에이전트 기반의 경험적 루프가 단일 라벨 방식보다 더 정확하고 검증 가능한 설명을 생성할 수 있음을 보여줍니다.
We introduce an autonomous multiagent framework for mechanistic interpretability that automates both explaining and finding internal features in large language models. The system runs two coupled loops: (1) explanation refinement, where an agent proposes competing hypotheses and iteratively tests them with targeted prompt controls and a multi-metric evaluation; and (2) feature discovery, where an agent generates prompt sets, constructs a k-nearest-neighbor graph in activation space, and retrieves candidate features using statistical separability and semantic coherence criteria. On Gemma-2 family models and MLP neurons in weight-sparse transformers, our agent improves over one-shot auto-interpretations, discovers language-specific and safety-relevant features, and produces auditable explanation traces, showing that agent-driven empirical loops yield sharper and more falsifiable explanations than one-shot labels.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.