2605.01555v1 May 02, 2026 cs.CL

에이전트를 활용한 언어 모델의 자동화된 해석 가능성 및 특징 발견

Automated Interpretability and Feature Discovery in Language Models with Agents

Javier Ferrando
Javier Ferrando
Citations: 777
h-index: 15
Arnau Marin-Llobet
Arnau Marin-Llobet
Citations: 51
h-index: 4

본 논문에서는 대규모 언어 모델의 메커니즘적 해석 가능성을 위한 자율적인 다중 에이전트 프레임워크를 소개합니다. 이 시스템은 두 가지 주요 기능을 자동화합니다: (1) 설명 개선, 여기서 에이전트는 경쟁적인 가설을 제안하고, 표적 프롬프트 제어 및 다중 지표 평가를 통해 반복적으로 이를 검증합니다; (2) 특징 발견, 여기서 에이전트는 프롬프트 세트를 생성하고, 활성화 공간에서 k-최근접 이웃 그래프를 구성하며, 통계적 분리 가능성 및 의미적 일관성 기준을 사용하여 후보 특징을 추출합니다. Gemma-2 모델 패밀리와 가중치 희소 트랜스포머의 MLP 뉴런에 대한 실험 결과, 저희의 에이전트는 단일 해석 방식보다 성능이 우수하며, 언어별 및 안전 관련 특징을 발견하고, 감사 가능한 설명 추적을 제공합니다. 이는 에이전트 기반의 경험적 루프가 단일 라벨 방식보다 더 정확하고 검증 가능한 설명을 생성할 수 있음을 보여줍니다.

Original Abstract

We introduce an autonomous multiagent framework for mechanistic interpretability that automates both explaining and finding internal features in large language models. The system runs two coupled loops: (1) explanation refinement, where an agent proposes competing hypotheses and iteratively tests them with targeted prompt controls and a multi-metric evaluation; and (2) feature discovery, where an agent generates prompt sets, constructs a k-nearest-neighbor graph in activation space, and retrieves candidate features using statistical separability and semantic coherence criteria. On Gemma-2 family models and MLP neurons in weight-sparse transformers, our agent improves over one-shot auto-interpretations, discovers language-specific and safety-relevant features, and produces auditable explanation traces, showing that agent-driven empirical loops yield sharper and more falsifiable explanations than one-shot labels.

1 Citations
0 Influential
7.5 Altmetric
38.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!