2606.26474v1 Jun 25, 2026 cs.LG

강화 학습 기반 도구 사용 능력을 단일 크로스코드 특징으로 제한

Localizing RL-Induced Tool Use to a Single Crosscoder Feature

Jessica Hullman
Jessica Hullman
Citations: 26
h-index: 2
A. Shportko
A. Shportko
Citations: 13
h-index: 3
Shubham Bhokare
Shubham Bhokare
Citations: 16
h-index: 1
Ahmed Zeyad A Alzahrani
Ahmed Zeyad A Alzahrani
Citations: 0
h-index: 0
Bowen Cheng
Bowen Cheng
Citations: 4,334
h-index: 3
G. Mercier
G. Mercier
Citations: 12
h-index: 2

강화 학습을 통한 미세 조정은 언어 모델의 내부 표현을 재구성하여 에이전트 행동, 특히 도구 사용 능력을 가능하게 하지만, 이러한 변화의 메커니즘적 기반은 아직 제대로 이해되지 못하고 있습니다. 강화 학습은 구조화된 도구 호출 생성 능력을 크게 향상시키지만, 어떤 특징들이 나타나는지, 어떤 특징들이 보존되는지, 그리고 식별된 특징들을 재학습 없이 행동 제어를 위해 활용할 수 있는지에 대한 명확한 정보는 부족합니다. 본 연구에서는 $ extit{전용 특징 크로스코드(DFC)}$가 $ exttt{Qwen2.5-3B}$ 모델에서 도구 호출 능력을 매개하는 일련의 작고 핵심적인 강화 학습 관련 특징들을 분리한다는 것을 보여줍니다. 48개의 크로스코드 하이퍼파라미터 조합을 사용하여 인코더-디코더 재구성 과정을 통해, 강화 학습 모델의 도구 정확도가 $+31.1 ext{ pp} ext{ } leq {9.7}$만큼 향상되었으며, 동시에 동결된 기본 모델에 도구 호출 능력이 $+6.8 ext{ pp} ext{ } leq 5.0$만큼 수동적으로 전달되는 현상, 즉 $ extit{능력 유출(capability spillover)}$이 발생했습니다. 이러한 결과는 DFC 분할을 통해 강화 학습으로 인해 도입된 기능이 최소한의, 제어 가능한 특징 집합으로 집중되어 있으며, 이를 통해 에이전트 역할을 수행하는 언어 모델의 런타임 행동 제어가 가능함을 보여줍니다.

Original Abstract

Fine-tuning through RL reshapes the internal representations of language models to enable agentic behaviors such as tool use, yet the mechanistic basis of these changes remains poorly understood. While RL substantially improves structured tool-call generation, it is unclear which features emerge, which are preserved, and whether identified features can be leveraged for retraining-free behavioral control. In this work, we show that $\textit{Dedicated Feature Crosscoders (DFC)}$ isolate a compact set of RL-specific features that mediate tool-calling capability in $\texttt{Qwen2.5-3B}$. Across a $48$-crosscoder hyperparameter sweep, encode-decode reconstruction improves the RL model's tool correctness by $+31.1 \pm {9.7}$ pp and passively transfers tool-calling ability to the frozen base model by $+6.8 \pm 5.0$ pp which we call a $\textit{capability spillover}$. Our findings show that DFC partitioning concentrates RL-introduced capability into a minimal, steerable feature set that enables runtime behavioral control of agentic LLMs.

0 Citations
0 Influential
1.5 Altmetric
7.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!