2607.05790v1 Jul 07, 2026 cs.AI

헤딩 정보 기반 활성화 제어를 통한 도구 사용 제어

Controlling Tool Use with Heading-Specific Activation Steering

D. Song
D. Song
Citations: 109
h-index: 6
Vincent Siu
Vincent Siu
Citations: 104
h-index: 6
Yuqiang Chen
Yuqiang Chen
Citations: 0
h-index: 0
Yang Liu
Yang Liu
Citations: 44
h-index: 4
Chenguang Wang
Chenguang Wang
Citations: 120
h-index: 7

도구 활용 대규모 언어 모델은 파라미터 지식을 넘어 외부 도구를 통해 기능을 확장하지만, 종종 불필요하게 도구를 호출하는 경향이 있습니다. 본 연구에서는 도구 사용 결정에 안정적인 내부 표현이 존재하는지, 즉 추출 및 조작이 가능한지를 조사합니다. 이는 도구가 추론 시점에 컨텍스트 내에서만 존재하고 모델 가중치에 직접적으로 인코딩되지 않기 때문에 매우 어려운 질문입니다. 우리는 헤딩 앵커 위치에서 추출된 스티어링 벡터가 다섯 가지 오픈 소스 모델과 세 가지 영역 전반에 걸쳐 도구 호출 행동에 양방향적인 인과적 제어를 행사하며, 파라미터 추론만으로 충분한 영역에서는 불필요한 도구 사용을 가장 효과적으로 억제한다는 것을 보여줍니다. 그러나 기하학적 분석 결과, 이러한 인과적 효과가 명확한 선형 구조와 일치하지 않는다는 사실이 밝혀졌습니다. 도구 호출 단계는 억제 벡터와의 일관된 부정적인 정렬보다는 확산되고 이중 모드를 보이는 경향을 보이며, 이는 선형 인코딩 모델에서 예측되는 패턴과 다릅니다. 또한, 서로 다른 유형의 도구는 대부분 뚜렷하게 구별되는 내부 특징을 사용하며, 이러한 특징 간에는 낮은 중복성을 나타냅니다. 우리는 이러한 기하학적 특성이 도구의 비파라미터적인 성격을 반영하는 것이라고 가정하고, 도구 사용 스티어링 벡터를 파라미터 기반 개념에 대한 스티어링 벡터와 구별합니다. 이러한 기하학적 불규칙성과 관찰된 인과적 효과 간의 관계는 여전히 연구 과제로 남아 있습니다.

Original Abstract

Tool-augmented large language models extend their capabilities beyond parametric knowledge through external tools, but tend to invoke them unnecessarily. We investigate whether tool-use decisions have any stable internal representation that can be extracted and manipulated, a question that is non-trivial given that tools exist entirely in context at inference time and have no direct encoding in model weights. We show that steering vectors extracted from heading-anchors positions exert bidirectional causal control over tool-invocation behavior across five open-source models and three domains, suppressing unnecessary tool use most effectively in domains where parametric reasoning suffices. However, geometric analysis reveals that this causal effectiveness does not correspond to clean linear structure: tool-invocation steps exhibit diffuse, bimodal alignment with the suppression vector rather than the consistent negative alignment a linear encoding account would predict, and different tool types recruit largely distinct internal signatures with low cross-tool feature overlap. We hypothesize these geometric properties are indicative of the non-parametric nature of tools, and distinguish tool-use steering vectors from those extracted for parametrically grounded concepts. The relationship between this geometric irregularity and the observed causal effectiveness remains an open question.

1 Citations
0 Influential
3.5 Altmetric
18.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!