방향 신호는 어디에서 오는가: 활성화 제어(Activation Steering)에서의 활성화 소스 선택
Where Steering Signals Come From: Activation Source Selection in Activation Steering
활성화 제어는 추론 시점에 벡터 또는 특징을 은닉 상태에 추가하여 언어 모델을 제어하는 기술이지만, 이러한 제어 신호의 상위 수준(upstream) 정보원은 종종 부수적인 세부 사항으로 간주됩니다. 본 연구에서는 이 정보원의 선택을 '활성화 소스 선택'이라고 정의하고, 이를 통해 제어 신호를 구성하는 데 사용되는 소스 컨텍스트와 활성화 판독 정책의 조합을 분석합니다. 하위 수준 개입은 고정된 상태에서, 세 가지 Instruction-tuned 모델과 네 가지 Steering Task를 사용하여 실험한 결과, 소스 활성화를 변경하는 것만으로도 제어 성공률에 상당한 변화가 있음을 확인했습니다. 또한, 효과적인 제어가 단순히 원하는 행동이 소스 텍스트에 나타나는 여부로 설명될 수 없다는 것을 발견했습니다. 대신, 강력한 신호는 모델이 목표 행동을 생성하거나 이어가기 직전의 '실행 경계 상태'에서 주로 발생합니다. 이러한 사전/사후 실현(pre-/post-realization) 구분이 답변 기반 소스가 때때로 효과적인 이유를 설명합니다. 왜냐하면 유용한 구성 요소는 단순히 목표 행동의 출현이 아니라 실행 경계 방향과 일치하기 때문입니다. 이 관점을 바탕으로, 본 연구에서는 '꼬리 제거(tail subtraction)' 기법을 제안합니다. 이는 경계 상태에서 공유되는 프롬프트 및 연속 의미를 제거하여 더 깨끗하고 안정적인 제어 신호를 얻는 방법입니다. 전반적으로, 본 연구 결과는 모델이 무엇을 하려고 하는지에 대한 표현에 의존하는 것이며, 단순히 이미 나타난 것에 의존하는 것이 아니라는 점을 시사합니다.
Activation steering controls language models by adding vectors or features to hidden states at inference time, but the upstream source of these steering signals is often treated as a secondary detail. We study this source choice as activation source selection: the combination of source context and activation readout policy used to collect the hidden states from which a steering signal is built. Holding the downstream intervention fixed, we show across three instruction-tuned models and four steering task families that changing only the source activations substantially changes steering success. We further find that effective steering is not explained simply by whether the desired behavior appears in the source text. Instead, strong signals come from execution-boundary states, where the model is about to produce or continue the target behavior. This pre-/post-realization distinction explains why answer-based sources sometimes work: their useful component aligns with execution-boundary directions rather than target appearance alone. Building on this view, we introduce tail subtraction, which removes shared prompt and continuation semantics from boundary states and yields cleaner, more stable steering signals. Overall, our results suggest that steering depends on representations of what the model is about to do, not merely on what has already appeared.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.