2608.05732v1 Aug 06, 2026 cs.LG

CircuitSteer: 희소 오토인코더 회로를 이용한 기하학적으로 정렬된 다층 제어 방법

CircuitSteer: Geometrically Aligned Multi-Layer Steering via Sparse Autoencoder Circuits

Parsa Razmara
Parsa Razmara
Citations: 108
h-index: 5
Seyedarmin Azizi
Seyedarmin Azizi
University of Southern California;Sharif University of Technology
Citations: 161
h-index: 7
Mehrshad Saadatinia
Mehrshad Saadatinia
Citations: 15
h-index: 1
Ardalan Aryashad
Ardalan Aryashad
Citations: 1
h-index: 1
Ali Abbasi
Ali Abbasi
Citations: 132
h-index: 6

대규모 언어 모델(LLM)의 행동을 제어하는 것은 AI 정렬을 위한 중요한 과제입니다. 기존의 제어 방법, 예를 들어 Contrastive Activation Addition (CAA)은 일반적으로 집계 활성화 차이에서 파생된 고정된 단일층 개입에 의존합니다. 이러한 방법은 의미적으로 다양한 입력에 대해 단일 개입을 적용하며, 종종 계층 전반에 걸쳐 일관된 행동 변화를 유지하지 못하여 제어의 효과를 제한합니다. 본 연구에서는 희소 오토인코더(SAE)를 활용하여 여러 계층에 분산된 응집적인 의미 회로를 식별하고 조작하는 새로운 프레임워크인 CircuitSteer를 소개합니다. 특징 공기성 및 디코더 방향의 기하학적 정렬을 기반으로 특징 흐름 회로를 구성함으로써, 특정 행동에 책임이 있는 특정 다층 하위 회로를 분리합니다. 그런 다음 이러한 희소 특징에서 밀집된 제어 벡터를 합성하고, 모델의 내부 의미 경로를 안내하기 위해 다중 지점 개입을 적용합니다. CircuitSteer는 독성, 감정 강도, 아첨, 거부 등 다양한 작업에 걸쳐 대조적인 예제를 사용하여 평가되었습니다. 모든 모델과 데이터 세트에서 CircuitSteer는 일관되게 유창성을 유지하는 제어 방법을 제공하며, 경쟁 방법은 텍스트 품질을 희생하거나 적용 범위가 부족하여 아첨 및 거부와 같은 복잡한 행동에서는 완전히 실패합니다. 이러한 결과는 선택된 특징 간의 기하학적 정렬을 강제함으로써 가능하게 된 다층 회로 제어가, 고정된 단일 지점 개입보다 훨씬 더 강력하고 효과적인 행동 제어를 제공한다는 것을 보여줍니다. 코드: https://github.com/mehrshad-sdtn/CircuitSteer

Original Abstract

Controlling the behavior of large language models (LLMs) remains a critical challenge for AI alignment. Existing steering methods, such as Contrastive Activation Addition (CAA), typically rely on fixed single-layer interventions derived from aggregate activation differences. These methods impose a single intervention across semantically diverse inputs and often fail to sustain consistent behavioral changes across layers, limiting the effectiveness of the steering. In this work, we introduce CircuitSteer, a novel framework that leverages Sparse Autoencoders (SAEs) to identify and manipulate coherent semantic circuits distributed across multiple layers. By constructing a feature flow circuit based on feature co-activation and the geometric alignment of decoder directions, we isolate the specific multi-layer subcircuits responsible for a target behavior. We then synthesize dense steering vectors from these sparse features and apply multi-point interventions to guide the model's internal semantic trajectory. We evaluate CircuitSteer using contrastive examples across a diverse set of tasks, including toxicity, emotion-intensity, sycophancy, and refusal, spanning two model families. Across all models and datasets, CircuitSteer is the only method to consistently produce fluency-preserving interventions; competing methods either sacrifice text quality or lack coverage, failing entirely on complex behaviors like sycophancy and refusal. These results demonstrate that multi-layer circuit steering, enabled by enforcing geometric alignment among selected features, yields strictly more robust and effective behavioral control than static single-point interventions. Code is available at https://github.com/mehrshad-sdtn/CircuitSteer.

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!