2608.03892v1 Aug 04, 2026 cs.AI

Qwen3 모델의 대조 학습 기반 활성화 추가를 통한 시간적 선호도 제어

Intertemporal Preference Steering in Qwen3 via Contrastive Activation Addition

Justin Shenk
Justin Shenk
Citations: 67
h-index: 4
Michal Mráz
Michal Mráz
Citations: 0
h-index: 0

본 연구에서는 Qwen3-32B라는 거대 언어 모델에서 시간 지평선에 대한 선형 표현을 분석하고, 이를 활용하여 모델의 시간 관련 선호도, 추천 및 능력을 변화시킵니다. 우리는 가이드된 답변을 기반으로 대조 학습 방식으로 선형 프로브를 훈련시켜 모델의 잔류 흐름 내에서 단기적인 것과 장기적인 것 사이의 방향성을 파악하고, 별도의 데이터셋인 이진 시간 선택 작업, 외부 데이터셋 기반의 금전적 시간 선택 작업 및 TravelPlanner 능력 평가 벤치마크를 사용하여 대조 학습 기반 활성화 추가 방식을 통해 모델을 제어하는 효과를 평가합니다. 주요 결과는 시간 지평선 방향성을 간단한 대조 학습 기반 선형 프로브를 통해 파악할 수 있으며, 이를 활용하여 큰 규모의 양방향 선호도 변화를 유도할 수 있다는 것입니다. 특히, 보상 크기와 지연 시간이 다른 외부 데이터셋 기반 금전적 선택 작업에서 모델의 무관심 임계값을 작은 것-빠른 것과 큰 것-나중 것 사이에서 양방향으로 크게 변화시킬 수 있었습니다. 또한, 적절한 수준의 시간 제어를 통해 계획 관련 능력 지표에서 개선 효과를 확인했습니다. 이러한 결과는 모델의 시간적 선호도가 측정 가능하고 제어 가능하다는 점을 시사하며, 이는 지연된 비용과 이점을 포함하는 조언을 제공하는 AI 시스템 및 장기적인 계획에 대한 안전 문제를 고려할 때 중요한 의미를 갖습니다.

Original Abstract

We study linear representations of temporal horizon in the large language model Qwen3-32B and use them to change the model's time-related preferences, recommendations, and capabilities. We train contrastive linear probes on teacher-forced temporal-choice answers to find a short-term versus long-term direction in the model's residual stream, and evaluate contrastive activation-addition steering on a held-out binary temporal-choice task, an out-of-distribution monetary intertemporal-choice task, and a TravelPlanner capability benchmark. The central result is that temporal-horizon directions can be identified with simple contrastive linear probes and then used for steering to induce large, bidirectional preference changes. On an out-of-distribution monetary choice task that varies reward size and delay, steering strongly shifts the model's indifference threshold between smaller-sooner and larger-later rewards in both directions. We further show improvements on a planning-related capability metric under moderate temporal steering. These results suggest that model intertemporal preferences are measurable and steerable, which is relevant for AI systems that give advice involving delayed costs and benefits, and for safety questions about long-horizon planning.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!