2607.20803v1 Jul 23, 2026 cs.CL

성격의 기하학: 융의 인지 기능을 활용한 활성화 제어

The Geometry of Personality: Activation Steering with Jungian Cognitive Functions

Junchen Fu
Junchen Fu
University of Glasgow
Citations: 708
h-index: 8
Liu Zai
Liu Zai
Citations: 134
h-index: 5
Joemon M. Jose University of Glasgow
Joemon M. Jose University of Glasgow
Citations: 0
h-index: 0
Leiden University
Leiden University
Citations: 0
h-index: 0
Yumeng Wang
Yumeng Wang
Citations: 0
h-index: 0

활성화 제어는 LLM을 제어하고 해석하는 데 유용하지만, 기존 연구에서는 주로 빅 파이브(Big Five)와 같은 정적인 특성 프레임워크를 사용하여 성격을 모델링합니다. 본 연구에서는 성격이 융의 여덟 가지 인지 기능을 활용한 일련의 인지 과정으로 표현되고 제어될 수 있는지 조사합니다. 이를 위해, 우리는 융 평가 프로토콜과 2,100개 이상의 역할극 캐릭터 내러티브 데이터 세트를 포함하는 프레임워크를 소개합니다. Llama-3.1-8B 모델에 대한 활성화 제어를 통한 벡터 추출 및 평가 실험 결과, 활성화 제어를 통해 여덟 가지 인지 기능 모두에 대해 효과적인 단조적인 제어가 가능함을 확인했습니다. 또한, 제어 가능성 외에도, 분석 결과 다음과 같은 사실이 밝혀졌습니다: 1. 성격 정보는 중간 Transformer 레이어에 집중되어 있습니다; 2. 제어 벡터는 합리적(rational) 기능과 비합리적(irrational) 기능 간의 구분에 일관된 구조적인 기하학적 관계를 나타냅니다; 3. 효과적인 다차원 제어 방향은 개별 기능 방향의 선형 조합으로 복구할 수 없습니다. 이러한 결과는 LLM 활성화 공간에서의 성격 표현에 대한 새로운 통찰력을 제공하며, 해석 가능하고 효과적이며 다차원적인 성격 제어를 연구하기 위한 프레임워크를 제시합니다.

Original Abstract

Activation steering enables control and interpretation of LLMs, yet existing work primarily models personality through static trait frameworks such as the Big Five. We investigate whether personality can instead be represented and controlled as a set of cognitive processes using the eight Jungian Cognitive Functions. To this end, we introduce a framework comprising a Jungian evaluation protocol and a dataset of over 2,100 role-playing character narrations. Activation steering vector extraction and evaluation experiments on Llama-3.1-8B demonstrate effective monotonic control over all eight cognitive functions through activation steering. Beyond controllability, our analysis reveals that: 1. personality information is concentrated in middle transformer layers; 2. steering vectors exhibit structured geometric relationships consistent with distinctions between rational and irrational functions; 3. effective multi-dimensional steering directions cannot be recovered as linear combinations of single-function directions. These findings provide new insights into the representation of personality in LLM activation space and establish a framework for studying interpretable, effective, and multi-dimensional personality control.

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!