표현에서 행동으로: LLM에서의 개인-상황-행동 삼위일체 탐구
From Representations to Behaviors: Exploring the Person-Situation-Behavior Triad in LLMs
인간의 성격 이론은 특성을 단일 점수가 나타내는 고립된 속성이 아니라, 개인, 상황 및 행동 간의 상호 작용을 통해 표현되는 안정적인 개별적 경향성으로 규정합니다. LLM에서 성격 관련 행동에 대한 기존 연구는 주로 성격 조건 하에서 생성된 결과물을 분석하며, 관찰 가능한 성격 관련 표현을 특징짓지만, 내부적으로 존재하는 성격 관련 표현의 존재 여부, 상황 간의 일관성 및 이러한 표현이 특정 행동을 어떻게 형성하는지에 대한 메커니즘적 증거는 부족합니다. 펀더(Funder)의 성격 삼위일체 프레임워크를 바탕으로, LLM 분석에 적합하도록 세 가지 구성 요소를 다음과 같이 적용했습니다. '개인'은 성격 관련 내부 표현을 의미하며, '상황'은 성격과 관련된 반응을 가능하게 하는 맥락을 의미하고, '행동'은 다양한 사회적 과제에서의 응답 패턴을 의미합니다. 본 연구는 LLM에서 특성과 유사한 표현을 발견, 제어 및 검증하기 위한 프레임워크를 제시합니다. 첫째, 공유된 상황에 기반한 대조적인 행동 쌍을 사용하여, SAE 분해(Spectral Analysis for Event decomposition)를 통해 반대되는 성격 극단과 관련된 희소한 내부 특징을 식별합니다. 이러한 특징이 실제로 특성 관련성을 가지는지 행동 변화, 토큰 수준의 활성화 패턴 및 패러프레이징에 대한 견고성 측면에서 검증합니다. 둘째, 특징 수준에서의 개입은 다양한 상황에서 양방향으로 성격 관련 변화를 유도하며, 응답의 타당성을 유지하여 문맥 간 일관성을 보여줍니다. 셋째, 동일한 개입을 사회적 지능 과제에 적용하면 인간 성격 연구 결과와 일치하는 이점-희생 패턴을 보이는 행동 변화가 나타나며, 이는 단순히 성격 점수 이상의 행동 수준에서의 검증을 제공합니다. 본 연구의 결과는 LLM이 내부 상태, 상황적 표현 및 행동 결과를 연결하는 제어 가능한 특성 유사한 표현을 포함하고 있음을 입증합니다.
Human personality theories characterize traits not as isolated attributes captured by a single score, but as stable individual tendencies expressed through the interplay among persons, situations, and behaviors. Existing studies of personality-related behavior in LLMs have primarily focused on outputs elicited under personality conditioning, characterizing observable trait-related expressions while lacking mechanistic evidence for the existence of internal personality-related representations, their cross-situational expression, and how these representations shape specific behaviors. Building on Funder's personality triad framework, we adapt its three components for LLM analysis: Person as personality-related internal representations, Situation as contexts that afford trait-relevant responses, and Behavior as response patterns on broader social tasks. We introduce a framework for discovering, controlling, and validating trait-like representations in LLMs. First, using contrastive behavior pairs grounded in shared situations, we identify sparse internal features associated with opposing poles of personality traits through SAE decomposition. We validate their trait relevance through effects on behavior to situation, token-level activation patterns, and robustness to paraphrasing. Second, feature-level interventions induce bidirectional trait-related shifts across a separate, diverse set of situations while preserving response validity, demonstrating consistent expression across contexts. Third, applying the same interventions to social intelligence tasks reveals behavioral changes with benefit-tradeoff patterns consistent with findings from human personality research, providing behavioral-level validation beyond personality scores. Our findings provide evidence that LLMs contain controllable trait-like representations linking internal states, situational expression, and behavioral outcomes.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.