2604.10031v1 Apr 11, 2026 cs.CL

CoSToM: 인과관계 지향적 제어를 통한 대규모 언어 모델의 내재적 정신 이론 정렬

CoSToM:Causal-oriented Steering for Intrinsic Theory-of-Mind Alignment in Large Language Models

Yang Deng
Yang Deng
Citations: 6
h-index: 1
Mengfan Li
Mengfan Li
Citations: 32
h-index: 3
Xuanhua Shi
Xuanhua Shi
Citations: 69
h-index: 5

정신 이론(Theory of Mind, ToM)은 타인의 정신 상태를 추론하는 능력으로, 사회적 지능의 중요한 특징입니다. 대규모 언어 모델(LLM)은 표준적인 ToM 평가에서 유망한 성능을 보이지만, 복잡한 특정 작업 시나리오에서는 종종 일반화에 실패하며, 추론을 모방하기 위해 프롬프트 지시(prompt scaffolding)에 크게 의존하는 경향이 있습니다. 내부 지식과 외부 행동 간의 이러한 근본적인 불일치는 다음과 같은 중요한 질문을 제기합니다: LLM이 진정으로 내재적인 인지 능력을 가지고 있으며, 이러한 내부 지식을 안정적이고 고품질의 행동으로 외부화할 수 있는가? 이에 대한 답변을 얻기 위해, 우리는 CoSToM(Causal-oriented Steering for ToM alignment)이라는 프레임워크를 제안합니다. CoSToM은 기계적인 해석에서 적극적인 개입으로 전환하는 방식입니다. 먼저, 인과적 추적을 사용하여 ToM 특징의 내부 분포를 매핑하고, 이를 통해 기본 ToM 의미를 인코딩하는 내부 계층의 특성을 경험적으로 밝힙니다. 이러한 통찰력을 바탕으로, 우리는 ToM에 중요한 계층 내에서 특정 활성화 제어(activation steering)를 통해 경량화된 정렬 프레임워크를 구현합니다. 실험 결과, CoSToM은 인간과 유사한 사회적 추론 능력과 대화 품질을 크게 향상시키는 것으로 나타났습니다.

Original Abstract

Theory of Mind (ToM), the ability to attribute mental states to others, is a hallmark of social intelligence. While large language models (LLMs) demonstrate promising performance on standard ToM benchmarks, we observe that they often fail to generalize to complex task-specific scenarios, relying heavily on prompt scaffolding to mimic reasoning. The critical misalignment between the internal knowledge and external behavior raises a fundamental question: Do LLMs truly possess intrinsic cognition, and can they externalize this internal knowledge into stable, high-quality behaviors? To answer this, we introduce CoSToM (Causal-oriented Steering for ToM alignment), a framework that transitions from mechanistic interpretation to active intervention. First, we employ causal tracing to map the internal distribution of ToM features, empirically uncovering the internal layers' characteristics in encoding fundamental ToM semantics. Building on this insight, we implement a lightweight alignment framework via targeted activation steering within these ToM-critical layers. Experiments demonstrate that CoSToM significantly enhances human-like social reasoning capabilities and downstream dialogue quality.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!