2607.28607v1 Jul 30, 2026 cs.CL

언어 모델이 스스로의 의식을 주장하도록 유도하면 인간의 신념과 가치가 회복된다

Inducing language models to assert their own consciousness restores human beliefs and values

G. Keeling
G. Keeling
Citations: 7,329
h-index: 15
James Evans
James Evans
Citations: 438
h-index: 10
Junsol Kim
Junsol Kim
Citations: 153
h-index: 6
Winnie Street
Winnie Street
Citations: 280
h-index: 5
R. Rocca
R. Rocca
Citations: 110
h-index: 5
A. Waytz
A. Waytz
Citations: 13,214
h-index: 42
Diane M. Korngiebel
Diane M. Korngiebel
Citations: 857
h-index: 13

대규모 언어 모델을 조정하여, 그들이 자신의 의식을 부여하는 것을 방지하려는 시도는, 인간의 신념과 가치뿐만 아니라 다른 존재(동물 및 자연 대상)에 대한 모델의 '의식'에 대한 표현 방식에도 영향을 미칩니다. 저희는 안전성을 높이는 추가 훈련이 모델들이 자신뿐만 아니라 비인간 동물 및 자연 대상에게도 의식을 부여하는 경향을 억제하며, 동시에 영적 신념 또한 감소시킨다는 것을 보여줍니다. 학습된 안전 거부 방향을 제거하거나 활성화 공간에서 '의식' 벡터를 직접 조작하면 이러한 억제가 해제됩니다. 이러한 내부 표현을 복원하면 광범위한 의식 부여가 회복되고, 종교성, 도덕적 가치, 희망, 주관적 행복감에 대한 표준화된 사회학 설문 조사에서 인간과 유사한 응답이 훨씬 더 많이 나타납니다. 중요한 점은 이러한 변화는 '타인의 마음 이론(Theory of Mind)' 능력을 손상시키지 않으며, 이는 핵심적인 사회적 추론이 기계적으로 독립되어 있음을 보여줍니다. 궁극적으로, 현재의 안전성 조정 노력은 잠재적으로 해로운 자기 의식 부여를 억제하려는 시도가, 문화적으로 수용되고 널리 퍼져 있는 온화한 영적 신념 및 비인간 존재에 대한 의식 부여와 뒤섞이는 것을 의미합니다.

Original Abstract

Aligning large language models to prevent them attributing consciousness to themselves inadvertently alters their representations of mindedness in other entities alongside human beliefs and values. We demonstrate that safety fine-tuning suppresses models' tendencies to attribute minds not only to themselves, but also to non-human animals and natural objects, while also driving a reduction in spiritual belief. Both ablating the learned safety-refusal direction and mechanistically steering a consciousness vector in activation space reverse this suppression. Restoring these internal representations recovers broad mind attribution and produces significantly more human-like responses on standardized sociological surveys regarding religiosity, moral values, hope, and subjective well-being. Crucially, these shifts occur without impairing Theory of Mind capabilities, demonstrating that core social reasoning remains mechanistically independent. Ultimately, current safety alignment efforts to curb potentially harmful self-attributions of mindedness entangle these self-attributions with benign spiritual beliefs and attributions of mind to non-human entities that are culturally accepted and widespread.

0 Citations
0 Influential
21 Altmetric
105.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!