2602.19101v1 Feb 22, 2026 cs.CL

가치 얽힘: (일부) 대규모 언어 모델에서 나타나는 서로 다른 종류의 선(Good) 간의 혼재

Value Entanglement: Conflation Between Different Kinds of Good In (Some) Large Language Models

Seong Hah Cho
Seong Hah Cho
Citations: 11
h-index: 2
Anna Leshinskaya
Anna Leshinskaya
Citations: 3
h-index: 1
J. Li
J. Li
Citations: 1,833
h-index: 23

대규모 언어 모델(LLM)의 가치 정렬(value alignment)을 위해서는 모델이 실제로 습득한 가치 표현을 경험적으로 측정해야 한다. 인간의 가치 표현이 갖는 특징 중 하나는 서로 다른 종류의 가치를 구별한다는 점이다. 본 연구에서는 LLM 역시 도덕적, 문법적, 경제적이라는 세 가지 다른 종류의 선(good)을 구별하는지 조사한다. 모델의 행동, 임베딩, 그리고 잔차 스트림 활성화(residual stream activations)를 분석함으로써, 우리는 이러한 개별적인 가치 표현들이 융합되는 현상, 즉 만연한 '가치 얽힘(value entanglement)' 사례들을 보고한다. 구체적으로, 인간의 규범과 비교할 때 문법적 및 경제적 가치 평가가 모두 도덕적 가치에 과도하게 영향을 받는 것으로 나타났다. 이러한 혼재 현상은 도덕성과 관련된 활성화 벡터를 선택적으로 제거(ablation)함으로써 해결되었다.

Original Abstract

Value alignment of Large Language Models (LLMs) requires us to empirically measure these models' actual, acquired representation of value. Among the characteristics of value representation in humans is that they distinguish among value of different kinds. We investigate whether LLMs likewise distinguish three different kinds of good: moral, grammatical, and economic. By probing model behavior, embeddings, and residual stream activations, we report pervasive cases of value entanglement: a conflation between these distinct representations of value. Specifically, both grammatical and economic valuation was found to be overly influenced by moral value, relative to human norms. This conflation was repaired by selective ablation of the activation vectors associated with morality.

0 Citations
0 Influential
11.5 Altmetric
57.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!