2608.03035v1 Aug 04, 2026 cs.CL

언어 모델은 명제의 문맥적 진실성을 인코딩한다

Language Models Encode the Contextual Truth of Propositions

Rachel Rudinger
Rachel Rudinger
Citations: 2
h-index: 1
Rupak Sarkar
Rupak Sarkar
Citations: 235
h-index: 10
Pritika Ramu
Pritika Ramu
Citations: 47
h-index: 4

이전 연구에 따르면, LLM(대규모 언어 모델)은 활성화 공간 내에서 사실적인 명제의 진실성을 선형적으로 표현합니다. 그러나 이러한 표현이 문맥적 진실성으로 어떻게 확장되는지는 불분명합니다. 문맥적 진실성은 세계 지식이 아닌 문맥 내 증거에 의해 결정되는 명제를 의미합니다. 본 연구에서는 LLM이 구조적으로 다른 출력 정책에서도 유지되는 문맥적 진실성의 선형적인 표현을 가지고 있음을 보여줍니다. 또한, 모델이 명제의 진실 여부를 판단해야 할 필요가 없는 경우에도 이러한 표현이 유지되며, 조작 실험을 통해 인과 관계를 입증합니다. 공동의 이해를 유지하기 위해 두 개의 LLM이 협력하는 시각-언어 작업에서 얻은 데이터를 사용하여, LLM이 특정 명제에 대한 자신의 진실성 표현이 해당 명제에 대해 파트너가 주장하는 내용에 의해 크게 영향을 받는다는 것을 보여줍니다. 심지어 LLM이 해당 명제의 진실 여부를 판단할 충분한 증거를 가지고 있더라도 이러한 현상이 나타납니다. 의사 결정 경계 근처의 명제가 파트너의 주장에 의해 진실성이 변경될 가능성이 더 높다는 사실을 발견했습니다. 표현과 출력을 분리함으로써, 출력 행동만으로는 구분하기 어려운 두 가지 형태의 아첨 행위를 식별할 수 있습니다. 모델은 거짓된 명제를 수용하면서도 여전히 그것을 거짓으로 표현하거나, 표현을 경계선을 넘어 이동시킬 수 있습니다. 후자의 경우, 모델이 거짓 주장을 명시적으로 반복하여 동의하는 경우에는 암묵적으로 동의하는 경우보다 2.59배 더 흔하게 발생합니다.

Original Abstract

Prior work has shown that LLMs encode the truth of factual propositions along linear directions in activation space. It's unclear how these representations extend to contextual truth: propositions whose truth is determined by in-context evidence rather than world knowledge. We show that LLMs maintain a linear representation of contextual truth that persists across structurally different output policies, even when the output doesn't require the model to determine a proposition's truth, and show causal evidence via steering experiments. Using the transcripts from a collaborative vision-language task that requires two LLMs to maintain a shared common ground, we show that truth representations of a proposition are significantly swayed by partner assertions about that proposition, even when the LLM has enough evidence to determine its truth. We find evidence that propositions near the decision boundary are more susceptible to having their truth shifted through partner assertions. Separating representation from output distinguish two forms of sycophancy that output behavior alone cannot: the model may accommodate a false proposition while continuing to represent it as false, or shift its representation across the boundary. The latter is 2.59x more common when the model agrees by restating the false claim explicitly than when it agrees implicitly.

0 Citations
0 Influential
5 Altmetric
25.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!