2602.08159v1 Feb 08, 2026 cs.LG

신뢰 공간: 언어 모델 내 정확성 표현의 기하학적 구조

The Confidence Manifold: Geometric Structure of Correctness Representations in Language Models

Seonglae Cho
Seonglae Cho
Citations: 18
h-index: 3
Zekun Wu
Zekun Wu
Citations: 169
h-index: 7
Adriano S. Koshiyama
Adriano S. Koshiyama
Citations: 218
h-index: 8
K. Costa
K. Costa
Citations: 8
h-index: 2

언어 모델이 "호주의 수도는 시드니이다"라고 주장할 때, 모델은 이것이 틀렸다는 것을 알고 있을까요? 본 연구에서는 5가지 아키텍처의 9개 모델에 걸쳐 정확성 표현의 기하학적 구조를 분석했습니다. 그 구조는 간단합니다. 판별력을 나타내는 신호는 3~8개의 차원을 차지하며, 추가적인 차원이 존재할수록 성능이 저하됩니다. 또한, 비선형 분류기는 선형 분리를 능가하지 못합니다. 저차원 부분 공간에서의 중심점 거리는 훈련된 탐지 성능(0.90 AUC)과 일치하며, 이를 통해 소량의 데이터만으로도 탐지가 가능합니다. 예를 들어, GPT-2 모델에서 25개의 레이블링된 데이터만으로 전체 데이터셋을 사용했을 때 달성할 수 있는 정확도의 89%에 도달할 수 있습니다. 활성화 조작을 통해 인과적인 검증을 수행한 결과, 학습된 방향은 오류율을 10.9%만큼 변화시키지만, 무작위 방향은 영향을 미치지 않습니다. 내부 탐지 방법은 0.80~0.97의 AUC를 달성하는 반면, 출력 기반 방법(P(True), 의미론적 엔트로피)은 0.44~0.64의 AUC에 그칩니다. 정확성 신호는 내부적으로 존재하지만, 출력으로는 명확하게 표현되지 않습니다. 중심점 거리가 탐지 성능과 일치한다는 사실은 클래스 분리가 평균 이동(mean shift)에 의한 것임을 시사하며, 이는 탐지가 학습된 것이 아니라 기하학적인 특성에 기반함을 의미합니다.

Original Abstract

When a language model asserts that "the capital of Australia is Sydney," does it know this is wrong? We characterize the geometry of correctness representations across 9 models from 5 architecture families. The structure is simple: the discriminative signal occupies 3-8 dimensions, performance degrades with additional dimensions, and no nonlinear classifier improves over linear separation. Centroid distance in the low-dimensional subspace matches trained probe performance (0.90 AUC), enabling few-shot detection: on GPT-2, 25 labeled examples achieve 89% of full-data accuracy. We validate causally through activation steering: the learned direction produces 10.9 percentage point changes in error rates while random directions show no effect. Internal probes achieve 0.80-0.97 AUC; output-based methods (P(True), semantic entropy) achieve only 0.44-0.64 AUC. The correctness signal exists internally but is not expressed in outputs. That centroid distance matches probe performance indicates class separation is a mean shift, making detection geometric rather than learned.

1 Citations
0 Influential
4 Altmetric
21.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!