2607.08046v1 Jul 09, 2026 cs.CL

LLM 예측 모델이 알고 있지만 말하지 못하는 것: 내부 표현을 활용한 교정 및 신뢰성 검증

What LLM Forecasters Know but Don't Say: Probing Internal Representations for Calibration and Faithfulness

E. Ho
E. Ho
Citations: 1
h-index: 1
Raphael Sarfati
Raphael Sarfati
Citations: 91
h-index: 5
Christopher J. Earls
Christopher J. Earls
Citations: 8
h-index: 2
S. Boppana
S. Boppana
Citations: 10
h-index: 2
P. Tiwari
P. Tiwari
Citations: 152
h-index: 7
Srikar Varadaraj
Srikar Varadaraj
Citations: 14
h-index: 1

예측 작업을 위해 미세 조정된 대규모 언어 모델은 정확할 수 있지만, 제대로 교정되지 않았을 수 있으며, 그들의 체인 오브 씽킹(Chain-of-Thought, CoT) 추론이 예측의 근거가 되는 증거를 충실히 반영하지 못하는 경우가 있습니다. 본 연구에서는 내부 표현이 이러한 측면에서 더 직접적인 정보를 제공할 수 있는지 탐구합니다. OpenForesight 데이터셋을 사용한 Eternis-Forecaster 8B 모델에 대해, 중간 활성화 값을 기반으로 representation pooling probes를 학습시킨 결과, 모델의 교정 능력이 크게 향상되었으며, 이 결과는 GLM-4.7-Flash 및 GLM-4.5-Air 모델에서도 동일하게 나타났습니다. 또한, evidence ablation (증거 제거) 및 diversionary injection (방향 전환 주입)을 통해 CoT의 신뢰성을 평가한 결과, 프롬프트 내에서 중요한 정보를 제거하면 모델의 예측이 종종 변경되는 반면, 추론 과정은 그대로 유지됨을 확인했습니다. 이러한 probes는 '거짓말 탐지기' 역할을 수행하며, 행동 변화를 추론 과정보다 훨씬 더 정확하게 파악하고, 84%의 경우에도 변화 방향을 예측합니다. 특히, CoT가 변수의 영향을 숨기는 경우에도 이를 감지할 수 있습니다. 마지막으로, 강제 답변 실험 결과, 예측은 대부분 추론이 시작되기 전에 결정되며, 하나의 사전 추론 단계만으로도 이미 결정된 답변과 신뢰도를 회복할 수 있으며, 이 사전 설정된 답변 분포를 활용하여 질문을 라우팅하면 생성되는 토큰 수를 30-47% 절약할 수 있습니다. 이러한 결과들을 종합적으로 고려했을 때, 내부 표현에 대한 probing은 언어 모델 예측 모델 및 추론 모델의 교정, 감사 및 분류를 위한 실용적인 도구로 활용될 수 있음을 시사합니다.

Original Abstract

Large language models fine-tuned for forecasting can be accurate yet poorly calibrated, and their chain-of-thought (CoT) reasoning may not faithfully reflect the evidence behind a forecast. We ask whether internal representations offer a more direct window into both. Working with Eternis-Forecaster 8B on OpenForesight, we train representation-pooling probes on intermediate activations and find they achieve substantially better calibration; a result that also holds for GLM-4.7-Flash and GLM-4.5-Air. We then assess CoT faithfulness through evidence ablation and diversionary injection: removing an influential source in the prompt often changes the model's forecast while leaving the reasoning trace untouched. The same probes function as lie detectors: their activations track behavioral shifts far better than the reasoning trace does, and they also predict the direction of change in 84% of cases, including when the CoT conceals the perturbation's influence. Finally, forced answering reveals that forecasts are largely fixed before reasoning begins: a single pre-reasoning pass recovers the committed answer and confidence, and routing questions by the spread of this pre-set answer distribution saves 30-47% of generated tokens, with no loss of accuracy. Together, these results establish probing internal representations as a practical tool for calibrating, auditing, and triaging language model forecasters and reasoning models more broadly.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!