불확실성 하의 사고: 언어 모델에서의 증거 활용 및 정보 탐색
Thinking Under Uncertainty: Evidence Use and Information-Seeking in Language Models
추론 시의 사고는 대규모 언어 모델의 성능을 향상시키지만, 전체적인 결과는 모델이 사용 가능한 증거를 더 효과적으로 활용하는지 또는 미래의 의사 결정을 개선할 수 있는 정보를 적극적으로 탐색하는지를 보여주지 않습니다. 본 연구에서는 이러한 반응들을 행동 선호도, 사고 시간, 그리고 보고된 확신 수준을 측정하여 구분했습니다. 10개의 공개 가중치 모델이 사고 모드와 비-사고 모드에서 동일한 조건의 두 가지 선택 옵션을 제시하는 실험을 수행했습니다. 인지 모델은 가치 기반 행동과 불확실성에 독립적인 선택 편향을, 그리고 탐색의 두 가지 특징인 UCB(Upper Confidence Bound)와 유사한 알려지지 않은 옵션에 대한 선호도, Thompson 샘플링과 유사한 전체 불확실성에 따른 선택 변동성을 분리했습니다. 평균적으로 사고는 가치 기반 행동을 강화하고 불확실성에 독립적인 선택 편향을 줄였지만, UCB와 유사한 탐색이나 Thompson 샘플링과 유사한 탐색 강화를 유발하지 않았습니다. 행동 외에도, 더 많은 관측값을 가진 정보 불균형 조건은 일관성 있게 균형 잡힌 조건보다 더 긴 사고 시간을 보였습니다. 보고된 확신 수준은 의사 결정의 어려움에 더욱 민감하게 반응했으며, 선택된 작업과 관련된 증거와 더 강한 상관관계를 나타냈습니다. 이러한 사고 시간 및 보고된 확신 수준 패턴을 메타인지 제어 및 메타인지 모니터링과 일치한다고 해석했지만, 두 가지 프로세스 모두를 명확히 입증하지는 못했습니다. 디코더 파라미터 조정(특히 온도)은 선택 편향과 사고 시간에 영향을 미쳤지만, 전체적인 상호 연관된 출력 패턴을 재현하지 못했습니다. 이러한 통제된 의사 결정 환경에서, 사고는 모델이 현재 증거에 기반하여 행동하는 방식을 개선했지만, 측정된 어떤 특징도 정보 탐색 정책으로의 전환을 뒷받침하지 않았습니다.
Inference-time thinking improves the performance of large language models, but aggregate outcomes do not reveal whether models use available evidence more effectively or seek information that could improve future decisions. We distinguish these responses by measuring action preference, thinking length, and reported confidence under matched uncertainty. Ten open-weight models completed matched horizon-style two-armed bandit trials in thinking and non-thinking modes. A cognitive model separated value-guided action and uncertainty-independent choice noise from two behavioral signatures of exploration: a UCB-like preference for the less-known arm and Thompson-like choice variability that increases with total uncertainty. On average, thinking strengthened value-guided action and reduced uncertainty-independent choice noise, without producing UCB-like exploration or strengthening Thompson-like exploration. Outside action, the information-imbalanced history condition, which also displayed more observations than the matched balanced condition, was associated with greater thinking length. Reported confidence became more sensitive to decision difficulty and more strongly associated with chosen task evidence. We interpret these thinking-length and reported-confidence patterns as consistent with metacognitive control and metacognitive monitoring, respectively, without establishing either process. Decoder sweeps, especially temperature, altered choice noise and thinking length but did not reproduce the joint cross-output pattern. In this controlled decision setting, thinking improved how models acted on current evidence, while neither measured signature supported a shift toward a more information-seeking policy.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.