2602.06286v1 Feb 06, 2026 cs.AI

LLM은 합리적 에이전트처럼 행동하는가? 확률적 의사결정에서의 신념 정합성 측정

Do LLMs Act Like Rational Agents? Measuring Belief Coherence in Probabilistic Decision Making

K. Yamin
K. Yamin
Citations: 30
h-index: 3
Jingjing Tang
Jingjing Tang
Citations: 5
h-index: 1
Santiago Cortes-Gomez
Santiago Cortes-Gomez
Citations: 43
h-index: 3
Amit Sharma
Amit Sharma
Citations: 41
h-index: 3
Eric Horvitz
Eric Horvitz
Citations: 245
h-index: 6
Bryan Wilder
Bryan Wilder
Citations: 21
h-index: 2

거대언어모델(LLM)은 최적의 행동을 위해 세계에 대한 불확실성과 다양한 결과의 효용을 모두 고려해야 하는 고위험 영역에서 에이전트로 점점 더 많이 활용되고 있지만, 그 의사결정 논리는 여전히 해석하기 어렵다. 본 연구는 LLM이 일관된 신념과 안정적인 선호를 가진 합리적 효용 극대화자인지를 조사한다. 우리는 진단 챌린지 문제에 대한 모델의 행동을 고찰한다. 연구 결과는 도출된 확률과 관찰된 행동에 대해, LLM의 추론과 이상적인 베이지안 효용 극대화 사이의 관계에 대한 통찰력을 제공한다. 우리의 접근 방식은 보고된 확률이 어떠한 합리적 에이전트의 참된 신념과도 일치할 수 없는 반증 가능한 조건을 제시한다. 우리는 이 방법론을 여러 LLM에 걸친 평가와 함께 다수의 의료 진단 영역에 적용한다. 마지막으로, 우리는 본 연구 결과의 시사점과 고위험 의사결정을 지원하는 데 있어 LLM 활용의 향후 방향에 대해 논의한다.

Original Abstract

Large language models (LLMs) are increasingly deployed as agents in high-stakes domains where optimal actions depend on both uncertainty about the world and consideration of utilities of different outcomes, yet their decision logic remains difficult to interpret. We study whether LLMs are rational utility maximizers with coherent beliefs and stable preferences. We consider behaviors of models for diagnosis challenge problems. The results provide insights about the relationship of LLM inferences to ideal Bayesian utility maximization for elicited probabilities and observed actions. Our approach provides falsifiable conditions under which the reported probabilities \emph{cannot} correspond to the true beliefs of any rational agent. We apply this methodology to multiple medical diagnostic domains with evaluations across several LLMs. We discuss implications of the results and directions forward for uses of LLMs in guiding high-stakes decisions.

3 Citations
0 Influential
3 Altmetric
18.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!