2607.26981v1 Jul 29, 2026 cs.CL

OptimismBench: 언어 모델 판단에서의 예측 편향 및 정렬 효과 분석

OptimismBench: Forecasting Bias and the Alignment Effect in Language Model Judgment

Seonglae Cho
Seonglae Cho
Citations: 18
h-index: 3
Adriano S. Koshiyama
Adriano S. Koshiyama
Citations: 218
h-index: 8

최근 대규모 언어 모델(LLM)은 의사 결정 지원 도구로 점점 더 많이 사용되고 있으며, 이들의 확률적 판단은 이후의 선택에 영향을 미칩니다. 그러나 이러한 판단이 체계적인 방향성을 갖는지는 밝혀내기 어려웠습니다. 기존의 교정 지표는 부호가 없는 오차를 집계하고, 실제 불확실성은 진정한 확률 값을 제공하지 않습니다. 예를 들어, LLM이 특정 스타트업의 성공 가능성을 70%로, 실패 가능성을 15%로 평가할 때, 누락된 15% 포인트는 전체 점수로 감지되지 않는 왜곡을 드러냅니다. 우리는 OptimismBench를 소개합니다. 이는 반전 쌍을 사용하여 방향성 편향을 탐지하며, 각 시나리오에서 P(성공)과 P(불패)를 모두 얻고, 두 프레임 간의 비대칭성을 통해 기준점 없이 서명된 편향 점수를 제공합니다. 8개 공급업체의 16개 모델을 분석한 결과, 14개의 모델이 낙관적인 경향을 보였으며, 비관적인 경향은 Anthropic사의 최상위 모델에서만 나타났습니다. 4가지 유형의 11쌍의 기본-챗 모델을 비교한 결과, 사후 학습 세트가 편향의 부호를 결정하며, 서로 다른 유형에서 반대 방향으로 변화하는 것을 확인했습니다. 이러한 패턴은 프롬프트, 온도, 관점 및 자체 편향 제거 방법을 적용해도 유지됩니다. 17개 모델과 6개 언어를 비교한 결과, 모델 식별자가 언어보다 더 큰 영향을 미치며, 모델 간의 변동성은 언어 간 변동성의 4.7배에 달합니다. 우리는 개별 모델의 방향성 편향 감사를 위한 3,870개의 데이터를 10개 언어로 공개했습니다. 정렬을 통해 모델이 더욱 유용해질 때, 이는 해당 모델의 확률 값에도 영향을 미치며, 이후 파이프라인은 기본적으로 이러한 편향을 상속합니다.

Original Abstract

Large language models are increasingly used as decision aids whose probability judgments shape downstream choices. Whether those judgments carry a systematic directional tilt has been hard to detect: calibration metrics aggregate unsigned errors, and naturalistic uncertainty offers no ground-truth probability. When an LLM rates a startup's success at 70% but its failure at 15%, the missing 15 points expose a distortion no aggregate score flags. We introduce OptimismBench, which detects directional bias with inverted pairs: each scenario elicits both P(success) and P(failure), and asymmetry between the two framings yields a signed bias score without ground truth. Across 16 models from 8 providers, fourteen are optimistic; pessimism appears only in Anthropic's frontier tier. Eleven matched base-versus-chat pairs across four families show post-training sets the sign of the bias, with opposite shifts in different families. The pattern survives prompt, temperature, perspective, and self-debiasing ablations. A seventeen-model six-language comparison further shows model identity dominates language, with inter-model variance at 4.7x inter-language variance. We release 3,870 items across 10 languages for per-model directional-bias auditing. When alignment makes a model more helpful, it also tilts its probabilities; downstream pipelines inherit the tilt by default.

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!