2602.22070v1 Feb 25, 2026 cs.AI

언어 모델은 알고리즘 에이전트와 인간 전문가에 대해 일관성 없는 편향을 보인다.

Language Models Exhibit Inconsistent Biases Towards Algorithmic Agents and Human Experts

Jessica Y. Bo
Jessica Y. Bo
Citations: 115
h-index: 6
Lillio Mok
Lillio Mok
Citations: 79
h-index: 5
Ashton Anderson
Ashton Anderson
Citations: 159
h-index: 5

대규모 언어 모델(LLM)은 다양한 정보원으로부터 데이터를 처리해야 하는 의사 결정 작업에 점점 더 많이 사용되고 있으며, 여기에는 인간 전문가와 다른 알고리즘 에이전트가 포함됩니다. LLM은 이러한 다양한 정보원을 어떻게 평가할까요? 우리는 인간 의사 결정자가 알고리즘의 예측에 대해 보이는 편향인 '알고리즘 회피'라는 잘 알려진 현상을 고려합니다. 행동 경제학의 실험적 방법을 활용하여, 8개의 서로 다른 LLM이 의사 결정 작업을 어떻게 위임하는지 평가합니다. 이때 위임 대상이 인간 전문가로 제시될 때와 알고리즘 에이전트로 제시될 때를 비교합니다. 다양한 평가 형식을 포괄하기 위해, 우리는 두 가지 방식으로 연구를 진행합니다. 첫째, LLM에게 직접적으로 신뢰도를 묻는 질문을 통해 선호도를 파악합니다. 둘째, 인간 전문가와 알고리즘 에이전트의 성능에 대한 예시를 제공하여, LLM이 어떤 선택을 하는지 '명시된 선호도'와 '실제 행동'을 모두 분석합니다. 다양한 작업에서 인간 전문가와 알고리즘의 신뢰도를 평가하도록 요청했을 때, LLM은 일반적으로 인간 전문가에게 더 높은 점수를 부여하며, 이는 인간 응답자의 이전 결과와 일치합니다. 그러나 인간 전문가와 알고리즘의 성능을 보여주고, 보상을 받을 수 있는 베팅을 통해 선택하도록 요청했을 때, LLM은 성능이 현저히 떨어지는 알고리즘을 과도하게 선택합니다. 이러한 상반된 결과는 LLM이 인간과 알고리즘에 대해 일관성 없는 편향을 내재하고 있을 수 있으며, 이는 고위험 상황에서 LLM이 사용될 때 신중하게 고려되어야 함을 시사합니다. 또한, 우리는 AI 안전성을 위한 평가의 견고성을 확보하기 위해, LLM이 작업 제시 방식에 얼마나 민감하게 반응하는지 논의합니다.

Original Abstract

Large language models are increasingly used in decision-making tasks that require them to process information from a variety of sources, including both human experts and other algorithmic agents. How do LLMs weigh the information provided by these different sources? We consider the well-studied phenomenon of algorithm aversion, in which human decision-makers exhibit bias against predictions from algorithms. Drawing upon experimental paradigms from behavioural economics, we evaluate how eightdifferent LLMs delegate decision-making tasks when the delegatee is framed as a human expert or an algorithmic agent. To be inclusive of different evaluation formats, we conduct our study with two task presentations: stated preferences, modeled through direct queries about trust towards either agent, and revealed preferences, modeled through providing in-context examples of the performance of both agents. When prompted to rate the trustworthiness of human experts and algorithms across diverse tasks, LLMs give higher ratings to the human expert, which correlates with prior results from human respondents. However, when shown the performance of a human expert and an algorithm and asked to place an incentivized bet between the two, LLMs disproportionately choose the algorithm, even when it performs demonstrably worse. These discrepant results suggest that LLMs may encode inconsistent biases towards humans and algorithms, which need to be carefully considered when they are deployed in high-stakes scenarios. Furthermore, we discuss the sensitivity of LLMs to task presentation formats that should be broadly scrutinized in evaluation robustness for AI safety.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!