2606.11635v1 Jun 10, 2026 cs.CY

LLM은 윤리적 추론에 능숙하지 않은가?

Are LLMs Bad at Moral Reasoning?

Menghang Zhu
Menghang Zhu
Citations: 274
h-index: 4
Seth Lazar
Seth Lazar
Citations: 45
h-index: 2

고도로 발전된 AI 시스템이 역동적이고 개방적인 환경에서 안전하게 작동하기 위해서는 행동의 윤리적 이유를 식별하고 이해하며, 그에 따라 행동을 제약할 수 있어야 합니다. 최근 연구들은 현재 가장 뛰어난 AI 시스템들의 이러한 능력, 즉 윤리적 역량을 평가하는 것을 목표로 하며, 대체로 비관적인 결론에 도달했습니다. 대표적인 연구 중 하나는 1,000개의 사례에 대한 인간이 작성한 윤리적 추론 평가 기준을 수집하고, 최첨단 AI 모델을 이러한 기준과 비교하여 실망스러운 결과를 얻었습니다. 본 논문에서는 MoReBench 데이터셋을 재활용하여 LLM의 윤리적 추론 능력(윤리적 역량의 필수적인 부분)에 대한 훨씬 더 긍정적인 그림을 제시하고자 합니다. 우리는 LLM이 제공된 사례에 대한 답변을 평가 기준과 비교하는 대신, 인간에게 주어진 동일한 과제 - 즉 특정 사례의 윤리적 분석을 위한 평가 기준을 생성하는 과제 - 를 LLM에게 부여했을 때, LLM이 생성한 평가 기준은 그들의 자유형 답변보다 인간의 평가 기준에 더 잘 부합하며, 차이가 있는 경우에도 대부분은 대부분의 윤리 문제의 방대한 다차원성을 반영할 뿐이며,

Original Abstract

For highly capable AI systems to operate safely in dynamic, open-ended environments, they must be able to identify, understand, and respond to moral reasons for action, and constrain their behaviour accordingly. A growing body of research aims to evaluate this capacity -- moral competence -- in today's most capable AI systems, recently reaching broadly pessimistic conclusions. One of the most ambitious such papers collects gold-standard human-authored rubrics for evaluating moral reasoning in 1,000 cases, and benchmarks frontier AI models against those rubrics, with underwhelming results. In this paper, we argue that the MoReBench dataset can be redeployed to give a much more optimistic picture of LLMs' moral reasoning (an essential part of moral competence). We show that if, instead of scoring LLMs' responses to these cases against these rubrics, we instead give the LLMs the same task given to humans -- to generate scoring rubrics for the moral analysis of particular cases -- the rubrics they generate are both better calibrated to the human rubrics than their open-ended responses, and, where they differ, plausibly reflect nothing more than the vast dimensionality of most moral problems, as well as highlighting some human departures from the "rubric for creating rubrics". Taking these points into consideration, the MoReBench dataset suggests that LLMs are significantly more capable at moral reasoning than was previously believed.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!