2601.13537v1 Jan 20, 2026 cs.CL

단어 선택이 평가에 미치는 영향: LLM 평가 시스템에서의 프레임 효과 편향

When Wording Steers the Evaluation: Framing Bias in LLM judges

Yerin Hwang
Yerin Hwang
Citations: 146
h-index: 7
Dongryeol Lee
Dongryeol Lee
Citations: 180
h-index: 7
Taegwan Kang
Taegwan Kang
Citations: 160
h-index: 7
Minwoo Lee
Minwoo Lee
Citations: 49
h-index: 3
Kyomin Jung
Kyomin Jung
Citations: 156
h-index: 6

대규모 언어 모델(LLM)은 프롬프트의 표현 방식에 따라 다양한 응답을 생성하는 것으로 알려져 있으며, 이는 프롬프트 표현 방식에 대한 미묘한 지시가 모델의 답변을 조작할 수 있음을 시사합니다. 그러나 이러한 프레임 효과 편향이 LLM 기반 평가에 미치는 영향은, LLM이 안정적이고 공정한 판단을 내릴 것으로 기대되는 상황에서, 아직 충분히 연구되지 않았습니다. 본 연구는 심리학의 프레임 효과 이론에서 영감을 받아, 의도적으로 설계된 프롬프트가 모델의 판단에 미치는 영향을 네 가지 중요한 평가 과제에 걸쳐 체계적으로 조사합니다. 우리는 '긍정적인 설명'과 '부정적인 설명'을 사용한 대칭적인 프롬프트를 설계하고, 이러한 프롬프트가 모델의 출력 결과에 상당한 차이를 유발한다는 것을 보여줍니다. 14개의 LLM 평가 모델을 분석한 결과, 모델들이 프레임 효과에 민감하게 반응하며, 모델 계열별로 동의 또는 거부 경향에 뚜렷한 차이가 있음을 확인했습니다. 이러한 결과는 프레임 효과 편향이 현재 LLM 기반 평가 시스템의 구조적인 문제임을 시사하며, 프레임 효과를 고려한 평가 프로토콜 개발의 필요성을 강조합니다.

Original Abstract

Large language models (LLMs) are known to produce varying responses depending on prompt phrasing, indicating that subtle guidance in phrasing can steer their answers. However, the impact of this framing bias on LLM-based evaluation, where models are expected to make stable and impartial judgments, remains largely underexplored. Drawing inspiration from the framing effect in psychology, we systematically investigate how deliberate prompt framing skews model judgments across four high-stakes evaluation tasks. We design symmetric prompts using predicate-positive and predicate-negative constructions and demonstrate that such framing induces significant discrepancies in model outputs. Across 14 LLM judges, we observe clear susceptibility to framing, with model families showing distinct tendencies toward agreement or rejection. These findings suggest that framing bias is a structural property of current LLM-based evaluation systems, underscoring the need for framing-aware protocols.

8 Citations
0 Influential
3.5 Altmetric
25.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!