2607.11871v1 Jul 13, 2026 cs.LG

불공정한 평가자 내부 탐구: LLM 평가 시스템의 편향에 대한 메커니즘적 해석 가능성 연구

Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias

Xiuying Chen
Xiuying Chen
Citations: 98
h-index: 6
Zirui Song
Zirui Song
Mohamed bin Zayed University of Artificial Intelligence;University of Technology Sydney,
Citations: 438
h-index: 13
Zixiang Xu
Zixiang Xu
Citations: 103
h-index: 7
Sixian Li
Sixian Li
Citations: 194
h-index: 6
Huaxing Liu
Huaxing Liu
Citations: 1
h-index: 1
Xiang Wang
Xiang Wang
Citations: 664
h-index: 8
Shuai Li
Shuai Li
Citations: 103
h-index: 3

기존의 LLM 기반 평가 시스템의 편향 연구는 주로 입력-출력 수준에서 이루어집니다. 즉, 입력을 변경하고 점수 변화를 측정하며 프롬프트 수준에서의 해결책을 제시합니다. 본 연구에서는 동일한 편향이 평가자의 숨겨진 상태(hidden state) 수준에서도 나타난다는 점을 주장하며, 이는 입력-출력 관점과 상호 보완적이며 실제적으로 유용한 정보를 제공할 수 있습니다. 본 연구는 7개의 평가 시스템, 7가지 편향 유형, 그리고 9개의 벤치마크를 대상으로 세 가지 주요 결과를 보고합니다. 첫째, 기본적인 평가 입력은 밀집된 활성화 공간(activation manifold)을 차지하지만, 편향된 입력은 특정 유형에 따라 낮은 차원의 부분 공간(subspace)으로 이동하며, 이 부분 공간은 깊이가 증가함에 따라 뚜렷해지는 경향이 있으며, 세 가지 추정 방법(estimator)에 의해 일관되게 회복됩니다. 둘째, 숨겨진 상태를 이러한 부분 공간을 따라 조작하면 점수를 양방향으로 변경할 수 있습니다. 즉, 순방향 이동은 깨끗한 입력에서 편향된 점수를 재현하는 반면, 역방향 이동은 편향된 입력에 대해 기본 점수를 복원합니다. 반면, 동일한 노름(norm)을 갖는 임의의 방향은 훨씬 작은 변화를 초래합니다. 셋째, 간단한 선형 투영(linear projection) 방법을 사용하여 특정 편향 방향의 특징 벡터를 추출하면, 완전히 새로운 3개의 벤치마크에서 평가 시스템의 실패를 예측할 수 있으며, 이는 텍스트 기반 대안보다 현저히 우수한 성능을 보입니다. 본 연구는 편향을 입력-출력 노이즈가 아닌 활성화 기하학(activation geometry)으로 해석함으로써, 기하학적 구조, 인과 관계 제어 및 실제적인 예측을 단일 프레임워크 내에서 통합합니다. 프로젝트 페이지는 https://xzx34.github.io/unfair-judge/ 에서 확인할 수 있습니다.

Original Abstract

Existing studies of LLM-as-judge scoring bias work predominantly at the input-output level: they perturb inputs, measure score deltas, and propose prompt-level mitigations. We argue that the same biases admit a representation-level account in the judge's hidden state, complementary to the input-output view and operationally useful in ways it does not afford. We report three findings, across seven judges, seven bias types, and nine benchmarks. Geometry: baseline judging inputs occupy a tight activation manifold while biased inputs are displaced along a low-dimensional, type-specific subspace that sharpens with depth and is recovered consistently by three families of estimators. Causal control: steering hidden states along this subspace drives scoring in both directions, forward shifts reproducing biased scoring on clean inputs and reverse shifts restoring baseline scoring on biased ones, while matched-norm random directions produce shifts an order of magnitude smaller. Operational: a simple linear projection onto the same bias-direction features anticipates judge failures on three entirely unseen benchmarks, substantially outperforming text-based alternatives. Reading bias as activation geometry, rather than as input-output noise, unifies geometric structure, causal control, and operational prediction within a single framework. The project page is available at https://xzx34.github.io/unfair-judge/

0 Citations
0 Influential
6.5 Altmetric
32.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!