평가자 교정: 확률적 교정이 LLM 에이전트 피드백 루프에서의 선호도 결합 현상을 완화하는가?
Calibrating the Evaluator: Does Probability Calibration Mitigate Preference Coupling in LLM Agent Feedback Loops?
대규모 언어 모델(LLM) 에이전트가 평가자의 피드백을 통해 행동을 조정할 때, 체계적인 평가자 편향이 에이전트의 학습된 전략 분포에 영향을 미쳐 '평가자 선호도 결합'이라는 현상이 발생합니다. 이전 연구에서는 이러한 결합 현상을 밝혀내고 이를 측정하기 위한 진단 프레임워크(EPC)를 제시했지만, 교정 기법이 이러한 효과를 완화할 수 있는지 여부는 조사하지 않았습니다. 본 연구는 평가자 교정을 완화 전략으로 사용하는 첫 번째 사례입니다. 즉, 평가자의 쌍대 비교 판단에 확률적 교정을 적용하여 부차적인 선호도 전파를 줄이는 것을 목표로 합니다. DeepSeek-V4-Pro를 실행기로 사용하고 GLM5.2를 평가자로 사용하여, 표준 이진 TTRL(승패)과 신뢰도 기반으로 교정된 TTRL(확률 가중 업데이트)을 비교하는 통제된 실험(N=5)을 수행했습니다. 그 결과, 교정을 적용하면 결합 계수 감마가 20-49% 감소하고, Jensen-Shannon 발산이 45-67% 감소하는 것을 확인했습니다. 대칭적인 LR 제어 그룹 분석 결과, 이러한 효과는 업데이트의 비대칭성 감소 때문이 아님을 확인했습니다. 본 연구에서 개발한 교정된 TTRL 프로토콜을 공개하며, LLM 기반 평가 시스템 구축 파이프라인에 대한 경량화된 완화 전략으로 활용할 것을 권장합니다.
When large language model (LLM) agents adapt their behavior through evaluator feedback, systematic evaluator biases propagate into the agent's learned strategy distribution - a phenomenon termed evaluator preference coupling. Prior work has documented this coupling and established a diagnostic framework (EPC) to measure it, but has not investigated whether calibration techniques can mitigate the effect. We present the first study of evaluator calibration as mitigation: applying probability calibration to the evaluator's pairwise judgments to reduce spurious preference propagation. In a controlled within-subjects experiment (N=5) comparing standard binary TTRL (win/loss) with confidence-calibrated TTRL (probability-weighted updates) using DeepSeek-V4-Pro as executor and GLM5.2 as evaluator, we find that calibration reduces the coupling coefficient gamma by 20-49% and Jensen-Shannon divergence by 45-67%. A symmetric-LR control confirms the effect is not due to reduced update asymmetry. We release the calibrated TTRL protocol and recommend it as a lightweight mitigation for LLM-as-judge deployment pipelines.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.