점수 규칙: 텍스트 평가 지표의 통계적 및 전략적 정렬
Scoring Rules! Statistical and Strategic Alignment for Text Evaluation Metrics
자연어 생성 시스템을 평가하는 데 널리 사용되는 참조 기반 텍스트 평가 지표는 후보 응답과 참조 응답을 비교하여 점수를 매깁니다. 평가지표의 신뢰성은 일반적으로 인간 평가와의 통계적 상관관계로 판단됩니다. 그러나 이러한 지표가 최적화 목표로 점점 더 많이 사용됨에 따라, 상관관계만으로는 충분하지 않습니다. 에이전트가 평가지표를 전략적으로 조작할 수 있기 때문입니다. 우리는 두 가지 상호 보완적인 정렬 개념을 통해 이 문제를 연구합니다. 평가지표는 인간 평가와 상관관계를 갖는 경우 통계적으로 정렬되고, 작업과 관련된 정보가 추가되지 않는 변경 사항에 저항하는 경우 전략적으로 정렬됩니다. 본 논문에서는 다음과 같은 두 가지 기여를 합니다. 첫째, 인간 평가와의 상관관계, 성능 저하 민감도 및 조작 방지 능력을 포함하는 참조 기반 지표에 대한 테스트 원칙을 제안합니다. 이러한 원칙은 지표가 인간 판단과 얼마나 일치하는지, 낮은 노력으로 인한 정보 손실을 얼마나 잘 감지하는지, 그리고 전략적인 점수 상승 시도를 얼마나 잘 저항하는지를 평가합니다. 둘째, 기존 및 새로운 지표를 정보 측정 방법, 추정 방법, 텍스트 표현 방식 및 예측 메커니즘의 네 가지 선택 요소로 분해하여 상호 정보 기반 지표에 대한 통합 설계 프레임워크를 개발했습니다. 동료 검토, 요약 및 질문 답변 작업에서 강력한 인간 평가 상관관계가 전략적 정렬을 의미하지 않는다는 것을 발견했습니다. LLM-as-a-Judge는 높은 상관관계를 달성하지만 조작에 취약합니다. 반면, 상호 정보 기반 지표는 조작 방지 능력을 크게 향상시킵니다. 또한, 우리의 프레임워크는 실험에서 가장 강력한 전체적인 견고성을 보이는 새로운 지표를 발견했으며, 이 지표는 인간 평가 상관관계에서도 경쟁력 있는 성능을 보입니다.
Reference-based text evaluation metrics, which are widely used to assess natural language generation systems, score a candidate response by comparing it with a reference response. The reliability of an evaluation metric is usually judged by its statistical correlation with human ratings. However, as these metrics are increasingly used as optimization objectives, correlation alone is no longer sufficient: agents may strategically game the evaluation metric. We study this issue through two complementary notions of alignment. A metric is statistically aligned if it correlates with human ratings and strategically aligned if it resists perturbations that do not add task-relevant information. We make two contributions. First, we propose test principles for reference-based metrics consisting of human-rating correlation, degradation sensitivity, and manipulation robustness. These principles evaluate whether a metric agrees with human judgments, penalizes low-effort information loss, and resists strategic score inflation. Second, we develop a unified design framework for mutual-information-based metrics that decomposes existing and new metrics into four choices: information measure, estimation method, text representation, and prediction mechanism. Across peer review, summarization, and question answering, we find that strong human-rating correlation does not imply strategic alignment: LLM-as-a-Judge achieves high correlation but is susceptible to manipulation. In contrast, mutual-information-based metrics substantially improve manipulation robustness. Our framework also uncovers a new metric that achieves the strongest overall robustness in our experiments while remaining competitive on human-rating correlation.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.