2601.01751v1 Jan 05, 2026 cs.IR

LLM 관련성 판단 편향 분석을 위한 쿼리-문서 밀집 벡터

Query-Document Dense Vectors for LLM Relevance Judgment Bias Analysis

Samaneh Mohtadi
Samaneh Mohtadi
Citations: 1
h-index: 1
Gianluca Demartini
Gianluca Demartini
Citations: 16
h-index: 2

대규모 언어 모델(LLM)은 인적 평가자에 비해 비용이 저렴하고 확장성이 뛰어나 정보 검색(IR) 평가 데이터셋 구축에 평가 도구로 활용되고 있습니다. 기존 연구에서는 LLM이 인적 평가자와 비교하여 얼마나 신뢰할 수 있는지에 초점을 맞추었지만, 본 연구에서는 LLM이 평균적으로 얼마나 잘 수행하는지뿐만 아니라, LLM이 관련성 판단 시 체계적인 오류를 범하는지 여부를 파악하고자 합니다. 이를 위해, 쿼리와 문서를 분석할 수 있도록 하는 새로운 표현 방식을 제안하여 관련성 레이블 분포를 분석하고, LLM과 인적 레이블을 비교하여 불일치 패턴을 파악하고 체계적인 불일치 영역을 특정합니다. 본 연구에서는 쿼리-문서(Q-D) 쌍을 통합 의미 공간에 임베딩하는 클러스터링 기반 프레임워크를 도입하여 관련성을 관계 속성으로 간주합니다. TREC Deep Learning 2019 및 2020 데이터셋에 대한 실험 결과, 인간과 LLM 간의 체계적인 불일치는 무작위적으로 분포하는 것이 아니라 특정 의미 클러스터에 집중되어 있음을 확인했습니다. 쿼리 수준 분석 결과, 정의 검색, 정책 관련 또는 모호한 맥락에서 반복적으로 발생하는 오류가 나타났습니다. 쿼리 클러스터 내에서 합의 수준이 크게 다른 쿼리는 불일치 발생 지점으로 나타났으며, LLM은 이러한 쿼리에서 관련 콘텐츠를 제대로 검색하지 못하거나 관련 없는 정보를 과도하게 포함하는 경향이 있었습니다. 본 프레임워크는 전반적인 진단과 지역 클러스터링을 연결하여 LLM 판단의 숨겨진 약점을 파악하고, 편향을 고려한 더욱 신뢰할 수 있는 IR 평가를 가능하게 합니다.

Original Abstract

Large Language Models (LLMs) have been used as relevance assessors for Information Retrieval (IR) evaluation collection creation due to reduced cost and increased scalability as compared to human assessors. While previous research has looked at the reliability of LLMs as compared to human assessors, in this work, we aim to understand if LLMs make systematic mistakes when judging relevance, rather than just understanding how good they are on average. To this aim, we propose a novel representational method for queries and documents that allows us to analyze relevance label distributions and compare LLM and human labels to identify patterns of disagreement and localize systematic areas of disagreement. We introduce a clustering-based framework that embeds query-document (Q-D) pairs into a joint semantic space, treating relevance as a relational property. Experiments on TREC Deep Learning 2019 and 2020 show that systematic disagreement between humans and LLMs is concentrated in specific semantic clusters rather than distributed randomly. Query-level analyses reveal recurring failures, most often in definition-seeking, policy-related, or ambiguous contexts. Queries with large variation in agreement across their clusters emerge as disagreement hotspots, where LLMs tend to under-recall relevant content or over-include irrelevant material. This framework links global diagnostics with localized clustering to uncover hidden weaknesses in LLM judgments, enabling bias-aware and more reliable IR evaluation.

2 Citations
0 Influential
1 Altmetric
7.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!