2608.10385v1 Aug 11, 2026 cs.IR

LLM 기반 정보 검색 평가에서 어조 조건부화를 통한 평가자 민감도 분석

Persona Conditioning as an Assessor-Sensitivity Probe for LLM-Based IR Evaluation

Samaneh Mohtadi
Samaneh Mohtadi
Citations: 1
h-index: 1
Pietro Bernardelle
Pietro Bernardelle
Citations: 35
h-index: 4
Gianluca Demartini
Gianluca Demartini
Citations: 150
h-index: 5
Joel Mackenzie
Joel Mackenzie
Citations: 80
h-index: 5

최근 대규모 언어 모델(LLM)은 정보 검색(IR) 평가에서 관련성 평가자로 점점 더 많이 사용되고 있으며, 이는 평가자의 프레임이 판단의 신뢰성과 하위 시스템 비교에 미치는 영향에 대한 질문을 제기합니다. 본 연구에서는 LLM 평가자의 민감도를 파악하기 위한 진단 메커니즘으로 어조 조건부화를 탐구합니다. PersonaHub 및 NVIDIA Nemotron-Personas-USA에서 가져온 작업 지향형 페르소나를 사용하여 의도 해석, 도메인 전문성, 대조적 판단, 증거 검증 및 전반적인 검색 품질 평가에 중점을 둔 다섯 가지 평가자 역할을 정의하고, 이를 표준 UMBRELA 기준과 비교합니다. TREC DL20 및 RAG24 데이터셋에서 여섯 가지 LLM 모델을 기반으로 분석한 결과, 평가자의 민감도는 균일하지 않고 구조화되어 있음을 확인했습니다. 일반적으로 판단은 기준선과 유사하게 유지되지만, 평가의 엄격성, 증거 임계값 또는 해석 강조점이 변경되는 경향이 있으며, 광범위한 관련성 반전은 발생하지 않습니다. 시스템 수준에서 고용량 모델은 시스템 순위 일치도를 유지하는 반면, 소형 모델은 페르소나에 의해 유발된 불안정성을 증폭시킵니다. 로컬 순위 변동 분석 결과, 민감도는 특정 검색 시스템 및 시스템 유형에 집중되어 있으며, 특히 DL20에서는 신경망 기반 랭킹/재순위 시스템에서, RAG24에서는 RAG 관련 파이프라인에서 두드러집니다. 페르소나 출처는 평가자 역할과 모델 용량보다 중요하지 않습니다. 이러한 결과는 어조 조건부 판단을 LLM 기반 IR 평가 파이프라인의 스트레스 테스트를 위한 제어된 민감도 분석 도구로 활용하고, 평가 결과가 평가자 프레임에 민감하게 반응하는 시스템을 식별하는 데 도움이 될 수 있음을 시사합니다.

Original Abstract

Large language models (LLMs) are increasingly used as relevance assessors in information retrieval (IR) evaluation, raising questions about how assessor framing affects judgment reliability and downstream system comparison. We study persona conditioning as a diagnostic mechanism for exposing LLM assessor sensitivity. Using task-oriented personas drawn from two complementary sources (PersonaHub and NVIDIA Nemotron-Personas-USA), we instantiate five assessor roles emphasizing intent interpretation, domain expertise, contrastive judgment, evidence verification, and global search-quality assessment, compared with a standard UMBRELA baseline. Across six LLM backbones on TREC DL20 and RAG24, our analyses reveal structured rather than uniform assessor sensitivity. Judgments usually remain close to the baseline while shifting assessment strictness, evidential threshold, or interpretation emphasis rather than producing widespread relevance reversals. At the system level, high-capacity models preserve system-ranking agreement, while smaller models amplify persona-induced instability. Local rank-displacement analysis shows sensitivity concentrates on particular retrieval systems and system types, especially neural ranking/reranking systems on DL20 and RAG-oriented pipelines on RAG24. Persona source matters less than assessor role and model capacity. These findings position persona-conditioned judging as a controlled sensitivity probe for stress-testing LLM-based IR evaluation pipelines and identifying systems whose evaluation outcomes are sensitive to assessor framing.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!