2605.28565v1 May 27, 2026 cs.DL

검증된 오도: 검색 기반 LLM에서 발생하는 구조적 인용 오류 측정

Verified Misguidance: Measuring Structural Citation Failures in Search-Augmented LLMs

Yongsik Seo
Yongsik Seo
Citations: 42
h-index: 4
Wooseok Jeong
Wooseok Jeong
Citations: 23
h-index: 2
Dongha Lee
Dongha Lee
Citations: 34
h-index: 3
Eunyoung Kim
Eunyoung Kim
Citations: 0
h-index: 0
Hyeonseo Jang
Hyeonseo Jang
Citations: 0
h-index: 0

검색 기반 LLM 사용자는 응답이 실제 출처에 근거하고 있다는 증거로 인용을 활용하지만, 대부분 인용된 페이지 자체를 확인하지 않습니다. 매일 수백만 건의 쿼리가 이러한 시스템을 거치므로, 인용 품질은 사용자가 정보에 입각한 결정을 내리는지, 아니면 오도되는지에 대한 중요한 결정 요인이 됩니다. 그러나 기존의 평가 기준들은 각각 개별적인 측면만을 다루고 있어, 인용의 신뢰성을 결정하는 전체적인 구조를 측정하지 못합니다. 우리는 CITETRACE라는 대규모 데이터셋을 구축했습니다. 이 데이터셋은 사용자 쿼리부터 검색된 출처, 그리고 생성된 답변까지의 전체 인용 연결망을 추적합니다. 여기에는 28개 커뮤니티에서 수집한 11,200개의 실제 쿼리와 5개 제공업체의 10개 모델로부터 얻은 112,000개의 응답이 포함되어 있으며, 이를 통해 761,495개의 평가 가능한 인용 쌍을 확보했습니다. 우리는 전문가 검증된 사전 정의된 매트릭스와 5단계 신뢰성 기준을 사용하여 각 인용의 의도-목적 일치성, 출처 적합성 및 답변-출처 충실성을 평가하는 3차원 평가 프레임워크를 설계했습니다. 이 프레임워크는 인용을 포함하는 응답을 생성하는 모든 시스템에 적용 가능합니다. 대규모로 이 프레임워크를 적용한 결과, 우리는 '검증된 오도(VERIFIED MISGUIDANCE; VM)'라는 체계적인 패턴을 발견했습니다. 즉, 모델은 실제 접근 가능한 출처를 인용하지만, 하나 이상의 측면에서 오류를 발생시켜 충실성이 높은 모델이 부적절한 출처를 선택하고, 그 반대의 경우도 발생하는 '충실성-적합성 균형' 문제를 야기합니다. 전체 데이터셋에서 30.6%의 인용이 출처 내용을 왜곡하고, 27.1%가 해당 분야에 적합하지 않은 출처에서 비롯됩니다. 응답 수준에서는 최대 96%의 사용자가 하나 이상의 구조적 오류를 포함하는 인용을 접하게 됩니다. 제공업체 간의 차이가 인용 품질 변동의 88~96%를 설명하며, 이는 출처 선택이 개별 모델의 능력보다는 LLM 자체의 요인보다 더 큰 영향을 받는다는 것을 시사합니다. CITETRACE와 그 평가 프레임워크는 배포된 검색 기반 시스템에서 발생하는 구조적 인용 오류를 진단할 수 있는 최초의 자료입니다.

Original Abstract

Users of search-augmented LLMs rely on citations as evidence that responses are grounded in real sources, and rarely verify the cited pages themselves. Millions of queries per day now pass through these systems, making citation quality a silent determinant of whether users are informed or misled-yet existing benchmarks each address one facet in isolation, leaving the joint structure that determines citation trustworthiness unmeasured. We construct CITETRACE, a large-scale dataset that traces the full citation chain from user query through retrieved source to generated answer: 11,200 real-world queries from 28 communities paired with 112,000 responses from ten models across five providers, yielding 761,495 evaluable citation pairs. We design a three-dimension evaluation framework that scores each citation on intent-purpose alignment, source suitability, and answer-source fidelity, using expert-validated predefined matrices and a five-level fidelity rubric; the framework applies to any system that produces citation-bearing responses. Applying this framework at scale, we identify a systematic pattern we call VERIFIED MISGUIDANCE (VM): models cite real, accessible sources yet fail along one or more dimensions, producing a fidelity-suitability trade-off in which faithful models select inappropriate sources and vice versa. Across our pool, 30.6% of citations distort their sources and 27.1% originate from domain-inappropriate sources; at the response level, up to 96% of users encounter at least one structurally misleading citation. Provider-level differences explain 88-96% of citation-quality variance, suggesting that source selection is governed more by factors beyond individual model capability than by the LLMs themselves. Together, CITETRACE and its evaluation framework provide the first resource for diagnosing structural citation failures in deployed search-augmented systems.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!