2605.28044v1 May 27, 2026 cs.AI

관련성은 정당성을 보장하지 않는다: 인용 RAG의 증거 강도 교정

Relevant Is Not Warranted: Evidence-Force Calibration for Cited RAG

Yihang Chen
Yihang Chen
Citations: 215
h-index: 3
Xinpeng Wei
Xinpeng Wei
Citations: 81
h-index: 5
Sipeng Zhang
Sipeng Zhang
Citations: 7
h-index: 2
Pinyan Qian
Pinyan Qian
Citations: 7
h-index: 2
Shuhua Lin
Shuhua Lin
Citations: 42
h-index: 3
Su Wang
Su Wang
Citations: 142
h-index: 5
Xiaoyuan Wang
Xiaoyuan Wang
Citations: 47
h-index: 3
Wenxuan Xu
Wenxuan Xu
Citations: 98
h-index: 6
Qi Yu
Qi Yu
Citations: 20
h-index: 3
Junxian You
Junxian You
Citations: 43
h-index: 3

인용된 RAG 평가에서, 종종 보이는 출처를 근거로 활용하지만, 실제로 주제적으로 관련이 있는 인용이라 할지라도, 제시된 문구에 대한 정당성을 충분히 제공하지 못하는 경우가 있습니다. 우리는 이러한 진단 오류를 '인용 세탁'이라고 부르며, 이는 관련된 출처가 과장된 주장을 뒷받침하기 위한 근거로 제시되는 현상을 의미합니다. 본 연구에서는 증거 강도 교정을 위한 대비 테스트 도구인 FORCEBENCH를 소개합니다. 각 항목은 인용된 부분을 고정하고, 다섯 가지 운영 축(관계, 양식, 범위, 시간적 유효성 및 수치적 구체성)에 따라 증거 강도가 조절된 주장과 해당 주장의 강도를 높인 변형을 쌍으로 구성합니다. 교정된 평가자는 증거 강도가 조절된 주장에 더 높은 점수를 부여해야 합니다. 주요 실험에서는 고정되고 지역적으로 필터링된 198쌍의 평가 데이터 세트를 사용했습니다. 인용 존재 여부에 대한 기본적인 검사는 의도적으로 정보 제공을 제한하며, 토큰 및 개체 중복은 여전히 32.8%에서 36.4%의 쌍에서 단조성을 위반합니다. 보고된 네 가지 모델 평가자 시스템에서, 일반적인 지원 프롬프팅은 이 증거 강도 교정 테스트에 충분하지 않았습니다(평균 MVR 47.2%). 반면, 명시적인 정당성 강도 프롬프팅을 사용하면 MVR이 24.5%로 낮아지지만, 여전히 완벽하지 않습니다. 우리는 이 벤치마크, 프롬프트, 출력 및 플러그인 파이프라인을 공개하여 인용 평가자들이 기존의 지원 메트릭과 함께 단조성 위반율 및 강도 민감도를 보고할 수 있도록 합니다.

Original Abstract

Cited RAG evaluation often treats visible sources as a grounding signal, but a real, topically relevant citation can still under-warrant the attached wording. We study this diagnostic failure as citation laundering: a related source is presented as warrant for an over-strong claim. We introduce FORCEBENCH, a contrastive stress test for evidence-force calibration. Each item holds a cited passage fixed and pairs an evidence-calibrated claim with a localized force-raised variant across five operational axes: relation, modality, scope, temporal validity, and numeric specificity. A calibrated evaluator should score the evidence-calibrated claim higher. Headline experiments use a fixed, locality-filtered 198-pair evaluation set. A citation-presence sanity check is uninformative by design; token and entity overlap still violate monotonicity on 32.8--36.4% of pairs. Across four reported model judges, standard generic support prompting is insufficient for this force-calibration stress test (aggregate MVR 47.2%), while explicit warrant-strength prompting lowers MVR to 24.5% but remains imperfect. We release the benchmark, prompts, outputs, and plug-in pipeline so citation evaluators can report monotonicity violation rate and force sensitivity alongside conventional support metrics.

15 Citations
0 Influential
3 Altmetric
30.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!