2605.29886v1 May 28, 2026 cs.CL

CRITIC-R1: 검색 기반 생성 모델을 위한 구조화된 비평 학습

CRITIC-R1: Learning Structured Critics for Retrieval-Augmented Generation

Xingcheng Fu
Xingcheng Fu
Guangxi Normal University
Citations: 921
h-index: 15
Jianxin Li
Jianxin Li
Citations: 326
h-index: 10
Wen Xiao
Wen Xiao
Citations: 151
h-index: 6
Chuanyue Yu
Chuanyue Yu
Citations: 12
h-index: 2
Qingyu Sun
Qingyu Sun
Citations: 39
h-index: 3
Runhua Xu
Runhua Xu
Citations: 12
h-index: 2
Ziwei Zhang
Ziwei Zhang
Citations: 1,724
h-index: 23

검색 기반 생성 (Retrieval-Augmented Generation, RAG)은 외부 증거를 활용하여 지식 집약적인 질문 응답 성능을 향상시킵니다. 그러나 기존의 RAG 방법은 여전히 환각 현상과 미묘한 추론 오류에 시달립니다. 최근 연구에서는 RAG 출력 결과를 개선하기 위해 외부 비평 모델을 도입했지만, 이러한 모델들은 종종 세부 사항이 부족하고 구조화되지 않은 피드백을 제공하며, 지나치게 공격적인 개입으로 인해 노이즈가 많고 신뢰성이 떨어지는 개선 결과를 초래하여 효과를 제한합니다. 이러한 문제점을 해결하기 위해, 우리는 강화 학습 (Reinforcement Learning, RL)을 사용하여 RAG 비평을 명시적인 오류 진단 문제로 정의하고 학습하는 구조화된 비평 프레임워크인 CRITIC-R1을 제안합니다. 우리의 프레임워크는 일반적인 RAG 오류를 판정, 오류 위치, 추론 분석 및 수정 생성 등 다양한 진단 차원으로 분류합니다. 이러한 능력을 학습하기 위해, 우리는 두 가지 보상 함수를 설계했습니다. Conservative Judgement Alignment (CJA)는 과도한 개입 현상을 완화하면서 신뢰할 수 있는 고수준 판단을 장려하는 반면, Diagnostic Quality Alignment (DQA)는 게이티드 보상을 통해 더욱 세밀한 진단 피드백을 개선합니다. 우리는 GRPO 기반의 강화 학습을 사용하여 외부 LLM(Large Language Model) 교사 모델로부터 수집된 프로세스 레벨의 감독 데이터를 활용하여 비평 모델을 훈련했습니다. 다섯 가지 질문 응답 벤치마크에서의 실험 결과, CRITIC-R1은 강력한 RAG 기본 모델보다 일관되게 답변 품질을 향상시키는 것으로 나타났습니다. 저희의 소스 코드는 다음 링크에서 확인할 수 있습니다: https://anonymous.4open.science/r/critic-r1-FCB0

Original Abstract

Retrieval-augmented generation (RAG) improves knowledge-intensive question answering by incorporating external evidence. However, existing RAG methods still suffer from hallucinations and subtle reasoning errors. Recent studies introduce external critics to refine RAG outputs, yet they often provide coarse-grained and weakly structured feedback, exhibit over-aggressive intervention, and lead to noisy and unreliable refinement, limiting their effectiveness for correction. To tackle these issues, we propose CRITIC-R1, a structured critic framework that formulates and learns RAG critique as an explicit error diagnosis problem using reinforcement learning (RL). Our framework categorizes common RAG errors into multiple diagnostic dimensions, including verdict, error location, reasoning analysis, and fix generation. To learn these capabilities, we design two reward functions: Conservative Judgement Alignment (CJA) first encourages calibrated high-level judgements while mitigating the over-aggressive phenomenon, whereas Diagnostic Quality Alignment (DQA) further improves fine-grained diagnostic feedback through gated rewards. We train the critic model using GRPO-based RL with process-level supervision collected from external LLM teacher models. Experiments across five QA benchmarks show that CRITIC-R1 consistently improves answer quality over strong RAG baselines. Our source code is available at https://anonymous.4open.science/r/critic-r1-FCB0

0 Citations
0 Influential
11.5 Altmetric
57.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!