C-MIG: 다중 관점 정보 이득 기반 검색 증강 생성 모델을 활용한 임상 진단 추론
C-MIG: Multi-view Information Gain-based Retrieval-Augmented Generation for Clinical Diagnosis Reasoning
검색 증강 생성(Retrieval-Augmented Generation, RAG)과 강화 학습의 결합은 대규모 언어 모델을 신뢰할 수 있는 의료 근거에 연결하는 데 유망한 결과를 보여주었습니다. 그러나 기존 방법은 정확히 일치하는 이진 보상에 의존하는데, 이는 임상 진단에서 다음과 같은 두 가지 문제를 야기합니다 (i) 의미적으로 관련이 있지만 문자 그대로 일치하지 않는 단계는 0의 신호를 받으므로 귀중한 학습 신호가 손실되고 (ii) 단일 차원의 보상은 다양한 추론 능력을 효과적으로 감독할 수 없습니다. 이러한 문제점을 해결하기 위해, 본 논문에서는 임상 진단을 위한 다중 관점 정보 이득 기반 검색 증강 생성 프레임워크인 C-MIG를 제안합니다. C-MIG는 검색된 문서와 문서 개선이라는 두 가지 상호 보완적인 관점에서 정보 이득을 추정하여, 무엇을 검색하고 어떻게 개선할지를 동시에 안내하며, 귀중한 보상 신호 손실 및 공헌도 분배 문제를 완화합니다. 또한, 임상 진단 시나리오에서 지식 재현 범위를 향상시키는 다중 하위 쿼리 검색 증강 전략을 설계했습니다. 네 가지 의료 벤치마크에 대한 종합적인 실험 결과는 C-MIG가 RAG-RL 방법 중에서 모든 영역(in-domain) 및 외부 영역(out-of-domain) 데이터셋에서 가장 우수한 성능을 보이며, 최첨단 범용 LLM보다 임상 진단 분야에서 더 뛰어난 성능을 달성한다는 것을 보여줍니다.
Retrieval-augmented generation combined with reinforcement learning has shown promise for grounding large language models in trustworthy medical evidence. However, existing methods rely on exact-match binary rewards, which in clinical diagnosis cause two issues: (i) semantically relevant but non-verbatim steps receive zero signal, discarding valuable learning signals; and (ii) uni-dimensional rewards cannot effectively supervise heterogeneous reasoning capabilities. To address these issues, we propose C-MIG, a Multi-view Information Gain-based retrieval-augmented generation framework for Clinical diagnosis. C-MIG estimates information gain under a frozen reference model from two complementary views, retrieved-document and document-refinement, to jointly guide what to retrieve and how to refine, alleviating the issues of valuable reward signal loss and credit assignment. We further design a multi-subquery retrieval augmentation strategy that improves knowledge recall coverage in clinical diagnostic scenarios. Comprehensive experiments on four medical benchmarks demonstrate that C-MIG achieves the best performance among all RAG-RL methods on both in-domain and out-of-domain sets, and outperforms state-of-the-art general-purpose LLMs for clinical diagnosis.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.