2607.18999v1 Jul 21, 2026 cs.CL

MedDDC-Eval: 진단 분리 평가 - 다중 대화 의료 상담 에이전트

MedDDC-Eval: Diagnosis-Decoupled Evaluation of Multi-Turn Medical Consultation Agents

Jianwei Lv
Jianwei Lv
Citations: 121
h-index: 4
Junfeng Wang
Junfeng Wang
Citations: 959
h-index: 6
Guofeng Zhang
Guofeng Zhang
Citations: 0
h-index: 0
Yizeng Quan
Yizeng Quan
Citations: 0
h-index: 0
Huaiyi Fang
Huaiyi Fang
Citations: 0
h-index: 0
Jinyao Liu
Jinyao Liu
Citations: 0
h-index: 0
Xu Duan
Xu Duan
Citations: 0
h-index: 0
Lening An
Lening An
Citations: 0
h-index: 0
Ouyang Yu
Ouyang Yu
Citations: 0
h-index: 0

다중 회차 의료 상담 에이전트는 질문할 내용을 결정하고, 환자의 응답에 적응하며, 수집된 정보가 충분한지 판단해야 합니다. 그러나 기존의 통합 평가 방식은 정책에서 추출된 병력의 품질과 정책 특유의 최종 진단 생성 능력을 혼동합니다. 강력한 진단 생성 능력은 부족한 병력을 보완할 수 있으며, 반대로 취약한 진단 생성 능력은 풍부한 병력을 가릴 수 있습니다. 우리는 MedDDC-Eval을 제안합니다. 이는 병력을 비교 대상으로 간주하고, 공유된 고정된 리더(reader)를 사용하여 병력과 진단의 매핑 관계를 일정하게 유지하는 진단 분리 평가 환경입니다. 두 개의 독립적인 데이터 세트를 활용하여, 사용 가능한 인터페이스와 감사 가능한 진단-경로-효율성(D/T/E) 지표를 통해 진단의 유용성, 정보 획득 및 효율성을 측정합니다. 방향적 의미론적 커버리지에 이어 결정론적인 일대일 매핑을 사용하여 개방형 항목에 대해 일관된 정밀도-재현율 값을 얻으며, 예측 또는 참조 항목당 최대 하나의 일치만 인정합니다. 병력을 고정하고 진단 리더만 변경하면 F1 점수가 2.2에서 19.0 포인트까지 변동하며, Record 및 Dialogue 데이터 세트에서 정책 순서의 18%와 36%가 뒤바뀌는 것을 확인했습니다. 또한, 진단 결과 및 경로 피드백을 사용하여 Qwen3-32B 모델을 그룹 상대적 정책 최적화(GRPO)를 통해 추가 학습시켰습니다. 100개의 사례로 구성된 Record 데이터 세트와 70개의 사례로 구성된 Dialogue 데이터 세트에서, 학습된 정책은 초기 상태보다 총 점수가 각각 9.7점과 4.6점 향상되었습니다. 주요 피드백 중 하나를 제거하면 검증 데이터 세트의 전체 성능이 저하되는 것을 확인했습니다. 이러한 결과는 MedDDC-Eval이 제어 가능한 속성 분석, 해석 가능한 병력 측정 및 평가 기반 증거 획득 정책 개발을 지원함을 보여줍니다.

Original Abstract

Multi-turn medical consultation agents must decide what to ask, adapt to patient responses, and determine when the collected evidence is sufficient. However, coupled evaluation conflates the quality of the policy-elicited history with policy-specific terminal diagnosis generation: strong generation can compensate for a thin history, while weaker generation can obscure a rich one. We introduce MedDDC-Eval, a diagnosis-decoupled testbed that treats elicited history as the comparison object and holds the history-to-diagnosis mapping constant through a shared frozen reader. Across two held-out sources, a grounded interface and an auditable diagnosis-trajectory-efficiency (D/T/E) harness measure diagnostic usefulness, information acquisition, and efficiency. Directional semantic coverage followed by deterministic one-to-one assignment yields coherent precision-recall counts for open-ended items, with at most one credited match per prediction or reference. Holding histories fixed, changing only the diagnostic reader shifts diagnosis F1 by 2.2-19.0 points and reverses 18% and 36% of pairwise policy orderings on the Record and Dialogue splits. We further apply standard Group Relative Policy Optimization (GRPO) over interactive multi-turn rollouts to post-train Qwen3-32B using diagnosis-result and trajectory feedback. On the 100-case Record and 70-case Dialogue splits, the trained policy improves over its initialization by 9.7 and 4.6 total-score points; removing either primary signal lowers held-out joint performance. These results show that MedDDC-Eval supports controlled attribution, interpretable elicited-history measurement, and evaluation-guided evidence-acquisition policy development.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!