2608.03028v1 Aug 04, 2026 cs.AI

약물 안전 추론에서의 환자 정보에 대한 반사실적 민감성 평가

Evaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Congkai Xie
Congkai Xie
Citations: 249
h-index: 6
Yuhang Liu
Yuhang Liu
Zhejiang University
Citations: 235
h-index: 5
Zhitian Hou
Zhitian Hou
Sun Yat Sen University
Citations: 11
h-index: 2
Zeyu Liu
Zeyu Liu
Citations: 27
h-index: 2
Zhengying Liu
Zhengying Liu
Citations: 0
h-index: 0
Shuo Cai
Shuo Cai
Citations: 10
h-index: 2
Zhijie Sang
Zhijie Sang
Citations: 43
h-index: 5
Kun Zeng
Kun Zeng
Citations: 0
h-index: 0

특정 환자의 상태를 충족하지 못함에도 불구하고 유효한 약물 안전 규칙을 적용하면 잘못된 결정을 내릴 수 있습니다. 기존의 의료 평가는 대부분 독립적이고 고정적인 시나리오를 사용합니다. 따라서 모델은 특정 약물과 위험 간의 연관성을 단순히 기억하여 올바른 답변을 할 수 있지만, 환자 정보를 활용하여 규칙이 적용되는지 판단했는지 여부를 보여주지는 않을 수 있습니다. 이러한 문제를 해결하기 위해, 우리는 환자별 약물 안전 추론에 대한 근거 기반 권장 사항과 전문가 검증 질문으로 구성된 벤치마크인 MedPIC-Bench를 소개합니다. 이 벤치마크는 지침 준수 질문과 함께, 환자 정보의 통제된 변경이 규칙 적용 여부에 영향을 미치는 쌍을 이루는 반사실적 질문을 포함합니다. 벤치마크에는 6가지 임상 및 추론 차원을 기준으로 주석이 달린 467개의 질문이 포함되어 있습니다. 28개의 의료 전문, 일반, 그리고 독점 LLM을 대상으로 실험한 결과, 모든 모델은 반사실적 질문에서 성능이 더 낮았으며, 평균 정확도는 63.6%에서 45.1%로 감소했습니다. 모델은 명시적인 환자 속성이 잘 알려진 금기 사항을 직접적으로 나타낼 때에는 좋은 성능을 보이지만, 환자 정보를 통해 안전 경고를 제한하거나 철회해야 할 때에는 어려움을 겪습니다. 모델의 설명에서는 종종 변경된 환자 정보가 언급되지만, 최종 답변은 이전의 안전 판단을 유지하는 경우가 많습니다. 이러한 취약점은 의료 전문 LLM에서도 나타나며, 그 평균 반사실적 성능은 일반 LLM보다 낮습니다. 따라서 MedPIC-Bench는 조건부 규칙 적용 여부를 측정 가능하게 만들고, 환자별 신뢰성을 평가하기 위한 정적인 약물 안전 정확도의 한계를 강조합니다.

Original Abstract

Applying a valid medication-safety rule when its patient-specific conditions are not met can produce an incorrect decision. Existing medical evaluations largely use isolated and fixed scenarios. A model may therefore answer correctly by recalling a drug-risk association without showing that it used patient information to decide whether the rule applies. To address this gap, we introduce MedPIC-Bench, a benchmark of source-verifiable recommendations and expert-validated questions for patient-specific medication-safety reasoning. It combines guideline-following questions with paired counterfactual questions in which a controlled change in patient information changes whether a rule applies. The benchmark contains 467 questions annotated along six clinical and reasoning dimensions. Across 28 medical-specific, general, and proprietary LLMs, every model performs worse on counterfactual questions, with mean accuracy falling from 63.6\% to 45.1\%. Models perform well when an explicit patient attribute directly signals a familiar contraindication, but struggle when patient information must narrow or withdraw a safety warning. Model rationales often acknowledge the changed patient information, yet the final answers retain the previous safety judgment. This vulnerability persists among medical-specific LLMs, whose average CF performance trails that of general LLMs. MedPIC-Bench therefore makes conditional rule application measurable and highlights the limitations of static medication-safety accuracy for assessing patient-specific reliability.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!