약물 안전 추론에서의 환자 정보에 대한 반사실적 민감성 평가
Evaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning
특정 환자의 상태를 충족하지 못함에도 불구하고 유효한 약물 안전 규칙을 적용하면 잘못된 결정을 내릴 수 있습니다. 기존의 의료 평가는 대부분 독립적이고 고정적인 시나리오를 사용합니다. 따라서 모델은 특정 약물과 위험 간의 연관성을 단순히 기억하여 올바른 답변을 할 수 있지만, 환자 정보를 활용하여 규칙이 적용되는지 판단했는지 여부를 보여주지는 않을 수 있습니다. 이러한 문제를 해결하기 위해, 우리는 환자별 약물 안전 추론에 대한 근거 기반 권장 사항과 전문가 검증 질문으로 구성된 벤치마크인 MedPIC-Bench를 소개합니다. 이 벤치마크는 지침 준수 질문과 함께, 환자 정보의 통제된 변경이 규칙 적용 여부에 영향을 미치는 쌍을 이루는 반사실적 질문을 포함합니다. 벤치마크에는 6가지 임상 및 추론 차원을 기준으로 주석이 달린 467개의 질문이 포함되어 있습니다. 28개의 의료 전문, 일반, 그리고 독점 LLM을 대상으로 실험한 결과, 모든 모델은 반사실적 질문에서 성능이 더 낮았으며, 평균 정확도는 63.6%에서 45.1%로 감소했습니다. 모델은 명시적인 환자 속성이 잘 알려진 금기 사항을 직접적으로 나타낼 때에는 좋은 성능을 보이지만, 환자 정보를 통해 안전 경고를 제한하거나 철회해야 할 때에는 어려움을 겪습니다. 모델의 설명에서는 종종 변경된 환자 정보가 언급되지만, 최종 답변은 이전의 안전 판단을 유지하는 경우가 많습니다. 이러한 취약점은 의료 전문 LLM에서도 나타나며, 그 평균 반사실적 성능은 일반 LLM보다 낮습니다. 따라서 MedPIC-Bench는 조건부 규칙 적용 여부를 측정 가능하게 만들고, 환자별 신뢰성을 평가하기 위한 정적인 약물 안전 정확도의 한계를 강조합니다.
Applying a valid medication-safety rule when its patient-specific conditions are not met can produce an incorrect decision. Existing medical evaluations largely use isolated and fixed scenarios. A model may therefore answer correctly by recalling a drug-risk association without showing that it used patient information to decide whether the rule applies. To address this gap, we introduce MedPIC-Bench, a benchmark of source-verifiable recommendations and expert-validated questions for patient-specific medication-safety reasoning. It combines guideline-following questions with paired counterfactual questions in which a controlled change in patient information changes whether a rule applies. The benchmark contains 467 questions annotated along six clinical and reasoning dimensions. Across 28 medical-specific, general, and proprietary LLMs, every model performs worse on counterfactual questions, with mean accuracy falling from 63.6\% to 45.1\%. Models perform well when an explicit patient attribute directly signals a familiar contraindication, but struggle when patient information must narrow or withdraw a safety warning. Model rationales often acknowledge the changed patient information, yet the final answers retain the previous safety judgment. This vulnerability persists among medical-specific LLMs, whose average CF performance trails that of general LLMs. MedPIC-Bench therefore makes conditional rule application measurable and highlights the limitations of static medication-safety accuracy for assessing patient-specific reliability.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.