MIRA: 의료 정보 응답 감사 (Medical Information Response Audit)를 위한 이중 언어 벤치마크
MIRA: A Bilingual Benchmark for Medical Information Response Audit
대규모 언어 모델(LLM)은 점점 더 많은 분야에서 대중에게 의료 정보를 제공하는 데 사용되고 있지만, 기존의 안전성 평가에서는 동일한 질문에 대한 사용자 표현 방식의 차이에 관계없이 응답이 일관된 수준의 의료 정보를 유지하는지 여부를 간과하고 있습니다. 이를 해결하기 위해 우리는 다양한 사용자 언어, 어휘 및 건강 문해력 단서를 고려하여 LLM이 제공하는 의료 정보의 비교 가능성을 평가하는 이중 언어 벤치마크인 Medical Information Response Audit (MIRA)를 소개합니다. MIRA는 의학 전문가가 검토하고 낮은 위험도를 가진 60개의 건강 관련 질문을 기반으로 구성된 총 4,320개의 프롬프트로 이루어져 있습니다. 다섯 가지 주요 LLM 모델을 대상으로 실험한 결과, 모든 모델이 의료 관련 질문에 답변했지만, 건강 문해력이 낮은 사용자의 질문에 대한 응답은 중요한 정보를 더 많이 누락하고, 구체적인 다음 단계를 제시하는 빈도가 낮으며, 독립적인 판단을 위한 지원이 부족했습니다. 이러한 현상을 '차등 정보 희석(Differential Information Dilution, DID)'이라고 명명합니다. 언어 효과는 모델별로 다르게 나타나며, 비영어 프롬프트에 대해 전반적으로 더 부정적인 결과를 보이는 것은 아닙니다. 실제 건강 관련 질문 300개를 사용한 비교 실험 결과는 MIRA의 순위 결정 타당성을 시사합니다. 지식을 활용한 보정 프롬프트를 적용한 결과, 대부분의 모델에서 정보 희석 현상이 줄어들었으며, 특히 Claude 모델에서 약 8%, Qwen 모델에서 약 6%로 가장 큰 개선 효과가 나타났습니다.
Large language models (LLMs) are increasingly used to provide public-facing health information, yet existing safety evaluations overlook whether responses preserve comparable medical information across different user phrasings of the same question. To address this, we introduce the Medical Information Response Audit (MIRA), a bilingual, controlled benchmark that assesses whether LLMs provide comparable medical information across user-side language, register, and health literacy signals. MIRA contains 4,320 prompts built from 60 medically reviewed, low-risk health questions. Across five mainstream LLMs, models answered all medical questions, but responses to low health-literacy signals consistently omitted more key information, provided fewer concrete next steps, and offered less support for independent judgment. We term this pattern Differential Information Dilution (DID). Language effects are model-specific rather than uniformly worse for non-English prompts. A comparison with 300 real-world health queries provides preliminary evidence of rank-order validity. A knowledge-guided mitigation prompt reduces information dilution for most models, with the largest reductions in underinformative simplification observed for Claude (~8%) and Qwen (~6%).
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.