2604.05051v1 Apr 06, 2026 cs.CL

이 치료법은 효과가 있나요? 의료 질문 답변(QA)에서 환자 질문 표현 방식에 따른 LLM의 민감성 평가

This Treatment Works, Right? Evaluating LLM Sensitivity to Patient Question Framing in Medical QA

Hye Sun Yun
Hye Sun Yun
Northeastern University
Citations: 193
h-index: 6
Geetika Kapoor
Geetika Kapoor
Citations: 75
h-index: 3
Michael Mackert
Michael Mackert
Citations: 35
h-index: 3
Ramez Kouzy
Ramez Kouzy
Citations: 857
h-index: 3
Wei Xu
Wei Xu
Citations: 47
h-index: 2
Byron C. Wallace
Byron C. Wallace
Citations: 48
h-index: 2
J. Li
J. Li
Citations: 1,833
h-index: 23

환자들은 점점 더 복잡하고 명확하게 표현하기 어려운 의료 관련 질문을 가지고 대규모 언어 모델(LLM)을 활용하고 있습니다. 그러나 LLM은 프롬프트의 표현 방식에 민감하며, 질문이 표현되는 방식에 따라 영향을 받을 수 있습니다. 이상적으로 LLM은 동일한 근거 자료에 기반할 때, 표현 방식에 관계없이 일관된 답변을 제공해야 합니다. 본 연구에서는 전문가가 선별한 문서를 활용한 검색 증강 생성(RAG) 환경에서 의료 질문 답변(QA) 시스템을 대상으로 체계적인 평가를 통해 이를 조사합니다. 우리는 환자 질문의 두 가지 변형 요인을 분석합니다. 즉, 질문의 긍정/부정적 표현 방식과 전문 용어 사용 여부입니다. 임상 시험 초록을 기반으로 구축된 6,614개의 질문 쌍 데이터셋을 활용하여 8개의 LLM에 대한 답변 일관성을 평가했습니다. 연구 결과, 긍정적 및 부정적 표현 방식의 질문 쌍은 동일한 표현 방식의 질문 쌍보다 모순되는 결론을 도출할 가능성이 훨씬 높다는 것을 확인했습니다. 이러한 표현 방식의 효과는 다중 턴 대화에서 더욱 두드러지며, 지속적인 설득 과정은 일관성 부족을 심화시킵니다. 표현 방식과 언어 스타일 간에는 유의미한 상호 작용이 없는 것으로 나타났습니다. 본 연구 결과는 의료 QA에서 LLM의 답변이 질문 표현 방식에 의해 체계적으로 영향을 받을 수 있으며, 동일한 근거 자료에도 불구하고 이러한 현상이 발생한다는 것을 보여줍니다. 이는 고위험 환경에서 RAG 기반 시스템을 평가할 때 표현 방식의 견고성이 중요한 평가 기준이 됨을 강조합니다.

Original Abstract

Patients are increasingly turning to large language models (LLMs) with medical questions that are complex and difficult to articulate clearly. However, LLMs are sensitive to prompt phrasings and can be influenced by the way questions are worded. Ideally, LLMs should respond consistently regardless of phrasing, particularly when grounded in the same underlying evidence. We investigate this through a systematic evaluation in a controlled retrieval-augmented generation (RAG) setting for medical question answering (QA), where expert-selected documents are used rather than retrieved automatically. We examine two dimensions of patient query variation: question framing (positive vs. negative) and language style (technical vs. plain language). We construct a dataset of 6,614 query pairs grounded in clinical trial abstracts and evaluate response consistency across eight LLMs. Our findings show that positively- and negatively-framed pairs are significantly more likely to produce contradictory conclusions than same-framing pairs. This framing effect is further amplified in multi-turn conversations, where sustained persuasion increases inconsistency. We find no significant interaction between framing and language style. Our results demonstrate that LLM responses in medical QA can be systematically influenced through query phrasing alone, even when grounded in the same evidence, highlighting the importance of phrasing robustness as an evaluation criterion for RAG-based systems in high-stakes settings.

1 Citations
0 Influential
11.5 Altmetric
58.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!