2606.12169v1 Jun 10, 2026 cs.CV

OpenMedReason: 의료 영상-언어 모델을 위한 과학적 추론 지도

OpenMedReason: Scientific Reasoning Supervision for Medical Vision-Language Models

Abeer Badawi
Abeer Badawi
Citations: 14
h-index: 1
Elham Dolatabadi
Elham Dolatabadi
Citations: 32
h-index: 2
Negin Baghbanzadeh
Negin Baghbanzadeh
Citations: 54
h-index: 4
Pritam Sarkar
Pritam Sarkar
Citations: 878
h-index: 12
Michael Colacci
Michael Colacci
Citations: 7
h-index: 1
Adibvafa Fallahpour
Adibvafa Fallahpour
Citations: 306
h-index: 6
Arash Afkanpour
Arash Afkanpour
Citations: 140
h-index: 6
Leonid Sigal
Leonid Sigal
Citations: 45
h-index: 3
Ali Etemad
Ali Etemad
Citations: 1,005
h-index: 15

대규모 시각-언어 모델(LVLM)의 임상 분야 활용은 정확한 최종 답변뿐만 아니라, 시각적 증거와 임상 지식에 기반한 추론 능력이 필요합니다. 본 연구에서는 OpenMedReason을 소개합니다. 이는 약 45만 개의 이미지-질문-답변 쌍으로 구성된 대규모 의료 추론 데이터셋이며, 주로 선별된 생의학 논문 및 인간이 작성한 과학적 자료에서 추출한 추론 과정을 포함합니다. OpenMedReason은 합성적인 사고 과정 외에도 고품질의 지도 정보를 제공하며, 방사선 영상, 현미경 이미지, 일반 사진, 차트 등 다양한 의료 분야의 시각 모달리티를 포괄합니다. 또한, OpenMedReason-Bench라는 별도의 평가 벤치마크를 함께 제공하여 LVLM의 인지 능력, 의료 지식, 추론 과정과 같은 세 가지 상호 보완적인 측면을 미세하게 평가할 수 있도록 지원합니다. 이를 통해 최종 답변 정확도 외에도 진단 능력을 종합적으로 평가할 수 있습니다. OpenMedReason은 풍부한 학습 자료로서, 지도 학습(SFT) 및 강화 학습 기반 정렬 모두에서 효과를 입증했습니다. OpenMedReason으로 학습했을 때, 기본 모델 대비 평균 20%의 VQA 정확도 향상을 보였으며, 유사 규모의 다른 의료 LVLM에 비해 4.2% 이내의 성능을 달성했습니다. 상세한 성능 분석 결과, 이러한 개선이 특정 측면에만 집중된 것이 아니라, 인지 능력, 의료 지식, 추론 과정을 종합적으로 향상시킨 것으로 확인되었습니다. 또한, OpenMedReason의 추론 과정은 기본 모델의 추론 과정보다 86.1%의 쌍대 비교에서 더 선호되는 것으로 나타났습니다. 코드 및 데이터셋은 huggingface.co/datasets/neginb/OpenMedReason 에서 확인할 수 있습니다.

Original Abstract

High-stakes clinical use of large vision-language models (LVLMs) requires reasoning that is grounded in visual evidence and clinical knowledge, not just correct final answers. We introduce OpenMedReason, a large-scale, open multimodal medical reasoning corpus comprising approximately 450K image-question-answer instances whose reasoning traces are primarily derived from curated biomedical, human-authored scientific articles. OpenMedReason provides high-fidelity supervision beyond synthetic chains of thought, covering diverse medical domain vision modalities such as radiological scans, microscopic images, visible light photographs, charts, and others. We complement it with OpenMedReason-Bench, a held-out benchmark that allows fine-grained evaluation of LVLMs along three complementary axes of capability, including perception, medical knowledge, and rationale, enabling diagnostic evaluation beyond final-answer accuracy. OpenMedReason is a rich training resource that exhibits its effectiveness in both supervised fine-tuning (SFT) and reinforcement-based alignment. Training with OpenMedReason yields a 20% average improvement in VQA accuracy over the base model and achieves performance within 4.2% of the strongest comparable-scale medical LVLMs. Fine-grained performance analysis confirms that the gains are not concentrated in any single axis: OpenMedReason improves perception, medical knowledge, and rationale jointly, and its reasoning traces are preferred over those of the base model in 86.1% of pairwise comparisons. We release the code and dataset at huggingface.co/datasets/neginb/OpenMedReason.

0 Citations
0 Influential
7.5 Altmetric
37.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!