2606.12953v1 Jun 11, 2026 cs.AI

OpenMedQ: 의료 영상-언어 모델을 위한 광범위한 공개 사전 학습

OpenMedQ: Broad Open Pretraining for Medical Vision-Language Models

Max Van Puyvelde
Max Van Puyvelde
Citations: 0
h-index: 0
Ibrahim Gulluk
Ibrahim Gulluk
Citations: 0
h-index: 0
Olivier Gevaert
Olivier Gevaert
Citations: 648
h-index: 11

본 논문에서는 OpenMedQ를 소개합니다. OpenMedQ는 현재까지 가장 방대한 규모의 공개 의료 데이터셋으로 사전 학습된 의료 영상-언어 모델입니다. 이 데이터셋은 병리학, 영상 진단, 현미경 이미지 및 텍스트 기반 임상 질의 응답을 포함하며, 총 14개의 데이터셋에 약 335만 개의 샘플이 사용되었습니다. OpenMedQ는 PathVQA에서 최고 수준인 BLEU-1 점수(75.9)를 달성했으며, 이는 최대 5620억 개의 파라미터를 가진 Med-PaLM M 모델보다 ~80배 더 큰 모델에서도 능가하는 결과입니다. 또한 VQA-MED 데이터셋에서 보고된 최고 BLEU-1 점수(64.5)와 일치합니다. OpenMedQ의 비전 인코더는 동일한 방식으로 8개의 의료 분류 벤치마크에 적용되었으며, BiomedCLIP (0.745), PMC-CLIP (0.745), PubMedCLIP (0.746) 및 처음부터 학습한 모델 (0.616)보다 높은 평균 macro-F1 점수(0.757)를 달성했습니다. 저희는 코드와 함께 인터랙티브 데모를 공개하여 커뮤니티의 재현 가능한 기준점으로 활용할 수 있도록 했습니다.

Original Abstract

We present OpenMedQ, a medical vision-language model pretrained on the broadest fully-open medical mix to date: 14 datasets totaling ~3.35M pretraining samples spanning pathology, radiology, microscopy, and text-only clinical QA. OpenMedQ reaches state-of-the-art BLEU-1 on PathVQA (75.9), beating Med-PaLM M variants up to 562B parameters (~80x larger), and matches the best reported VQA-MED BLEU-1 (64.5). Its vision encoder, transferred to 8 unseen medical classification benchmarks under an identical downstream recipe, obtains the highest average macro-F1 (0.757) among BiomedCLIP (0.745), PMC-CLIP (0.745), PubMedCLIP (0.746), and a from-scratch baseline (0.616). We release our code and an interactive demo is publicly available as a reproducible baseline for the community.

1 Citations
0 Influential
5.5 Altmetric
28.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!