OpenMedQ: 의료 영상-언어 모델을 위한 광범위한 공개 사전 학습
OpenMedQ: Broad Open Pretraining for Medical Vision-Language Models
본 논문에서는 OpenMedQ를 소개합니다. OpenMedQ는 현재까지 가장 방대한 규모의 공개 의료 데이터셋으로 사전 학습된 의료 영상-언어 모델입니다. 이 데이터셋은 병리학, 영상 진단, 현미경 이미지 및 텍스트 기반 임상 질의 응답을 포함하며, 총 14개의 데이터셋에 약 335만 개의 샘플이 사용되었습니다. OpenMedQ는 PathVQA에서 최고 수준인 BLEU-1 점수(75.9)를 달성했으며, 이는 최대 5620억 개의 파라미터를 가진 Med-PaLM M 모델보다 ~80배 더 큰 모델에서도 능가하는 결과입니다. 또한 VQA-MED 데이터셋에서 보고된 최고 BLEU-1 점수(64.5)와 일치합니다. OpenMedQ의 비전 인코더는 동일한 방식으로 8개의 의료 분류 벤치마크에 적용되었으며, BiomedCLIP (0.745), PMC-CLIP (0.745), PubMedCLIP (0.746) 및 처음부터 학습한 모델 (0.616)보다 높은 평균 macro-F1 점수(0.757)를 달성했습니다. 저희는 코드와 함께 인터랙티브 데모를 공개하여 커뮤니티의 재현 가능한 기준점으로 활용할 수 있도록 했습니다.
We present OpenMedQ, a medical vision-language model pretrained on the broadest fully-open medical mix to date: 14 datasets totaling ~3.35M pretraining samples spanning pathology, radiology, microscopy, and text-only clinical QA. OpenMedQ reaches state-of-the-art BLEU-1 on PathVQA (75.9), beating Med-PaLM M variants up to 562B parameters (~80x larger), and matches the best reported VQA-MED BLEU-1 (64.5). Its vision encoder, transferred to 8 unseen medical classification benchmarks under an identical downstream recipe, obtains the highest average macro-F1 (0.757) among BiomedCLIP (0.745), PMC-CLIP (0.745), PubMedCLIP (0.746), and a from-scratch baseline (0.616). We release our code and an interactive demo is publicly available as a reproducible baseline for the community.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.