2608.03322v1 Aug 04, 2026 cs.CV

LocAnyMed: 다중 모드 의료 이미지에 대한 시각-언어 연계 기술

LocAnyMed: Vision-Language Grounding for Multimodal Medical Images

Tao Huang
Tao Huang
Citations: 15
h-index: 3
Wentao Jiang
Wentao Jiang
Citations: 99
h-index: 4
Shanshan Ye
Shanshan Ye
Citations: 0
h-index: 0
Jing Zhang
Jing Zhang
Citations: 45
h-index: 2
Zhiwei Wang
Zhiwei Wang
Citations: 10
h-index: 1
Xiaohui Yang
Xiaohui Yang
Citations: 0
h-index: 0
Sihan Ma
Sihan Ma
Citations: 437
h-index: 10
Tong Liu
Tong Liu
Citations: 0
h-index: 0
Zi-hao Wang
Zi-hao Wang
Citations: 0
h-index: 0

의료 영상 기반 객체 위치 파악(visual grounding)은 임상적 질문을 의료 이미지 내의 특정 영역과 연결하여 설명 가능한 의료 인공지능 시스템 구축에 중요한 역할을 합니다. 하지만, 대부분의 일반적인 객체 위치 파악 모델들은 자연 이미지 데이터로 훈련되며, 기존의 의료 관련 위치 정보 자원들은 다양한 영상 촬영 방식, 데이터셋 및 작업 정의 방식으로 분산되어 있습니다. 이러한 격차를 해소하기 위해, 저희는 LocAnyMed-200K라는 다중 모드 의료 영상 기반 객체 위치 파악 데이터셋을 구축했습니다. 이 데이터셋은 CT, 광학 의료 영상, 초음파 및 X선 이미지 약 200만 개의 이미지-질문-답변 예제를 포함합니다. 저희는 다양한 객체 감지 및 위치 정보 자원을 통합하여, 하나 이상의 경계 상자, 좌표 또는 '없음(no-target)' 출력을 지원하는 통일된 자유 형식 지침 형식을 사용했습니다. LocateAnything-3B 모델을 LocAnyMed-200K 데이터셋으로 전파라미터 미세 조정했을 때, 검증 데이터셋에서 F1@IoU 값이 10.64에서 85.59로 향상되었으며, 이는 대규모의 특정 분야에 대한 지도 학습이 일반적인 객체 위치 파악 모델에게 효과적인 의료 관련 위치 정보 기능을 부여할 수 있음을 보여줍니다. 또한, 임상적으로 해석 가능한 객체 위치 파악 시스템은 예측을 뒷받침하는 근거를 제공해야 합니다. 따라서, 저희는 해부학적 맥락, 시각적 관찰 및 공간적 결론을 구조화된 추론을 통해 연결하는 추가 데이터인 LocAnyMed-CoT-20K를 개발하여 교차 소스 일반화 성능을 더욱 향상시켰습니다. 이러한 자원들은 다양한 의료 영상 촬영 방식에서 위치 정확도와 근거 품질을 연구하기 위한 통합적인 기반을 제공합니다. 관련 코드는 다음 GitHub 주소에서 공개적으로 이용할 수 있습니다: https://github.com/MiliLab/LocAnyMed.

Original Abstract

Medical visual grounding connects free-form clinical queries to spatial evidence in medical images and is an important component of interpretable medical artificial intelligence. However, general-purpose grounding models are predominantly trained on natural images, while existing medical localization resources remain fragmented across imaging modalities, datasets, and task formulations. To address this gap, we construct LocAnyMed-200K, a multimodal medical visual grounding dataset containing approximately 200K image-query-answer examples across computed tomography, optical medical imaging, ultrasound, and X-ray. We harmonize heterogeneous detection and localization resources into a unified free-form instruction format that supports one or multiple bounding boxes, point coordinates, and no-target outputs for negative queries. Full-parameter fine-tuning of LocateAnything-3B on LocAnyMed-200K improves F1@IoU 0.50 from 10.64 to 85.59 on a held-out evaluation split, demonstrating that large-scale domain-specific supervision can equip a general grounding model with effective medical localization capabilities. Beyond spatial coordinates, a clinically interpretable grounding system should also communicate the evidence supporting its prediction. We therefore derive LocAnyMed-CoT-20K, a rationale-augmented subset that connects anatomical context, visual observations, and spatial conclusions through structured reasoning and further improves cross-source generalization through fine-tuning. Together, these resources provide a unified foundation for studying both localization accuracy and rationale quality across heterogeneous medical imaging modalities. The code is publicly available at https://github.com/MiliLab/LocAnyMed.

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!