2607.24743v1 Jul 27, 2026 cs.CV

ClinFusion: 포괄적인 의료 이해를 위한 비전 중심의 다중 모드 대규모 언어 모델 시스템

ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding

Hangjie Yuan
Hangjie Yuan
Citations: 301
h-index: 5
Tao Feng
Tao Feng
Citations: 63
h-index: 4
Yi Yang
Yi Yang
Citations: 427
h-index: 5
Sicheng Yang
Sicheng Yang
Citations: 6
h-index: 2
Jiahong Dong
Jiahong Dong
Citations: 37
h-index: 3
Zhitao Zeng
Zhitao Zeng
Citations: 57
h-index: 3
Weihua Chen
Weihua Chen
Citations: 65
h-index: 4
Fan Wang
Fan Wang
Citations: 45
h-index: 3
Zuozhu Liu
Zuozhu Liu
Citations: 271
h-index: 9
Yanqiu Xing
Yanqiu Xing
Citations: 0
h-index: 0
Yizeng Han
Yizeng Han
Citations: 100
h-index: 5
Jitao Wang
Jitao Wang
Citations: 10
h-index: 2
Yichen Qian
Yichen Qian
Citations: 94
h-index: 3
Shengxuan Luo
Shengxuan Luo
Citations: 0
h-index: 0
Qing Xie
Qing Xie
Citations: 0
h-index: 0
Weigen Yao
Weigen Yao
Citations: 0
h-index: 0
Zhiwei Tang
Zhiwei Tang
Citations: 0
h-index: 0
Xianzhe Xu
Xianzhe Xu
Citations: 846
h-index: 11
Lirong Wu
Lirong Wu
Citations: 0
h-index: 0
Jinwang Wang
Jinwang Wang
Citations: 331
h-index: 10
Pengju Wang
Pengju Wang
Citations: 0
h-index: 0
Jiasheng Tang
Jiasheng Tang
Citations: 29
h-index: 2
Shaochen Wang
Shaochen Wang
Citations: 0
h-index: 0
Feng Xu
Feng Xu
Citations: 0
h-index: 0

다중 모드 대규모 언어 모델(MLLM)은 임상 실무에 혁신을 가져올 엄청난 잠재력을 가지고 있지만, 의료 분야에 이를 적용하는 것은 근본적으로 비전 중심적인 과제입니다. 모델은 다양한 2차원 및 3차원 의료 이미지를 통해 지식을 습득해야 하며, 평가 프로토콜은 방사선 전문의의 임상 실무와 일치하고 정확하고 세밀하며 사실성에 기반한 평가를 제공해야 합니다. 본 논문에서는 이러한 제한 사항을 체계적으로 해결하기 위해 설계된 포괄적인 의료 이해를 위한 비전 중심 MLLM인 ClinFusion을 소개합니다. 우리는 Cascade Spatial-Aware Locality Fusion 연산자를 특징으로 하는 구조적이고 캐스케이드 방식으로 구성된 비전 인코더 아키텍처를 제안하여 다양한 2차원 및 원시 3차원 의료 이미지 이해를 통합된 인코더 내에서 결합합니다. 또한, 지침 준수 평가를 위한 MedIF-Bench와 임상적으로 일치하고 사실성을 기반으로 한 보고서 생성 평가를 위한 관심 영역(Region-of-Interest) 기반 방법을 포함하는 비전 기반 평가 프레임워크를 소개합니다. 실험 결과, ClinFusion은 시각적 질의 응답, 보고서 생성 및 지침 준수를 포괄하는 다양한 2차원 및 3차원 다중 모드 의료 벤치마크에서 새로운 최고 성능을 달성했으며, 텍스트 기반 의료 작업에서도 선도적인 오픈 소스 의료 MLLM(예: Hulu-Med, Lingshu)보다 24개의 벤치마크 중 20개에서 더 나은 성능을 보였습니다. 또한 GPT-5.2 및 Gemini-3-Flash와 같은 강력한 독점 모델보다 16개의 벤치마크 중 13개에서 더 우수한 다중 모드 기능을 보여주며, 검색 증강 및 도구 지원 임상 워크플로우를 위한 에이전트 기반 도구 사용으로 추가 확장될 수 있습니다. 보이지 않는 상태로 진행된 방사선 전문의 평가 결과, ClinFusion이 가장 높은 순위를 차지한 보고서를 생성했으며, 제안하는 RoI 기반 지표가 조사된 모든 자동 평가 지표 중에서 전문가 판단과 가장 강한 상관관계를 갖는다는 것을 입증했습니다.

Original Abstract

Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundamentally a vision-centric challenge: models must absorb knowledge from heterogeneous 2D and 3D medical images, and evaluation protocols must align with radiologists' clinical practice and provide an accurate, fine-grained and factualness-driven assessment. In this paper, we introduce ClinFusion, a vision-centric MLLM designed for holistic medical understanding that systematically addresses these limitations. We propose a compositional and cascaded vision encoder architecture featuring a Cascade Spatial-Aware Locality Fusion operator that unifies diverse 2D and native 3D medical image understanding within a fused encoder. We further introduce a vision-grounded evaluation framework, including MedIF-Bench for instruction-following assessment and a region-of-interest-grounded method for clinically aligned and factualness-driven report generation evaluation. We show that ClinFusion sets a new state-of-the-art across a comprehensive suite of 2D and 3D multimodal medical benchmarks---spanning visual question answering, report generation, and instruction following---as well as textual medical tasks, outperforming leading open-source medical MLLMs (\textit{e.g.}, Hulu-Med, Lingshu) on 20 out of 24 benchmarks and demonstrating multimodal capabilities better than powerful proprietary models such as GPT-5.2 and Gemini-3-Flash on 13 out of 16 benchmarks, and can be further augmented with agentic tool use for retrieval-augmented and tool-assisted clinical workflows. A blinded evaluation by board-certified radiologists confirms that ClinFusion produces the highest-ranked reports, and validates our RoI-grounded metric as achieving the strongest correlation with expert judgment among all automatic evaluation metrics examined.

0 Citations
0 Influential
5.5 Altmetric
27.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!