2607.27806v1 Jul 30, 2026 cs.CV

LoMeVQA: 장기 의료 시각 질의응답을 위한 종합적인 벤치마크

LoMeVQA: A Comprehensive Benchmark for Longitudinal Medical VQA

Yihang Liu
Yihang Liu
Citations: 110
h-index: 6
Cheng Yang
Cheng Yang
Citations: 0
h-index: 0
Zhilin Wu
Zhilin Wu
Citations: 0
h-index: 0
Zhangkai Ni
Zhangkai Ni
Citations: 284
h-index: 2
Longzhen Yang
Longzhen Yang
Citations: 62
h-index: 5
Ying Wen
Ying Wen
Citations: 57
h-index: 3
Lianghua He
Lianghua He
Citations: 22
h-index: 3

임상 환경에서 환자들은 종종 여러 차례 진료를 받으면서 다양한 영상 검사를 진행하게 되며, 이는 시간의 흐름에 따른 데이터를 생성합니다. 이러한 시간 정보를 모델링하는 것은 질병의 진행 및 치료 반응을 정확하게 평가하는 데 매우 중요합니다. 그러나 다중 모드 대규모 언어 모델(MLLM)의 빠른 발전에도 불구하고, 장기 의료 시각 추론은 아직 충분히 연구되지 않았습니다. 이러한 간극을 메우기 위해, 우리는 206,000개의 장기 시각 질의응답(VQA) 쌍으로 구성된 종합적인 벤치마크인 LoMeVQA를 제안합니다. LoMeVQA는 진행 분류, 진행 설명, 진행 보고서 생성, 차등 영역 지정 및 차등 영역 설명을 포함한 다섯 가지 작업을 다룹니다. 데이터셋을 구축하기 위해, 우리는 다음과 같은 자동화 파이프라인을 개발했습니다 (1) 환자 기록을 시간 순으로 정리하고, (2) 의료 지식 그래프를 통해 임상적으로 의미 있는 개체를 추출하며, (3) 이러한 개체의 시간적 변화를 모델링하여 대규모 언어 모델이 고품질의 장기 VQA 쌍을 생성하도록 안내합니다. 광범위한 실험 결과는 범용 및 의료 분야 MLLM 모두 LoMeVQA에서 낮은 성능을 보이며, 이는 시간 추론에 상당한 제한 사항이 있음을 보여줍니다. 이러한 제한 사항을 해결하기 위해, 우리는 모든 작업에서 최첨단 성능을 달성하는 MedLong-8B를 소개합니다. 벤치마킹 외에도, 우리는 상세한 분석을 통해 주요 실패 요인을 밝혀내고 장기 의료 시각 추론을 개선하는 방법에 대한 통찰력을 제공합니다. 당사의 데이터는 다음 주소에서 이용 가능합니다: https://github.com/pepperbubble/LoMeVQA

Original Abstract

In clinical practice, patients often undergo multiple imaging examinations over successive visits, yielding longitudinal data. Modeling such temporal information is crucial for reliable assessment of disease progression and treatment response. However, despite the rapid advancement of multimodal large language models (MLLMs), longitudinal medical visual reasoning remains largely underexplored. To fill this gap, we propose LoMeVQA, a comprehensive benchmark consisting of 206K longitudinal visual question answering (VQA) pairs for temporal medical image analysis. LoMeVQA covers five tasks: progress classification, progress description, progress report generation, differential region grounding, and differential region description. To construct the dataset, we develop an automated pipeline that (1) organizes patient records chronologically, (2) extracts clinically meaningful entities via a medical knowledge graph, and (3) models their temporal evolution to guide large language models in generating high-quality longitudinal VQA pairs. Extensive evaluations demonstrate that both general-purpose and medical-domain MLLMs perform poorly on LoMeVQA, revealing substantial limitations in temporal reasoning. To address these limitations, we introduce MedLong-8B, which achieves state-of-the-art performance across all tasks. Beyond benchmarking, we conduct detailed analyses that uncover key failure modes and shed light on how to improve longitudinal medical visual reasoning. Our data is available at: https://github.com/pepperbubble/LoMeVQA

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!