2607.01751v1 Jul 02, 2026 cs.CV

MedStreamBench: 스트리밍 및 사전 대응 의료 영상 분석을 위한 시간 인지 벤치마크

MedStreamBench: A Time-Aware Benchmark for Streaming and Proactive Medical Video Understanding

Shujian Gao
Shujian Gao
Citations: 24
h-index: 2
Yuan Wang
Yuan Wang
Citations: 3
h-index: 1
Songtao Jiang
Songtao Jiang
Citations: 239
h-index: 7
Zuozhu Liu
Zuozhu Liu
Citations: 271
h-index: 9
Zhe Hu
Zhe Hu
Citations: 0
h-index: 0

기존의 의료 영상 벤치마크는 모델이 올바른 답변을 생성하는지를 평가하는 데 주로 사용되지만, 답변을 적절한 시점에 제공하는지는 거의 평가하지 않습니다. 실제 임상 환경에서 AI 시스템은 무엇을 예측할 것인지뿐만 아니라 언제 답변하고, 판단을 유보하거나, 사전 경고를 발령해야 하는지 결정해야 합니다. 이는 벤치마크 평가와 실제 적용 요구 사항 간의 중요한 격차를 야기합니다. 본 논문에서는 시간 인지를 고려한 의료 영상 분석 벤치마크인 MedStreamBench를 제시합니다. MedStreamBench는 22개의 의료 데이터 세트와 5,419개의 질의응답(QA) 사례를 포함하며, 회고적, 현재, 미래, 사전 대응이라는 네 가지 시간적 설정을 지원합니다. 기존 벤치마크와 달리 MedStreamBench는 모델이 전체 영상을 사용하는 것을 제한하고, 시간적으로 제한된 증거 창 내에서만 작동하도록 설계되었으며, 단일 라운드 평가뿐만 아니라 스트리밍 평가도 지원합니다. 또한, 임상적으로 중요한 경고를 언제 발동해야 하는지 판단해야 하는 사전 모니터링 설정을 도입했습니다. MedStreamBench는 답변의 정확성 외에도 응답성과 증거 이후의 안정성을 통해 시간적 행동을 평가합니다. 최첨단 범용 및 의료 영상-언어 모델에 대한 실험 결과, 오프라인 인식과 시간적으로 맥락화된 의사 결정 간에는 상당한 격차가 있으며, 특히 스트리밍 및 사전 대응 환경에서 성능이 현저히 저하되는 것으로 나타났습니다. 본 벤치마크는 https://huggingface.co/datasets/Venn2024/MedStreamBench 에서 확인할 수 있습니다.

Original Abstract

Existing medical video benchmarks primarily evaluate whether a model produces the correct answer, but rarely assess whether it answers at the right time. In real clinical settings, AI systems must decide not only what to predict, but also when to answer, defer judgment, or proactively raise alerts. This creates a critical gap between benchmark evaluation and deployment requirements. We present MedStreamBench, a benchmark for time-aware medical video understanding. MedStreamBench integrates 22 medical datasets and 5,419 QA instances across four temporal settings: retrospective, present, future, and proactive. Unlike conventional benchmarks that assume full-video access, MedStreamBench restricts models to temporally bounded evidence windows and supports both single-turn and streaming evaluation. We further introduce a proactive monitoring setting that requires models to determine whether and when clinically relevant alerts should be triggered. Beyond answer correctness, MedStreamBench evaluates temporal behavior through responsiveness and post-evidence stability. Experiments on leading general-purpose and medical vision-language models reveal a substantial gap between offline recognition and temporally grounded decision-making, with performance dropping markedly in streaming and proactive settings. Our benchmark is available at https://huggingface.co/datasets/Venn2024/MedStreamBench.

0 Citations
0 Influential
24.5 Altmetric
122.5 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!