심초음파 영상의 표준 뷰 분류를 위한 시공간 융합 모델
Spatio-Temporal Fusion Model for Standard View Classification of Echocardiographic Videos
표준 심초음파 뷰 자동 분류는 효율적인 임상 워크플로우에 매우 중요하지만, 세 가지 주요 과제를 안고 있습니다. 첫째, 공개적으로 이용 가능한 데이터셋은 규모가 작고 제한적이며 다양한 뷰를 포함하지 않습니다. 둘째, 심초음파 뷰 분류를 위한 일부 최신 비디오 기반 아키텍처의 성능이 충분히 연구되지 않았습니다. 셋째, 일부 뷰 카테고리는 시각적으로 매우 유사하여 단일 프레임 특징만으로는 구분이 어렵고, 불균일한 프레임 품질은 강력한 시간 정보 융합을 복잡하게 만듭니다. 이러한 과제에 대응하기 위해, 저희는 5,138개의 비디오, 910,579개의 프레임을 포함하며 9가지 표준 뷰를 담고 있는 '심초음파 영상 아홉 뷰 (EV9V)' 데이터셋을 공개합니다. EV9V 데이터셋은 현재까지 공개된 가장 큰 심초음파 비디오 데이터셋이라고 생각됩니다. EV9V를 사용하여, 컨볼루션 신경망(CNN), 순환 신경망(RNN) 및 트랜스포머와 같은 대표적인 비디오 분류 아키텍처에 대한 체계적인 성능 비교 분석을 수행했습니다. 또한, 저희는 공간 해부학적 구조와 시간 기반 심장 움직임을 동시에 학습하는 효율적인 이중 스트림 CNN-LSTM 프레임워크인 '시공간 융합 모델 (STFM)'을 제안합니다. 제안된 프레임워크는 학습 시 대표적인 비디오 세그먼트를 우선적으로 샘플링하고, 추론 시 근거 기반 융합을 활용하여 심초음파 비디오의 프레임 품질 변화에 대한 강건성을 향상시킵니다. 광범위한 실험 결과는 저희 방법이 다양한 비디오 분류 모델에서 경쟁력 있는 성능을 달성하며, 불확실성 인지 시공간 학습이 심초음파 뷰 분류에 효과적임을 입증합니다. 관련 코드는 https://github.com/bgx666/stfm 에서 확인할 수 있습니다.
Automated classification of standard echocardiographic views is crucial for efficient clinical workflow but faces three main challenges. First, publicly available datasets are scarce and limited in scale and view coverage. Second, the performance of some modern video-level architectures for echocardiographic view classification remains underexplored. Third, some view categories exhibit highly similar spatial appearances, making single-frame features insufficient for discrimination, while heterogeneous frame quality complicates robust temporal information fusion. To address these challenges, we release the Echocardiographic Videos of Nine Views (EV9V) dataset, comprising 5,138 videos, 910,579 frames, and 9 standard views, which is, to the best of our knowledge, the largest publicly available echocardiography video dataset. Using EV9V, we systematically benchmark representative video classification architectures, including Convolutional Neural Networks (CNNs), Recurrent Neural Networks (RNNs), and Transformers. Furthermore, we propose a Spatio-Temporal Fusion Model (STFM), an efficient dual-stream CNN-LSTM (Long Short-Term Memory) framework that jointly captures spatial anatomical structures and temporal cardiac dynamics. The proposed framework leverages uncertainty-aware learning to preferentially sample representative video segments during training and evidence-based fusion during inference, improving robustness to variations in frame quality across echocardiographic videos. Extensive experiments demonstrate that our method achieves competitive performance across diverse video classification models, validating the effectiveness of uncertainty-aware spatio-temporal learning for echocardiographic view classification. The code is available at https://github.com/bgx666/stfm.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.