MMA-Former: 다중 창 혼합 헤드 어텐션 트랜스포머를 이용한 3차원 MRI 영상에서의 적응적 주변 신경 침윤 예측
MMA-Former: Multi-Window Mixture-of-Head Attention Transformer for Adaptive PNI Prediction in 3D MRI
주변 신경 침윤(PNI)은 담관암의 중요한 예후 지표입니다. 3차원 MRI 영상을 이용한 비침습적인 PNI 예측은 어렵고, 이를 위해서는 미세한 세부 정보와 전체 맥락을 효율적으로 포착하는 모델이 필요합니다. 본 연구에서는 다중 창 혼합 헤드 어텐션 트랜스포머(MMA-Former)라는 새로운 3차원 아키텍처를 제안합니다. 이 아키텍처는 병렬적인 다단계 특징 추출을 위한 코스-파인 트랜스포머(CFT) 구조를 포함하고 있습니다. 우리는 더욱 발전된 형태의 윈도우별 혼합 헤드 어텐션(WS-MoH) 메커니즘을 통합하여 CFT 구조에 적용했습니다. WS-MoH는 표준 멀티 헤드 셀프 어텐션(MSA)과 달리, 각 3차원 창에 대한 표현을 생성하고, 전체 창을 전문화된 또는 일반적인 어텐션 헤드로 동적으로 연결합니다. 이를 통해 각 창의 지역적 맥락에 맞춰 공간적으로 적응적인 특징 추출이 가능하며, 전문성을 향상시키고 중복을 줄여 파라미터 수를 증가시키지 않습니다. 168개의 T1 가중 MRI 스캔으로 구성된 과거 데이터를 사용하여 MMA-Former를 평가한 결과, AUC 값이 0.752로 나타났으며, 이는 최고 성능의 CNN(AUC 0.708) 및 트랜스포머 기반 모델(AUC 0.681)을 능가하는 성능입니다.
Perineural invasion (PNI) is a critical prognostic factor in cholangiocarcinoma. Non-invasive prediction from 3D MRI is challenging, demanding models that efficiently capture both fine-grained details and global context. We propose the Multi-window Mixture-of-Head Attention Transformer (MMA-Former), a novel end-to-end 3D architecture featuring a Coarse-Fine Transformer (CFT) structure for parallel multi-scale feature extraction. We advance this structure by integrating a novel Window-Specific Mixture-of-Head attention (WS-MoH) mechanism. Unlike standard Multi-Head Self Attention (MSA), WS-MoH generates a representation for each 3D window and dynamically routes the entire window to specialized or common attention heads. This enables spatially adaptive feature extraction tailored to the local context of each window, enhancing specialization and reducing redundancy without increasing parameters. Evaluated on a retrospective dataset of 168 T1-weighted MRI scans, MMA-Former achieved an AUC of 0.752, outperforming other 3D architectures, including the best CNN (AUC of 0.708) and Transformer baselines (AUC of 0.681).
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.