ProgFormer: 계층적 복셀 확산 트랜스포머를 이용한 경시 뇌 MRI 예측
ProgFormer: Hierarchical Voxel Diffusion Transformer for Longitudinal Brain MRI Prediction
뇌의 미래 구조 MRI를 예측하는 것은 어려운 과제입니다. 왜냐하면 경시적인 변화는 종종 미묘하며 특정 해부학적 영역에 국한되는 반면, 대부분의 개인별 뇌 구조는 시간이 지남에 따라 안정적으로 유지되기 때문입니다. 따라서 효과적인 모델은 전체적인 뇌 구조 일관성을 유지하면서도 미세한 질병 진행에 민감해야 합니다. 기존의 잠재 공간 기반 방법은 계산 효율성을 향상시키지만, 압축-복원 과정에서 정보 손실이 발생합니다. 반면, 직접 복셀 공간 기반 방법은 잠재 공간 복원을 피하지만, 일반적으로 뇌 구조와 진행 관련 변화를 모델링하기 위해 단일 예측 경로를 사용합니다. 따라서 미묘한 국소적 변화는 우세한 안정적인 뇌 구조에 의해 가려질 수 있습니다. 이러한 문제점을 해결하기 위해, 우리는 경시 뇌 MRI 예측을 위한 계층적 복셀 공간 확산 트랜스포머인 ProgFormer를 제안합니다. ProgFormer는 3D 패치 토큰으로부터 주요 부피 예측을 수행하는 거친 경로를 사용하며, 이 경로는 전체 뇌 구조와 경시 맥락을 모델링합니다. 그런 다음 미세 경로는 거친 표현을 개별 패치 내의 복셀 수준 정제를 위한 시공간적 기준으로 활용합니다. 두 경로는 조건부 플로우 매칭을 통해 직접 복셀 공간에서 속도장을 공동으로 추정하며, 별도로 학습된 이미지 오토인코더 없이 엔드 투 엔드 예측을 가능하게 합니다. 예측된 미래 스캔은 추정된 속도장을 일련의 Euler 단계에 걸쳐 통합하여 가우시안 노이즈로부터 생성됩니다. ADNI, AIBL 및 OASIS를 포함한 세 가지 널리 사용되는 벤치마크 데이터셋에서 쌍별 및 시퀀스 설정 모두에서 수행한 광범위한 실험 결과는 여러 최첨단 방법과 비교했을 때 우수한 성능을 보여줍니다.
Predicting future structural MRI of a brain is challenging because longitudinal changes are often subtle and confined to specific anatomical regions, while most subject-specific brain structure remains stable over time. An effective model should therefore preserve global brain structural consistency while remaining sensitive to fine-grained disease progression. Existing latent-space-based methods improve computational efficiency, but suffer from information loss during their compression-reconstruction procedure. In contrast, direct voxel-space methods avoid latent reconstruction but commonly use a unified prediction pathway to model brain structure and progression-related changes. Subtle local changes may therefore be overshadowed by the dominant stable brain structure. To address these challenges, we propose ProgFormer, a hierarchical voxel-space Diffusion Transformer for longitudinal brain MRI prediction. ProgFormer uses a coarse pathway to perform the primary volumetric prediction from 3D patch tokens. This pathway models overall brain structure and longitudinal context. The fine pathway then uses the coarse representations as spatio-temporal grounding for voxel-level refinement within individual patches. The two pathways jointly estimate a velocity field directly in voxel space through conditional flow matching, enabling end-to-end prediction without a separately learned image autoencoder. The predicted future scan is then generated from Gaussian noise by integrating the estimated velocity field over a sequence of Euler steps. Extensive experimental results on three widely used benchmarks, ADNI, AIBL, and OASIS, under both pairwise and trajectory settings demonstrate favourable performance compared against several state-of-the-art methods.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.