2606.16234v1 Jun 15, 2026 cs.CV

구조적 지침을 활용한 혈관 영상 합성: 망막 사진과 희소 광간섭 단층 촬영 영상을 이용한 형광혈관조영술 생성

Propagating Structural Guidance: Synthesizing Fluorescein Angiography from Fundus Images and Sparse OCT Scans

Yi Zhou
Yi Zhou
Citations: 1,115
h-index: 14
Tao Zhou
Tao Zhou
Citations: 285
h-index: 8
Tengfei Ma
Tengfei Ma
Citations: 11
h-index: 2
Ruiqi Wu
Ruiqi Wu
Citations: 89
h-index: 4
Chenran Zhang
Chenran Zhang
Citations: 72
h-index: 3
Y. Geng
Y. Geng
Citations: 5
h-index: 1
Na Su
Na Su
Citations: 11
h-index: 2
Xiangyuan Duanmu
Xiangyuan Duanmu
Citations: 0
h-index: 0
Wen Fan
Wen Fan
Citations: 6
h-index: 1

형광혈관조영술(FFA)은 망막 혈관 이상을 평가하는 데 매우 중요하지만, 획득 과정이 침습적이며 항상 가능한 것은 아닙니다. 반면, 색깔 망막 사진(CFP)은 비침습적이고 널리 활용 가능하여, CFP에서 FFA를 생성하는 연구가 진행되어 왔습니다. 그러나 기존 연구들은 CFP 표면 질감에만 의존하므로 기능적인 혈관 정보와 미묘한 병리학적 변화를 재구성하는 데 근본적인 한계가 있습니다. 이러한 문제를 해결하기 위해, 우리는 광간섭 단층 촬영(OCT)에서 제공하는 구조적 지침을 활용하여 CFP로부터 FFA를 합성하는 새로운 프레임워크를 제안합니다. 우리는 3,676명의 환자 눈으로부터 얻은 CFP, FFA, 그리고 OCT 영상을 쌍으로 구성한 다중 모드 망막 영상 데이터셋을 구축했습니다. 이는 망막 영상 분야에서 처음으로 만들어진 삼모드 정렬 데이터셋입니다. OCT와 망막 영상 간의 공간적 격차를 해소하기 위해, 우리는 깊이 정보를 기반으로 OCT 특징을 망막 평면에 투영하고, 적응형 레이어 정규화를 통해 이를 CFP 인코더에 주입하는 '공간적으로 정렬된 다중 모드 융합(Spatially Aligned Cross-Modal Fusion, SACMF)' 모듈을 제안합니다. 특징 융합 외에도, 우리는 '토큰 기반 다중 모드 정렬(Token-wise Cross-Modality Alignment, TCMA)'이라는 토큰 수준의 대조 학습 전략을 도입하여, CFP와 FFA 표현을 해당 공간 위치에서 명시적으로 정렬합니다. 우리의 방법은 최첨단 방법과 비교했을 때 우수한 합성 성능을 보입니다. 또한, 광범위한 실험 결과는 우리 접근 방식으로 생성된 FFA 이미지가 기존 방법보다 질병 진단 성능을 향상시킨다는 것을 보여주며, 이는 임상 환경에서 비침습적인 의사 결정 지원 도구로서의 잠재력을 강조합니다. 코드 및 관련 정보는 다음 링크에서 확인할 수 있습니다: https://github.com/while-plus/OCT-guide-FFA-Syn.

Original Abstract

Fundus fluorescein angiography (FFA) is critical for assessing retinal vascular abnormalities, but its acquisition is invasive and not always feasible. In contrast, color fundus photography (CFP) is non-invasive and widely accessible, which has motivated studies on CFP-to-FFA synthesis. However, prior works rely solely on CFP surface texture, fundamentally limiting the ability to reconstruct functional vascular information and subtle pathological changes. To address this, we propose a novel framework that synthesizes FFA from CFP with structural guidance provided by optical coherence tomography (OCT). We construct a multi-modal retinal imaging dataset with paired CFP, FFA, and OCT from 3,676 patient eyes--the first tri-modally aligned dataset in retinal imaging. To bridge the spatial gap between OCT and fundus modalities, we propose a Spatially Aligned Cross-Modal Fusion (SACMF) module that projects depth-resolved OCT features onto the fundus plane and injects them into the CFP encoder via adaptive layer normalization. Beyond feature fusion, we further introduce Token-wise Cross-Modality Alignment (TCMA), a token-level contrastive learning strategy that explicitly aligns CFP and FFA representations at corresponding spatial positions. Our method achieves superior synthesis performance compared to state-of-the-art methods. Moreover, extensive experiments demonstrate that the FFA images synthesized by our approach bring greater improvements in downstream disease diagnosis performance than existing methods, highlighting the clinical potential of our approach as a non-invasive decision-support tool in routine workflows. The code is available at https://github.com/while-plus/OCT-guide-FFA-Syn.

0 Citations
0 Influential
30.4657359028 Altmetric
0.0 Score
Original PDF
1

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!