검색 증강 기반 해부학적 지침을 활용한 텍스트-CT 이미지 생성
Retrieval-Augmented Anatomical Guidance for Text-to-CT Generation
볼륨 기반 의료 이미징을 위한 텍스트 기반 생성 모델은 의미론적 제어를 제공하지만, 명시적인 해부학적 지침이 부족하여 종종 공간적으로 모호하거나 해부학적으로 일관성이 없는 결과를 초래합니다. 반면, 구조 기반 방법은 강력한 해부학적 일관성을 보장하지만, 일반적으로 대상 이미지를 합성할 때 사용할 수 없는 실제 주석 정보가 필요합니다. 본 연구에서는 현실적인 추론 환경에서 의미론적 및 해부학적 정보를 통합하는 검색 증강 기반 텍스트-CT 이미지 생성 방법을 제안합니다. 제안하는 방법은 방사선 보고서가 주어지면, 3D 비전-언어 인코더를 사용하여 의미적으로 관련된 임상 사례를 검색하고, 해당 사례의 해부학적 주석을 구조적 프록시로 활용합니다. 이 프록시는 텍스트 기반 잠재 확산 모델의 ControlNet 브랜치를 통해 주입되어, 의미론적 유연성을 유지하면서도 거친 해부학적 지침을 제공합니다. CT-RATE 데이터 세트에 대한 실험 결과, 검색 증강 생성이 텍스트만 사용한 기존 방식에 비해 이미지 충실도와 임상적 일관성을 향상시키는 것으로 나타났으며, 이러한 방식에서 본질적으로 부족한 명시적인 공간 제어 기능을 제공합니다. 추가 분석 결과, 검색 품질이 중요하며, 의미론적으로 일치하는 프록시는 모든 평가 지표에서 일관된 성능 향상을 가져오는 것으로 나타났습니다. 본 연구는 볼륨 기반 의료 이미지 합성에서 의미론적 조건부 제어와 해부학적 타당성을 연결하는 원칙적이고 확장 가능한 메커니즘을 제시합니다. 코드 공개 예정입니다.
Text-conditioned generative models for volumetric medical imaging provide semantic control but lack explicit anatomical guidance, often resulting in outputs that are spatially ambiguous or anatomically inconsistent. In contrast, structure-driven methods ensure strong anatomical consistency but typically assume access to ground-truth annotations, which are unavailable when the target image is to be synthesized. We propose a retrieval-augmented approach for Text-to-CT generation that integrates semantic and anatomical information under a realistic inference setting. Given a radiology report, our method retrieves a semantically related clinical case using a 3D vision-language encoder and leverages its associated anatomical annotation as a structural proxy. This proxy is injected into a text-conditioned latent diffusion model via a ControlNet branch, providing coarse anatomical guidance while maintaining semantic flexibility. Experiments on the CT-RATE dataset show that retrieval-augmented generation improves image fidelity and clinical consistency compared to text-only baselines, while additionally enabling explicit spatial controllability, a capability inherently absent in such approaches. Further analysis highlights the importance of retrieval quality, with semantically aligned proxies yielding consistent gains across all evaluation axes. This work introduces a principled and scalable mechanism to bridge semantic conditioning and anatomical plausibility in volumetric medical image synthesis. Code will be released.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.