2603.08305v1 Mar 09, 2026 cs.CV

검색 증강 기반 해부학적 지침을 활용한 텍스트-CT 이미지 생성

Retrieval-Augmented Anatomical Guidance for Text-to-CT Generation

P. Soda
P. Soda
Citations: 3,667
h-index: 31
V. Guarrasi
V. Guarrasi
Citations: 772
h-index: 16
C. M. Caruso
C. M. Caruso
Citations: 214
h-index: 6
Daniele Molino
Daniele Molino
Citations: 19
h-index: 3

볼륨 기반 의료 이미징을 위한 텍스트 기반 생성 모델은 의미론적 제어를 제공하지만, 명시적인 해부학적 지침이 부족하여 종종 공간적으로 모호하거나 해부학적으로 일관성이 없는 결과를 초래합니다. 반면, 구조 기반 방법은 강력한 해부학적 일관성을 보장하지만, 일반적으로 대상 이미지를 합성할 때 사용할 수 없는 실제 주석 정보가 필요합니다. 본 연구에서는 현실적인 추론 환경에서 의미론적 및 해부학적 정보를 통합하는 검색 증강 기반 텍스트-CT 이미지 생성 방법을 제안합니다. 제안하는 방법은 방사선 보고서가 주어지면, 3D 비전-언어 인코더를 사용하여 의미적으로 관련된 임상 사례를 검색하고, 해당 사례의 해부학적 주석을 구조적 프록시로 활용합니다. 이 프록시는 텍스트 기반 잠재 확산 모델의 ControlNet 브랜치를 통해 주입되어, 의미론적 유연성을 유지하면서도 거친 해부학적 지침을 제공합니다. CT-RATE 데이터 세트에 대한 실험 결과, 검색 증강 생성이 텍스트만 사용한 기존 방식에 비해 이미지 충실도와 임상적 일관성을 향상시키는 것으로 나타났으며, 이러한 방식에서 본질적으로 부족한 명시적인 공간 제어 기능을 제공합니다. 추가 분석 결과, 검색 품질이 중요하며, 의미론적으로 일치하는 프록시는 모든 평가 지표에서 일관된 성능 향상을 가져오는 것으로 나타났습니다. 본 연구는 볼륨 기반 의료 이미지 합성에서 의미론적 조건부 제어와 해부학적 타당성을 연결하는 원칙적이고 확장 가능한 메커니즘을 제시합니다. 코드 공개 예정입니다.

Original Abstract

Text-conditioned generative models for volumetric medical imaging provide semantic control but lack explicit anatomical guidance, often resulting in outputs that are spatially ambiguous or anatomically inconsistent. In contrast, structure-driven methods ensure strong anatomical consistency but typically assume access to ground-truth annotations, which are unavailable when the target image is to be synthesized. We propose a retrieval-augmented approach for Text-to-CT generation that integrates semantic and anatomical information under a realistic inference setting. Given a radiology report, our method retrieves a semantically related clinical case using a 3D vision-language encoder and leverages its associated anatomical annotation as a structural proxy. This proxy is injected into a text-conditioned latent diffusion model via a ControlNet branch, providing coarse anatomical guidance while maintaining semantic flexibility. Experiments on the CT-RATE dataset show that retrieval-augmented generation improves image fidelity and clinical consistency compared to text-only baselines, while additionally enabling explicit spatial controllability, a capability inherently absent in such approaches. Further analysis highlights the importance of retrieval quality, with semantically aligned proxies yielding consistent gains across all evaluation axes. This work introduces a principled and scalable mechanism to bridge semantic conditioning and anatomical plausibility in volumetric medical image synthesis. Code will be released.

0 Citations
0 Influential
15.5 Altmetric
77.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!