2506.00633v3 May 31, 2025 cs.CV

정렬에서 합성으로: 텍스트-CT 생성 모델을 위한 대비 학습 기반의 볼륨 데이터 정합

From Alignment to Synthesis Contrastive Volumetric Grounding for Text-to-CT Generation

P. Soda
P. Soda
Citations: 3,667
h-index: 31
V. Guarrasi
V. Guarrasi
Citations: 772
h-index: 16
Filippo Ruffini
Filippo Ruffini
Citations: 158
h-index: 6
C. M. Caruso
C. M. Caruso
Citations: 214
h-index: 6
Daniele Molino
Daniele Molino
Citations: 19
h-index: 3

방사선 보고서로부터 의미적으로 제어 가능한 3차원 CT 이미지를 생성하기 위해서는 풍부한 텍스트 인코더뿐만 아니라, 볼륨 공간에 기반한 시각-언어 정합이 필수적입니다. 기존의 텍스트-CT 생성 방식은 언어 데이터 또는 2차원 시각-언어 데이터를 사용하여 사전 학습된 인코더를 활용하여 생성 과정을 제어합니다. 이러한 방식은 언어적으로는 풍부하지만, 볼륨에 대한 정보가 부족하다는 구조적인 한계를 가지고 있습니다. 본 연구에서는 3차원 시각-언어 정합의 품질이, 텍스트 인코더의 풍부함보다 의미적 제어 가능성에 더 큰 영향을 미친다는 점을 강조합니다. 이를 해결하기 위해, 텍스트 수준에서만 작동하는 구조화된 어려운 예시 데이터를 활용하여 학습된 생성 지향적인 3D-CLIP 인코더를 제안합니다. 이러한 설계는 추가적인 3차원 메모리 비용 없이 대비 학습의 난이도를 높여, 볼륨 데이터 인코더에 내재된 작은 배치 크기 제약을 극복합니다. 결과적으로, 제안하는 인코더는 완전한 엔드-투-엔드 잠재 확산 모델을 조건부로 사용하여 작동하며, 이를 통해 초해상도 파이프라인에서 발생하는 공간적 왜곡 및 슬라이스 간 불일치를 제거합니다. 체계적인 실험을 통해 정합 품질과 생성 제어 가능성 사이의 명확한 경험적 연관성을 확인했습니다. 18가지 병리 상태에 대한 CT-RATE 데이터셋으로 평가 결과, 본 연구는 이미지 충실도와 사실 정확성 모두에서 최첨단 성능을 달성했으며, 동시에 경쟁 모델보다 짧은 추론 시간과 GPU 메모리를 사용합니다. 코드: https://github.com/danielemolino/Text2CT

Original Abstract

Generating semantically controllable 3D CT volumes from radiology reports requires more than a rich text encoder, it requires vision-language alignment grounded in volumetric space. Existing Text-to-CT approaches condition generation on encoders pretrained with language only or 2D vision-language objectives, providing conditioning signals that are linguistically expressive but volumetrically blind. We argue this is a structural limitation: the quality of 3D vision-language alignment, not the richness of the text encoder, is the primary bottleneck for semantic controllability in volumetric diffusion models. To address this, we propose a generation-oriented 3D-CLIP encoder trained with structured hard negatives that operate exclusively at the text level. This design increases contrastive difficulty without any additional 3D memory cost, overcoming the small-batch constraints inherent to volumetric encoders. The resulting encoder conditions a fully end-to-end latent diffusion model that operates directly in 3D latent space, eliminating the spatial artifacts and cross-slice inconsistencies introduced by super-resolution pipelines. Through systematic ablations, we establish a clear empirical link between grounding quality and downstream generative controllability. Evaluated on CT-RATE across 18 pathological conditions, our method achieves state-of-the-art performance on both image fidelity and factual correctness, while requiring less inference time and GPU memory than all competing methods. Code is at https://github.com/danielemolino/Text2CT.

8 Citations
2 Influential
0 Altmetric
22.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!