2607.27779v1 Jul 30, 2026 cs.CV

CXR-Retrieve: 흉부 방사선 영상의 구성적 텍스트-이미지 검색

CXR-Retrieve: Compositional Text-to-Image Retrieval in Chest Radiography

Chaim Baskin
Chaim Baskin
Ben-Gurion University of the Negev
Citations: 1,165
h-index: 15
M. Kimhi
M. Kimhi
Citations: 62
h-index: 5
Ehud Rivlin
Ehud Rivlin
Citations: 13
h-index: 2
Tom Erez
Tom Erez
Citations: 126
h-index: 4

대규모 흉부 방사선 영상 데이터베이스는 대부분 연구가 정형화된 임상 정보 주석이 아닌 자유 형식 보고서와 함께 제공되기 때문에 검색하기 어렵습니다. 시각-언어 모델은 텍스트-이미지 검색을 위한 자연스러운 인터페이스를 제공하지만, 현재의 생물의학 모델은 주로 보고서-이미지 매칭에 최적화되어 있으며 짧은 임상 검색 쿼리에 대한 만족도를 높이는 데는 어려움이 있습니다. 이는 객관적인 불일치를 야기합니다. 즉, 모델은 쿼리의 단어와 관련된 이미지를 검색할 수 있지만, 전체 임상 제약을 충족하지 못하는 경우가 발생하며, 특히 '폐렴 없이 폐 경화'와 같은 접속사와 부정 표현에서 이러한 현상이 두드러집니다. 저희는 구성적 흉부 X선 텍스트-이미지 검색을 위한 표준 데이터셋인 CXR-Retrieve를 소개합니다. 이 데이터셋은 MIMIC-CXR-JPG의 공식 테스트 세트에서 추출한 5,159개의 테스트 이미지와 단일 및 접속 형태의 긍정적/부정적 진단을 모두 포함하는 145개의 텍스트 쿼리로 구성되어 있습니다. 관련성은 검색된 이미지가 명시된 모든 병리 제약을 충족하는지 여부에 따라 정의되며, 단순히 페어링된 보고서와 일치하는지에 대한 기준은 아닙니다. 더 나아가, 임상 검색을 위한 레이블 기반 대비 학습 방법을 제안합니다. 저희의 방법은 공유된 확인된 부재를 포함하여 호환 가능한 명시된 병리 제약을 가진 이미지-텍스트 쌍을 서로 끌어당기는 동시에 모순되는 쌍을 명시적으로 밀어냅니다. 사전에 훈련된 CXR-CLIP 모델을 기반으로, 저희의 방법은 두 가지 병리가 동시에 존재하는 경우 Precision@5를 8.5% 포인트 향상시키고, 부정 쿼리의 경우 Precision@5를 22.0% 포인트 향상시켰습니다. 이러한 결과는 신뢰할 수 있는 흉부 X선 검색을 위해서는 어떤 진단이 언급되었는지뿐만 아니라, 그러한 진단이 어떻게 임상적으로 명시되는지를 모델링하는 학습 목표가 필요하다는 것을 보여줍니다.

Original Abstract

Large chest radiography archives are difficult to search because most studies are paired only with free-text reports rather than structured clinical annotations. Vision-language models offer a natural interface for text-to-image retrieval, but current biomedical models are primarily optimized for report-to-image matching rather than for satisfying short clinical search queries. This creates an objective mismatch: a model may retrieve images related to words in the query while failing to satisfy the full clinical constraint, especially for conjunctions and negations such as ``atelectasis and no pneumonia.'' We introduce CXR-Retrieve, a structured benchmark for compositional chest X-ray text-to-image retrieval. The benchmark contains 5,159 test images from the official test-split of MIMIC-CXR-JPG and 145 textual queries spanning single and conjunction findings, both positive and negative. Relevance is defined by whether a retrieved image satisfies all asserted pathology constraints, rather than by whether it matches a paired report. We further propose a label-aware contrastive fine-tuning objective for clinical retrieval. Our method attracts image-text pairs with compatible asserted pathology constraints, including shared confirmed absences, while explicitly repelling contradictory pairs. Starting from the in-domain CXR-CLIP checkpoint, our method improves Precision@5 over CXR-CLIP by 8.5 percentage points on two-pathology conjunctions and by 22.0 percentage points on negation queries. These results show that reliable chest X-ray retrieval requires training objectives that model not only which findings are mentioned, but also how they are clinically asserted.

0 Citations
0 Influential
7.5 Altmetric
37.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!