언어 기반 약한 지도 학습을 이용한 CT 영상 부피 분할
Open-Ended CT Volume Segmentation with Weak Supervision from Language
본 논문에서는 CT 스캔 영상을 분석하는 텍스트 기반 분할 모델 학습 방법을 제시합니다. 이 방법은 픽셀 단위의 정확한 지도 정보와 함께, 보고서에서 얻은 대략적이지만 확장 가능한 슬라이스 수준의 정보를 결합하여 사용합니다. 방대한 양의 영상-보고서 데이터베이스에서, 해당 발견 사항이 나타나는 슬라이스의 위치 정보를 포함하는 설명을 추출합니다. 그런 다음, 일반적인 2D 이미지 분할 모델인 SAM3을, 정밀하게 레이블링된 데이터를 활용한 표준 분할 손실 함수와 함께, 추출된 약한 지도 정보를 기반으로 한 슬라이스 수준의 분류 손실 함수를 사용하여 미세 조정합니다. ReXGroundingCT 데이터셋에 대한 실험 결과는 제안하는 방법이 분할 Dice 점수를 향상시킨다는 것을 보여줍니다. 완전하게 레이블링된 1000개의 부피 데이터를 사용할 경우 8%의 상대적인 성능 향상이 나타났으며, 250개의 완전하게 레이블링된 부피 데이터를 사용할 경우 22%의 성능 향상이 관찰되었습니다.
We introduce a method for training a text-conditioned segmentation model for CT scans, which combines voxel-level supervision with coarse but scalable slice-level supervision from reports. We extract, from a large database of scan-report pairs, descriptions of findings with indices of slices where those findings occur. We then finetune a general-purpose 2D image segmentation model, SAM3, with standard segmentation losses from strongly labeled data and with a slice-level classification loss from the extracted weak supervision. Our results on the ReXGroundingCT dataset illustrate that this strategy improves the segmentation dice score: from an 8% relative gain when there are 1000 fully labeled volumes to 22% when there are 250 fully labeled volumes.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.