EasyLens: 의료용 시각-언어 모델을 위한 학습 불필요한 플러그 앤 플레이 방식의 미세 병변 표현 증폭기
EasyLens: A Training-Free Plug-and-Play Subtle-Lesion Representation Amplifier for Medical Vision-Language Models
의료용 시각-언어 모델(VLMs)은 임상 이미지 해석, 특히 병변 감지 및 보고서 생성에 상당한 잠재력을 보여왔습니다. 그러나 이들의 실제 유용성은 종종 희소하고 대비가 낮으며 복잡한 해부학적 맥락 내에 존재하는 미세 병변에 대한 민감도가 부족하여 제한됩니다. 로컬 시각 토큰이 집계됨에 따라, 이러한 약한 병변 신호는 전역 이미지 표현에서 과소 대표될 수 있으며, 이는 의료용 VLM이 인식하기 어렵게 만듭니다. 병변 감지 능력을 향상시키기 위한 기존 노력은 주로 의료 분야의 비전 인코더 사전 훈련, 임상 용어 기반 정렬 또는 학습 가능한 병리학적 표현 강화에 의존합니다. 이러한 접근 방식은 효과적이지만 일반적으로 추가적인 학습이나 모델별 적응이 필요하며 특정 질병 형태에 과적합될 수 있으므로, 기존 의료용 VLM에 적용하기 어렵습니다. 이러한 제한 사항을 해결하기 위해, 저희는 의료용 VLM을 위한 학습 불필요한 플러그 앤 플레이 방식의 미세 병변 표현 증폭기인 EasyLens를 제안합니다. EasyLens는 먼저 EasyBank라는 병리학-해부학 프로토타입 공간을 구축하여, 의심스러운 영역을 병리학적 패턴과 해부학적으로 인식된 정상 기준과 비교할 수 있도록 합니다. EasyTag는 정상 조직에 대한 맹목적인 증폭을 방지하기 위해 반사실적 프로토타입 추론을 통해 병변과 관련된 영역을 선택합니다. 또한, 미세 병변 신호가 전역 이미지 표현에서 희석되는 것을 막기 위해, EasyAmplifier는 형태학 기반 잔차 강화 방법을 사용하여 선택된 병변 관련 영역의 표현을 강화하여 전역 이미지 임베딩에 대한 기여도를 높입니다. 여러 의료 이미지 데이터 세트 및 기존 의료용 VLM 백본에서의 실험 결과, EasyLens는 미세 병변 감지를 개선하고 기존 인코더 기반 방법보다 우수한 성능을 보였습니다.
Medical vision-language models (VLMs) have shown increasing potential for clinical image interpretation, including lesion detection and report generation. However, their practical utility remains limited by insufficient sensitivity to subtle lesions, whose visual evidence is often sparse, low-contrast, and embedded within complex anatomical context. As local visual tokens are aggregated, these weak lesion cues can become underrepresented in global image representations, making them difficult for medical VLMs to recognize. Existing efforts to improve lesion sensitivity mainly rely on medical-domain vision-encoder pre-training, clinical-term-guided alignment, or trainable pathological representation enhancement. Although effective, these approaches usually require additional training or model-specific adaptation and may overfit to particular disease morphologies, limiting their applicability to frozen medical VLMs. To address these limitations, we propose EasyLens, a training-free plug-and-play subtle-lesion representation amplifier for medical VLMs. EasyLens first constructs EasyBank, a pathology-anatomy prototype space that provides lesion-related prototypes and anatomy-aware normal references for comparing suspicious patches against both pathological and normal anatomical patterns. To avoid blindly amplifying normal tissues, EasyTag selects lesion-relevant patches through counterfactual prototype reasoning. To counteract the dilution of subtle lesion cues in global image representations, EasyAmplifier strengthens the selected lesion-relevant patch representations through morphology-guided residual enhancement, thereby increasing their contribution to the global image embedding. Experiments on multiple medical image datasets and frozen medical VLM backbones show that EasyLens improves subtle-lesion detection and outperforms existing encoder-enhancement baselines.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.