STAIL: 대규모 언어 모델을 활용한 의료 영상 분석에서의 의미 기반 텍스트 연동 점진적 학습
STAIL: Semantic Text-Anchored Incremental Learning for Medical Imaging via Large Language Models
의료 영상 분석에 적용되는 심층 학습 모델은 동적인 환경에서 새로운 임상 작업에 지속적으로 적응할 때 심각한 재앙적 망각 현상을 겪습니다. 기존의 점진적 학습 방법들은 일반적으로 과거 데이터를 반복 학습하여 이러한 문제를 완화하지만, 이 방식은 상당한 저장 공간을 필요로 하고, 개인 정보 보호 문제를 야기하며, 희소한 데이터 샘플로는 실제 데이터 분포를 제대로 반영하지 못합니다. 인간 인지 메커니즘에서 영감을 받아, 우리는 순차적인 임상 작업에 적합한 새로운 프레임워크인 의미 기반 텍스트 연동 점진적 학습(Semantic Text-Anchored Incremental Learning, STAIL)을 제안합니다. STAIL은 재학습의 한계를 극복하기 위해 비대칭 의미 통합 버퍼(Asymmetric Semantic Consolidation Buffer, SCB)를 도입했습니다. 최소한의 이미지 앵커와 광범위한 텍스트 설명을 포함하는 SCB는 저장 공간 비용을 최소화하면서 이전 작업의 풍부한 의미 정보를 재구성할 수 있도록 합니다. 또한, 우리는 개발 초기 단계에서 안정적인 의미 공간을 활용하는 대규모 언어 모델 기반 의미 앵커링 메커니즘(LLM-derived Semantic Anchoring Mechanism, LSAM)을 설계했습니다. 이 메커니즘은 진화하는 시각적 특징을 텍스트 표현에 명시적으로 연결하여 거시적 및 미시적 수준에서 가소성과 안정성을 안내하고 제어합니다. 망막 검사, 초음파, X-ray 이미지를 포함한 세 가지 다양한 의료 데이터셋에 대한 광범위한 실험 결과, STAIL은 매우 효과적인 플러그 앤 플레이 모듈로 작용한다는 것을 보여줍니다. STAIL은 다양한 기존 모델의 성능을 종합적으로 향상시키며, 지속적인 성능 측면에서 평균 2.24%의 AAA-AUC 향상과 망각 감소 측면에서 3.55%의 BWT-AUC 향상을 달성했습니다. 코드 공개 예정입니다.
Deep learning models applied to medical image analysis suffer from severe catastrophic forgetting when continually adapting to new clinical tasks in dynamic environments. Mainstream incremental learning methods typically mitigate this by rehearsing raw historical images. However, this pixel-level rehearsal incurs significant storage overhead, raises privacy concerns, and fails to adequately capture the true data distribution with sparse exemplars. Inspired by human cognitive mechanisms, we propose a novel framework termed Semantic Text-Anchored Incremental Learning (STAIL) for sequential clinical tasks. To overcome the rehearsal bottleneck, STAIL introduces an asymmetric semantic consolidation buffer (SCB). By incorporating a minimal set of image anchors and extensive textual descriptions, the SCB enables dense semantic reconstruction of old tasks at a minimal storage cost. Furthermore, we design an LLM-derived Semantic Anchoring Mechanism (LSAM) that leverages the stable semantic space of frozen large language models as developmental priors. This mechanism explicitly anchors evolving visual features to textual representations, guiding and constraining plasticity and stability at both macroscopic and microscopic levels. Extensive experiments across three heterogeneous medical datasets, covering fundus, ultrasound, and X-ray imaging, demonstrate that STAIL acts as a highly effective plug-and-play module. It comprehensively enhances the performance of various existing baselines, achieving average gains of 2.24\% in AAA-AUC for sustained performance and 3.55\% in BWT-AUC for reduced forgetting. Code is available.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.