검색 증강 해석 가능 학습: 의료 분야의 작업별 제로샷 모델을 향하여
Retrieval-Augmented Interpretable Learning: Towards Task-Specific Zero-Shot Models in Healthcare
본 논문에서는 검색 증강 해석 가능 학습(RAIL)이라는 확률적 메타 학습 프레임워크를 소개합니다. RAIL은 자연어 기반의 작업 설명과 이전에 학습된 작업별 예측 모델들의 정보를 활용하여, 작업에 특화된 해석 가능한 모델을 제로샷 방식으로 생성합니다. RAIL은 관련 작업들을 검색하고, 계수 공간을 통해 구조를 전달하며, 원래 진단-특징 공간에서 새로운 예측 모델을 생성합니다. 이를 통해 특징 수준에서의 설명과 함께 제로샷 및 소량 데이터 기반의 임상 절차 예측이 가능합니다. RAIL의 확률적 형식은 검색 과정, 모델 계수 및 예측에 대한 불확실성을 제공하여 신뢰성 확보를 위한 배포를 지원합니다. 즉, 불확실한 예측이나 불안정한 설명은 자동 결정으로 간주되기보다는 추가적인 임상 검토를 위해 표시될 수 있습니다. 이는 RAIL이 의료 환경에서 특히 유용하며, 여기서 예측 작업은 데이터 분포가 매우 편향되어 있고, 새로운 임상 목표가 빈번하게 발생하며, 모델은 검사가 가능하고 불확실성을 고려하며 인간의 감독과 호환되어야 합니다. 다양한 임상 절차 예측 작업에서 RAIL은 데이터 가용성 수준에 관계없이 안정적인 성능을 유지합니다. 제로샷 환경에서는 73.4%의 정확도를 달성하며, 이는 지도 학습 방식으로 특정 작업을 위한 모델을 학습할 수 없는 경우입니다. 또한, 단 2~4개의 예시만 있는 극히 소량 데이터 환경에서도 약 73.2%의 정확도를 유지하며, 이 수준은 일반적인 지도 학습 방식의 작업별 모델 성능에 가깝습니다. RAIL은 임상적 지식을 반영한 작업 표현을 활용하여 검색, 불확실성 및 계수 수준의 진단 정보를 제공함으로써 모델의 동작 방식을 더욱 투명하게 만듭니다. 이러한 결과는 새로운 작업에 적응하면서도 해석 가능성과 신뢰성을 유지할 수 있는 확장 가능한 임상 예측 시스템 개발을 위한 중요한 방향을 제시합니다.
We introduce Retrieval-Augmented Interpretable Learning (RAIL), a probabilistic meta-learning framework for zero-shot generation of task-specific interpretable models that synthesizes coefficient-space structure from natural-language task descriptions and a memory of previously learned task-specific predictors. RAIL retrieves related source tasks, transfers structure through coefficient space, and generates a new predictor in the original diagnostic-feature space, enabling zero-shot and few-shot clinical procedure prediction with feature-level explanations. Its probabilistic formulation provides uncertainty over retrieval, model coefficients, and predictions, supporting reliability-aware deployment: uncertain predictions or unstable explanations can be flagged for additional clinical review rather than treated as automatic decisions. This makes RAIL particularly suited for healthcare settings, where prediction tasks are highly long-tailed, new clinical targets arise frequently, and models must remain inspectable, uncertainty-aware, and compatible with human oversight. Across long-tailed clinical procedure prediction tasks, RAIL maintains reliable performance across data-availability regimes: it achieves 73.4% accuracy in the held-out zero-shot settings, where no supervised task-specific model can be trained, and remains near 73.2% accuracy in the extreme few-shot regime with only 2-4 examples, where supervised task-specific models perform close to chance. RAIL further benefits from clinically informed task representations and yields retrieval, uncertainty, and coefficient-level diagnostics that make model behavior more transparent. These results suggest a path toward scalable clinical prediction systems that can adapt to new tasks while preserving interpretability and reliability.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.