2607.28538v1 Jul 30, 2026 cs.CV

ScaFE: LLM 기반 임상 특징 프로그램으로 구현된 데이터 효율적인 흉터 분류

ScaFE: Data-Efficient Scar Classification with LLM-Generated Clinical Feature Programs

Hangting Ye
Hangting Ye
Citations: 148
h-index: 7
Ruman Wang
Ruman Wang
Citations: 0
h-index: 0

임상 사진에서 병리적 흉터를 분류하는 것은 제한적인 전문가 레이블 데이터를 가지고 있고, 병원 간 이미지 획득 과정의 상당한 차이가 존재함에도 불구하고, 케로이드와 과도성 흉터를 구별해야 합니다. 엔드 투 엔드 이미지 모델은 여전히 데이터에 의존하는 반면, 사진을 호스팅된 시각-언어 모델(VLM)로 보내는 것은 현지 데이터 거버넌스 요구 사항과 충돌할 수 있으며, 결과에 대한 재현성과 감사 가능성을 저해합니다. 본 연구에서는 LLM으로부터 얻은 임상 지식을 이미지 진단을 수행하는 대신, 결정적이고 실행 가능한 특징 프로그램으로 변환하는 ScaFE(Scar Feature Engineering)를 소개합니다. 웹 기반 LLM은 임상 증거를 검색하고 시각적으로 평가할 수 있는 흉터 속성을 측정하는 프로그램을 생성합니다. 후보 프로그램은 제한된 로컬 환경에서 실행되며, 반복적인 개선 및 보정을 위해 집계된 검증 통계 및 특징 수준의 SHAP 요약 정보만 반환됩니다. 원본 이미지와 환자 레벨의 출력 결과는 현지 시스템에 유지됩니다. 경량화된 랜덤 포레스트는 이러한 구조화된 표현을 기반으로 작동합니다. 세 병원에서 수집된 600장의 사진을 사용하여, leave-one-site-out 방식으로 평가한 결과, ScaFE는 81.0%의 사이트별 가중 평균 정확도를 달성하여, 가장 강력한 기준 모델인 BiomedCLIP보다 10.0%p 더 높은 성능을 보였습니다. 개발 데이터의 10%만을 사용하여 ScaFE는 72.0%의 가중 평균 정확도를 유지했으며, 이는 11.8%p 더 높은 수치입니다. 반복적인 개선 과정을 통해 실행 가능한 프로그램의 비율은 66.7%에서 95.0%로 향상되었으며, 최종 특징의 91.7%에 대해 검증된 증거가 제공되었습니다. 이러한 결과는 LLM 지식이 현지 시스템에서 운영되며 감사 가능성이 높은 특징 프로그램을 통해 데이터 효율적인 의료 이미지 분류를 지원할 수 있음을 보여줍니다.

Original Abstract

Classifying pathological scars from clinical photographs requires distinguishing keloids from hypertrophic scars despite limited expert-labeled data and substantial acquisition variation across hospitals. End-to-end image models remain data-dependent, whereas sending photographs to a hosted vision-language model (VLM) may conflict with local data-governance requirements and yields decisions that are difficult to reproduce and audit. We introduce ScaFE (Scar Feature Engineering), which transfers clinical knowledge from a large language model (LLM) into deterministic, executable feature programs instead of asking the model to diagnose images. A web-enabled LLM retrieves clinical evidence and synthesizes programs that measure visually assessable scar attributes. Candidate programs execute in a restricted local environment, and only aggregate validation statistics and feature-level SHAP summaries are returned for iterative repair and refinement; raw images and patient-level outputs remain local. A lightweight Random Forest then operates on the resulting structured representation. On 600 photographs from three hospitals under leave-one-site-out evaluation, ScaFE achieves 81.0% site-macro balanced accuracy, exceeding the strongest baseline, BiomedCLIP, by 10.0 percentage points. With only 10% of the development data, ScaFE retains 72.0% balanced accuracy and an 11.8-point lead. Iterative refinement also raises the executable-program rate from 66.7% to 95.0%, with verified evidence for 91.7% of the final features. These results show that LLM knowledge can support data-efficient, cross-site medical image classification through local and auditable feature programs rather than direct VLM decisions.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!