2606.18063v1 Jun 16, 2026 cs.CV

LLM이 상처를 분석할 때: 이미지에서 임상적으로 의미 있는 특징으로

When LLMs Analyze Scars: From Images to Clinically-Meaningful Features

Hangting Ye
Hangting Ye
Citations: 148
h-index: 7
Ruman Wang
Ruman Wang
Citations: 0
h-index: 0

의료 영상 분류는 근본적인 난제에 직면합니다. 딥러닝 모델은 뛰어난 성능을 보이지만, 실제 임상 시나리오에서는 어노테이션 비용, 개인 정보 보호 제약 및 질병 희귀성 등으로 인해 심각한 데이터 부족 문제가 발생합니다. 이러한 문제는 특히 병리적 상처 분류에서 두드러지는데, 케로이드와 과도증식 흉터를 구별하려면 미묘한 전문 지식이 필요하며, 레이블이 지정된 이미지는 극히 제한적입니다. 본 연구에서는 대규모 언어 모델(LLM)을 종단 간 분류기가 아닌 지식 기반 특징 엔지니어링 도구로 재배치하는 새로운 패러다임을 제안합니다. 우리는 이 프레임워크를 ScaFE (Scar Feature Engineering)라고 부릅니다. 핵심적인 아이디어는 LLM이 풍부한 의료 지식을 포함하고 있으며, 이러한 지식이 실행 가능한 특징 추출 코드로 외부화될 수 있어 고차원 이미지를 저차원의 임상적으로 해석 가능한 표현으로 변환할 수 있다는 것입니다. 구체적으로, 우리는 확립된 상처 평가 기준을 사용하여 LLM에 프롬프트를 제공하여 밴쿠버 흉터 평가 척도와 같은 임상적 점수 시스템과 일치하는 특징을 추출하는 결정론적인 Python 코드를 생성하도록 합니다. 본 연구의 접근 방식은 세 가지 주요 장점을 제공합니다: (1) 데이터 효율성, 즉 지식 습득과 통계 학습을 분리하여 제한된 훈련 샘플로 강력한 성능을 달성합니다; (2) 개인 정보 보호, 즉 원본 이미지는 외부 LLM에 노출되지 않고 로컬에서 처리됩니다; (3) 명확한 특징을 통해 임상적 추론에 기반한 해석 가능성을 제공합니다. 상처 분류에 대한 광범위한 실험 결과, 본 연구의 방법은 제한된 데이터 조건 하에서 기존의 종단 간 딥러닝 모델 또는 LLM을 블랙박스 분류기로 사용하는 것보다 일관되게 우수한 성능을 보이며, 이는 데이터 효율적이고 임상적으로 투명한 의료 AI 시스템에 LLM을 통합하는 유망한 방향을 제시합니다.

Original Abstract

Medical image classification faces a fundamental dilemma: while deep learning models achieve remarkable performance at scale, real-world clinical scenarios often suffer from severe data scarcity due to annotation costs, privacy constraints, and disease rarity. This challenge is particularly pronounced in pathological scar classification, where differentiating keloids from hypertrophic scars requires subtle expert knowledge and labeled images are extremely limited. We propose a novel paradigm that repositions large language models (LLMs) as knowledge-driven feature engineers rather than end-to-end classifiers. We call this framework ScaFE (Scar Feature Engineering). Our key insight is that LLMs encode rich medical knowledge that can be externalized as executable feature extraction code, enabling the transformation of high-dimensional images into low-dimensional, clinically interpretable representations. Specifically, we prompt an LLM with established scar assessment criteria to generate deterministic Python code that extracts features aligned with clinical scoring systems such as the Vancouver Scar Scale. Our approach offers three key advantages: (1) data efficiency, achieving robust performance with limited training samples by decoupling knowledge acquisition from statistical learning; (2) privacy preservation, as raw images are processed locally without exposure to external LLMs; and (3) interpretability, through explicit features grounded in clinical reasoning. Extensive experiments on scar classification demonstrate that our method consistently outperforms end-to-end deep learning baselines or using LLMs as black-box classifiers under limited data conditions, establishing a promising direction for integrating LLMs into data-efficient and clinically transparent medical AI systems.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!