Pocket-Dentist: 효율적인 다중 모드 대규모 언어 모델을 활용한 온디바이스 치과 영상 이해
Pocket-Dentist: On-Device Dental Image Understanding via Efficient Multimodal Large Language Models
치과 시각-언어 모델의 평가는 데이터셋, 작업 정의 및 지표 측면에서 분산되어 있으며, 종종 계산 비용을 간과합니다. 이는 전문 의료 센터 외 지역에서의 치과 검진 적용을 제한하는데, 왜냐하면 신속한 추론, 제한된 하드웨어, 그리고 환자 이미지에 대한 로컬 처리 기능이 실용적이고 개인 정보 보호를 위한 사전 임상 검사에 필수적이기 때문입니다. 본 연구에서는 효율성을 고려한 치과 다중 모드 질의응답 벤치마크인 Pocket-Dentist를 소개합니다. 이 벤치마크는 약 1,159명의 환자를 포함하는 세 가지 데이터셋, 다섯 가지 작업 유형 및 일곱 가지 지표를 통합합니다. 일반적인 14개의 시각-언어 모델(VLM)에 대한 분석 결과, 흥미로운 사실을 발견했습니다. 즉, 컴팩트한 VLM(예: 20억 파라미터 모델)이 더 큰 VLM보다 정확도가 높으면서도 치과 영상 이해에서 훨씬 낮은 계산 비용을 요구한다는 것입니다. iPhone 17 Pro에 로컬로 배포된 미세 조정된 컴팩트 VLM인 Pocket-Dentist-2B는 각 샘플을 4.31초 만에 처리하여, 기준이 되는 70억 파라미터 모델보다 지연 시간을 4.9배 줄이고 메모리 사용량을 2.3배 감소시켰습니다.
Evaluations of dental vision-language models remain fragmented across datasets, task definitions and metrics, and often ignore their computational cost. This limits their widespread deployment for dental screening outside specialist centres, where timely inference, limited hardware, and local handling of patient images are vital for practical, privacy-preserving clinical prescreening. Here we present Pocket-Dentist, an efficiency-aware benchmark for dental multimodal question answering that brings together three datasets spanning approximately 1,159 patients, five task types and seven metrics. Across typical 14 VLMs, our results reveals an interesting observation: compact VLMs (e.g., 2B-parameter models) outperform larger VLMs in accuracy while requiring substantially lower computational costs in dental image understanding. Deployed locally on an iPhone 17 Pro, our finetuned compact VLM Pocket-Dentist-2B processed each sample in 4.31 s, reducing latency by 4.9-fold and memory use by 2.3-fold compared with a 7B baseline.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.