2605.29299v1 May 28, 2026 cs.CV

Pocket-Dentist: 효율적인 다중 모드 대규모 언어 모델을 활용한 온디바이스 치과 영상 이해

Pocket-Dentist: On-Device Dental Image Understanding via Efficient Multimodal Large Language Models

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Ting Dang
Ting Dang
Citations: 134
h-index: 6
Hong Jia
Hong Jia
Citations: 100
h-index: 5
Kai Bian
Kai Bian
Citations: 5
h-index: 1
Xucheng Guo
Xucheng Guo
Citations: 0
h-index: 0
Bin Chen
Bin Chen
Citations: 14
h-index: 1
Lingyan Ruan
Lingyan Ruan
Citations: 239
h-index: 7

치과 시각-언어 모델의 평가는 데이터셋, 작업 정의 및 지표 측면에서 분산되어 있으며, 종종 계산 비용을 간과합니다. 이는 전문 의료 센터 외 지역에서의 치과 검진 적용을 제한하는데, 왜냐하면 신속한 추론, 제한된 하드웨어, 그리고 환자 이미지에 대한 로컬 처리 기능이 실용적이고 개인 정보 보호를 위한 사전 임상 검사에 필수적이기 때문입니다. 본 연구에서는 효율성을 고려한 치과 다중 모드 질의응답 벤치마크인 Pocket-Dentist를 소개합니다. 이 벤치마크는 약 1,159명의 환자를 포함하는 세 가지 데이터셋, 다섯 가지 작업 유형 및 일곱 가지 지표를 통합합니다. 일반적인 14개의 시각-언어 모델(VLM)에 대한 분석 결과, 흥미로운 사실을 발견했습니다. 즉, 컴팩트한 VLM(예: 20억 파라미터 모델)이 더 큰 VLM보다 정확도가 높으면서도 치과 영상 이해에서 훨씬 낮은 계산 비용을 요구한다는 것입니다. iPhone 17 Pro에 로컬로 배포된 미세 조정된 컴팩트 VLM인 Pocket-Dentist-2B는 각 샘플을 4.31초 만에 처리하여, 기준이 되는 70억 파라미터 모델보다 지연 시간을 4.9배 줄이고 메모리 사용량을 2.3배 감소시켰습니다.

Original Abstract

Evaluations of dental vision-language models remain fragmented across datasets, task definitions and metrics, and often ignore their computational cost. This limits their widespread deployment for dental screening outside specialist centres, where timely inference, limited hardware, and local handling of patient images are vital for practical, privacy-preserving clinical prescreening. Here we present Pocket-Dentist, an efficiency-aware benchmark for dental multimodal question answering that brings together three datasets spanning approximately 1,159 patients, five task types and seven metrics. Across typical 14 VLMs, our results reveals an interesting observation: compact VLMs (e.g., 2B-parameter models) outperform larger VLMs in accuracy while requiring substantially lower computational costs in dental image understanding. Deployed locally on an iPhone 17 Pro, our finetuned compact VLM Pocket-Dentist-2B processed each sample in 4.31 s, reducing latency by 4.9-fold and memory use by 2.3-fold compared with a 7B baseline.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!