2604.14866v1 Apr 16, 2026 cs.CV

MetaDent: 치과용 시각-언어 모델을 위한 임상 이미지 라벨링

MetaDent: Labeling Clinical Images for Vision-Language Models in Dentistry

Wenlong Deng
Wenlong Deng
Citations: 205
h-index: 7
Meng-Xun Li
Meng-Xun Li
Citations: 8
h-index: 1
Zhijian Wu
Zhijian Wu
Citations: 7
h-index: 1
Jiamin Wu
Jiamin Wu
Citations: 10
h-index: 2
Yue Han
Yue Han
Citations: 41
h-index: 2
James K. H. Tsoi
James K. H. Tsoi
Citations: 26
h-index: 3
Gui-Song Xia
Gui-Song Xia
Citations: 41
h-index: 4
Cui Huang
Cui Huang
Citations: 8
h-index: 1
C. Jin
C. Jin
Citations: 63
h-index: 3

시각-언어 모델(VLMs)은 의료 영상 분석 분야에서 상당한 잠재력을 보여주고 있지만, 세분화된 주석이 포함된 데이터셋과 종합적인 벤치마크의 부족으로 인해 구강 내 사진 분야에서의 활용은 아직 미흡한 실정입니다. 이러한 문제를 해결하기 위해, 우리는 MetaDent라는 종합적인 리소스를 제시합니다. MetaDent은 다음과 같은 내용을 포함합니다. (1) 임상, 공개 및 웹 소스에서 수집된 새로운 대규모 치과 이미지 데이터셋; (2) 치과 사진의 계층적이고 임상적으로 미묘한 특성을 반영하도록 설계된 반정형 주석 프레임워크; (3) 최첨단 VLM을 임상 이미지 이해 능력에 대해 평가하기 위한 종합적인 벤치마크 세트. 우리의 라벨링 방식은 고수준의 이미지 요약과 함께, 이상 징후에 대한 점별, 자유 텍스트 설명을 결합합니다. 이 방법은 풍부하고 확장 가능하며 작업에 독립적인 표현을 가능하게 합니다. 우리는 다양한 소스에서 60,669개의 치과 이미지를 수집하고, 이 중 대표적인 2,588개의 이미지를 위와 같은 메타-라벨링 방식으로 주석을 달았습니다. 대규모 언어 모델(LLM)을 활용하여, 약 15,000개의 시각적 질의 응답(VQA) 쌍과 18개 클래스의 멀티-라벨 분류 데이터셋과 같은 표준 벤치마크를 생성했습니다. 우리는 인간 검토 및 오류 분석을 통해 LLM 기반 전환이 신뢰성을 유지하고 의미적 정확성을 보장한다는 것을 확인했습니다. 그런 다음, 우리는 VQA, 분류 및 이미지 캡셔닝 작업에서 최첨단 VLM을 평가했습니다. 정량적 결과는 가장 발전된 모델조차도 구강 내 장면의 세분화된 이해에 어려움을 겪으며, 적당한 정확도를 달성하고 이미지 캡셔닝에서 일관성 없거나 불완전한 설명을 생성한다는 것을 보여줍니다. 우리는 데이터셋, 주석 및 도구를 공개하여 재현 가능한 연구를 촉진하고 치과 응용 분야를 위한 시각-언어 시스템 개발을 가속화하고자 합니다.

Original Abstract

Vision-Language Models (VLMs) have demonstrated significant potential in medical image analysis, yet their application in intraoral photography remains largely underexplored due to the lack of fine-grained, annotated datasets and comprehensive benchmarks. To address this, we present MetaDent, a comprehensive resource that includes (1) a novel and large-scale dentistry image dataset collected from clinical, public, and web sources; (2) a semi-structured annotation framework designed to capture the hierarchical and clinically nuanced nature of dental photography; and (3) comprehensive benchmark suites for evaluating state-of-the-art VLMs on clinical image understanding. Our labeling approach combines a high-level image summary with point-by-point, free-text descriptions of abnormalities. This method enables rich, scalable, and task-agnostic representations. We curated 60,669 dental images from diverse sources and annotated a representative subset of 2,588 images using this meta-labeling scheme. Leveraging Large Language Models (LLMs), we derive standardized benchmarks: approximately 15K Visual Question Answering (VQA) pairs and an 18-class multi-label classification dataset, which we validated with human review and error analysis to justify that the LLM-driven transition reliably preserves fidelity and semantic accuracy. We then evaluate state-of-the-art VLMs across VQA, classification, and image captioning tasks. Quantitative results reveal that even the most advanced models struggle with a fine-grained understanding of intraoral scenes, achieving moderate accuracy and producing inconsistent or incomplete descriptions in image captioning. We publicly release our dataset, annotations, and tools to foster reproducible research and accelerate the development of vision-language systems for dental applications.

1 Citations
0 Influential
3.5 Altmetric
18.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!