2601.14084v1 Jan 20, 2026 cs.CV

DermaBench: 피부과 시각적 질의응답 및 추론을 위한 임상의 주석 벤치마크 데이터셋

DermaBench: A Clinician-Annotated Benchmark Dataset for Dermatology Visual Question Answering and Reasoning

Abdurrahim Yilmaz
Abdurrahim Yilmaz
Citations: 43
h-index: 3
Ozan Erdem
Ozan Erdem
Citations: 1
h-index: 1
Ece Gokyayla
Ece Gokyayla
Citations: 1
h-index: 1
Burc Bugra Dagtas
Burc Bugra Dagtas
Citations: 0
h-index: 0
Dilara İlhan Erdil
Dilara İlhan Erdil
Citations: 171
h-index: 3
G. Gencoglan
G. Gencoglan
Citations: 652
h-index: 15
Burak Temelkuran
Burak Temelkuran
Citations: 459
h-index: 11
A. Acar
A. Acar
Citations: 141
h-index: 7

비전-언어 모델(VLM)은 의료 애플리케이션에서 점점 더 중요해지고 있지만, 피부과 분야에서의 평가는 주로 병변 인식과 같은 이미지 수준의 분류 작업에 초점을 맞춘 데이터셋으로 인해 여전히 제한적입니다. 이러한 데이터셋은 인식에는 유용하지만, 멀티모달 모델의 완전한 시각적 이해, 언어 그라운딩 및 임상적 추론 능력을 평가할 수는 없습니다. 모델이 피부과 이미지를 해석하고, 미세한 형태학적 특징을 추론하며, 임상적으로 의미 있는 설명을 생성하는 방식을 평가하기 위해서는 시각적 질의응답(VQA) 벤치마크가 필요합니다. 이에 우리는 Diverse Dermatology Images (DDI) 데이터셋을 기반으로 구축된, 임상의가 주석을 단 피부과 VQA 벤치마크인 DermaBench를 소개합니다. DermaBench는 피츠패트릭 피부 타입 I-VI에 걸친 570명의 환자로부터 얻은 656장의 임상 이미지로 구성되어 있습니다. 전문 피부과 의사들은 22개의 주요 질문(단일 선택, 다중 선택, 개방형)으로 구성된 계층적 주석 스키마를 사용하여 각 이미지에 대해 진단, 해부학적 위치, 병변 형태, 분포, 표면 특징, 색상 및 이미지 품질을 주석 처리하였으며, 개방형 서술 설명 및 요약을 포함하여 약 14,474개의 VQA 스타일 주석을 생성했습니다. DermaBench는 원천 라이선스를 준수하기 위해 메타데이터 전용 데이터셋으로 공개되며 Harvard Dataverse에서 이용할 수 있습니다.

Original Abstract

Vision-language models (VLMs) are increasingly important in medical applications; however, their evaluation in dermatology remains limited by datasets that focus primarily on image-level classification tasks such as lesion recognition. While valuable for recognition, such datasets cannot assess the full visual understanding, language grounding, and clinical reasoning capabilities of multimodal models. Visual question answering (VQA) benchmarks are required to evaluate how models interpret dermatological images, reason over fine-grained morphology, and generate clinically meaningful descriptions. We introduce DermaBench, a clinician-annotated dermatology VQA benchmark built on the Diverse Dermatology Images (DDI) dataset. DermaBench comprises 656 clinical images from 570 unique patients spanning Fitzpatrick skin types I-VI. Using a hierarchical annotation schema with 22 main questions (single-choice, multi-choice, and open-ended), expert dermatologists annotated each image for diagnosis, anatomic site, lesion morphology, distribution, surface features, color, and image quality, together with open-ended narrative descriptions and summaries, yielding approximately 14.474 VQA-style annotations. DermaBench is released as a metadata-only dataset to respect upstream licensing and is publicly available at Harvard Dataverse.

2 Citations
0 Influential
7.5 Altmetric
39.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!