FDARxBench: FDA 제네릭 의약품 평가를 위한 규제 및 임상 추론 벤치마킹
FDARxBench: Benchmarking Regulatory and Clinical Reasoning on FDA Generic Drug Assessment
본 논문에서는 미국 식품의약국(FDA)의 의약품 설명서 문서를 활용하여, 문서 기반 질의응답(QA) 시스템을 평가하기 위한 전문가가 선별한 실제 기반 벤치마크인 FDARxBench를 소개합니다. 의약품 설명서에는 풍부하지만 이질적인 임상 및 규제 정보가 포함되어 있어, 현재의 언어 모델이 정확한 질의응답을 수행하기 어렵습니다. FDA 규제 평가 전문가와의 협력을 통해, FDARxBench를 개발하고, 사실 기반, 다단계 추론, 그리고 거부(refusal) 작업을 포괄하는 고품질의 전문가가 선별한 질의응답 예제를 생성하는 다단계 파이프라인을 구축했습니다. 또한, 개방형(open-book) 및 폐쇄형(closed-book) 추론 능력을 평가하기 위한 평가 프로토콜을 설계했습니다. 독점 모델과 공개 모델에 대한 실험 결과, 사실 기반 지식, 장문 맥락 검색, 그리고 안전한 거부 행동 측면에서 상당한 격차가 존재하는 것을 확인했습니다. 본 벤치마크는 FDA 제네릭 의약품 평가의 필요성에 의해 개발되었지만, 의약품 설명서 이해에 대한 규제 수준의 평가를 위한 중요한 기반을 제공합니다. 본 벤치마크는 LLM이 의약품 설명서 관련 질문에 대해 어떻게 작동하는지 평가하는 데 사용될 수 있도록 설계되었습니다.
We introduce an expert curated, real-world benchmark for evaluating document-grounded question-answering (QA) motivated by generic drug assessment, using the U.S. Food and Drug Administration (FDA) drug label documents. Drug labels contain rich but heterogeneous clinical and regulatory information, making accurate question answering difficult for current language models. In collaboration with FDA regulatory assessors, we introduce FDARxBench, and construct a multi-stage pipeline for generating high-quality, expert curated, QA examples spanning factual, multi-hop, and refusal tasks, and design evaluation protocols to assess both open-book and closed-book reasoning. Experiments across proprietary and open-weight models reveal substantial gaps in factual grounding, long-context retrieval, and safe refusal behavior. While motivated by FDA generic drug assessment needs, this benchmark also provides a substantial foundation for challenging regulatory-grade evaluation of label comprehension. The benchmark is designed to support evaluation of LLM behavior on drug-label questions.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.