SafeMed-R1: 의료 대규모 언어 모델의 안전성 및 윤리적 정렬을 위한 임상의 검증 시스템
SafeMed-R1: Clinician-Audited Safety and Ethics Alignment for Medical Large Language Models
대규모 언어 모델(LLM)은 점점 더 전문가 수준의 시험 성과를 보이고 있지만, 감사 가능한 추론, 안전 및 윤리적 정렬, 그리고 악의적인 오용에 대한 강건성을 요구하는 규제 때문에 실제 임상 환경에서의 활용은 제한적입니다. 본 연구에서는 추적 가능한 임상 신뢰성(CTS) 파이프라인을 통해 각 추론 과정을 임상의의 평가 기준 점수 및 편집 기록과 연결하고, 안전 및 윤리적 감독과 레드팀 스트레스 테스트를 통해 정렬된 SafeMed-R1 모델을 제시합니다. SafeMed-R1은 다양한 임상 벤치마크에서 평균 79.6%의 정확도를 달성했습니다. 적대적인 안전성 테스트에서는 가장 낮은 위험 점수를 보이며, 기준 모델에 비해 안전하지 않은 출력 결과를 약 3~5% 감소시켰습니다. 또한, 30개의 의약품 안전 관련 사례 연구에서 SafeMed-R1은 PGY1 및 PGY2 레지던트와 유사한 수준의 의료적 정확도를 보였으며, 의약품 안전성, 지침 일관성 및 임상 유용성 측면에서 더 높은 점수를 받았습니다. 종합적으로 볼 때, 본 연구 결과는 임상의 검증을 통한 감독 데이터 출처 정보와 함께, 특정 도메인에 맞춘 안전 및 윤리적 정렬이 추론 시 검색이나 인용 기반 접근 방식을 사용하지 않고도 규제 관련 증거를 강화할 수 있음을 시사합니다.
Large language models(LLMs) increasingly match expert performance on licensing examinations, yet routine clinical use remains limited because governance requires auditable reasoning, safety and ethics alignment, and resilience to adversarial misuse. Here we present SafeMed-R1, trained with a traceable Clinical Trust Signals(CTS) pipeline that links each reasoning instance to clinician rubric scores and edit histories, and aligned through safety and ethics supervision and red team stress testing. SafeMed-R1 attains a macro-averaged accuracy of 79.6% across clinical benchmarks. Under adversarial safety testing, it shows the lowest aggregated risk and reduces unsafe outputs by about 3 to 5% relative to its baseline. In a paired expert study of 30 medication safety vignettes, SafeMed-R1 matches PGY1 and PGY2 residents on medical correctness and scores higher for medication safety, guideline consistency, and clinical usefulness. Collectively, these results suggest that clinician-audited supervision provenance, together with domain-tailored safety and ethics alignment, can strengthen governance-relevant evidence without relying on inference-time retrieval or citation grounding.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.