2606.18203v1 Jun 16, 2026 cs.CL

RubricsTree: 건강 정보 및 의료 기술 분야에서 개인 맞춤형 건강 관리 시스템의 확장 가능하고 발전적인 개방형 평가 방법

RubricsTree: Scalable and Evolving Open-Ended Evaluation of Personal Health Agents across Health Memory and Medical Skills

Philip S. Yu
Philip S. Yu
Citations: 149
h-index: 7
Simon A. Lee
Simon A. Lee
UCLA
Citations: 246
h-index: 10
Zechen Li
Zechen Li
University of New South Wales
Citations: 144
h-index: 7
Weizhi Zhang
Weizhi Zhang
Citations: 443
h-index: 9
Ahmed A. Metwally
Ahmed A. Metwally
Citations: 116
h-index: 5
Mark Malhotra
Mark Malhotra
Citations: 397
h-index: 11
D. McDuff
D. McDuff
Citations: 3,743
h-index: 11
Salman Rahman
Salman Rahman
Citations: 346
h-index: 3
R. Luo
R. Luo
Citations: 320
h-index: 8
Shwetak Patel
Shwetak Patel
Citations: 1
h-index: 1
Hamid Palangi
Hamid Palangi
Citations: 59
h-index: 4
Benjamin Graef
Benjamin Graef
Citations: 29
h-index: 2
A. Heydari
A. Heydari
Citations: 0
h-index: 0
Zeinab Esmaeilpour
Zeinab Esmaeilpour
Citations: 16
h-index: 2
Erik Schenck
Erik Schenck
Citations: 116
h-index: 4
Chloe Zhang
Chloe Zhang
Citations: 0
h-index: 0
Yamin Li
Yamin Li
Citations: 0
h-index: 0
Menglian Zhou
Menglian Zhou
Citations: 73
h-index: 5
L. Sunden
L. Sunden
Citations: 9
h-index: 2

사용자 건강 데이터(센서)를 활용하는 LLM 기반 개인 맞춤형 건강 관리 시스템은 의료 접근성의 불평등을 완화할 수 있는 유망한 솔루션입니다. 그러나 대규모 임상 적용은 개방형 평가의 어려움 때문에 제한적입니다. 의사 주석은 신뢰성이 높지만 비용이 많이 들고 확장 가능하지 않으며, LLM 기반 평가 시스템은 확장 가능하지만 주관적이고 일관성이 부족하며 때로는 임상적으로 잘못된 결과를 초래할 수 있습니다. 본 연구에서는 4,000건의 실제 사용자 질의 데이터를 기반으로 전문가 패널이 참여하는 반복적인 인간-루프 관리 프로토콜을 통해 얻은 인사이트를 바탕으로 개발된 RubricsTree라는 확장 가능한 평가 프레임워크를 소개합니다. RubricsTree는 100개 이상의 세분화된, 임상적으로 검증 가능한 부울(Boolean) 항목으로 구성된 계층적 분류 체계를 가지고 있으며, 컨텍스트 인지 적응형 라우터는 각 질의에 따라 관련 있는 자동 가중치 항목 집합만 활성화하여 전문가 수준의 품질을 유지하면서 확장 가능한 평가를 가능하게 합니다. 체계적인 메타평가를 통해 RubricsTree가 (i) 어려운 개방형 질의에서 기존의 대규모 평가 기준보다 높은 수준의 전문가 일관성을 보여주며, (ii) 컨텍스트에 따라 저하된 응답을 정확하게 감지하고, (iii) 구조화된 지침, 텍스트 피드백 또는 성능 최적화를 위한 보상으로 사용될 때 Gemini, GPT 및 Qwen 모델 패밀리에 대해 HealthBench에서 최대 약 66%의 상대적인 성능 향상을 가져옴을 보여줍니다. 따라서 RubricsTree는 제품 수준의 개인 맞춤형 의료 AI를 지속적으로 개선하기 위해 필요한 확장 가능하고 감사 가능하며 발전 가능한 평가 인프라를 제공합니다.

Original Abstract

The LLM-empowered personal health agents with user health (sensor) metrics have offered a promising pathway to alleviate global disparities in healthcare access. However, large-scale clinical deployment remains constrained by an open-ended evaluation bottleneck: physician annotation is reliable but costly and unscalable, while LLM-as-a-judge evaluators are scalable but subjective, inconsistent, and sometimes clinically misaligned. We introduce RubricsTree, a scalable evaluation framework with an expert-aligned hierarchical taxonomy of over 100 atomic, clinically-verifiable Boolean rubrics, evolving from the insights of 4,000 real user queries through an iterative human-in-the-loop curation protocol with an expertise panel led by an experienced physician. A context-aware adaptive router activates only the relevant auto-weighted rubric subset per query, providing the throughput needed for scalable evaluation with expert-aligned quality. Through a systematic meta-evaluation, we show that RubricsTree (i) substantially exceeds a strong large-scale evaluation baseline in expert alignment on challenging open-ended queries; (ii) reliably penalizes contextually degraded responses; and (iii) when used as structured instructions, text feedback, or training rewards for performance optimization, yields up to ~66% relative gains on HealthBench for Gemini, GPT, and Qwen model families. RubricsTree thus provides a scalable, auditable, and evolving evaluation infrastructure required for the continuous optimization of product-level personal healthcare AI.

0 Citations
0 Influential
5.5 Altmetric
27.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!