2606.18147v1 Jun 16, 2026 cs.AI

WEQA: 웨어러블 건강 데이터 질의 응답 시스템 - 쿼리 적응형 에이전트 기반 추론

WEQA: Wearable hEalth Question Answering with Query-Adaptive Agentic Reasoning

Y. Wu
Y. Wu
Citations: 97
h-index: 5
Tong Xia
Tong Xia
Citations: 1,357
h-index: 15
Xin Liu
Xin Liu
Citations: 23
h-index: 2
D. McDuff
D. McDuff
Citations: 3,743
h-index: 11
Yuwei Zhang
Yuwei Zhang
Citations: 93
h-index: 4
Bianca Emmerich
Bianca Emmerich
Citations: 0
h-index: 0
Dimitris Spathis
Dimitris Spathis
Citations: 44
h-index: 3
Cecilia Mascolo
Cecilia Mascolo
Citations: 49
h-index: 1

언어 모델은 의료 분야 질의 응답에서 놀라운 성능을 보이며, 때로는 일반 의사의 정확도를 능가하기도 합니다. 그러나 웨어러블 기기에서 생성되는 건강 데이터에 대한 질의 응답은 여전히 어려운 과제이며 연구가 부족한 영역입니다. 이러한 웨어러블 센서는 지속적이고 고차원적인 데이터를 생성하며, 이는 LLM(Large Language Model)의 사전 학습 과정에서 사용되는 텍스트 기반 데이터 분포와 일치시키기 어렵습니다. 다양한 센서 모드 및 사용자 의도를 효과적으로 처리하기 위해서는 고정된 추론 흐름이나 단일의 사전 학습된 모델로는 한계가 있습니다. 이러한 문제점을 해결하기 위해, 우리는 LLM 추론과 특수 웨어러블 분석 및 모델링 도구를 통합하는 쿼리 적응형 에이전트 프레임워크인 WEQA를 제안합니다. LLM 컨트롤러는 실행 계획을 생성하고 각 질의를 적절한 센서 분석 및 사전 학습된 모델 조합으로 동적으로 라우팅하며, 외부 지식을 활용하여 답변의 정확성을 검증합니다. 또한, 우리는 세 가지 건강 분야에서 분석 및 예측 작업을 수행하는 네 개의 공개 웨어러블 데이터셋으로 구성된 벤치마크를 구축했습니다. 실험 결과, 제안하는 프레임워크는 LLM 및 에이전트 기반 모델보다 24% 더 높은 정확도를 보였으며, 12명의 의료 전문가와 8명의 사용자를 대상으로 실시한 블라인드 테스트에서 유용성 및 임상적 타당성 측면에서 상당한 개선을 확인했습니다.

Original Abstract

Language models are remarkably capable at medical question answering, in some cases surpassing the accuracy of general physicians. However, answering questions about wearable health data remains challenging and understudied, as these ubiquitous sensors produce continuous, high-dimensional, and longitudinal data, which is non-trivial to align with text-centric distributions in LLM pretraining. The diversity of sensor modalities and user intents cannot be effectively handled by a fixed reasoning workflow or a single pretrained foundation model. To address these challenges, we propose WEQA, a query-adaptive agent framework that unifies LLM reasoning with specialized wearable analytical and modeling tools. An LLM controller is employed to synthesize execution plans and dynamically route each query to the appropriate combination of sensor analysis and pretrained models, and perform grounded response auditing with external knowledge. We also curate a benchmark spanning four open wearable datasets comprising analytic and predictive tasks in three different health domains. Experiments show that our framework is 24% more accurate than LLM and agentic baselines, and a blinded study with 12 medical experts and 8 users shows substantial gains in usefulness and clinical soundness.

0 Citations
0 Influential
7.5 Altmetric
37.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!