장기적인 개인 건강 관리를 위한 자체 진화 에이전트
A Self-Evolving Agent for Longitudinal Personal Health Management
개인 건강 관리는 반복적인 상호 작용을 통해 이루어지지만, 대부분의 건강 AI 시스템은 각 요청을 독립적으로 처리합니다. 본 연구에서는 사용자의 일상생활, 선호도, 측정값 및 위험 요인이 변화함에 따라 지원 내용을 업데이트하는 오픈 소스 에이전트 아키텍처인 HealthClaw를 개발했습니다. HealthClaw는 공유된 안전 규칙 및 의료 지식을 개인의 장기적인 기억(프로필 정보, 재사용 가능한 절차 및 사건 기록 포함)과 분리합니다. 각 상호 작용 후에 유도 추론을 통해 프로필 업데이트, 절차 수정, 사건 기록 유지 또는 제외할 내용을 결정합니다. 우리는 HealthClaw를 합성된 1년 데이터 세트와 9개의 200개 사례 생물 의학 작업으로 평가했습니다. 900건의 장기적인 지원 요청에서, HealthClaw는 현재 쿼리 방식에 비해 응답 정확도를 0.2%에서 45.7%로 향상시켰으며, 프롬프트 측면에서의 컨텍스트 노출량은 전체 기록 방식을 사용할 때보다 71.7% 더 낮았습니다. 100건의 개인 정보 보호 테스트에서 HealthClaw는 기준 모델보다 높은 수준의 개인 정보 보호를 제공하는 답변 품질을 생성하고 안전하지 않은 정보 공개 건수가 적었습니다. 생물 의학 작업 전반에 걸쳐, 특정 작업의 주요 지표에서 평균 절대적 성능 향상은 27.0%p였으며, 그 중 7건은 거짓 발견율 보정 후에도 유의미한 개선을 보였습니다. 이러한 오프라인 벤치마크 결과는 장기적인 개인 건강 에이전트를 위한 관리되고 자체 진화하는 메모리 시스템의 가능성을 보여주지만, 임상적 효능은 전향적 평가를 통해 검증되어야 합니다. HealthClaw는 https://github.com/HC-Guo/HealthClaw 에서 공개적으로 이용할 수 있습니다.
Personal health management unfolds over repeated encounters, yet most health AI systems treat each request in isolation. We developed HealthClaw, an open-source agent architecture that updates support as a person's routines, preferences, measurements and risks change. It separates shared safety rules and medical knowledge from private longitudinal memory containing profile facts, reusable procedures and episodic traces. After each episode, induction determines what should update the profile, revise a procedure, remain episodic or be excluded. We evaluated HealthClaw with a synthetic year-long benchmark and nine 200-case biomedical tasks. Across 900 longitudinal support probes, answer accuracy increased from 0.2% with current-query prompting to 45.7% with HealthClaw, while prompt-side context exposure was 71.7% lower than with full-history prompting. In 100 privacy probes, HealthClaw produced higher privacy-aware answer quality and fewer unsafe disclosures than both baselines. Across the biomedical tasks, the mean absolute gain in the task-specific primary metric was 27.0 percentage points, and seven gains remained significant after false-discovery-rate correction. These offline benchmarks support governed, self-evolving memory for longitudinal personal health agents, although clinical effectiveness requires prospective evaluation. HealthClaw is publicly available at https://github.com/HC-Guo/HealthClaw.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.