ArogyaSutra: 인도어 기반 다중 모드 의료 추론을 위한 멀티 에이전트 프레임워크
ArogyaSutra: A Multi-Agent Framework for Multimodal Medical Reasoning in Indic Languages
멀티모달 대규모 언어 모델(MLLM)은 일반 영역에서 유망한 추론 능력을 보여주었지만, 특히 다국어 및 저자원 환경에서의 의료와 같은 전문 분야에서는 성능이 제한적입니다. 이러한 격차는 인도 농촌 지역과 같이 환자들이 복잡한 의료 질문을 자국어인 인도어로 표현하고 의료 이미지를 포함한 다양한 정보를 사용하는 곳에서 중요한 문제입니다. 기존의 영어 중심 MLLM은 이러한 활용 사례를 지원하기 어렵기 때문에 AI 기반 의료 지원에 대한 공정한 접근성을 제한합니다. 이 문제를 해결하기 위해, 우리는 8가지 이질적인 소스로부터 구축된 대규모 다국어 멀티모달 의료 질의응답 데이터셋인 ArogyaBodha를 소개합니다. 이 데이터셋은 영어와 함께 인도 주요 7개 언어를 포함하며, 31개의 인체 시스템, 6가지 이미징 모드 및 21개의 임상 영역을 포괄합니다. 또한, 우리는 단계별 추론 기반 의사 결정을 위해 도구 연결(tool grounding)과 이중 메모리 메커니즘을 통합하고, 저장된 액터-크리틱 시뮬레이션 경로를 활용하여 지식 전달을 수행하는 액터-크리틱 기반 멀티 에이전트 프레임워크인 ArogyaSutra를 제안합니다. 실험 결과, 우리 데이터셋과 프레임워크는 모든 인도어에서 다국어 의료 추론 정확도를 향상시키며, 각 구성 요소의 기여도를 검증하는 분석을 통해 효과를 입증했습니다. 소스 코드와 데이터셋은 다음 URL에서 제공됩니다: https://iitp-cse.github.io/ArogyaSutra/
Multimodal Large Language Models (MLLMs) have shown promising reasoning capabilities in general domains, yet their performance remains limited in specialized settings such as healthcare, especially in multilingual and low-resource scenarios. This gap is critical in regions like rural India, where patients often express complex medical queries in native Indic languages and rely on multimodal inputs such as medical images. Existing English-centric MLLMs struggle to support such use cases, limiting equitable access to AI-driven healthcare assistance. To address this challenge, we introduce ArogyaBodha, a large-scale multilingual multimodal medical question-answer dataset constructed from eight heterogeneous sources, covering 31 body systems, six imaging modalities, and 21 clinical domains across English and seven major Indian languages. We further propose ArogyaSutra, an actor-critic-based multi-agent framework that integrates tool grounding with dual-memory mechanisms for step-wise, reasoning-aware decision making, and uses stored actor-critic simulation trajectories for distillation. Experiments show that our dataset and framework improve multilingual medical reasoning accuracy across all Indic languages, with ablations validating the contribution of each component. The source code and dataset are available at: https://iitp-cse.github.io/ ArogyaSutra/
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.