2603.06542v1 Mar 06, 2026 cs.SD

RAMoEA-QA: 강력한 호흡 음성 질의 응답을 위한 계층적 전문화

RAMoEA-QA: Hierarchical Specialization for Robust Respiratory Audio Question Answering

G. Bertolino
G. Bertolino
Citations: 1
h-index: 1
Yuwei Zhang
Yuwei Zhang
Citations: 65
h-index: 2
Tong Xia
Tong Xia
Citations: 1,357
h-index: 15
Domenico Talia
Domenico Talia
Citations: 145
h-index: 7
Cecilia Mascolo
Cecilia Mascolo
Citations: 3
h-index: 1

대화형 생성형 AI는 급속히 의료 분야로 진출하고 있으며, 범용 모델은 다양한 환자 데이터를 통합하고 다양한 상호 작용 방식을 지원하는 동시에 임상적으로 의미 있는 결과를 생성해야 합니다. 호흡기 질환 관리 분야에서, 모바일 마이크를 통해 녹음된 비침습적 오디오 데이터는 대규모 선별 및 장기 모니터링을 가능하게 하지만, 데이터의 이질성 문제가 특히 심각합니다. 녹음 데이터는 장치, 환경, 획득 프로토콜에 따라 크게 다르며, 질문은 다양한 의도와 형식을 포함합니다. 기존의 생체의학 오디오-언어 질의 응답 시스템은 일반적으로 단일 구조로 되어 있으며, 다양한 호흡기 데이터와 질문 의도를 처리하기 위한 전문화 메커니즘이 없습니다. 또한, 이러한 시스템은 제한된 환경에서만 검증되었기 때문에, 실제 환경에서 발생하는 변화에 얼마나 안정적으로 대응하는지는 불분명합니다. 이러한 제한 사항을 해결하기 위해, 우리는 호흡기 오디오 질의 응답을 위한 계층적으로 라우팅되는 생성 모델인 RAMoEA-QA를 소개합니다. RAMoEA-QA는 여러 유형의 질문을 통합하고 단일 다중 모달 시스템 내에서 이산적 및 연속적 목표를 모두 지원합니다. RAMoEA-QA는 두 단계의 조건부 전문화를 적용합니다. 오디오 Mixture-of-Experts는 각 녹음 데이터를 적합한 사전 훈련된 오디오 인코더로 라우팅하고, Language Mixture-of-Adapters는 공유된 동결된 LLM에서 LoRA 어댑터를 선택하여 질문 의도와 응답 형식을 일치시킵니다. RAMoEA-QA는 각 예제에 대한 음향 표현과 생성 동작을 전문화함으로써, 최소한의 파라미터 오버헤드로 강력한 기준 모델 및 라우팅 제거 실험에서 일관되게 더 우수한 성능을 보입니다. 또한, 특정 분야에서의 테스트 정확도를 0.72로 향상시키고 (최첨단 기준 모델의 경우 0.61 및 0.67), 도메인, 모달리티 및 작업 변화에 대한 가장 강력한 일반화 성능을 보여줍니다.

Original Abstract

Conversational generative AI is rapidly entering healthcare, where general-purpose models must integrate heterogeneous patient signals and support diverse interaction styles while producing clinically meaningful outputs. In respiratory care, non-invasive audio, such as recordings captured via mobile microphones, enables scalable screening and longitudinal monitoring, but the heterogeneity challenge is particularly acute: recordings vary widely across devices, environments, and acquisition protocols, and questions span multiple intents and question formats. Existing biomedical audio-language QA systems are typically monolithic, without any specialization mechanisms for tackling diverse respiratory corpora and query intents. They are also only validated in limited settings, leaving it unclear how reliably they handle the shifts encountered in real-world settings. To address these limitations, we introduce RAMoEA-QA, a hierarchically routed generative model for respiratory audio question answering that unifies multiple question types and supports both discrete and continuous targets within a single multimodal system. RAMoEA-QA applies two-stage conditional specialization: an Audio Mixture-of-Experts routes each recording to a suitable pre-trained audio encoder, and a Language Mixture-of-Adapters selects a LoRA adapter on a shared frozen LLM to match the query intent and answer format. By specializing both acoustic representations and generation behaviour per example, RAMoEA-QA consistently outperforms strong baselines and routing ablations with minimal parameter overhead, improving in-domain test accuracy to 0.72 (vs. 0.61 and 0.67 for state-of-the-art baselines) and exhibiting the strongest generalization for diagnosis under domain, modality, and task shifts.

0 Citations
0 Influential
7.5 Altmetric
37.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!