2602.15909v2 Feb 16, 2026 eess.AS

Resp-Agent: 다중 모드 호흡 음 생성 및 질병 진단을 위한 에이전트 기반 시스템

Resp-Agent: An Agent-Based System for Multimodal Respiratory Sound Generation and Disease Diagnosis

Pengfei Zhang
Pengfei Zhang
Citations: 64
h-index: 3
Tianxin Xie
Tianxin Xie
Citations: 90
h-index: 5
Minghao Yang
Minghao Yang
Citations: 73
h-index: 4
Li Liu
Li Liu
Citations: 3,394
h-index: 8

딥 러닝 기반 호흡 청진은 현재 다음과 같은 두 가지 근본적인 문제에 직면해 있습니다. (i) 신호를 분광 이미지로 변환하는 과정에서 일시적인 음향 이벤트와 임상적 맥락이 손실되는 정보 손실 현상, (ii) 심각한 클래스 불균형으로 인해 데이터 가용성이 제한되는 문제. 이러한 문제점을 해결하기 위해, 우리는 Active Adversarial Curriculum Agent (Thinker-A$^2$CA)라는 새로운 방식으로 작동하는 자율적인 다중 모드 시스템인 Resp-Agent를 제안합니다. 기존의 정적인 파이프라인과 달리, Thinker-A$^2$CA는 중앙 제어기로 작동하며, 진단상의 약점을 적극적으로 파악하고, 닫힌 루프 방식으로 표적 합성을 수행합니다. 표현 격차 문제를 해결하기 위해, 우리는 EHR (전자의무기록) 데이터를 오디오 토큰과 Strategic Global Attention 및 희소 오디오 앵커를 통해 연결하는 Modality-Weaving Diagnoser를 도입하여, 장기적인 임상적 맥락과 밀리초 단위의 일시적인 현상을 모두 포착합니다. 데이터 격차 문제를 해결하기 위해, 우리는 텍스트 기반의 대규모 언어 모델 (LLM)에 모달리티 주입을 통해 텍스트 기반 모델을 조정하는 Flow Matching Generator를 설계하여, 병리적 내용과 음향 스타일을 분리하고 진단하기 어려운 샘플을 생성합니다. 이러한 연구를 위한 기반으로, 우리는 229,000개의 녹음 파일과 LLM을 통해 추출된 임상적 설명을 포함하는 벤치마크 데이터셋인 Resp-229k를 공개합니다. 광범위한 실험 결과, Resp-Agent는 다양한 평가 환경에서 기존 방법보다 일관되게 우수한 성능을 보이며, 데이터 부족 및 클래스 불균형 상황에서 진단의 안정성을 향상시킵니다. 저희의 코드와 데이터는 https://github.com/zpforlove/Resp-Agent 에서 확인할 수 있습니다.

Original Abstract

Deep learning-based respiratory auscultation is currently hindered by two fundamental challenges: (i) inherent information loss, as converting signals into spectrograms discards transient acoustic events and clinical context; (ii) limited data availability, exacerbated by severe class imbalance. To bridge these gaps, we present Resp-Agent, an autonomous multimodal system orchestrated by a novel Active Adversarial Curriculum Agent (Thinker-A$^2$CA). Unlike static pipelines, Thinker-A$^2$CA serves as a central controller that actively identifies diagnostic weaknesses and schedules targeted synthesis in a closed loop. To address the representation gap, we introduce a Modality-Weaving Diagnoser that weaves EHR data with audio tokens via Strategic Global Attention and sparse audio anchors, capturing both long-range clinical context and millisecond-level transients. To address the data gap, we design a Flow Matching Generator that adapts a text-only Large Language Model (LLM) via modality injection, decoupling pathological content from acoustic style to synthesize hard-to-diagnose samples. As a foundation for these efforts, we introduce Resp-229k, a benchmark corpus of 229k recordings paired with LLM-distilled clinical narratives. Extensive experiments demonstrate that Resp-Agent consistently outperforms prior approaches across diverse evaluation settings, improving diagnostic robustness under data scarcity and long-tailed class imbalance. Our code and data are available at https://github.com/zpforlove/Resp-Agent.

0 Citations
0 Influential
29.493061443341 Altmetric
0.0 Score
Original PDF
2

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!