듣고, 호출하고, 이해하기: 대규모 오디오 언어 모델을 위한 기술 기반의 다중 모드 에이전트
Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models
복잡한 음향 문제는 모델이 직접적인 오디오 입력에서 답을 찾는 것보다, 음향 연산을 수행하고, 외부 도구와 상호 작용하며, 결과적으로 생성된 텍스트 또는 처리된 오디오 정보를 바탕으로 추론해야 할 수 있습니다. 본 연구에서는 이러한 문제를 도구 기반의 음성 추론 문제로 정의하고, SpeechAgent-R이라는 음성 에이전트를 개발하여, 내재적인 다중 모드 이해 능력을 외부 기술 및 도구와 연계하도록 합니다. 이 기능을 지원하기 위해, 24개의 작업, 8개의 기술, 9개의 도구를 포함하는 65,492개의 상호 작용 경로와 507.6시간의 오디오 데이터로 구성된 HIU-Corpus를 구축했습니다. SpeechAgent-R은 먼저 경로 기반의 지도 학습을 통해 구조화된 상호 작용 행동을 학습하고, 다중 단계 강화 학습을 통해 의사 결정 능력을 향상시킵니다. 또한, 작업 성능, 상호 작용 품질 및 다양한 작업 환경으로의 일반화 능력을 종합적으로 평가하기 위한 HIU-Bench를 소개합니다. HIU-Bench는 56개의 작업을 포함하는 1,395개의 샘플로 구성되어 있으며, 여기에는 도구 사용 방식과 워크플로우 구성을 크게 변경한 in-distribution (ID) 및 out-of-distribution (OOD) 데이터가 포함됩니다. SpeechAgent-R은 ID 작업에서 84.17점, OOD 작업에서 70.94점을 달성하여, 동일한 에이전트 구조를 갖는 기본 모델 대비 각각 15.40점과 14.23점의 성능 향상을 보였습니다. 이러한 결과는 기술 및 도구 연계 학습이 음성 에이전트가 다양한 작업 환경에 적응하고, 가변적인 도구 상호 작용을 처리하는 능력을 향상시킨다는 것을 보여줍니다.
Complex acoustic problems may require models to perform acoustic operations, interact with external tools and reason over the resulting textual or processed-audio observations rather than answer directly from a fixed audio input. We study such problems as tool-interactive audio reasoning and develop SpeechAgent-R, an audio agent that coordinates its intrinsic multimodal understanding with external skills and tools. To support this capability, we construct HIU-Corpus, comprising 65,492 interaction trajectories and 507.6 hours of audio across 24 tasks, 8 skills and 9 tools. SpeechAgent-R first learns structured interaction behaviors through trajectory-based supervised fine-tuning and then improves its decisions through multi-turn reinforcement learning. We further introduce HIU-Bench to jointly evaluate task performance, interaction quality and generalization to diverse task settings. It contains 1,395 samples across 56 tasks, including in-distribution (ID) and out-of-distribution (OOD) splits with substantial shifts in tool usage and workflow composition. SpeechAgent-R achieves 84.17 on ID tasks and 70.94 on OOD tasks, improving over the base model under the same agent harness by 15.40 and 14.23 points. These results demonstrate that learning skill and tool coordination improves audio agents' ability to handle diverse task settings and adaptive tool interactions.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.