자유 형식 텍스트 프롬프트를 이용한 통합 음성 및 효과음 생성
Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts
오디오 생성이 상당한 발전을 이루었지만, 음성과 효과음을 자연스럽게 결합하여 일관된 오디오를 생성하는 것은 여전히 어려운 과제입니다. 현재의 방법들은 세밀한 상호작용을 포착하지 못하는 분리된 파이프라인에 의존하거나, 구조화된 입력과 외부 텍스트 재작성을 필요로 하여 자유 형식 텍스트 프롬프트의 유연성을 제한합니다. 본 논문에서는 새로운 과제인 '자유 형식 텍스트 프롬프트에서 통합 오디오 생성'을 제시하며, 이는 제약 없이 자연어로부터 음성, 효과음 및 이들의 결합을 직접 합성하는 것을 목표로 합니다. 이러한 과제를 해결하기 위해, 우리는 통합적이고 자기 회귀적인 LLM 기반 프레임워크인 PlanAudio를 제안합니다. 첫째, PlanAudio는 기존 텍스트 인코더 대신 LLM의 고유한 추론 능력을 활용하여 모델 아키텍처를 단순화합니다. 둘째, 고수준의 의미 이해와 저수준의 음향 합성을 연결하는 암시적 계획 메커니즘인 '의미 기반 연쇄적 사고(semantic latent chain-of-thought)'를 도입합니다. 또한, 복합 오디오 시나리오 평가를 위한 특화된 벤치마크인 PlanAudio-Bench를 제작했습니다. 우리는 음성, 효과음 및 이들의 결합 시나리오에서 실험을 수행했으며, 그 결과 PlanAudio는 기존 파이프라인 및 통합 모델과 비교하여 일반적으로 더 우수한 성능을 보였으며, 단일 시나리오에 특화된 모델들과 경쟁력 있는 결과를 얻었습니다. 추가 분석 결과, 의미 기반 연쇄적 사고가 다른 연쇄적 사고 메커니즘보다 우수하며, 지속적인 다중 시나리오 훈련 커리큘럼의 중요성을 강조합니다.
Audio generation has made significant progress, yet synthesizing unified audio where speech and sounds are naturally composited remains a challenge. Current methods either rely on disjoint pipelines, which fail to capture fine-grained interactions, or require structured inputs and external text rewriting, which limits the flexibility of free-form text prompts. In this paper, we introduce a new task: Free-Form-Text-Prompt-to-Unified-Audio generation, which aims to directly synthesize unified audio containing speech, sound, and their composites from unconstrained natural language. To address this task, we propose PlanAudio, a unified, autoregressive LLM-based framework. First, it simplifies the model architecture by leveraging intrinsic LLM reasoning capability instead of traditional text encoders. Second, it introduces a semantic latent chain-of-thought mechanism, an implicit planning mechanism that bridges high-level semantic understanding and low-level acoustic synthesis. Furthermore, we create PlanAudio-Bench, a specialized benchmark for evaluating composite audio scenarios. We perform evaluations in the scenarios of speech, sound, and their composites. The results demonstrate that PlanAudio generally outperforms the existing pipeline and unified baselines, while staying competitive with models designed for a single scenario. Our analysis further reveals the superiority of semantic latent CoT over other CoT mechanisms and highlights the importance of continuous multi-scenario training curricula.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.