2607.02119v1 Jul 02, 2026 eess.AS

통합 오디오 이해 및 생성을 위한 효율적인 vLLM 기반 추론 파이프라인

An Efficient vLLM-Based Inference Pipeline for Unified Audio Understanding and Generation

Jinchuan Tian
Jinchuan Tian
Citations: 928
h-index: 14
Shinji Watanabe
Shinji Watanabe
Citations: 82
h-index: 6
Siddhant Arora
Siddhant Arora
Citations: 1,464
h-index: 18
Haoran Wang
Haoran Wang
Citations: 2
h-index: 1

대규모 다중 모드 모델은 이해 능력에서 뛰어난 성능을 보이지만, 높은 처리량을 제공하는 추론 엔진은 다중 모드 생성 기능을 기본적으로 지원하지 않습니다. 특히 음성 언어 모델의 경우, AR+NAR 또는 동기식 멀티 토큰 예측(MTP)과 같은 분리된 방식으로 다층 오디오 토큰을 생성하는 것은 표준 단일 스트림 루프와 충돌합니다. 본 연구에서는 통합 음성 이해 및 생성을 위한 vLLM 기반 추론 파이프라인을 제시합니다. 우리는 자기 회귀 디코딩을 확장하여 지연 패턴 제거 및 조정된 멀티 스트림 샘플링을 기본적으로 실행할 수 있도록 하고, 엔드 투 엔드 웨이브폼 합성을 위해 GPU 내부에 음향 디코더를 통합했습니다. 중요한 점은 기존의 통념과는 달리, 분류기 자유 가이드(CFG)가 처리량을 절반으로 줄인다는 생각에 반하는 결과를 얻었습니다. 연속적인 배치 내에서 페어링된 조건부 및 무조건부 요청을 동시에 실행함으로써, 우리의 CFG 구현은 일반적인 방식 대비 80%의 처리량을 유지하며, 이중 요청 및 로짓 병합 오버헤드를 효과적으로 흡수합니다. 저희는 개발한 프레임워크를 오픈 소스로 공개합니다.

Original Abstract

While Large Multimodal Models excel in comprehension, high-throughput inference engines lack native support for multimodal generation. This is severe in Speech Language Models, where generating multi-layered audio tokens via decoupled AR+NAR or synchronous Multi-Token Prediction (MTP) with delay-pattern interleaving conflicts with standard single-stream loops. We present a vLLM-based inference pipeline for unified speech understanding and generation. We extend autoregressive decoding to natively execute delay-pattern de-interleaving and coordinated multi-stream sampling, integrating an on-GPU acoustic decoder for end-to-end waveform synthesis. Crucially, we overcome the shared intuition that Classifier-Free Guidance (CFG) halves throughput. By co-scheduling paired conditional and unconditional requests within a continuous batch, our CFG implementation sustains 80% of non-CFG throughput, absorbing dual-request and logit merging overheads. We open-source our framework.

0 Citations
0 Influential
9 Altmetric
45.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!