2607.13408v1 Jul 15, 2026 eess.AS

오디오 인지 능력을 갖춘 대규모 언어 모델로부터 얻은 세부적인 피드백을 활용한 텍스트-음성 명령어 추종 성능 향상

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models

Chun-Yi Kuan
Chun-Yi Kuan
Citations: 534
h-index: 13
Bo-Ru Lu
Bo-Ru Lu
Citations: 13
h-index: 2
Qinming Tang
Qinming Tang
Citations: 0
h-index: 0
Ankur Gandhe
Ankur Gandhe
Citations: 764
h-index: 18
Chieh-Chi Kao
Chieh-Chi Kao
Citations: 866
h-index: 14
Siwon Kim
Siwon Kim
Citations: 0
h-index: 0
Byeonggeun Kim
Byeonggeun Kim
Citations: 14
h-index: 2
Suyoun Kim
Suyoun Kim
Citations: 45
h-index: 3
Hung-yi Lee
Hung-yi Lee
Citations: 136
h-index: 6
Chao Wang
Chao Wang
Citations: 15
h-index: 2

최근의 텍스트-음성 변환 모델들은 고품질 오디오를 생성하지만, 여러 음향 이벤트와 시간 순서를 포함하는 명령어에 대한 이해도가 부족한 경우가 많습니다. 이러한 문제는 기존의 평가 및 학습 신호가 주로 전반적인 유사성이나 지각적 품질을 강조하고, 명령어 수준에서의 정확성에 대한 감독이 제한적이기 때문에 발생합니다. 본 연구에서는 오디오 인지 능력을 갖춘 대규모 언어 모델(ALLM)을 사용하여 생성된 오디오에서 목표 이벤트의 존재 여부와 시간 관계를 검증하는 세부적인 판단 기준을 제공하는 명령어 수준 프레임워크를 제안합니다. ALLM의 판단 결과를 벤치마크 테스트 및 인간 평가를 통해 검증한 후, 이를 활용하여 직접 선호도 최적화를 위한 선호 쌍을 구성합니다. 또한, 다중 이벤트의 시간 순서 추종 능력을 평가하기 위한 스토리 기반 벤치마크인 S3Bench를 소개합니다. 실험 결과는 제안하는 방법이 기존 벤치마크와 S3Bench 모두에서 이벤트 완성도, 시간 정렬 및 전체적인 명령어 추종 정확도를 향상시키면서 오디오 품질을 유지함을 보여줍니다.

Original Abstract

Recent text-to-audio models generate high-quality audio, but often fail to follow instructions involving multiple sound events and temporal order. This gap arises because existing evaluation and training signals mainly emphasize global similarity or perceptual quality, with limited supervision on instruction-level correctness. We propose an instruction-level framework that uses audio-aware large language models (ALLMs) as fine-grained judges to verify target event presence and temporal relations in generated audio. After validating ALLM judgments on benchmarks and through human verification, we use their feedback to construct preference pairs for direct preference optimization. We further introduce S3Bench, a narrative benchmark for evaluating multi-event temporal instruction following. Experiments show that our method improves event completeness, temporal ordering, and joint instruction-following accuracy across existing benchmarks and S3Bench, while maintaining audio quality.

0 Citations
0 Influential
9 Altmetric
45.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!