2605.28035v1 May 27, 2026 cs.AI

MTAVG-Bench 2.0: 다중 화자 오디오-비디오 생성 모델의 영화적 표현 실패 요인 진단

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation

Heyan Huang
Heyan Huang
Citations: 146
h-index: 7
Xian-Ling Mao
Xian-Ling Mao
Citations: 58
h-index: 4
Yue Liu
Yue Liu
Citations: 2
h-index: 1
Haitian Li
Haitian Li
Citations: 13
h-index: 2
Yanghao Zhou
Yanghao Zhou
Citations: 10
h-index: 1
Liang-Jin Chen
Liang-Jin Chen
Citations: 40
h-index: 1
Yiming Cheng
Yiming Cheng
Citations: 54
h-index: 2
Xu Liu
Xu Liu
Citations: 12
h-index: 3
Dian Jin
Dian Jin
Citations: 34
h-index: 3
Jiajun Xu
Jiajun Xu
Citations: 12
h-index: 2
Jing Liao
Jing Liao
Citations: 127
h-index: 4
Tian Lan
Tian Lan
Citations: 6
h-index: 2
Ziqi Zhou
Ziqi Zhou
Citations: 24
h-index: 3
Yu Bai
Yu Bai
Beijing Academy of Artificial Intelligence (BAAI)
Citations: 286
h-index: 9
Changsen Yuan
Changsen Yuan
Citations: 231
h-index: 7
Jinxing Zhou
Jinxing Zhou
Citations: 45
h-index: 5
Xuefeng Chen
Xuefeng Chen
Citations: 45
h-index: 2
Yousheng Feng
Yousheng Feng
Citations: 7
h-index: 1

최근 몇 년 동안, 다중 화자 오디오-비디오 생성(MTAVG) 모델은 동기화 및 오디오-시각 정렬과 같은 기본적인 지표에서 유망한 성능을 보여주었습니다. 그러나 이러한 지표는 장면 수준의 생생한 표현력을 평가하기에는 충분하지 않습니다. 다중 캐릭터 장면에서는 생성 모델이 오디오-시각적 사실성을 넘어 일관성 있는 캐릭터 연기와 다른 고차원의 영화적 요소를 전달해야 합니다. 이러한 격차를 해소하기 위해, 우리는 다중 화자 오디오-비디오 생성에서 영화적 표현의 실패 요인을 진단하는 벤치마크인 MTAVG-Bench 2.0을 소개합니다. 기존 설정이 주로 기본적인 다회 대화 품질에 초점을 맞춘 것과는 달리, MTAVG-Bench 2.0은 단편 드라마 및 장면 수준의 생성을 목표로 하며, 연기, 내러티브, 분위기 및 오디오-시각 언어를 포괄하는 고차원의 실패 분류 체계를 확립합니다. 이 분류 체계에 기반하여, 우리는 1만 개 이상의 질의응답 평가 사례를 구축했으며, 단편 드라마 수준의 평가 및 실패 요인의 시간적 위치 추정을 위한 하위 집합도 포함되어 있습니다. 이를 통해, 거대 언어 모델이 고차원의 오디오-시각적 오류를 진단하는 능력을 체계적으로 평가합니다. 실험 결과는 Gemini와 같은 상용 거대 모델이 다른 평가 도구보다 훨씬 뛰어난 성능을 보이지만, 가장 강력한 모델조차도 여전히 우리 벤치마크의 복잡한 실패 요인에 어려움을 겪고 있음을 보여줍니다. 이러한 결과는 MTAVG-Bench 2.0이 영화적 다중 화자 오디오-비디오 생성에서 실패 진단을 위한 체계적인 벤치마크를 제공한다는 것을 입증합니다.

Original Abstract

In recent years, Multi-Talker Audio-Video Generation (MTAVG) models have shown promising performance on fundamental metrics such as lip-sync and audio-visual alignment. However, these metrics remain insufficient for assessing cinematic expressiveness in scene-level generation. In multi-character scenes, generation models must go beyond audio-visual realism to convey coherent character performance and other higher-level cinematic qualities. To fill this gap, we introduce MTAVG-Bench 2.0, a benchmark for diagnosing failure modes of cinematic expressiveness in multi-talker audio-video generation. Unlike prior settings that mainly focus on the quality of basic multi-turn dialogue, MTAVG-Bench 2.0 targets short-drama and scene-level generation, and establishes a high-level failure taxonomy spanning acting, narrative, atmosphere, and audio-visual language. Based on this taxonomy, we construct more than 10,000 question-answering evaluation instances, together with subsets for short-drama-level assessment and temporal localization of failure modes, to systematically evaluate the ability of omni large language models to diagnose high-level audio-visual failures. Experimental results show that commercial omni models such as Gemini substantially outperform other evaluators, yet even the strongest models continue to struggle with complex failures in our benchmark. These results demonstrate that MTAVG-Bench 2.0 provides a systematic benchmark for failure diagnosis in cinematic multi-talker audio-video generation.

0 Citations
0 Influential
4.5 Altmetric
22.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!