2606.19256v1 Jun 17, 2026 cs.AI

X+Slides: 청중 특성을 고려한 슬라이드 자동 생성 성능 평가

X+Slides: Benchmarking Audience-Conditioned Slide Generation

Xuanhe Zhou
Xuanhe Zhou
Citations: 14
h-index: 2
Wei Zhou
Wei Zhou
Citations: 128
h-index: 6
Fan Wu
Fan Wu
Citations: 54
h-index: 4
Jiawei Hong
Jiawei Hong
Citations: 553
h-index: 4
Haodong Chen
Haodong Chen
Citations: 9
h-index: 1
Xinyu Shao
Xinyu Shao
Citations: 21
h-index: 2
Yanbin Zhu
Yanbin Zhu
Citations: 13
h-index: 2
Bojun Wang
Bojun Wang
Citations: 0
h-index: 0
Anya Jia
Anya Jia
Citations: 209
h-index: 1

소스 문서로부터 자동으로 프레젠테이션 슬라이드를 생성하는 것은 대규모 언어 모델(LLM)의 중요한 응용 분야입니다. 기존 벤치마크는 주로 슬라이드의 완결성과 기술적 깊이를 평가하지만, 실제 환경에서 중요한 요소인 '청중'을 간과합니다. 예를 들어, 전문가들은 엄밀한 증명을 요구하는 반면, 의사 결정권자들은 실행 가능한 결론을 우선시합니다. 이러한 격차를 해소하기 위해, 우리는 청중의 특성을 고려한 슬라이드 생성에 특화된 벤치마크인 X+Slides를 소개합니다. X+Slides는 113개의 주제와 7가지 프레젠테이션 시나리오를 아우르는 다양한 데이터셋을 기반으로 구축되었으며, 8,133개의 중복되지 않은, 소스 문서에 기반한 평가 데이터를 활용하는 동적 평가 프레임워크를 사용합니다. X+Slides는 동일한 소스 문서 기반의 평가 데이터에 대해 청중별 유용성 가중치를 부여하여 다음과 같은 네 가지 상호 보완적인 지표를 제공합니다: '청중 포괄성'(Audience Coverage)은 청중에 필수적인 정보가 얼마나 전달되는지를 측정하며, '도메인별 포괄성'(Domain-wise Coverage)은 어떤 유형의 정보가 다루어지는지를 보여줍니다. '효율성'(Efficiency)은 단위 주의 비용당 제공되는 유용성을 측정하고, '정확성'(Correctness)은 슬라이드 내용이 소스 문서에 의해 뒷받침되는지 확인합니다. DeepPresenter, SlideTailor 및 NotebookLM에 대한 실험 결과, 현재 시스템들은 상당한 양이지만 여전히 불완전한 수준의 청중에 필수적인 정보를 복원할 수 있음을 보여줍니다. 특정 기준($τ_A=0.7$)에서 DeepPresenter는 0.714의 최고 '청중 포괄성'을 달성하고, SlideTailor는 0.594를, NotebookLM은 0.853을 기록했으며, 이는 소스 문서 기반 평가 없이 시각적 품질과 광범위한 주제 보장이 얼마나 중요한 근거가 될 수 있는지 보여줍니다.

Original Abstract

Automatically generating slide decks from source documents is an important application of large language models (LLMs). Existing benchmarks primarily assess slide completeness and technical depth, while overlooking the target audience as a critical real-world factor. For instance, specialists demand rigorous proofs, whereas decision-makers prioritize actionable conclusions. To bridge this gap, we introduce X+Slides, a benchmark specifically designed for audience-conditioned slide generation. Built on a diverse corpus spanning 113 topics and seven presentation scenes, X+Slides employs a dynamic evaluation framework constructed from 8,133 deduplicated, source-grounded probes. By assigning audience-specific utility weights to the same source-grounded probes, X+Slides reports four complementary metrics: Audience Coverage measures how much audience-essential information is conveyed, Domain-wise Coverage shows which information types are covered, Efficiency measures delivered utility per unit of attention cost, and Correctness verifies whether slide claims are supported by the source. Experiments on DeepPresenter, SlideTailor, and NotebookLM show that current systems can recover a substantial but still incomplete part of audience-essential information: at $τ_A=0.7$, DeepPresenter reaches a best Audience Coverage of 0.714, SlideTailor reaches 0.594, and the NotebookLM ablation reaches 0.853 while showing clear grounding differences. These results indicate that visual quality and broad topic coverage should not be treated as evidence support without source-grounded evaluation.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!