2606.31154v1 Jun 30, 2026 cs.LG

PPT-Eval: 파워포인트 작업을 위한 컴퓨터 사용 에이전트 평가 벤치마크

PPT-Eval: A Benchmark for Computer-Use Agents on PowerPoint Tasks

Graham Neubig
Graham Neubig
Citations: 966
h-index: 16
Apurva Gandhi
Apurva Gandhi
Citations: 191
h-index: 5
Thong Q. Nguyen
Thong Q. Nguyen
Citations: 19
h-index: 2
Vishwas Suryanarayanan
Vishwas Suryanarayanan
Citations: 11
h-index: 2
Raja Hasnain Anwar
Raja Hasnain Anwar
Citations: 16
h-index: 2
Firoz Shaik
Firoz Shaik
Citations: 11
h-index: 2
Shubhang Desai
Shubhang Desai
Citations: 4
h-index: 1
Muhammad Taqi Raza
Muhammad Taqi Raza
Citations: 0
h-index: 0
Vishal Chowdhary
Vishal Chowdhary
Citations: 6
h-index: 2

슬라이드 생성 및 편집은 전문적, 교육적 환경에서 광범위하게 사용되는 다중 모드 활동이며, 실제 환경에서의 컴퓨터 사용 에이전트를 테스트하기에 이상적인 플랫폼입니다. Microsoft PowerPoint는 프레젠테이션 제작을 위한 가장 널리 사용되고 기능이 풍부한 환경 중 하나입니다. 본 논문에서는 난이도별로 구성된 12개의 파일에 걸쳐 슬라이드 콘텐츠 생성 및 프레젠테이션 편집 시나리오를 포함하는 총 120개의 파워포인트 작업으로 구성된 벤치마크인 PPT-Eval을 소개합니다. 이 분야의 주요 과제는 평가입니다. 작업은 복잡하고 다중 모드이며, 종종 여러 개의 유효한 솔루션을 허용합니다. 또한, 현재 에이전트는 종종 부분적인 진행만 이루는 경우가 많으며, 이는 단순 성공/실패 지표로는 제대로 반영하기 어렵습니다. 이를 해결하기 위해, 파워포인트 작업에 대한 작업별 평가 기준을 설계하는 강력한 평가 프레임워크를 개발했습니다. 이 프레임워크는 기존의 기준 기반 평가 연구에서 영감을 얻고 확장되었습니다. 이러한 기준은 중간 단계에 부분 점수를 부여하고 불필요한 변경 및 미흡한 디자인 요소에 대해 감점을 부과하며, 자연어 피드백을 제공합니다. 이러한 세밀한 접근 방식은 매우 효과적이며, 인간의 판단과의 Kendall's τ-b 상관관계는 0.77입니다. 연구 결과, 현재 최첨단 에이전트조차도 파워포인트 작업을 해결하는 데 어려움을 겪고 있으며, Claude-4.5-Opus와 같은 강력한 모델조차도 45%의 성공률과 평균 부분 점수 57%를 기록했습니다. 본 벤치마크는 다음 위치에서 확인할 수 있습니다: https://microsoft.github.io/ppteval.

Original Abstract

Creating and editing slides is a rich, multimodal activity that is ubiquitous in professional and educational settings, making it an ideal testbed for real-world computer-use agents. Microsoft PowerPoint is among the most widely adopted and feature-rich environments for presentation creation. We introduce PPT-Eval, a benchmark of 120 PowerPoint tasks across 12 files that cover both content creation and presentation editing scenarios, organized by difficulty. A central challenge in this domain is evaluation: tasks are complex, multimodal, and often admit many valid solutions. Moreover, today's agents frequently make only partial progress, which binary success metrics fail to capture. To address this, we design a robust evaluation framework to help create task-specific rubrics for PowerPoint tasks, taking inspiration from and building on past works for rubric-based evaluation. These rubrics award partial credit for intermediate steps, penalize unnecessary changes and poor aesthetics, and provide natural language feedback. This nuanced approach proves highly effective, achieving a Kendall's τ-b correlation of 0.77 with human judgments. We find that existing frontier agents still struggle with solving PowerPoint tasks, with strong models like Claude-4.5-Opus achieving only a 45% success rate and an average partial score of 57%. The benchmark is located at: https://microsoft.github.io/ppteval.

3 Citations
1 Influential
8 Altmetric
45.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!