PPT-Eval: 파워포인트 작업을 위한 컴퓨터 사용 에이전트 평가 벤치마크
PPT-Eval: A Benchmark for Computer-Use Agents on PowerPoint Tasks
슬라이드 생성 및 편집은 전문적, 교육적 환경에서 광범위하게 사용되는 다중 모드 활동이며, 실제 환경에서의 컴퓨터 사용 에이전트를 테스트하기에 이상적인 플랫폼입니다. Microsoft PowerPoint는 프레젠테이션 제작을 위한 가장 널리 사용되고 기능이 풍부한 환경 중 하나입니다. 본 논문에서는 난이도별로 구성된 12개의 파일에 걸쳐 슬라이드 콘텐츠 생성 및 프레젠테이션 편집 시나리오를 포함하는 총 120개의 파워포인트 작업으로 구성된 벤치마크인 PPT-Eval을 소개합니다. 이 분야의 주요 과제는 평가입니다. 작업은 복잡하고 다중 모드이며, 종종 여러 개의 유효한 솔루션을 허용합니다. 또한, 현재 에이전트는 종종 부분적인 진행만 이루는 경우가 많으며, 이는 단순 성공/실패 지표로는 제대로 반영하기 어렵습니다. 이를 해결하기 위해, 파워포인트 작업에 대한 작업별 평가 기준을 설계하는 강력한 평가 프레임워크를 개발했습니다. 이 프레임워크는 기존의 기준 기반 평가 연구에서 영감을 얻고 확장되었습니다. 이러한 기준은 중간 단계에 부분 점수를 부여하고 불필요한 변경 및 미흡한 디자인 요소에 대해 감점을 부과하며, 자연어 피드백을 제공합니다. 이러한 세밀한 접근 방식은 매우 효과적이며, 인간의 판단과의 Kendall's τ-b 상관관계는 0.77입니다. 연구 결과, 현재 최첨단 에이전트조차도 파워포인트 작업을 해결하는 데 어려움을 겪고 있으며, Claude-4.5-Opus와 같은 강력한 모델조차도 45%의 성공률과 평균 부분 점수 57%를 기록했습니다. 본 벤치마크는 다음 위치에서 확인할 수 있습니다: https://microsoft.github.io/ppteval.
Creating and editing slides is a rich, multimodal activity that is ubiquitous in professional and educational settings, making it an ideal testbed for real-world computer-use agents. Microsoft PowerPoint is among the most widely adopted and feature-rich environments for presentation creation. We introduce PPT-Eval, a benchmark of 120 PowerPoint tasks across 12 files that cover both content creation and presentation editing scenarios, organized by difficulty. A central challenge in this domain is evaluation: tasks are complex, multimodal, and often admit many valid solutions. Moreover, today's agents frequently make only partial progress, which binary success metrics fail to capture. To address this, we design a robust evaluation framework to help create task-specific rubrics for PowerPoint tasks, taking inspiration from and building on past works for rubric-based evaluation. These rubrics award partial credit for intermediate steps, penalize unnecessary changes and poor aesthetics, and provide natural language feedback. This nuanced approach proves highly effective, achieving a Kendall's τ-b correlation of 0.77 with human judgments. We find that existing frontier agents still struggle with solving PowerPoint tasks, with strong models like Claude-4.5-Opus achieving only a 45% success rate and an average partial score of 57%. The benchmark is located at: https://microsoft.github.io/ppteval.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.