SpreadsheetBench 2: 실제 비즈니스 스프레드시트 워크플로우에서의 에이전트 평가
SpreadsheetBench 2: Evaluating Agents on End-to-End Business Spreadsheet Workflows
스프레드시트는 비즈니스 분석, 재무 모델링, 보고 및 의사 결정에 널리 사용됩니다. 그러나 대부분의 기존 스프레드시트 벤치마크는 단일 수식 생성 또는 로컬 셀 편집과 같은 격리된 작업을 평가하며, 따라서 실제 비즈니스 환경에서의 전체 워크플로우를 제대로 반영하지 못합니다. 본 논문에서는 세 가지 작업 범주(생성, 디버깅 및 시각화)를 포함하는 스프레드시트 에이전트를 위한 워크플로우 수준 벤치마크인 extsc{SpreadsheetBench 2}를 소개합니다. 이 벤치마크는 재무 보고서 및 기업 공시 자료를 포함한 실제 비즈니스 데이터를 기반으로 구축되었으며, 해당 분야 전문가에 의해 주석이 달리고 검증되었습니다. 벤치마크에는 총 321개의 작업이 포함되어 있으며, 각 작업은 평균적으로 11.8개의 워크시트를 사용하며 593.5개의 셀 수정 작업을 필요로 하여, 여러 시트 간의 의존성을 가진 대규모 워크북을 반영합니다. 본 논문에서는 통일된 다단계 에이전트 프레임워크 하에서 8개의 최첨단 거대 언어 모델(LLM)을 평가하고, 추가적으로 LLM 기반 스프레드시트 제품들을 보완적인 기준선으로 포함했습니다. 결과는 현재 시스템들이 실제 워크플로우에서 여전히 신뢰성이 떨어진다는 것을 보여줍니다. 최고 성능 모델의 전체 작업 정확도는 34.89%에 불과하며, 디버깅 정확도는 12.00%로 매우 낮습니다. 트래jectory 분석 및 오류 분류 결과, 스프레드시트 검사 부족과 잘못된 대상 셀 선택이 주요 문제점으로 지적됩니다. 이러한 결과를 종합적으로 고려할 때, extsc{SpreadsheetBench 2}는 신뢰성 있는 스프레드시트 자동화를 발전시키기 위한 도전적인 테스트 환경으로 자리매김합니다. 프로젝트 페이지: https://spreadsheetbench.github.io/
Spreadsheets are widely used for business analysis, financial modeling, reporting, and decision-making. However, most existing spreadsheet benchmarks evaluate isolated operations such as single-formula generation or local cell edits, and therefore fail to capture end-to-end workflows in realistic business settings. We introduce \textsc{SpreadsheetBench 2}, a workflow-level benchmark for spreadsheet agents that covers three task categories: generation, debugging, and visualization. The benchmark is constructed from authentic business data, including financial reports and corporate filings, and is annotated and validated by domain experts. The benchmark contains 321 tasks; each instance averages 11.8 worksheets and requires 593.5 cell modifications, reflecting large multi-sheet workbooks with cross-sheet dependencies. We evaluate eight frontier large language models under a unified multi-turn agent scaffold, and additionally include several LLM-based spreadsheet products as complementary baselines. Results show that current systems remain far from reliable on real-world workflows: the best model achieves 34.89\% overall task accuracy, and debugging accuracy is as low as 12.00\%. Trajectory analysis and a failure taxonomy further indicate that insufficient spreadsheet inspection and incorrect target-cell selection are the dominant bottlenecks. Together, these findings position \textsc{SpreadsheetBench 2} as a challenging testbed for advancing reliable spreadsheet automation. Project page: https://spreadsheetbench.github.io/
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.