2605.05811v1 May 07, 2026 cs.AI

시트 토큰: 다중 시트 스프레드시트 이해를 위한 그래프 기반 강화 표현

Sheet as Token: A Graph-Enhanced Representation for Multi-Sheet Spreadsheet Understanding

Boyuan Guan
Boyuan Guan
Citations: 94
h-index: 2
Yiming Lei
Yiming Lei
Citations: 39
h-index: 2
D. Zhu
D. Zhu
Citations: 6
h-index: 1
Tianyu Shi
Tianyu Shi
Citations: 9
h-index: 2
Chunhui Wang
Chunhui Wang
Citations: 1
h-index: 1
Yujia Zhang
Yujia Zhang
Citations: 18
h-index: 1
Zhuo Hao
Zhuo Hao
Citations: 15
h-index: 2
Yiqin Wang
Yiqin Wang
Citations: 316
h-index: 8

워크북 단위의 스프레드시트 이해는 언어 모델 기반 데이터 분석 에이전트에 점점 더 중요해지고 있지만, 관련 정보가 종종 이질적인 스키마, 레이아웃 및 암묵적인 관계를 가진 여러 시트에 분산되어 있어 여전히 어려운 과제입니다. 기존의 검색 증강 방식은 확장성을 높이기 위해 일반적으로 스프레드시트를 행, 열 또는 블록으로 분해하지만, 이러한 청크 중심 표현 방식은 워크시트를 고립된 텍스트 범위로 분리하고 전역적인 시트 수준의 의미를 약화시킬 수 있습니다. 본 논문에서는 각 워크시트를 단일한 의미 단위로 취급하여 다중 시트 스프레드시트 검색을 위한 그래프 기반 강화 프레임워크인 '시트 토큰(Sheet as Token)'을 제안합니다. 저희 방법은 시트 이름, 열 머리글, 대표 값 및 레이아웃 특징으로부터 스키마 정보를 활용하여 워크시트를 추출하고, 각 워크시트를 컴팩트한 밀집 토큰으로 인코딩합니다. 자연어 쿼리가 주어지면, 그래프 검색기는 의미, 쿼리 의존성, 스키마 일관성 및 모양 호환성 관계를 사용하여 쿼리에 특화된 후보 그래프를 시트 토큰 위에서 구성하고, 이러한 채널들을 다단계 그래프 변환기를 통해 결합하여 관련 시트 집합을 검색합니다. 저희가 구축한 다중 시트 스프레드시트 데이터셋에 대한 실험 결과, 시트 수준 토큰화는 안정적인 표현을 학습하며, 그래프 기반의 시트 간 추론은 제한적인 추가적인 그래프 계산으로 얕은 그래프 기반의 성능을 능가하는 목록 기반 검색 성능을 향상시킬 수 있음을 보여줍니다. 이러한 결과는 시트 수준 토큰화가 확장 가능한 다중 시트 스프레드시트 이해를 위한 유망한 추상화 방법임을 시사합니다.

Original Abstract

Workbook-scale spreadsheet understanding is increasingly important for language-model-based data analysis agents, but remains challenging because relevant information is often distributed across multiple sheets with heterogeneous schemas, layouts, and implicit relationships. Existing retrieval-augmented approaches typically decompose spreadsheets into rows, columns, or blocks to improve scalability; however, such chunk-centric representations can fragment worksheets into isolated text spans and weaken global sheet-level semantics. We propose Sheet as Token, a graph-enhanced framework that treats each worksheet as a unified semantic unit for multi-sheet spreadsheet retrieval. Our method extracts schema-aware records from sheet names, column headers, representative values, and layout features, and encodes each worksheet into a compact dense token. Given a natural-language query, a Graph Retriever constructs a query-specific candidate graph over sheet tokens using semantic, query-conditioned, schema-consistency, and shape-compatibility relations, and composes these channels through a multi-stage graph transformer to retrieve supporting sheet sets. Experiments on a constructed multi-sheet spreadsheet corpus show that sheet-level tokenization learns stable representations, and that graph-enhanced cross-sheet reasoning improves listwise retrieval over a shallow graph baseline with limited additional graph-side computation. These results suggest that sheet-level tokenization is a promising abstraction for scalable multi-sheet spreadsheet understanding.

1 Citations
0 Influential
4 Altmetric
21.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!