2607.27654v1 Jul 30, 2026 cs.CL

단일 문서에서 교차 문서로: 대규모 언어 모델의 다중 수준 이벤트 분석 벤치마킹

From Single- to Cross-Document: Benchmarking Multi-Granularity Event Analysis of Large Language Models

Jie Zou
Jie Zou
Citations: 132
h-index: 7
Pei Ke
Pei Ke
Citations: 7
h-index: 2
Tao Wen
Tao Wen
Citations: 0
h-index: 0
Shuai Shao
Shuai Shao
Citations: 0
h-index: 0
Xu Han
Xu Han
Citations: 8
h-index: 1
Guannan Li
Guannan Li
Citations: 0
h-index: 0
Tao Tian
Tao Tian
Citations: 8
h-index: 1
Jinjie Qiu
Jinjie Qiu
Citations: 0
h-index: 0
Lan Wang
Lan Wang
Citations: 36
h-index: 3
Ke Qin
Ke Qin
Citations: 41
h-index: 4

이벤트 분석은 정보 추출의 필수적이고 기본적인 분야이며, 다양한 이벤트 중심 작업을 서로 다른 문서 세분성 수준에서 수행합니다. 대규모 언어 모델(LLM)은 일부 작업에서 유망한 성능을 보였지만, 기존 벤치마크의 제한된 문서 세분성, 작업 설계 및 데이터 소스로 인해 LLM의 이벤트 분석 능력에 대한 종합적인 이해는 부족합니다. 이러한 한계를 해결하기 위해, 우리는 LLM의 다중 수준 이벤트 분석 성능을 평가하는 체계적인 벤치마크인 MiGUE-Bench를 소개합니다. 대규모 평가를 지원하기 위해, 우리는 먼저 LLM 기반의 자체 수정 주석 프레임워크인 MiGUE-Pipeline을 개발하여 자동 레이블링을 통해 고품질의 이벤트 관련 데이터를 확장 가능하게 확보합니다. 그런 다음, 우리는 벤치마크의 핵심 작업 네 가지(이벤트 감지, 관계 추론, 구조 유도 및 미래 예측)를 설계하여 원자 수준의 이벤트 세부 정보부터 복잡한 교차 문서 내러티브에 이르기까지 다양한 수준에서 모델의 역량을 평가합니다. 최첨단 LLM과 검색 증강 생성(RAG) 방법에 대한 광범위한 실험은 현재 LLM의 성능 한계를 명확히 하고 중요한 결점을 파악하여, 어려운 이벤트 분석 작업에서 LLM을 개선하는 데 필요한 통찰력을 제공합니다.

Original Abstract

Event analysis is an essential and fundamental direction of information extraction, involving various event-centric tasks at different granularity of documents. While large language models (LLMs) have preliminarily achieved promising performance in part of these tasks individually, their capability in event analysis still lacks comprehensive understanding due to restricted document granularity, task designs, and data source of existing benchmarks. To address these limitations, we introduce MiGUE-Bench, a systematic benchmark for assessing the performance of LLMs in multi-granularity event analysis. To support large-scale evaluation, we first develop an LLM-driven self-correcting annotation framework called MiGUE-Pipeline, enabling scalable acquisition of high-quality source data of events with automatic labels. Then, we design four core tasks in our benchmark, i.e., event detection, relation reasoning, structure induction, and future prediction, to probe model competence at different levels, from atomic event details to complex cross-document narratives. Extensive experiments on state-of-the-art LLMs and retrieval-augmented generation (RAG) methods delineate the current capability boundary and identify critical deficiencies, providing insights into the future improvement of LLMs in challenging event analysis tasks.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!