2601.08510v2 Jan 13, 2026 cs.CL

STAGE: 영화 시나리오 기반 지식 그래프 구축, 질의 응답 및 인물 역할 수행을 위한 벤치마크

STAGE: A Benchmark for Knowledge Graph Construction, Question Answering, and In-Script Role-Playing over Movie Screenplays

Qiuyu Tian
Qiuyu Tian
Citations: 10
h-index: 1
Feng Chen
Feng Chen
Citations: 14
h-index: 3
Zequn Liu
Zequn Liu
Citations: 38
h-index: 2
Youyong Kong
Youyong Kong
Citations: 7
h-index: 1
Fan Guo
Fan Guo
Citations: 993
h-index: 9
Jinjing Shen
Jinjing Shen
Citations: 1
h-index: 1
Yiyun Luo
Yiyun Luo
Citations: 18
h-index: 2
Xin Zhang
Xin Zhang
Citations: 63
h-index: 4
Zhijing Xie
Zhijing Xie
Citations: 47
h-index: 4
Yiding Li
Yiding Li
Citations: 4
h-index: 1
Yin Xia
Yin Xia
Citations: 90
h-index: 2
Yuyao Li
Yuyao Li
Citations: 12
h-index: 2

영화 시나리오는 복잡한 인물 관계, 시간 순서에 따른 사건, 대화 기반 상호작용이 얽혀 있는 풍부한 장편 서사 구조를 가지고 있습니다. 기존 벤치마크는 질의 응답이나 대화 생성과 같은 개별적인 하위 작업에 초점을 맞추는 경우가 많지만, 모델이 일관성 있는 이야기 세계를 구축하고 다양한 추론 및 생성 과정에서 이를 활용하는지 평가하는 경우는 드뭅니다. 본 논문에서는 전체 길이의 영화 시나리오에 대한 서사 이해를 위한 통합 벤치마크인 STAGE (Screenplay Text, Agents, Graphs and Evaluation)를 소개합니다. STAGE는 지식 그래프 구축, 장면 레벨 이벤트 요약, 장문 시나리오 질의 응답, 그리고 시나리오 내 인물 역할 수행이라는 네 가지 작업을 정의하며, 모든 작업은 공유된 서사 세계 표현을 기반으로 합니다. 본 벤치마크는 영어 및 중국어 영화 150편에 대한 정제된 시나리오, 큐레이션된 지식 그래프, 이벤트 및 인물 중심의 어노테이션을 제공하여, 모델이 세계 표현을 구축하고, 서사 이벤트의 추상화 및 검증을 수행하며, 장문의 서사에 대해 추론하고, 인물 일관성을 유지하는 응답을 생성하는 능력을 종합적으로 평가할 수 있도록 지원합니다.

Original Abstract

Movie screenplays are rich long-form narratives that interleave complex character relationships, temporally ordered events, and dialogue-driven interactions. While prior benchmarks target individual subtasks such as question answering or dialogue generation, they rarely evaluate whether models can construct a coherent story world and use it consistently across multiple forms of reasoning and generation. We introduce STAGE (Screenplay Text, Agents, Graphs and Evaluation), a unified benchmark for narrative understanding over full-length movie screenplays. STAGE defines four tasks: knowledge graph construction, scene-level event summarization, long-context screenplay question answering, and in-script character role-playing, all grounded in a shared narrative world representation. The benchmark provides cleaned scripts, curated knowledge graphs, and event- and character-centric annotations for 150 films across English and Chinese, enabling holistic evaluation of models' abilities to build world representations, abstract and verify narrative events, reason over long narratives, and generate character-consistent responses.

2 Citations
0 Influential
4.5 Altmetric
24.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!