STAGE: 영화 시나리오 기반 지식 그래프 구축, 질의 응답 및 인물 역할 수행을 위한 벤치마크
STAGE: A Benchmark for Knowledge Graph Construction, Question Answering, and In-Script Role-Playing over Movie Screenplays
영화 시나리오는 복잡한 인물 관계, 시간 순서에 따른 사건, 대화 기반 상호작용이 얽혀 있는 풍부한 장편 서사 구조를 가지고 있습니다. 기존 벤치마크는 질의 응답이나 대화 생성과 같은 개별적인 하위 작업에 초점을 맞추는 경우가 많지만, 모델이 일관성 있는 이야기 세계를 구축하고 다양한 추론 및 생성 과정에서 이를 활용하는지 평가하는 경우는 드뭅니다. 본 논문에서는 전체 길이의 영화 시나리오에 대한 서사 이해를 위한 통합 벤치마크인 STAGE (Screenplay Text, Agents, Graphs and Evaluation)를 소개합니다. STAGE는 지식 그래프 구축, 장면 레벨 이벤트 요약, 장문 시나리오 질의 응답, 그리고 시나리오 내 인물 역할 수행이라는 네 가지 작업을 정의하며, 모든 작업은 공유된 서사 세계 표현을 기반으로 합니다. 본 벤치마크는 영어 및 중국어 영화 150편에 대한 정제된 시나리오, 큐레이션된 지식 그래프, 이벤트 및 인물 중심의 어노테이션을 제공하여, 모델이 세계 표현을 구축하고, 서사 이벤트의 추상화 및 검증을 수행하며, 장문의 서사에 대해 추론하고, 인물 일관성을 유지하는 응답을 생성하는 능력을 종합적으로 평가할 수 있도록 지원합니다.
Movie screenplays are rich long-form narratives that interleave complex character relationships, temporally ordered events, and dialogue-driven interactions. While prior benchmarks target individual subtasks such as question answering or dialogue generation, they rarely evaluate whether models can construct a coherent story world and use it consistently across multiple forms of reasoning and generation. We introduce STAGE (Screenplay Text, Agents, Graphs and Evaluation), a unified benchmark for narrative understanding over full-length movie screenplays. STAGE defines four tasks: knowledge graph construction, scene-level event summarization, long-context screenplay question answering, and in-script character role-playing, all grounded in a shared narrative world representation. The benchmark provides cleaned scripts, curated knowledge graphs, and event- and character-centric annotations for 150 films across English and Chinese, enabling holistic evaluation of models' abilities to build world representations, abstract and verify narrative events, reason over long narratives, and generate character-consistent responses.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.