2606.17391v1 Jun 16, 2026 cs.CL

NarrativeWorldBench: 최첨단 모델의 한계점을 보여주는 벤치마크와 장기 협업 오디오 드라마를 위한 잠재 세계 모델

NarrativeWorldBench: A Frontier-Saturated Benchmark and a Latent World Model for Long-Horizon Co-Creative Audio Drama

Vasu Sharma
Vasu Sharma
Citations: 919
h-index: 7
Logan Mann
Logan Mann
University of California, Santa Barbara
Citations: 2
h-index: 1
Abdur Rahman
Abdur Rahman
Indian Institute of Technology Delhi
Citations: 15
h-index: 2
Mohammad Saifullah
Mohammad Saifullah
Citations: 61
h-index: 3
Taaha Kazi
Taaha Kazi
Citations: 13
h-index: 2

200~800부작으로 구성되는 장편 연재형 오디오 드라마는 중요한 창작 매체이며, 현재 가장 뛰어난 성능을 보이는 대규모 언어 모델(LLM)조차도 이 분야에서 어려움을 겪습니다. 본 연구에서는 다양한 유형의 21개 모델(기존 모델, 미세 조정 모델, 개방형 최첨단 모델, 폐쇄형 최첨단 모델, 추론 모델)을 대상으로 구조적 내러티브 지표를 활용한 균일한 성능 평가를 진행했습니다. 폐쇄형 시스템은 모든 모델에서 플롯 구성 요소 F1 점수가 [0.78, 0.81] 범위에 도달하며, 에피소드 수(horizon)가 h=200이 되면 약 -0.20의 F1 점수 감소를 보입니다. 본 연구에서는 NarrativeWorldBench라는 오픈형 벤치마크를 제안합니다. 이 벤치마크는 9가지 내러티브 구조 지표를 활용하여 {10, 20, 50, 100, 200} 에피소드 수에 따른 성능을 평가하며, 힌디어, 타밀어, 텔루구어, 마라티어를 포함한 네 가지 인도 지역 언어로 다국어 평가를 수행합니다. 또한 N-VSSM이라는 내러티브 변동 상태 공간 모델을 소개합니다. 이 모델은 Mamba-2 아키텍처를 기반으로 하며, 이벤트에 조건화된 사후 확률과 80억 개의 파라미터를 가진 디코더를 사용하여 200편 이상의 에피소드 동안 구조화된 256차원의 잠재 세계 상태를 유지합니다. N-VSSM은 폐쇄형 시스템보다 4배 낮은 컴퓨팅 자원으로 모든 에피소드 수에서 플롯 구성 요소 F1 점수가 0.84 이상을 달성합니다. 학습된 문화적 전이 함수는 다국어 정확도를 Likert 점수 기준 +0.20 ~ +0.23만큼 향상시킵니다. 또한, 12명의 전문 작가(총 240개 시도)를 대상으로 한 사용자 연구에서 N-VSSM은 Claude Opus 4.5보다 장편의 일관성 측면에서 71% 더 높은 선호도를 받았으며, 제어 가능성 측면에서도 Likert 점수 기준 +1.3점이 더 높게 평가되었습니다.

Original Abstract

Long-form serialized audio drama, with arcs that run for 200 to 800 episodes, is a major creative medium and a setting where frontier large language models (LLMs) fail. We benchmark 21 models, spanning classical, fine-tuned, open-frontier, closed-frontier, and reasoning tiers, on a uniform set of structural narrative metrics. All closed-frontier systems saturate at a plot-beat F1 in the band [0.78, 0.81] and collapse by about -0.20 F1 at horizon h=200. We introduce NarrativeWorldBench, an open benchmark of nine narrative-structure metrics evaluated across horizons h in {10, 20, 50, 100, 200}, with cross-lingual evaluation across four Indic languages (Hindi, Tamil, Telugu, Marathi). We introduce N-VSSM, a Narrative Variational State-Space Model that maintains a structured 256-dimensional latent world state over more than 200 episodes via a Mamba-2 backbone with an event-conditioned posterior and an 8B decoder. N-VSSM holds plot-beat F1 >= 0.84 across all horizons at 4x lower compute than the closed-frontier band. A learned Cultural Transfer Function lifts cross-language fidelity by +0.20 to +0.23 Likert points. In a within-subjects writer study (n = 12 professional authors, 240 trials), N-VSSM is preferred over Claude Opus 4.5 on long-arc consistency 71% of the time and rated +1.3 Likert points higher on controllability.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!