2603.05890v1 Mar 06, 2026 cs.CL

이야기 속에서 길을 잃다: LLM으로 생성된 장편 스토리에서의 일관성 오류

Lost in Stories: Consistency Bugs in Long Story Generation by LLMs

Junjie Li
Junjie Li
Citations: 286
h-index: 4
Xinru Guo
Xinru Guo
Citations: 39
h-index: 4
Yuhao Wu
Yuhao Wu
Citations: 100
h-index: 6
Roy Ka-Wei Lee
Roy Ka-Wei Lee
Citations: 10
h-index: 1
Hongzhi Li
Hongzhi Li
Citations: 4
h-index: 1
Yutao Xie
Yutao Xie
Citations: 4
h-index: 1

이야기꾼이 스스로의 이야기를 잊어버린다면 어떤 일이 벌어질까요? 대규모 언어 모델(LLM)은 이제 수만 단어에 달하는 이야기를 생성할 수 있지만, 종종 이야기 전체의 일관성을 유지하는 데 실패합니다. 장편 이야기를 생성할 때, 이러한 모델들은 스스로 설정한 사실, 등장인물의 특징, 그리고 세계관의 규칙에 대해 모순되는 내용을 생성할 수 있습니다. 기존의 스토리 생성 벤치마크는 주로 줄거리의 품질과 유창성에 초점을 맞추고 있어, 일관성 오류는 대부분 탐구되지 않았습니다. 이러한 격차를 해소하기 위해, 우리는 장편 스토리 생성에서 이야기의 일관성을 평가하기 위한 벤치마크인 ConStory-Bench를 제시합니다. 이 벤치마크는 2,000개의 프롬프트를 포함하며, 네 가지 작업 시나리오를 다루고, 5가지 오류 범주와 19가지 세부 유형으로 분류된 분류 체계를 정의합니다. 또한, 모순을 감지하고 각 판단의 근거를 명시적인 텍스트 증거로 제시하는 자동화된 파이프라인인 ConStory-Checker를 개발했습니다. 다섯 가지 연구 질문을 통해 다양한 LLM을 평가한 결과, 일관성 오류는 명확한 경향성을 보입니다. 즉, 사실적 및 시간적 측면에서 가장 흔하게 발생하며, 이야기의 중간 부분에서 나타나는 경향이 있고, 토큰 수준의 엔트로피가 높은 텍스트 세그먼트에서 발생하며, 특정 유형의 오류는 함께 발생하는 경향이 있습니다. 이러한 결과는 장편 이야기 생성에서 일관성을 향상시키기 위한 향후 노력에 도움이 될 수 있습니다. 프로젝트 페이지는 https://picrew.github.io/constory-bench.github.io/ 에서 확인할 수 있습니다.

Original Abstract

What happens when a storyteller forgets its own story? Large Language Models (LLMs) can now generate narratives spanning tens of thousands of words, but they often fail to maintain consistency throughout. When generating long-form narratives, these models can contradict their own established facts, character traits, and world rules. Existing story generation benchmarks focus mainly on plot quality and fluency, leaving consistency errors largely unexplored. To address this gap, we present ConStory-Bench, a benchmark designed to evaluate narrative consistency in long-form story generation. It contains 2,000 prompts across four task scenarios and defines a taxonomy of five error categories with 19 fine-grained subtypes. We also develop ConStory-Checker, an automated pipeline that detects contradictions and grounds each judgment in explicit textual evidence. Evaluating a range of LLMs through five research questions, we find that consistency errors show clear tendencies: they are most common in factual and temporal dimensions, tend to appear around the middle of narratives, occur in text segments with higher token-level entropy, and certain error types tend to co-occur. These findings can inform future efforts to improve consistency in long-form narrative generation. Our project page is available at https://picrew.github.io/constory-bench.github.io/.

4 Citations
2 Influential
3 Altmetric
23.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!