2607.11798v1 Jul 13, 2026 cs.CV

StoryTeller: 학습 없이 스토리 정보를 활용하는 장편 오디오 설명 시스템

StoryTeller: Training-Free Narrative Grounding for Long-Form Audio Description

SouYoung Jin
SouYoung Jin
Citations: 37
h-index: 2
Minh T. Dinh
Minh T. Dinh
Citations: 0
h-index: 0
Seung-Yeon Hahm
Seung-Yeon Hahm
Citations: 32
h-index: 3

장편 오디오 설명(AD)은 단순히 보이는 행동을 묘사하는 것 이상으로, 시각 장애 또는 저시력(BLV) 청중이 영화를 따라갈 수 있도록 등장인물, 사건, 관계 및 스토리 맥락을 장면 전체에 걸쳐 보존해야 합니다. 최신 비디오-언어 모델(VLM)은 짧은 클립에서는 효과적이지만, 종종 각 순간을 독립적으로 처리하여 누가 등장인물인지, 어떤 사건이 중요한지, 그리고 현재 장면이 이전 스토리 맥락과 어떻게 연결되는지에 대한 설명을 놓치는 경우가 많습니다. 우리는 스토리 정보를 고려한 장편 AD를 위한 학습이 필요 없는 프레임워크인 StoryTeller를 제안합니다. StoryTeller는 지역적인 시각적 단서에만 의존하는 대신, 검증된 스토리 메모리를 유지하여 장면 전체에 걸쳐 스토리와 관련된 정보를 전달하고, 이를 통해 후속 설명이 일관되고 맥락을 기반으로 정보 제공하도록 합니다. StoryTeller는 원본 비디오와 영화 제목만 입력받아 공개 영화 메타데이터를 활용하여 이름과 스토리 맥락을 파악할 수 있으며, 의미 분석 필터링 및 VLM 검증을 통해 비디오에 의해 뒷받침되는 사실만을 사용합니다. 이 방법은 자막, 시나리오, AD 스크립트, 정렬된 캡션, 등장인물 목록, 미리 계산된 얼굴 정보 또는 특정 작업에 대한 미세 조정이 필요하지 않습니다. 생성된 AD가 스토리 정보를 얼마나 잘 보존하는지 평가하기 위해, 우리는 생성된 설명만으로 언어 모델이 스토리 맥락 관련 질문에 답변할 수 있는지 테스트하는 질의응답 벤치마크인 StoryAD-QA를 소개합니다. 표준 AD 벤치마크 및 다양한 장편 비디오에 대한 실험 결과, StoryTeller는 자동 평가, 질의응답 기반 평가 및 인간 평가 모두에서 강력한 기준 모델보다 일관되게 스토리 일관성, 사실적 정확성 및 스토리 이해도를 향상시키는 것으로 나타났습니다.

Original Abstract

Long-form audio description (AD) requires more than describing visible actions: it must preserve characters, events, relationships, and story context across scenes so that blind and low-vision (BLV) audiences can follow a film. Modern video-language models (VLMs) are effective on short clips, but they often treat each moment independently, producing descriptions that miss who characters are, why events matter, and how the current scene connects to earlier narrative context. We propose StoryTeller, a training-free framework for story-aware long-form AD. Instead of relying only on local visual cues, StoryTeller maintains a verified narrative memory that carries forward story-relevant information across scenes, enabling later descriptions to remain coherent, grounded, and contextually informative. Given only raw video and a movie title, StoryTeller can optionally retrieve public movie metadata to resolve names and story context, while accepting only facts that are supported by the video through semantic filtering and VLM verification. The method requires no subtitles, scripts, AD transcripts, aligned captions, character banks, precomputed face identities, or task-specific fine-tuning. To evaluate whether generated AD preserves narrative information, we introduce StoryAD-QA, a question-answering benchmark that tests whether a language model can answer story-context questions using only the generated descriptions. Experiments on standard AD benchmarks and diverse long-form videos show that StoryTeller consistently improves narrative coherence, factual grounding, and story comprehension over strong baselines in automatic, QA-based, and human evaluations.

0 Citations
0 Influential
1.5 Altmetric
7.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!