2605.26525v1 May 26, 2026 cs.CV

ReCA: 재귀적 컨텍스트 할당을 통한 멀티샷 장편 비디오 확장

ReCA: Multi-Shot Long Video Extrapolation via Recursive Context Allocation

Akide Liu
Akide Liu
Citations: 18
h-index: 3
Bohan Zhuang
Bohan Zhuang
Citations: 6,351
h-index: 37
Weijie Wang
Weijie Wang
Citations: 271
h-index: 4
Gholamreza Haffari
Gholamreza Haffari
Citations: 10
h-index: 1
Jinbo Xing
Jinbo Xing
Citations: 42
h-index: 4
Chaojie Mao
Chaojie Mao
Citations: 2,528
h-index: 10
Yefei He
Yefei He
Citations: 96
h-index: 5
Zeyu Zhang
Zeyu Zhang
Citations: 19
h-index: 2
Ye Li
Ye Li
Citations: 69
h-index: 4
Zihan Wang
Zihan Wang
Citations: 25
h-index: 2
Y. Liu
Y. Liu
Citations: 499
h-index: 7

분 단위의 영화 같은 비디오 생성은 생성형 비디오 모델의 핵심적인 과제입니다. 기존 방법들은 이 과제의 일부만을 해결합니다. 싱글샷 확장은 기준 프레임을 유지하지만, 영화적 구조가 부족하며, 멀티샷 스토리텔링은 구조를 부여하지만 관찰된 내용을 이어가는 것이 아니라 시각적으로 새로운 상태를 창조하는 경향이 있습니다. 본 연구에서는 Multi-Shot Video Extrapolation (MSVE)이라는 과제를 정의합니다. MSVE는 주어진 프레임 또는 클립을 영화적 구조로 구성된 일련의 샷으로 확장하면서, 기준 상태를 유지하고 스토리 전개 의도를 발전시키는 것입니다. 이 설정은 짧은 비디오 모델의 제한적인 생성 예산 하에서 작동합니다. 우리는 세 가지 상호 연관된 병목 현상을 확인했습니다: (1) 글로벌 플래너는 전체 시나리오에서 지원되지 않는 상세 정보를 과도하게 지정하고, (2) 샷 수준의 프롬프트는 전체 스토리를 담고 있을 때 중요한 상태 정보가 희석되며, (3) 시간적 연결은 생성된 프레임을 손실이 발생하는 메모리로 만들어, 개체, 장면, 객체 및 동작 상태가 점차 사라집니다. MSVE 연구 결과는 장편 비디오 실패가 단순히 컨텍스트 길이의 제한 때문이 아니라, 컨텍스트 할당 방식의 문제라는 것을 보여줍니다. 우리는 Recursive Context Allocation (ReCA)이라는 추론 시간 프레임워크를 제안합니다. ReCA는 계획 및 생성 과정에서 계층적으로 컨텍스트를 할당합니다. 재귀적으로 MSVE 문제를 컨텍스트 제한적인 하위 문제로 분해하고, 리프 노드에서 고정된 생성기를 사용하며, 시간에 따른 구조화된 상태 업데이트를 전파합니다. 이 과제를 평가하기 위해, 기존의 짧은 클립 벤치마크에서는 다루지 않는 3분에서 5분 길이의 장편 비디오 생성을 위한 프롬프트를 특별히 설계한 MSVE-Bench와 NB-Q라는 새로운 프로토콜을 제안합니다. 실험 결과, ReCA는 가장 강력한 기존 모델보다 평균 정규화된 성능이 8~16% 향상되었으며, 멀티샷 일관성 지표가 28~43% 향상되었습니다. 프로젝트 페이지는 https://reca.vmv.re 에서 확인할 수 있습니다.

Original Abstract

Minute-scale cinematic video generation is a central challenge for generative video models. Existing paradigms address only fragments of this challenge: single-shot extrapolation preserves an anchor but lacks cinematic structure, while multi-shot storytelling imposes structure yet remains free to invent its visual states rather than continue an observed one. We define Multi-Shot Video Extrapolation (MSVE), a task that extends an observed frame or clip into a sequence of cinematically structured shots while preserving anchor state and advancing narrative intent. This setting operates under the finite per-call generation budget of short-video models. We identify three coupled bottlenecks: (1) global planners over-specify unsupported details from full screenplays; (2) shot-level prompts dilute task-relevant state when carrying the complete story; and (3) temporal chaining turns generated frames into a lossy memory in which identity, scene, object, and action state decay. MSVE reveals that long-video failure is not merely a limitation of context length, but a failure of context allocation. We propose Recursive Context Allocation (ReCA), an inference-time framework that allocates context hierarchically across planning and generation. ReCA recursively decomposes MSVE into context-bounded subproblems, invokes frozen generators at leaf nodes, and propagates structured state updates across time. To evaluate this setting, we further propose MSVE-Bench and NB-Q, a source-grounded protocol with prompts purpose-built for 3 to 5 minute long-video generation, a regime not addressed by existing short-clip benchmarks. Compared to previous methods, ReCA improves average normalized score by 8 to 16 percent over the strongest competing controller and improves multi-shot consistency metrics by 28 to 43 percent. View the project page at https://reca.vmv.re.

2 Citations
0 Influential
18.5 Altmetric
94.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!