2606.26964v1 Jun 25, 2026 cs.AI

먼저 보고 움직이기: 동적인 3차원 스토리 월드에서 서사 기반의 시각적 주의 집중

Look-Before-Move: Narrative-Grounded World Visual Attention in Dynamic 3D Story Worlds

Zhi Wang
Zhi Wang
Citations: 198
h-index: 3
Zhenhong Sun
Zhenhong Sun
Citations: 39
h-index: 4
Jiaming Bian
Jiaming Bian
Citations: 24
h-index: 1
Yueh-Hua Wu
Yueh-Hua Wu
Citations: 203
h-index: 5
Huadong Mo
Huadong Mo
Citations: 3
h-index: 1
Bingliang Li
Bingliang Li
Citations: 75
h-index: 4
Pichao Wang
Pichao Wang
Citations: 63
h-index: 3
Hailan Ma
Hailan Ma
Citations: 83
h-index: 3

구체화된 인공지능과 세계 모델이 점점 더 역동적인 3차원 환경에서 작동함에 따라, 시각적 인식은 단순히 주어진 관찰 내용을 해석하는 것을 넘어 적극적으로 무엇을 관찰할지를 결정해야 합니다. 본 연구에서는 동적인 3차원 스토리 월드에서의 카메라 계획 문제를 통해 이 문제를 탐구합니다. 여기서 카메라는 부드러운 움직임을 생성할 뿐만 아니라 이동하기 전에 어떤 시각적 증거를 확보해야 할지 결정해야 합니다. 우리는 이러한 능력을 '서사 기반의 세계 시각적 주의 집중(Narrative-Grounded World Visual Attention)'으로 정의하며, 카메라는 서사 의도와 물리적인 3차원 제약 조건 하에서 무엇을 관찰할지, 어떻게 관찰 내용을 구성할지, 그리고 시간에 따른 주의 집중을 어떻게 전환할지를 결정하는 구체화된 관찰자 역할을 합니다. 이러한 능력을 구현하기 위해, 우리는 '먼저 보고 움직이기(Look-Before-Move)'라는 카메라 계획 프레임워크를 제안합니다. 이 프레임워크는 관찰 사양과 동작 실행을 분리합니다. 먼저, 지시자의 의도를 실행 가능한 시각적 제약 조건으로 변환하는 '의미론적 관찰 계약(Semantic Observation Contract)'을 구축하고, 그 다음에는 '몬테 카를로 뷰포인트 탐색(Monte Carlo Viewpoint Search)'을 통해 서사에 부합하고 기하학적으로 실현 가능한 뷰포인트를 찾습니다. 마지막으로, 선택된 뷰포인트를 연결하여 연속적이고 충돌 방지되며 시간적으로 일관된 카메라 움직임을 생성하기 위해 '의미론적 경로 고정(Semantic Trajectory Grounding)'을 적용합니다. 또한, 우리는 StoryBlender를 기반으로 애니메이션 캐릭터, 의미론적 장면 구성 및 실행 가능한 3차원 환경을 포함하는 50개의 스토리, 457개의 장면, 1585개의 샷으로 구성된 동적인 3차원 스토리 월드 벤치마크를 구축했습니다. 실험 결과는 우리 프레임워크가 대표적인 기준 모델보다 주관적 인식, 의도 일관성 및 경로 품질을 향상시킨다는 것을 보여주며, 이는 카메라 움직임을 생성하기 전에 시각적 주의 집중을 조직하는 것이 중요하다는 점을 입증합니다.

Original Abstract

As embodied AI and world models increasingly operate in dynamic 3D environments, visual perception must move beyond passively interpreting given observations toward actively deciding what to observe. We study this problem through camera planning in dynamic 3D story worlds, where the camera must not only generate smooth motion, but also decide what visual evidence should be acquired before it moves. We formulate this capability as Narrative-Grounded World Visual Attention, where the camera acts as an embodied observer that determines what to observe, how to compose the observation, and how to shift attention over time under narrative intent and physical 3D constraints. To realize this capability, we propose Look-Before-Move, a camera planning framework that separates observation specification from motion execution. It first builds a Semantic Observation Contract to convert directorial intent into executable visual constraints, then performs Monte Carlo Viewpoint Search to find narrative-compliant and geometrically feasible viewpoints, and finally applies Semantic Trajectory Grounding to connect selected viewpoints into continuous, collision-aware, and temporally coherent camera motion. We further construct a dynamic 3D Story World Benchmark based on StoryBlender, covering 50 stories, 457 scenes, and 1585 shots with animated characters, semantic scene configurations, and executable 3D environments. Experiments show that our framework improves subject perception, intent consistency, and trajectory quality over representative baselines, demonstrating the importance of organizing visual attention before generating camera motion.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!