2607.26910v1 Jul 29, 2026 cs.CV

CinemaTraj: LLM 에이전트를 활용한 3D 장면의 원자적 카메라 트래커토리 생성

CinemaTraj: Composing Atomic Camera Trajectories for 3D Scenes with LLM Agents

Yanfeng Zhang
Yanfeng Zhang
Citations: 15
h-index: 3
Liqiu Meng
Liqiu Meng
Citations: 10
h-index: 2
Qianru Li
Qianru Li
Citations: 1
h-index: 1
Xuyang Chen
Xuyang Chen
Citations: 11
h-index: 2
Erkin Türköz
Erkin Türköz
Citations: 1
h-index: 1
Lu Liu
Lu Liu
Citations: 20
h-index: 3
Tao Wu
Tao Wu
Citations: 15
h-index: 2
Xuqin Wang
Xuqin Wang
Citations: 10
h-index: 2

자연어 설명을 기반으로 3D 공간 내에서 시네마틱하게 표현력 있는 카메라 트래커토리를 자동으로 생성하는 것은 매우 중요하고 실용적인 과제이며, 부동산 광고부터 가상 투어 제작에 이르기까지 다양한 응용 분야를 가지고 있습니다. 기존 방법들은 대부분 2D 이미지 사전 지식에 의존하여 진정한 3차원 공간 인지 능력이 부족하거나, 트래커토리 생성을 시네마틱 의미와 동떨어진 기하학적 경로 계획 문제로 취급합니다. 본 논문에서는 CinemaTraj라는 프레임워크를 제시하며, 카메라 트래커토리 계획을 언어 기반의 공간 추론 문제로 재구성합니다. CinemaTraj는 RGB-D 이미지 세트와 사용자 프롬프트를 입력으로 받아, LLM 에이전트에 구조화된 3D 장면 그래프를 제공합니다. 이 에이전트는 프롬프트를 일련의 원자적 시네마틱 동작(도로리, 오빗, 크레인, 팬, 틸트, 줌, 아크)으로 분해합니다. 각 동작은 시네마틱 표현력이 뛰어나고 충돌 회피를 위한 최적화가 가능한 새로운 파라메트릭 트래커토리 표현 방식을 통해 구현됩니다. 장면 그래프는 구조적인 공간 사전 지식 역할을 하며, 에이전트의 추론을 환경에 대한 정확한 기하학적 및 의미적 정보에 기반하도록 합니다. CinemaTraj는 또한 카메라 움직임과 동기화된 내레이션과 자막을 생성하여, 나레이션이 포함된 시네마틱 비디오 결과물을 만듭니다. 실제 ScanNet++ 환경에서 CinemaTraj를 평가한 결과, 프롬프트의 충실도, 충돌 방지 능력, 높은 시네마틱 품질을 갖는 트래커토리를 생성하며, 기존 방식보다 프롬프트 일치성, 트래커토리 품질 및 안전성 측면에서 우수한 성능을 보였습니다.

Original Abstract

Automatically generating cinematically expressive camera trajectories through 3D scenes from natural language descriptions is a challenging task of high practical value, with applications ranging from real-estate advertising to virtual tour creation. Existing methods either lack true 3D spatial awareness by relying on 2D image priors, or treat trajectory generation as a geometric path planning problem divorced from cinematographic semantics. We present CinemaTraj, a framework that reframes camera trajectory planning as a language-grounded spatial reasoning problem. Given a set of RGB-D images and a user prompt, CinemaTraj equips an LLM agent with a structured 3D scene graph: the agent decomposes the prompt into a sequence of atomic cinematographic movements (dolly, orbit, crane, pan, tilt, zoom, arc). Each movement is instantiated via a novel parametric trajectory representation that is both cinematographically expressive and optimizable for collision avoidance. The scene graph acts as a structured spatial prior, grounding the agent's reasoning in accurate geometric and semantic knowledge of the environment. CinemaTraj further generates synchronized voiceover and subtitles aligned with camera motion, producing narrated cinematic video outputs. We evaluate CinemaTraj on real-world ScanNet++ environments, and show that it produces prompt-faithful, collision-free trajectories with high cinematographic quality, outperforming existing approaches on prompt alignment, trajectory quality, and safety metrics.

0 Citations
0 Influential
1.5 Altmetric
7.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!