CinemaTraj: LLM 에이전트를 활용한 3D 장면의 원자적 카메라 트래커토리 생성
CinemaTraj: Composing Atomic Camera Trajectories for 3D Scenes with LLM Agents
자연어 설명을 기반으로 3D 공간 내에서 시네마틱하게 표현력 있는 카메라 트래커토리를 자동으로 생성하는 것은 매우 중요하고 실용적인 과제이며, 부동산 광고부터 가상 투어 제작에 이르기까지 다양한 응용 분야를 가지고 있습니다. 기존 방법들은 대부분 2D 이미지 사전 지식에 의존하여 진정한 3차원 공간 인지 능력이 부족하거나, 트래커토리 생성을 시네마틱 의미와 동떨어진 기하학적 경로 계획 문제로 취급합니다. 본 논문에서는 CinemaTraj라는 프레임워크를 제시하며, 카메라 트래커토리 계획을 언어 기반의 공간 추론 문제로 재구성합니다. CinemaTraj는 RGB-D 이미지 세트와 사용자 프롬프트를 입력으로 받아, LLM 에이전트에 구조화된 3D 장면 그래프를 제공합니다. 이 에이전트는 프롬프트를 일련의 원자적 시네마틱 동작(도로리, 오빗, 크레인, 팬, 틸트, 줌, 아크)으로 분해합니다. 각 동작은 시네마틱 표현력이 뛰어나고 충돌 회피를 위한 최적화가 가능한 새로운 파라메트릭 트래커토리 표현 방식을 통해 구현됩니다. 장면 그래프는 구조적인 공간 사전 지식 역할을 하며, 에이전트의 추론을 환경에 대한 정확한 기하학적 및 의미적 정보에 기반하도록 합니다. CinemaTraj는 또한 카메라 움직임과 동기화된 내레이션과 자막을 생성하여, 나레이션이 포함된 시네마틱 비디오 결과물을 만듭니다. 실제 ScanNet++ 환경에서 CinemaTraj를 평가한 결과, 프롬프트의 충실도, 충돌 방지 능력, 높은 시네마틱 품질을 갖는 트래커토리를 생성하며, 기존 방식보다 프롬프트 일치성, 트래커토리 품질 및 안전성 측면에서 우수한 성능을 보였습니다.
Automatically generating cinematically expressive camera trajectories through 3D scenes from natural language descriptions is a challenging task of high practical value, with applications ranging from real-estate advertising to virtual tour creation. Existing methods either lack true 3D spatial awareness by relying on 2D image priors, or treat trajectory generation as a geometric path planning problem divorced from cinematographic semantics. We present CinemaTraj, a framework that reframes camera trajectory planning as a language-grounded spatial reasoning problem. Given a set of RGB-D images and a user prompt, CinemaTraj equips an LLM agent with a structured 3D scene graph: the agent decomposes the prompt into a sequence of atomic cinematographic movements (dolly, orbit, crane, pan, tilt, zoom, arc). Each movement is instantiated via a novel parametric trajectory representation that is both cinematographically expressive and optimizable for collision avoidance. The scene graph acts as a structured spatial prior, grounding the agent's reasoning in accurate geometric and semantic knowledge of the environment. CinemaTraj further generates synchronized voiceover and subtitles aligned with camera motion, producing narrated cinematic video outputs. We evaluate CinemaTraj on real-world ScanNet++ environments, and show that it produces prompt-faithful, collision-free trajectories with high cinematographic quality, outperforming existing approaches on prompt alignment, trajectory quality, and safety metrics.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.