2603.06140v1 Mar 06, 2026 cs.CV

Place-it-R1: 멀티모달 대규모 언어 모델(MLLM)의 환경 인식 추론 잠재력을 활용한 비디오 객체 삽입

Place-it-R1: Unlocking Environment-aware Reasoning Potential of MLLM for Video Object Insertion

Bohai Gu
Bohai Gu
Citations: 37
h-index: 2
Taiyi Wu
Taiyi Wu
Citations: 4
h-index: 1
Dazhao Du
Dazhao Du
Citations: 378
h-index: 8
Jian Liu
Jian Liu
Citations: 345
h-index: 3
Shuai Yang
Shuai Yang
Citations: 336
h-index: 6
Xiaotong Zhao
Xiaotong Zhao
Citations: 133
h-index: 4
Alan Zhao
Alan Zhao
Citations: 29
h-index: 3
Song Guo
Song Guo
Citations: 31
h-index: 3

최신 비디오 편집 기술은 비디오 객체를 삽입할 때 높은 시각적 충실도를 달성하지만, 시각적 충실도를 최적화하는 데 집중하여 물리적 인과 관계를 고려하지 못하는 경우가 많아 환경과 물리적으로 일관되지 않은 편집 결과가 발생합니다. 본 연구에서는 멀티모달 대규모 언어 모델(MLLM)의 환경 인식 추론 잠재력을 활용하여 비디오 객체 삽입을 위한 엔드투엔드 프레임워크인 Place-it-R1을 제안합니다. 저희의 프레임워크는 MLLM의 연쇄적 사고(Chain-of-Thought, CoT) 추론을 활용하여 비디오 확산을 조정하며, '생각-후-삽입(Think-then-Place)' 패러다임을 따릅니다. 인지적 추론과 생성 실행 간의 간극을 해소하기 위해 세 가지 주요 혁신을 도입했습니다. 첫째, MLLM은 물리적 장면 이해 및 상호 작용 추론을 수행하여 환경 인식 연쇄적 사고 토큰을 생성하고, 물리적으로 타당한 삽입 영역을 추론하여 확산 과정을 명시적으로 안내합니다. 둘째, MLLM 가이드형 공간 직접 선호 최적화(Spatial Direct Preference Optimization, DPO)를 도입하여 확산 모델의 출력을 MLLM에 피드백하여 점수를 매김으로써 시각적 자연스러움을 향상시킵니다. 추론 과정에서 MLLM은 반복적인 개선 단계를 유발하고 확산 모델로부터 적응적 조정을 이끌어내어, 점진적으로 편집 품질을 향상시키는 폐루프 시스템을 형성합니다. 또한, 사용자가 선택할 수 있는 두 가지 모드를 제공합니다. 첫 번째는 물리적 타당성을 높이기 위해 환경 수정(예: 지지 구조 생성)을 허용하는 유연한 모드이며, 두 번째는 장면의 무결성을 유지하여 최대 충실도를 제공하는 표준 모드입니다. 이를 통해 사용자는 물리적 타당성과 충실도 간의 균형을 명시적으로 제어할 수 있습니다. 광범위한 실험 결과는 Place-it-R1이 최첨단 솔루션 및 상용 모델과 비교하여 물리적으로 일관된 비디오 객체 삽입을 달성함을 보여줍니다.

Original Abstract

Modern video editing techniques have achieved high visual fidelity when inserting video objects. However, they focus on optimizing visual fidelity rather than physical causality, leading to edits that are physically inconsistent with their environment. In this work, we present Place-it-R$1$, an end-to-end framework for video object insertion that unlocks the environment-aware reasoning potential of Multimodal Large Language Models (MLLMs). Our framework leverages the Chain-of-Thought (CoT) reasoning of MLLMs to orchestrate video diffusion, following a Think-then-Place paradigm. To bridge cognitive reasoning and generative execution, we introduce three key innovations: First, MLLM performs physical scene understanding and interaction reasoning, generating environment-aware chain-of-thought tokens and inferring valid insertion regions to explicitly guide the diffusion toward physically plausible insertion. Then, we introduce MLLM-guided Spatial Direct Preference Optimization (DPO), where diffusion outputs are fed back to the MLLM for scoring, enabling visual naturalness. During inference, the MLLM iteratively triggers refinement cycles and elicits adaptive adjustments from the diffusion model, forming a closed-loop that progressively enhances editing quality. Furthermore, we provide two user-selectable modes: a plausibility-oriented flexible mode that permits environment modifications (\eg, generating support structures) to enhance physical plausibility, and a fidelity-oriented standard mode that preserves scene integrity for maximum fidelity, offering users explicit control over the plausibility-fidelity trade-off. Extensive experiments demonstrate Place-it-R1 achieves physically-coherent video object insertion compared with state-of-the-art solutions and commercial models.

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!