EffectLearner: 실세계 비디오 객체 제거를 위한 환경 인지 객체-영향 추론
EffectLearner: World-Aware Object-Effect Reasoning for Real-World Video Object Removal
비디오 객체 제거는 대상 객체를 제거하는 것뿐만 아니라, 고품질의 시각적 충실도를 유지하면서 공간 및 시간적으로 일관된 방식으로 해당 객체가 유발하는 영향까지 제거해야 합니다. 기존 방법들은 주로 미리 정의된 영향 범주와 고정된 데이터 분포로부터 객체-영향 관계를 암묵적으로 학습하며, 이는 복합적인 영향을 포함하는 실제 환경에서의 일반화 성능을 제한합니다. 본 연구에서는 의미론적 추론을 강화한 프레임워크인 EffectLearner를 제안합니다. EffectLearner는 VLM 기반의 객체-영향 추론 모듈과 DiT 기반의 비디오 제거 모듈을 결합합니다. 구조화된 영향 분석 프롬프트를 통해 추론 모듈은 대상이 강조된 비디오에 대한 교차 모드 추론을 수행하고, 객체와 관련된 영향을 고려한 간결한 컨텍스트 정보를 추출합니다. 이 정보는 비디오 제거 모듈이 포괄적인 객체-영향 제거를 수행하도록 안내합니다. 또한, 움직임 인지 마스크 가이드 및 움직임 일관성 감독을 통해 객체의 움직임과 변화하는 장면 환경에서도 제거 범위와 시간적 안정성을 향상시킵니다. EffectLearner 프레임워크의 실제 환경에서의 성능을 극대화하기 위해, 복잡한 객체 유발 효과를 위한 특수 데이터셋인 EffectWorld를 구축하고, 일반적인 감독 방법과 복합적인 효과 데이터를 결합한 점진적인 학습 방법을 도입했습니다. 표준 벤치마크인 ROSE-Bench에서 EffectLearner는 대부분의 지표에서 기존 모델보다 우수한 성능을 보였으며, EffectWorld-Eval 및 어려운 EffectWorld-Wild 데이터셋에서도 명확한 장점을 보여주어 복잡한 실제 환경에서의 고품질 비디오 객체 제거 능력을 입증했습니다.
Video object removal must eliminate not only the target object but also its induced effects while maintaining high-fidelity and spatiotemporally coherent restoration. Existing methods mainly learn object-effect correspondences implicitly from predefined effect categories and fixed data distributions, limiting their generalization to complex real-world scenes involving compositional effects, spatially detached or weakly correlated effects, long-tail physical phenomena, and dynamically evolving interactions. We propose EffectLearner, a semantic-reasoning-enhanced framework that combines a VLM-based Object-Effect Reasoner with a DiT-based Video Eraser. Guided by a structured effect-analysis prompt, the Reasoner performs cross-modal reasoning over a target-highlighted video and extracts compact effect-aware context, which guides the Video Eraser toward comprehensive object-effect removal. Motion-aware mask guidance and motion-consistency supervision further improve removal coverage and spatiotemporal stability under object motion and evolving scene dynamics. To fully exploit the framework in challenging real-world scenarios, we further construct EffectWorld, a paired video dataset specifically designed for complex object-induced effects, and introduce a progressive training curriculum that combines common supervision with complex-effect data. On the standard ROSE-Bench, EffectLearner outperforms existing baselines on most metrics and achieves clear advantages on both EffectWorld-Eval and the challenging EffectWorld-Wild, demonstrating its ability to deliver high-quality video object removal in complex real-world scenes.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.