2607.21580v1 Jul 23, 2026 cs.CV

GraphVid: 상호 작용 그래프 기반의 대화형 비디오 생성

GraphVid: Interactive Graph-Controllable Video Generation

Ismini Lourentzou
Ismini Lourentzou
Citations: 1,446
h-index: 19
Onkar Susladkar
Onkar Susladkar
Citations: 257
h-index: 8
Tushar Prakash
Tushar Prakash
Citations: 4
h-index: 1
A. Juvekar
A. Juvekar
Citations: 71
h-index: 4
Kiet A. Nguyen
Kiet A. Nguyen
Citations: 54
h-index: 5
Tianjiao Yu
Tianjiao Yu
Citations: 91
h-index: 5
Muntasir Wahed
Muntasir Wahed
Citations: 443
h-index: 6
Vedant r Shah
Vedant r Shah
University of Illinois Urbana-Champaign
Citations: 19
h-index: 2

정교한 다중 객체 상호 작용을 텍스트 프롬프트나 모션 제어 입력으로 명시하는 어려움 때문에, 제어 가능한 비디오 생성은 여전히 어려운 과제입니다. 실제로 경로 기반 제어는 종종 사용자가 여러 객체의 정확한 이동 경로를 그려야 하며, 이는 장면의 복잡도가 증가함에 따라 성능 저하를 일으키고 가려짐이나 겹침 현상에서는 모호해지는 경향이 있습니다. 본 논문에서는 유연하면서도 정교한 다중 주체 제어를 가능하게 하기 위해, 구조화된 상호 작용 그래프를 통해 대화형 제어를 지원하는 이미지-비디오 생성 모델인 $ extbf{GraphVid}$를 소개합니다. 또한, 상호 작용에 대한 정보를 담은 구조적 관계 주석이 포함된 대규모 비디오 데이터셋인 $ extbf{GraphVid-Bench}$를 구축하여 상호 작용을 인식하는 비디오 생성 모델의 학습을 지원합니다. 기존의 모션 제어 방법보다 훨씬 적은 훈련 데이터와 파라미터를 사용하면서도, GraphVid는 뛰어난 제어 가능성과 비디오 품질을 제공합니다. Motion-I2V와 비교했을 때, GraphVid는 FID를 최대 39.9% 감소시키고 FVD를 37.6% 감소시키는 동시에 PSNR (9.87=>15.98) 및 SSIM (0.38=>0.61)을 향상시켰습니다. 본 연구 결과는 구조화된 의미 인터페이스가 제어 가능한 비디오 생성에 있어 강력한 패러다임이 될 수 있음을 보여줍니다.

Original Abstract

Controllable video generation remains challenging due to the difficulty of specifying precise multi-object interactions using text prompts or motion-control inputs that primarily constrain pixel movement. In practice, trajectory-based control often requires users to draw accurate tracks for multiple objects, which scales poorly with scene complexity and becomes ambiguous under occlusion or overlap. To enable flexible yet precise multi-subject control, we introduce $\textbf{GraphVid}$, a graph-conditioned image-to-video generation model that enables interactive control through structured interaction graphs. We further curate $\textbf{GraphVid-Bench}$, a large-scale interaction-centric video dataset with structured relational annotations to enable training of interaction-aware video generation models. Despite using substantially less training data and fewer trainable parameters than prior motion-control methods, GraphVid delivers strong controllability and video quality. Compared with Motion-I2V, GraphVid reduces FID by up to 39.9% and FVD by 37.6%, while improving PSNR (9.87=>15.98) and SSIM (0.38=>0.61). Our results highlight the potential of structured semantic interfaces as a powerful paradigm for controllable video generation.

0 Citations
0 Influential
9.5 Altmetric
47.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!