2607.24560v1 Jul 27, 2026 cs.CV

EgoPlay: 이벤트 기반 비디오 편집을 위한 개인 시점 스트림 시스템

EgoPlay: Event-Triggered Video Editing for Egocentric Streams

Peter Wonka
Peter Wonka
Citations: 788
h-index: 14
Rameen Abdal
Rameen Abdal
Citations: 1
h-index: 1
Sergey Tulyakov
Sergey Tulyakov
Citations: 1,064
h-index: 17
Runjia Li
Runjia Li
University of Oxford
Citations: 308
h-index: 8
Chaoyang Wang
Chaoyang Wang
Citations: 2,189
h-index: 20
Jinjie Mai
Jinjie Mai
Citations: 2,080
h-index: 9
G. Qian
G. Qian
Citations: 1
h-index: 1
W. Menapace
W. Menapace
Citations: 2,265
h-index: 22
Arpit Sahni
Arpit Sahni
Citations: 59
h-index: 4
Ashkan Mirzaei
Ashkan Mirzaei
Citations: 922
h-index: 13
Bernard Ghanem
Bernard Ghanem
Citations: 60
h-index: 3

본 논문에서는 EgoPlay를 소개합니다. EgoPlay는 개인 시점(egocentric) 스트림 데이터를 대상으로 하는, 이벤트 트리거 기반의 비디오-투-비디오 편집 시스템입니다. 이는 사전 훈련된 V2V diffusion transformer 모델을 미세 조정하여 개발되었으며, 주로 Ego4D 데이터셋에서 구축된 이벤트 조건 데이터를 활용합니다. EgoPlay는 단안(monocular) 비디오와 "X가 발생하면 Y를 수행한다"라는 형태의 이벤트 트리거 프롬프트를 입력받아, 이벤트 X가 언제 발생하는지 추론하고, 이벤트 발생 전 프레임을 보존하며, 이벤트 발생 후 부분에만 편집 Y를 적용합니다. 기존 방식과는 달리, 별도의 이벤트 검출기와 편집기를 연결하는 대신, EgoPlay는 이벤트 인식, 시간 제약 및 픽셀 단위 편집을 하나의 통합 모델에서 공동으로 학습하며, 음수(negative) 및 다중 이벤트 프롬프트도 처리할 수 있습니다. 이를 위해, 양성 트리거, 가짜 트리거 부정 예제 및 다중 이벤트 프롬프트를 포함하는 106,000개의 이벤트 트리거 클립-프롬프트 쌍으로 구성된 대규모 데이터셋을 구축했습니다. 또한, 이벤트 트리거 기반의 지도 학습을 통해 양방향 비디오 diffusion 편집기를 훈련하고, 청크 단위로 스트리밍 가능한 추론을 위한 인과 관계 모델을 개발했습니다. 더 나아가, 이벤트 인식 능력을 평가하기 위해, 트리거 이후 편집 품질, 트리거 이전 프레임 보존 및 오탐(false-trigger)에 대한 강건성 측정을 분리하는 새로운 평가 프로토콜을 도입했습니다. Ego4D 벤치마크에서 EgoPlay는 최첨단인 지침 기반의 개인 시점 비디오 편집 시스템인 EgoEdit를 크게 능가했으며, 편집 품질, 시각적 품질 및 배경 일관성 측면에서 각각 17.7%, 16.9% 및 16.4%의 상대적인 성능 향상을 보였습니다. 또한, VLM(Vision-Language Model) 기반 검출기-편집 시스템보다 동일한 지표에서 15.7%, 14.5% 및 13.5% 더 높은 성능을 보이면서도 GPU 메모리 사용량은 절반 이하입니다.

Original Abstract

We introduce EgoPlay, an event-triggered video-to-video editor for egocentric streams, obtained by fine-tuning a pretrained V2V diffusion transformer on event-conditioned data built primarily from Ego4D. Given a monocular video and an event-triggered prompt of the form "when X happens, do Y," EgoPlay infers whether and when event X occurs, preserves pre-event frames, and applies edit Y only to the post-event continuation. Rather than cascading a separate event detector with an editor, EgoPlay learns event recognition, temporal restraint, and pixel-level editing jointly in a single end-to-end model, while also handling negative and multi-event prompts. To support this, we construct a large-scale dataset of 106K event-triggered clip-prompt pairs spanning positive triggers, fabricated-trigger negatives, and multi-event prompts. We then train a bidirectional video diffusion editor with event-triggered supervision and derive a causal variant for chunk-by-chunk streamable inference. We further introduce an event-aware evaluation protocol that separately measures post-trigger editing quality, pre-trigger preservation, and false-trigger robustness. On the Ego4D benchmark, EgoPlay substantially outperforms EgoEdit, the state-of-the-art instruction-based egocentric video editing baseline, with relative gains of 17.7%, 16.9%, and 16.4% in editing quality, visual quality, and background consistency. It also surpasses a VLM-guided detector-editor baseline by 15.7%, 14.5%, and 13.5% on the same metrics, while using less than half the GPU memory.

0 Citations
0 Influential
11 Altmetric
55.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!