2606.23256v1 Jun 22, 2026 cs.CV

P-JEPA: 절차 기반 비디오 표현 학습을 위한 공동 임베딩 예측 아키텍처

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture

N. Navab
N. Navab
Citations: 11
h-index: 2
Felix Tristram
Felix Tristram
Citations: 52
h-index: 3
Stefano Gasperini
Stefano Gasperini
Technical University of Munich (TUM)
Citations: 445
h-index: 10
Benjamin D. Killeen
Benjamin D. Killeen
Citations: 0
h-index: 0
Marcel Walch
Marcel Walch
Citations: 1,468
h-index: 20
Christiane Benz
Christiane Benz
Citations: 0
h-index: 0
Ghazal Ghazaei
Ghazal Ghazaei
Citations: 274
h-index: 6

임베디드 AI 플랫폼의 발전으로 인해 복잡하고 다단계 작업을 지원하는 지능형 시스템을 위해 절차 기반 비디오 표현 학습에 대한 관심이 높아지고 있습니다. 대규모 잠재적 예측 훈련을 활용하여, 비디오 모델은 비디오 동역학을 파악하여 활동 이해, 시공간 위치 추정 및 예측 제어와 같은 하위 작업에서 사용됩니다. 그러나 절차 기반 비디오에는 자기 주의 메커니즘의 이중 제곱 복잡성으로 인해 이러한 모델이 지원하지 않는 장거리 종속성이 있는 동작들이 포함될 수 있습니다. 예를 들어, 동일한 시각적 특징을 가질 수 있지만 절차의 다른 지점에서 발생하는 서로 다른 동작 (예: 가스 레인지 켜기 대 끄기)이 있을 수 있습니다. 본 논문에서는 백본 아키텍처에 독립적인 접근 방식을 제안하며, 이 방법은 문제를 밀집된, 프레임 정렬된 동작 공간으로 줄이고 마스크 처리된 잠재 벡터를 예측하여 장시간 비디오 표현을 학습합니다. 이러한 접근 방식 덕분에, 우리의 절차 기반 공동 임베딩 예측 아키텍처 (P-JEPA)는 30분 이상 길이의 비디오를 처리할 수 있으며, 이를 통해 절차 단계에 대한 효과적인 장기 이해가 가능합니다. 우리는 VJEPA2.1, TSM 및 I3D로 추출된 특징을 사용하여 EgoExo4D, EgoProceL 및 Assembly101 데이터 세트에서 P-JEPA를 평가한 결과, 일관되게 선형 분리성, 스트리밍 추론 및 시간 액션 분할 성능이 향상되는 것을 확인했습니다. 특히, P-JEPA는 EgoExo4D의 미세 분류 작업에서 최첨단 결과를 달성했으며, LLM 기반 방법보다 10배 적은 파라미터를 사용하고 실시간으로 실행됩니다.

Original Abstract

The increasing maturity of embodied AI platforms has driven a growing interest in procedural video representation learning to support intelligent assistance systems for complex, multi-step tasks. Leveraging large-scale latent predictive training, video foundation models capture video dynamics, enabling downstream tasks such as activity understanding, spatiotemporal localization, and predictive control. However, procedural videos include actions with long-range dependencies that these models do not support, due to the quadratic complexity of self-attention. Distinct actions, for example, may be visually similar despite appearing at different points in the procedure, such as turning the stove on versus off. Here, we propose a backbone-agnostic approach that learns long-duration video representations by reducing the problem to a dense, frame-aligned action space and predicting pooled masked latent vectors. This approach allows our Procedural Joint Embedding Predictive Architecture (P-JEPA) to ingest videos over 30 minutes long, enabling effective long-form understanding of procedural steps. We evaluate P-JEPA using features extracted with VJEPA2.1, TSM, and I3D over the EgoExo4D, EgoProceL, and Assembly101 datasets, finding that it consistently improves linear separability, streaming inference, and temporal action segmentation performance, achieving state-of-the-art results on EgoExo4D fine-grained action classification while using an order of magnitude fewer parameters than LLM-based methods and running in real time.

0 Citations
0 Influential
10 Altmetric
50.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!