2608.04426v1 Aug 05, 2026 cs.CV

예측 후 검색: 비디오 프레임 기반의 인스턴스 간 미래 상태 검색

Predict, Then Retrieve: Cross-Instance Future-State Retrieval from Video Prefixes

A. Luu
A. Luu
Citations: 5,973
h-index: 34
Cong-Duy Nguyen
Cong-Duy Nguyen
Citations: 300
h-index: 10
Quynh T. N. Vo
Quynh T. N. Vo
Citations: 9
h-index: 1
Thong Nguyen
Thong Nguyen
Citations: 432
h-index: 11
Vinh-Hien Do
Vinh-Hien Do
Citations: 0
h-index: 0

본 논문에서는 Predictive State Retrieval (PSR)이라는 새로운 태스크를 소개합니다. PSR은 모델이 짧은 비디오 프레임과 객체의 미래 상태에 대한 시간적 질문을 입력받아, 다른 비디오 또는 이미지에서 해당 상태를 나타내는 인스턴스를 검색하는 방식으로 작동합니다. 기존의 행동 예측(라벨 예측), 특정 시점 검색(비디오 내 이벤트 위치 추정), 비디오 생성(픽셀 합성)과는 달리, PSR은 예측과 인스턴스 간 검색을 결합하여 여러 시간적 지평선을 고려합니다. 우리는 네 개의 데이터 세트를 기반으로 인간이 검증한 정답 데이터를 포함하는 벤치마크를 구축했으며, 난이도 수준 및 성능 상한선(oracle ceiling)을 정의했습니다. 또한 LFTR이라는 경량 검색 모델을 제안하며, 이 모델은 동결된 인코더를 사용하여 질문과 시간적 지평선에 조건화된 미래 잠재 변수를 예측하고, 이를 의미론적 및 시각적 공간에서 매칭합니다. 성능 분석 결과, 실제 미래 상태는 명시되면 쉽게 검색될 수 있지만, 프레임 접근 권한을 가진 대규모 멀티모달 언어 모델을 포함한 모든 평가 모델이 성능 상한선에 훨씬 미치지 못하는 것으로 나타났습니다. 따라서 핵심 학습 과제는 인식(perception)보다는 예측(forecasting)이라는 것을 알 수 있습니다. LFTR은 상당한 성능 향상을 보였으며, 추론 비용도 낮습니다. 추가 분석 결과, 성능 향상은 잠재 변수 롤아웃(latent rollout)보다 교차 공간 융합(cross-space fusion) 및 어려운 부정 샘플링 학습(hard-negative training)에 기인합니다. 우리는 본 논문에서 구축한 벤치마크 데이터 세트, 코드, 그리고 평가 스크립트를 공개합니다.

Original Abstract

We introduce Predictive State Retrieval (PSR), a task in which a model observes a short video prefix and a temporal question about an object's future state, then retrieves instances from other videos or images that depict that state. Unlike action anticipation, which predicts a label, moment retrieval, which localizes an observed event within a video, or video generation, which synthesizes pixels, PSR combines anticipation with cross-instance retrieval across multiple temporal horizons. We construct a benchmark from four datasets with graded, human-validated ground truth, difficulty tiers, and an oracle ceiling. We also propose LFTR, a lightweight retriever with frozen encoders that predicts a question- and horizon-conditioned future latent and matches it in complementary semantic and visual spaces. A ceiling decomposition reveals a clear bottleneck: the true future state is highly retrievable once specified, whereas every predictor we evaluate, including a large multimodal language model with access to the prefix frames, remains far below the oracle. Thus, forecasting rather than perception is the central learnable challenge. LFTR narrows this gap at substantially lower inference cost, and ablations attribute its gains to cross-space fusion and hard-negative training rather than latent rollout. We release the benchmark, code, and evaluation scripts.

0 Citations
0 Influential
17 Altmetric
85.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!