2607.29401v1 Jul 31, 2026 cs.CV

OSEF: 단일 단계 증거 융합을 통한 다중 비디오 장면 절차 계획

OSEF: One-Step Evidence Fusion for Cross-Video Scene Procedure Planning

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Bin Li
Bin Li
Citations: 8
h-index: 2
Lei Zhang
Lei Zhang
Citations: 0
h-index: 0
Zhentong Ye
Zhentong Ye
Citations: 0
h-index: 0
Sijia Zhou
Sijia Zhou
Citations: 0
h-index: 0
Jiaqi Xuan
Jiaqi Xuan
Citations: 232
h-index: 6
Shuaiwu Dong
Shuaiwu Dong
Citations: 0
h-index: 0
G. Tong
G. Tong
Citations: 0
h-index: 0

비디오 장면 절차 계획(VSPP)은 목표 시작-종료 관찰 정보를 미리 제공하지만, 증거 자체를 검색해야 할 때 플래너가 어떻게 작동해야 하는지에 대한 문제는 여전히 남아 있습니다. 본 연구에서는 다중 비디오 장면 절차 계획(CVSPP)을 소개합니다. CVSPP는 질문에 대한 답변이 가려진 시작-종료 쿼리와 K개의 후보 비디오를 입력으로 받아, 모델이 관련 비디오를 검색하고, 관련된 영역을 찾고, 행동 시퀀스를 예측해야 합니다. 여기에는 두 가지 어려움이 있습니다. 동일한 작업의 데모는 단계와 영역을 공유하며, 초기 단계에서 잘못된 장면 체인을 플래너에 전달하면 성능 저하가 발생합니다. 본 연구에서는 유형화된 부정적인 역할을 포함하는 11개의 소스 데이터셋 기반 벤치마크를 구축하고, 오류 방지 기능과 답변 유출 방지 장치를 추가했으며, 증거 및 계획 축에 대한 별도의 평가 지표를 사용했습니다. 14개의 소스-호라이즌 셀에서 9가지 플래너 패밀리를 사용하여 대부분의 시퀀스 수준을 기준으로 성능을 평가했습니다. 그런 다음, 본 연구에서는 One-Step Evidence Fusion (OSEF)을 제안합니다. OSEF는 모든 후보에 대한 질문에 조건화된 셀 및 스팬 격자를 생성하고, 미리 특정 영역을 잘라내지 않고 전체 소프트 격자를 토큰-글로벌 어댑터를 통해 플래너에 전달합니다. OSEF는 벤치마크에서 순위 결정이 가능한 것으로 평가되는 6개의 모든 셀에서 가장 높은 성능을 보였습니다. 또한, 동일한 작업의 COIN 및 CrossTask 셀에서 기존 최고 성능 모델보다 정확한 비디오 및 계획 성공률을 각각 2.9~10.7 포인트 향상시켰으며, 토큰-글로벌 인터페이스가 가장 큰 성능 향상을 가져왔습니다. 5개의 변환된 소스 셀은 벤치마크의 기준선에 도달하거나 거의 도달했으며, 이는 남은 잠재력을 나타냅니다. 추가 자료로는 모델 생성 코드와 평가 코드를 제공합니다.

Original Abstract

Video Scene Procedure Planning (VSPP) supplies the target start-goal observations in advance, leaving open how a planner should act when the evidence must itself be retrieved. We introduce Cross-Video Scene Procedure Planning (CVSPP): given an answer-redacted start-goal query and K candidate videos, a model must retrieve the supporting video, localize the relevant window, and predict the action sequence. Two obstacles couple here. Same-task demonstrations share stages and windows, and an early hard selection passes the wrong scene chain to the planner. We build an eleven-source benchmark with typed negative roles, a fail-closed answer-leakage gate, and separate Evidence- and Plan-axis metrics. On its 14 source-horizon cells we adapt nine planner families against a majority-sequence floor. We then present One-Step Evidence Fusion (OSEF), which scores a query-conditioned cell-and-span lattice over all candidates and feeds the full soft lattice to the planner through a token-global adapter, cropping no window beforehand. OSEF ranks first on all six cells the benchmark certifies as method-rankable. On four matched same-task COIN and CrossTask cells it improves exact-video-and-plan success by 2.9-10.7 points over an enhanced hard-selection SOTA, and a component study assigns the largest single increment to the token-global interface. Five converted-source cells sit at or near the majority-sequence floor, the benchmark's remaining headroom. The supplementary package includes model constructors and evaluation code.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!