DynaPix: 시각-언어 모델이 정확한 미래를 식별할 수 있을까?
DynaPix: Can Vision-Language Models Identify the Exact Future?
물리적 환경에서 행동하려면, 가능성 있는 상태가 아닌 실제 후속 상태를 알아야 합니다. 현재의 평가 방법은 종종 단어나 현실적으로 보이는 이미지를 허용하지만, 예측된 상태가 실제 상태와 비교 검증되지 않는 경우가 많습니다. 우리는 예측 결과를 검증할 수 있도록 설계된 벤치마크인 DynaPix (Dynamic Pixels)를 소개합니다. 이 벤치마크에서는 특정 이벤트 직전에 종료되는 비디오 클립과 이후 시점에 대한 질문이 주어지면, 모델은 후보 이미지 목록 또는 대규모 이미지 데이터베이스에서 실제 미래 이미지를 선택해야 합니다. 장면은 물리 시뮬레이터에서 생성되므로 정확한 이미지와 시간 정보가 알려져 있으며, 오답은 의도적으로 유사하게 설계되었습니다. 모델은 종종 눈에 띄는 이벤트가 목표 시점에 있을 때 성공하지만, 경과된 시간이 목표 시점을 나타낼 때는 거의 무작위 수준의 성능을 보입니다. 대규모 이미지 데이터베이스 검색은 더욱 어렵습니다. 왜냐하면 실제 이미지가 거의 첫 번째 순위에 랭크되지 않기 때문입니다. 사람들은 경과 시간만으로 판단해야 하는 항목을 잘 처리하지만, 모델은 이러한 질문에 어려움을 겪는 것으로 보입니다. 시뮬레이터의 실제 기록에서 추출한 장면 데이터를 사용하여 학습하면 이 문제를 상당 부분 해결할 수 있지만, 특히 긴 경과 시간을 포함하는 경우에는 여전히 문제가 남아 있습니다. 따라서 DynaPix는 시간 정보와 예측 간의 연관성 부족 (temporal-anchoring gap)을 보여줍니다. 즉, 모델은 특정 이벤트에 대한 예측은 잘 하지만, 단순히 시간에 기반한 예측은 어려워합니다.
Acting in a physical scene requires knowing its real later state, not a plausible one. Current evaluations often accept words or a realistic-looking image, so the predicted state is never checked against the true one. We introduce DynaPix (Dynamic Pixels), a benchmark that makes prediction checkable. Given a video clip that stops before a key event and a question about a later moment, a model must pick the true future image from close candidates or a large gallery. The scenes come from a physics simulator, so the correct image and its time are known exactly and the wrong options are deliberately similar. Models often succeed when a visible event marks the target moment, but are near chance when only elapsed time marks it. Gallery search is harder still, as the true image rarely ranks first. People handle the elapsed-time items well, so the difficulty lies with the models, not the questions. Training on scene accounts drawn from the simulator's true record, not a teacher's guess, repairs much of this but not the longer elapsed time case. DynaPix thus exposes a temporal-anchoring gap: models attach a prediction to an event far better than to time itself.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.