2608.05631v1 Aug 06, 2026 cs.CV

ChronoVision: 잠재 상태 재구성을 통한 시간 추론

ChronoVision: Temporal Reasoning via Latent State Reconstruction

Bingxuan Li
Bingxuan Li
Citations: 86
h-index: 3
Yifan Shen
Yifan Shen
Citations: 88
h-index: 5
Xu Cao
Xu Cao
Citations: 24
h-index: 3
Rushi Wang
Rushi Wang
Citations: 15
h-index: 3
Tianjiao Yu
Tianjiao Yu
Citations: 91
h-index: 5
Jian Xu
Jian Xu
Citations: 73
h-index: 2
Boyi Li
Boyi Li
Citations: 348
h-index: 5
Yuner Zhang
Yuner Zhang
Citations: 13
h-index: 1
Houze Yang
Houze Yang
Citations: 4
h-index: 1

다중 모드 대규모 언어 모델은 수동적인 인지 능력에서는 뛰어난 성능을 보이지만, 다단계 시간적 추론이 필요한 복잡한 시각적 인지 작업에서는 어려움을 겪습니다. 이러한 성능 저하는 주로 언어 기반 추론의 고유한 모호성에서 비롯되며, 이는 종종 연속적인 시각적 변화를 정확하게 표현하지 못합니다. 이를 해결하기 위해, 우리는 시각적 논리를 잠재 이미지와 일치시키는 다중 모드 프레임워크인 ChronoVision을 제안합니다. 지도 학습 과정에서, 재구성 시각 헤드는 최종 변환된 상태의 잠재 표현을 예측하고, 관심 영역(ROI) 주의 집중 모듈은 의미 기반 범위를 활용하여 모델이 중요한 시각적 증거에 집중하도록 합니다. 훈련 후에는, 우리는 암묵적인 프로세스 연결 메커니즘과 함께 강화 학습을 적용하며, 결과 정확성, 잠재 프로세스 일치도 및 비지도 시각적 집중도를 평가하는 복합 보상 함수를 통해 모델을 안내합니다. 또한, 우리는 시간 추적 능력을 엄격한 이미지 순서화 작업으로 재구성하여 평가하는 새로운 데이터셋인 Vbvr-VQA를 소개합니다. 실험 결과는 ChronoVision이 Vbvr-VQA에서 74.8%의 인내역 정확도와 71.6%의 외부 영역 정확도를 달성했으며, 매우 어려운 교차 도메인 벤치마크인 IntPhys2에서도 55.0%의 높은 정확도를 보임을 보여줍니다.

Original Abstract

Multimodal large language models excel at passive perception but struggle with complex visual cognitive tasks requiring multi-step temporal reasoning. This degradation largely stems from the inherent ambiguity of language-based reasoning, which often fails to accurately articulate continuous visual transformations. To address this, we propose ChronoVision, a multimodal framework designed to align visual logic with latent imagery. During supervised fine-tuning, a Reconstructive Visual Head predicts the latent representation of the final transformed state, while an ROI Attention Locating module focuses the model on key visual evidence via semantic span queries. In post-training, we apply reinforcement learning with an implicit process grounding mechanism, guided by a composite reward function that evaluates outcome correctness, latent process alignment, and unsupervised visual focus. Furthermore, we introduce Vbvr-VQA, a novel dataset that evaluates temporal tracking by reformulating video reasoning into a strict image-ordering task. Experiments demonstrate that ChronoVision achieves state-of-the-art performance on Vbvr-VQA with 74.8% in-domain and 71.6% out-of-domain accuracy, alongside a strong 55.0% accuracy on IntPhys2, a highly challenging cross-domain benchmark.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!