2606.05677v1 Jun 04, 2026 cs.CV

LongSpace: 비디오를 통해 인지부터 회상까지의 장기 공간 기억 탐구

LongSpace: Exploring Long-Horizon Spatial Memory from Perception to Recall in Video

Honggang Zhang
Honggang Zhang
Citations: 38
h-index: 3
Longteng Guo
Longteng Guo
Citations: 1,611
h-index: 18
Shiqiang Lang
Shiqiang Lang
Citations: 30
h-index: 2
Peiwen Sun
Peiwen Sun
Citations: 94
h-index: 4
Jing Liu
Jing Liu
Citations: 21
h-index: 3
Haoyang He
Haoyang He
Citations: 32
h-index: 2
Yuanteng Chen
Yuanteng Chen
Citations: 39
h-index: 3
Tao Liu
Tao Liu
Citations: 15
h-index: 2
Lan Yang
Lan Yang
Citations: 24
h-index: 2

다중 모달 대규모 언어 모델(MLLM)은 이미지 및 비디오 이해 능력이 향상되었으며, 점차 더 긴 시각적 입력을 처리할 수 있게 되었습니다. 자율 주행 및 로봇 내비게이션과 같은 장기 과제는 현재 장면을 인식하는 것뿐만 아니라, 모델이 이전에 관찰한 공간 구조, 경로, 시점 변화 및 객체 상태를 기억하고 검색해야 하기 때문에 더 높은 수준의 능력을 요구합니다. 이러한 능력을 평가하기 위해, 우리는 장면 인지, 공간 관계 및 공간 기억을 포괄하는 장기 공간 기억 비디오 벤치마크인 LongSpace-Bench를 소개합니다. 본 연구에서는 또한, 장기 비디오 공간 추론을 위한 메모리 프레임워크인 LongSpace를 제안합니다. LongSpace는 긴 비디오를 연속적인 청크로 처리하고, 초기 디코더 레이어에 3차원 구조적 단서를 통합하며, 질문 기반 검색을 위한 레이어별 메모리를 구축합니다. 여러 공간 추론 벤치마크에서의 실험 결과는 LongSpace가 장기 비디오의 공간 이해 능력을 향상시키며, 명시적인 공간 기억이 장기 과제에 대한 비디오 MLLM의 핵심 능력임을 입증한다는 것을 보여줍니다.

Original Abstract

Multimodal Large Language Models (MLLMs) have advanced image and video understanding and can increasingly handle longer visual inputs. Long-horizon tasks such as autonomous driving and robotic navigation require more than recognizing the current view, as models must remember and retrieve previously observed spatial layouts, routes, viewpoint changes, and object states. To evaluate this capability, we introduce LongSpace-Bench, a room-tour video benchmark for long-horizon spatial memory, covering scene perception, spatial relations, and spatial memory. In this work, we further propose LongSpace, a memory framework for long-video spatial reasoning. LongSpace models long videos as sequential chunks, incorporates 3D structural cues into early decoder layers, and constructs layer-aware memory for question-guided retrieval. Experiments on multiple spatial reasoning benchmarks show that LongSpace improves long-video spatial understanding, further demonstrating explicit spatial memory as a key capability for long-horizon video MLLMs.

1 Citations
0 Influential
9 Altmetric
46.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!