RSVideo: 귀하의 비전-언어 모델이 원격 감지 영상에 적합한가?
RSVideo: Are Your Vision-Language Models Ready for Remote Sensing Videos?
원격 감지 영상은 대상 속성의 실시간 변화, 단기 활동 및 장면 진화를 관찰할 수 있도록 합니다. 이들은 이미지로는 캡처할 수 없는 움직임, 행동, 상호 작용 및 장면 변화를 기록합니다. 기존 모델은 주로 개별 이미지 또는 광범위한 시간 범위를 포괄하는 이산적인 시퀀스 데이터에 초점을 맞춥니다. 그러나 원격 감지 영상 이해에 대한 비전-언어 모델을 평가하기 위한 통일된 평가 환경은 아직 부족합니다. 우리는 10,773개의 인스턴스, 147만 프레임 및 17.02시간의 영상을 포함하는 원격 감지 영상 데이터 세트인 RSVideo-10K를 소개합니다. 이 데이터 세트는 무인 항공기 및 위성 플랫폼 데이터를 모두 포함하고 있습니다. 고정된 평가 벤치마크인 RSVideo-Bench는 2,731개의 테스트 인스턴스를 포함하며, 원격 감지 영상 이해의 두 가지 상호 보완적인 측면인 L1 인식 및 L2 추론을 평가합니다. 이 평가는 7가지 능력 그룹과 17가지 작업을 포괄합니다. 결과는 현재 비전-언어 모델이 여전히 작은 지역적 증거를 복구하고, 짧은 기간의 상태를 추적하고, 장면 제약 조건 내에서의 공간 관계를 사용하는 데 어려움을 겪고 있음을 보여줍니다. 이러한 분석을 바탕으로, 우리는 작은 대상의 시공간 초점에 대한 강화 학습 프레임워크인 RSVideo를 제안합니다. RSVideo는 질문과 관련된 영역을 선택하고 불필요한 배경 토큰을 억제하여 여러 프레임을 가로질러 관련 영역에 집중합니다. RSVideo는 InternVL3.5-14B 모델에서 최대 9.01%의 성능 향상을 달성했으며, Qwen3.6-27B 모델에서는 26개의 오픈 소스 비전-언어 모델 아키텍처 중에서 가장 높은 정확도인 40.63%를 달성했습니다. 코드 및 데이터는 https://github.com/HongjieZhou0329/RSVideo 에서 확인할 수 있습니다.
Remote-sensing videos enable real-time observation of changes in target attributes, short-term activities, and scene evolution. They record motion, actions, interactions, and scene changes that cannot be captured by isolated images. Existing models primarily target single images or discrete temporal observations spanning a long time range. However, a unified evaluation setting for assessing vision-language models on continuous remote-sensing video understanding remains lacking. We introduce RSVideo-10K, a remote-sensing video dataset comprising 10,773 instances, 1.47 million frames, and 17.02 hours of footage, containing both unmanned aerial vehicles and satellite platforms. Its fixed evaluation benchmark, RSVideo-Bench, contains 2,731 test instances and evaluates two complementary aspects of remote-sensing video understanding: L1 Perception and L2 Reasoning, spanning seven capability groups and 17 tasks. Evaluations show that current vision-language models still struggle to recover small local evidence, track short-lived states, and use scene-constrained spatial relations. Based on this analysis, we further propose RSVideo, a reinforcement learning framework for small-target spatiotemporal focusing that selects question-relevant regions across frames and suppresses redundant background tokens. RSVideo achieves a maximum absolute improvement of 9.01% with InternVL3.5-14B and attains the highest accuracy of 40.63% with Qwen3.6-27B across 26 open-source vision-language backbones.Codes will be available at https://github.com/HongjieZhou0329/RSVideo.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.