SUMO: 비선형 상태 공간 모델을 이용한 모든 움직임의 분할 및 추적
SUMO: Segment and Track Any Motion with Nonlinear State Space Models
시각 객체 추적 (VOT) 과 이동 객체 분할 (MOS) 은 컴퓨터 비전 분야에서 시공간적인 객체 동역학을 모두 다루는 두 가지 핵심적인 작업입니다. 기존 방법들은 주로 시각적인 정보에 의존하기 때문에, 실제 환경에서 객체의 움직임이 복잡하고 비선형적인 경우 종종 실패합니다. 이러한 한계를 극복하기 위해, 우리는 정확하고 일관된 VOT 및 MOS를 위한 비전 기반 분할과 비선형 동역학을 통합하는 제로샷, 학습 불필요한 통합 프레임워크인 SUMO를 제안합니다. 구체적으로, 로봇 공학 원리에서 영감을 받은 비선형 상태 공간 모델 (SSM) 을 개발하여 복잡한 객체 동역학을 포착합니다. 이 모델을 기반으로, 정확한 상태 추정을 위한 선택적 언센티드 필터 (SUF)를 제안하며, 이는 공동 점수 부여 메커니즘과 다중 소스 예측의 동적 융합 기능을 통해 시간이 지남에 따라 가장 가능성 높은 객체 상태를 식별합니다. 또한, 메모리 선택 메커니즘을 적용하여 메모리 프레임의 신뢰성을 평가합니다. 광범위한 실험 결과는 SUMO가 VOT 및 MOS 작업 모두에서 최첨단 성능을 달성함을 보여줍니다.
Visual Object Tracking (VOT) and Moving Object Segmentation (MOS) are two fundamental tasks in computer vision that involve both spatial and temporal object dynamics. Existing methods rely predominantly on visual cues and thus often falter in real-world scenarios where object motions are inherently complex and nonlinear. To address this limitation, we propose SUMO, a zero-shot, training-free, unified framework integrating nonlinear dynamics with vision-based segmentation for accurate and consistent VOT and MOS. Specifically, we develop a nonlinear State Space Model (SSM) inspired by robotics principles to capture the complex object dynamics. Building on this model, we propose a Selective Unscented Filter (SUF) for accurate state estimation, which features a joint scoring mechanism and dynamically fuses multi-source predictions to identify the most plausible object state over time. Furthermore, we apply a memory selection mechanism to evaluate the reliability of memory frames. Our extensive experimental results show that SUMO achieves state-of-the-art performance on both VOT and MOS tasks.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.