2606.12826v1 Jun 11, 2026 cs.CV

DIMOS: 인스턴스 수준의 움직이는 객체 분할을 위한 특징 분리

DIMOS: Disentangling Instance-level Moving Object Segmentation

Zeke Xie
Zeke Xie
Citations: 7
h-index: 2
Hongxiang Huang
Hongxiang Huang
Citations: 36
h-index: 3
Hong Ren
Hong Ren
Citations: 45
h-index: 2
Xiaopeng Lin
Xiaopeng Lin
Citations: 201
h-index: 7
Yulong Huang
Yulong Huang
Citations: 196
h-index: 6
Bo-Xun Cheng
Bo-Xun Cheng
Citations: 185
h-index: 8

움직이는 객체 인스턴스 분할(MIS)은 교통 감시, 자율 주행 및 동물 추적과 같은 다양한 응용 분야에서 그 중요성이 점점 더 커지고 있습니다. 이벤트 카메라는 비동기적인 밝기 변화를 기록하여 높은 시간 해상도와 동적 범위를 제공하며, 이는 움직임 정보에 매우 민감합니다. 이미지 특징과 이벤트를 결합함으로써, 이벤트로부터 얻은 움직임 정보를 활용하여 이미지로부터 얻은 공간 정보를 보완하고 MIS의 성능을 향상시킬 수 있습니다. 그러나 현재의 다중 모드 MIS 방법은 종종 낮은 해상도에서 발생하는 희소한 이벤트 특징으로 인해 작은 움직이는 객체 인스턴스를 분할하는 데 어려움을 겪습니다. 또한, 이벤트 특징은 외관 속성과 움직임 정보를 함께 포함하고 있어 효과적인 교차 모드 융합을 더욱 어렵게 만듭니다. 이러한 문제점을 해결하기 위해, 먼저 이미지 및 이벤트 양쪽의 데이터에서 외관과 움직임 정보를 분리하여 추출하는 이중 분리 특징 추출 프레임워크를 제안합니다. 이를 통해 특징 밀도를 향상시킵니다. 또한, 분포적으로나 의미적으로 일관된 특징을 다양한 수준에서 정렬하는 교차 모드 정렬 방법을 도입하여 풍부한 공간적 및 시간적 세부 정보와 함께 보다 효과적인 융합을 가능하게 합니다. 실험 결과는 제안된 방법이 다중 모드 MIS 분야에서 최첨단 성능을 달성하며, 특히 빠른 움직임이나 저조도 환경과 같은 어려운 조건에서 작은 객체 인스턴스를 분할하는 데 탁월함을 보여줍니다.

Original Abstract

Moving instance segmentation (MIS) attracts increasing attention due to its broad applications in traffic surveillance, autonomous driving, and animal tracking. Event cameras record asynchronous brightness changes, providing high temporal resolution and dynamic range, which makes them highly sensitive to motion information. By fusing event and image features, motion cues from events can complement spatial details from images, enhancing the performance of MIS. However, current multimodal MIS methods still struggle to segment small moving instances, as event cameras often yield sparse features under limited resolution. Moreover, event features entangle appearance attributes with motion cues, which further restricts effective cross-modal fusion. To address these challenges, we first propose a dual-disentangling feature extraction framework that separates and extracts appearance and motion information within both image and event modalities, thereby improving feature density. Subsequently, a multi-granularity cross-modal alignment is introduced to align distributionally and semantically consistent features across modalities, enabling more effective fusion with rich spatial and temporal details. The experiment results demonstrate that our method achieves state-of-the-art performance in multimodal MIS, especially for small instances under challenging conditions such as fast motion and low-light settings.

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!