RDVSv2: RGB-D 비디오에서 눈에 띄는 객체 검출을 위한 대규모 벤치마크
RDVSv2: A Large-scale Benchmark for RGB-D Video Salient Object Detection
본 논문에서는 프레임 단위의 상세한 어노테이션이 포함된 RGB-D 비디오에서 눈에 띄는 객체를 검출하기 위한 대규모 벤치마크인 RDVSv2를 소개합니다. 이 분야의 기존 데이터셋은 규모나 어노테이션 품질 면에서 제한적이며, 또한 기하학적으로 일관성 없는 깊이 정보를 사용하는 경향이 있습니다. 이러한 한계를 극복하기 위해, RDVSv2는 공개된 스테레오 온라인 비디오로부터 구축되었으며, 249개의 비디오 시퀀스와 29,077개의 어노테이션된 프레임을 포함합니다. 이 데이터셋은 스테레오 비디오에서 추출한 깊이 맵과 함께, 아이 트래킹 가이드에 의해 어노테이션된 프레임 단위의 눈에 띄는 객체 마스크를 제공합니다. 기존 데이터셋과 비교했을 때, RDVSv2는 규모가 훨씬 크며 더 다양하고 어려운 시나리오를 포함합니다. 또한, Segment Anything Model 2 (SAM2)를 기반으로 RGB-D VSOD 분야의 강력한 초기 성능을 제시합니다. 특히, 파라미터 효율적인 미세 조정(PEFT) 전략을 사용하여 SAM2 인코더를 RGB, 깊이 및 광학 흐름 정보를 함께 인코딩하도록 적용했습니다. 다양한 실험 결과, RDVSv2는 기존의 RGB-D VSOD 방법론에 대해 훨씬 더 높은 수준의 어려움을 제시함을 보여줍니다. 동시에, 제안된 초기 성능 모델은 RDVSv2와 기존의 RGB-D VSOD 벤치마크에서 최고 수준의 결과를 달성했습니다. 본 연구에서는 RDVSv2와 제공되는 초기 성능 모델이 향후 RGB-D VSOD 및 관련 다중 모달 비디오 이해 연구에 유용한 자료로 활용될 수 있기를 바랍니다. 데이터셋과 코드는 https://github.com/ltynick/RDVSv2 에서 확인할 수 있습니다.
We introduce RDVSv2, a large-scale benchmark for RGB-D video salient object detection (RGB-D VSOD) with dense frame-level annotations. Existing datasets in this emerging field are often limited in scale and annotation quality, while also relying on less geometry-consistent depth cues. To address these limitations, RDVSv2 is built from publicly accessible stereoscopic online videos and contains 249 video sequences with 29,077 annotated frames. It includes depth maps derived from stereoscopic videos, together with frame-wise salient object masks annotated with eye-tracking guidance. Compared with existing datasets, RDVSv2 is much larger in scale and covers more diverse and challenging scenarios. In addition, we establish a strong baseline for RGB-D VSOD based on Segment Anything Model 2 (SAM2). Specifically, we employ a parameter-efficient fine-tuning (PEFT) strategy to adapt the SAM2 encoder to jointly encode RGB, depth, and optical flow cues. Extensive experiments show that RDVSv2 is substantially more challenging for existing RGB-D VSOD methods. Meanwhile, the proposed baseline achieves state-of-the-art results on RDVSv2 and existing RGB-D VSOD benchmarks. We hope that RDVSv2 and the provided baseline will serve as useful resources for future research on RGB-D VSOD and related multi-modal video understanding tasks. Our dataset and code will be available at https://github.com/ltynick/RDVSv2.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.