ForecastOcc: 비전 기반 의미 있는 공간 점유 예측
ForecastOcc: Vision-based Semantic Occupancy Forecasting
자율 주행은 미래 환경 상태에 대해 효과적으로 추론하기 위해 시간의 흐름에 따른 공간 정보와 의미 정보를 모두 예측해야 합니다. 기존의 비전 기반 공간 점유 예측 방법은 주로 정적 및 동적 객체와 같은 움직임과 관련된 범주에 초점을 맞추고 있으며, 의미 정보는 대부분 고려되지 않습니다. 최근의 의미 기반 공간 점유 예측 방법은 이러한 격차를 해소하려 시도하지만, 별도의 네트워크에서 얻은 과거 공간 점유 예측 결과에 의존합니다. 이는 현재 방법이 오차 누적에 민감하게 반응하게 하고, 이미지로부터 직접적인 시공간 특징을 학습하는 것을 방해합니다. 본 연구에서는 비전 기반 의미 있는 공간 점유 예측을 위한 최초의 프레임워크인 ForecastOcc를 제시합니다. ForecastOcc는 과거 카메라 이미지로부터 직접적으로 미래의 공간 점유 상태와 의미 범주를 동시에 예측하며, 외부적으로 추정한 지도를 사용하지 않습니다. 우리는 ForecastOcc를 두 가지 상호 보완적인 환경에서 평가했습니다. 첫째는 Occ3D-nuScenes 데이터셋에서의 다중 뷰 예측, 둘째는 SemanticKITTI 데이터셋에서의 단안 예측이며, 이 분야의 첫 번째 벤치마크를 설정했습니다. 우리는 프레임워크 내에서 두 가지 2차원 예측 모듈을 활용하여 첫 번째 기본 모델을 제시했습니다. 더욱 중요하게는, 우리는 시간적 크로스 어텐션 예측 모듈, 2D-to-3D 뷰 트랜스포머, 공간 점유 예측을 위한 3D 인코더, 그리고 여러 시점에서 3차원 격자 수준의 예측을 수행하는 의미 기반 공간 점유 예측 헤드를 통합한 새로운 아키텍처를 제안합니다. 두 데이터셋에 대한 광범위한 실험 결과, ForecastOcc는 기본 모델보다 일관되게 뛰어난 성능을 보이며, 자율 주행에 중요한 장면의 역학 및 의미를 포착하는 의미적으로 풍부하고 미래에 대한 인식을 갖춘 예측 결과를 제공합니다.
Autonomous driving requires forecasting both geometry and semantics over time to effectively reason about future environment states. Existing vision-based occupancy forecasting methods focus on motion-related categories such as static and dynamic objects, while semantic information remains largely absent. Recent semantic occupancy forecasting approaches address this gap but rely on past occupancy predictions obtained from separate networks. This makes current methods sensitive to error accumulation and prevents learning spatio-temporal features directly from images. In this work, we present ForecastOcc, the first framework for vision-based semantic occupancy forecasting that jointly predicts future occupancy states and semantic categories. Our framework yields semantic occupancy forecasts for multiple horizons directly from past camera images, without relying on externally estimated maps. We evaluate ForecastOcc in two complementary settings: multi-view forecasting on the Occ3D-nuScenes dataset and monocular forecasting on SemanticKITTI, where we establish the first benchmark for this task. We introduce the first baselines by adapting two 2D forecasting modules within our framework. Importantly, we propose a novel architecture that incorporates a temporal cross-attention forecasting module, a 2D-to-3D view transformer, a 3D encoder for occupancy prediction, and a semantic occupancy head for voxel-level forecasts across multiple horizons. Extensive experiments on both datasets show that ForecastOcc consistently outperforms baselines, yielding semantically rich, future-aware predictions that capture scene dynamics and semantics critical for autonomous driving.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.