ObsDriveBench: 관찰 가능성 인지 기반의 악천후 환경에서의 다중 모드 이해 성능 평가
ObsDriveBench: Benchmarking Multimodal Understanding under Adverse Weather with Observability Awareness
악천후 조건에서 자율 주행은 여전히 중요한 과제이지만, 기존의 시각-언어 벤치마크는 주로 표준 조건, 합성 데이터 또는 단일 모드 입력에 대한 평가를 수행합니다. 따라서 실제 악천후 환경에서의 다중 모드 입력을 사용하는 시각-언어 모델의 성능이 어떻게 나타나는지에 대한 명확한 이해가 부족합니다. 우리는 이러한 어려움이 '환경 관찰 가능성'의 저하에서 비롯된다고 주장합니다. 안개, 비, 눈, 그리고 낮은 조명 조건에서는 다중 모드 센서 데이터가 신뢰성이 떨어지고 서로 일관되지 않아 장면 이해 및 후속 의사 결정에 어려움을 초래합니다. 이러한 문제를 해결하기 위해, 실제 악천후 자율 주행 환경을 위한 다중 모드 벤치마크인 **ObsDriveBench**를 제안합니다. 저희의 벤치마크는 '관찰 가능성 인지', '공간적 신뢰성', 그리고 '위험 인식 의사 결정'이라는 세 가지 능력 차원을 중심으로 설계되었으며, 이를 통해 모델이 열악한 환경에서 어떻게 동작하는지에 대한 상세한 분석을 가능하게 합니다. 저희는 카메라, LiDAR, 레이더 센서로부터 수집된 동기화된 데이터를 사용하여 관찰 가능성 메타-주석, 장면 설명 및 능력을 평가하기 위한 객관식 문제를 구성하여 총 14,000개의 학습 데이터와 13,000개의 테스트 데이터를 포함하는 벤치마크를 구축했습니다. 실험 결과 기존의 시각-언어 모델은 성능 저하가 발생하는 것을 확인했습니다. 또한, 정상 기상 환경에서의 지도 학습과 악천후 환경에서의 강화 학습을 결합한 **ObsDrive** 모델을 제안하여 세 가지 능력 측면에서 모델의 견고성을 향상시켰습니다. 데이터셋 및 평가 코드는 다음 GitHub 링크(https://github.com/russellyq/ObsDriveBench)에서 제공됩니다.
Autonomous driving under adverse weather remains a critical challenge, yet existing vision-language benchmarks mainly evaluate under standard conditions, synthetic corruptions, or single modality. As a result, it remains unclear how vision-language models behave under real-world adverse weather with multi-modal inputs. We argue that a key difficulty lies in degraded environmental observability: under fog, rain, snow, and low illumination, multi-modal observations become unreliable and cross-modally inconsistent, posing challenges to scene understanding, and subsequent decision-making. To study this, we introduce \textbf{ObsDriveBench}, a real-world multi-modal benchmark for adverse-weather autonomous driving. Our benchmark is designed with three capability dimensions: \textbf{observability awareness}, \textbf{spatial reliability}, and \textbf{risk-aware decision-making}, enabling fine-grained diagnosis of model behavior under degraded observations. We construct the benchmark through observability meta-annotation, scene description, and capability oriented multiple-choice tasks over synchronized camera, LiDAR, and radar inputs, forming a benchmark with over 14k training and 13k test questions. Experiments reveal consistent performance degradation of existing vision-language models. We further introduce \textbf{ObsDrive} model with normal-weather supervised fine-tuning and adverse-weather reinforcement learning, improving robustness across all three capabilities. The dataset and evaluation code will be released at \href{https://github.com/russellyq/ObsDriveBench}{\texttt{ObsDriveBench}}.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.