숨겨진 정보의 중요성: 비전-언어 모델을 활용하여 계획에 중요한 영향을 미치는 가려진 객체 식별
What's Hidden Matters: Identifying Planning-Critical Occluded Agents using Vision-Language Models
자율 주행 차량은 계획 수립에 중요한 역할을 하는 객체가 시야에서 가려질 수 있는 복잡한 환경에서 안전하게 운행해야 합니다. 현재의 접근 방식은 종종 모든 가려짐 현상을 동일한 수준의 보수적으로 처리하여 불필요하게 방어적인 운전을 유발하거나, 숨겨진 공간을 추론하지만 계획에 미치는 영향을 정확히 평가하지 못합니다. 본 연구는 비전-언어 모델(VLM)이 자율 주행 차량의 경로에 가장 중요한 영향을 미치는 특정 가려진 객체를 식별하고 분석할 수 있도록 하여, 인식과 계획 간의 중요한 격차를 해소합니다. 우리는 정보 이론적 지표인 Planning KL-divergence (PKL)을 사용하여 가려진 객체가 자율 주행 차량의 계획에 미치는 영향에 따라 체계적으로 식별하고 순위를 매기는 새로운 프레임워크를 제안합니다. 이 계획 수립에 대한 인식을 갖춘 순위 정보를 활용하여, GPT-5와 같은 전문 VLM을 사용하여 시각적 증거와 해당 작업에 필요한 추론을 담은 풍부하고 구조화된 주석을 생성합니다. 본 프레임워크를 nuScenes 데이터셋에 적용하여 고 영향 시나리오에 초점을 맞춘 새로운 벤치마크를 구축했습니다. 다양한 범용 및 도메인 특화 VLM에 대한 포괄적인 실험을 수행한 결과, PKL 기반 데이터로 미세 조정하면 모든 모델에서 성능이 크게 향상되는 것을 확인했습니다. 특히, 본 연구의 결과는 작은 크기의 미세 조정된 모델이 훨씬 더 큰 규모의 제로샷 모델보다 우수한 성능을 보이며, PKL 기반 데이터 선택 전략이 무작위 샘플링에 비해 약 30% 정도 성능을 향상시킨다는 것을 보여줍니다. 본 연구는 VLM을 훈련하여 계획 수립에 중요한 가려짐 현상에 집중하도록 하는 최초의 체계적인 접근 방식을 제시하며, 이를 통해 자율 주행에서 보다 의미 기반적이고 효율적인 위험 평가를 가능하게 합니다.
Autonomous vehicles must safely navigate complex environments where planning-critical agents may be hidden from view. Current approaches often treat all occlusions with uniform conservatism, yielding needlessly defensive driving, or they infer hidden spaces without estimating the impact on the planner. This work bridges the critical gap between perception and planning by enabling Vision-Language Models (VLMs) to identify and reason about the specific hidden agents that are most critical to the ego-vehicle's trajectory. We introduce a novel framework that uses Planning KL-divergence (PKL), an information-theoretic metric, to systematically identify and rank occluded agents based on their impact on the ego vehicle's plan. Using this planning-aware ranking, we employ an expert VLM (GPT-5) to generate rich, structured annotations that capture the visual evidence and reasoning required for this task. We apply this framework to the nuScenes dataset to create a new benchmark focused on high-impact scenarios. We conduct comprehensive experiments on a wide range of general-purpose and domain-adapted VLMs, demonstrating that fine-tuning on our PKL-guided data yields dramatic performance improvements across all models. Notably, our results show that smaller, fine-tuned models significantly outperform their much larger zero-shot counterparts, and that our PKL-guided data selection strategy improves performance by approximately 30\% over random sampling. Our work presents the first systematic approach for training VLMs to focus on planning-critical occlusions, enabling more semantically grounded and efficient risk assessment in autonomous driving.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.