자율 주행에 대한 탐색 연구: 내부 예측 오류와 자가 계획 간의 연관성
What Probing Reveals about Autonomous Driving: Linking Internal Prediction Errors to Ego Planning
대규모 데이터셋과 빠른 시뮬레이션은 안전하고 견고해 보이는 운전 정책 개발을 가능하게 했지만, 이상적인 환경에서의 뛰어난 성능은 여전히 결함 있는 추론 및 위험한 휴리스틱스를 숨길 수 있습니다. 폐루프 시뮬레이터에서 얻는 요약 점수는 정책에 대한 중요한 정보를 제공하지 않아, 해당 정책이 실제로 주변 차량의 움직임을 예측하는지, 자율차가 어떻게 미래 계획을 생성하는지에 대한 이해를 어렵게 하며, 단순히 이상적인 환경에서만 작동하는 취약한 휴리스틱스에 의존하는 것은 아닌지 판단하기 어렵습니다. 운전 정책의 한계와 약점을 더 잘 이해하기 위해, 우리는 주변 차량이 다음에 어디로 움직일 것인지에 대한 예측 능력과 안전한 경로를 생성하는 방법에 대한 계획 능력을 탐색하는 데 중점을 둡니다. 이러한 두 가지 능력은 효과적인 운전 정책에서 기대되는 행동을 반영하며, 이를 통해 데이터 기반 행동 복제 및 시뮬레이션 기반 강화 학습 정책의 품질을 평가합니다. 이러한 능력의 존재 여부를 평가하기 위해, 규모에 따른 변화를 조사하여 더 큰 데이터셋과 더 긴 시뮬레이션 훈련이 예측 및 계획 능력이 향상되었는지, 아니면 단순히 더 나은 행동 휴리스틱스를 통해 이루어진 것인지 확인합니다. 우리는 선형 탐색 및 목표 지향적 교란을 사용하여 모방 학습 및 강화 학습 모델에서 이러한 내부 신호가 언제 나타나는지, 최고점에 도달하는지 또는 실패하는지를 추적합니다. 우수한 폐루프 성능에도 불구하고, 정책은 종종 충돌 직전 상황에서 적절한 시기에 주변 차량의 움직임을 예측하지 못하며, 이는 자가 계획을 위한 예측 신호에 대한 한계를 드러냅니다. 마지막으로, 인과 관계 분석 결과, 잘못된 예측을 수정하면 더 안전한 경로를 향한 자가 계획이 개선되는 것으로 나타났습니다.
Large-scale datasets and fast simulators have enabled improvements in driving policies that appear safe and robust, yet strong performance in nominal scenarios can still mask flawed reasoning and unsafe heuristics. Summary scores from closed-loop simulators do not give significant insight into the policy, making it difficult to determine whether they truly predict the motion of surrounding vehicles, how the ego vehicle generates future plans, or whether they merely rely on brittle heuristics that happen to succeed in nominal scenarios. To better understand the limits and weaknesses of driving policies, we focus on probing for forms of prediction, i.e., where surrounding vehicles will move next, and planning, i.e., understanding how to generate safe trajectories. We focus on these two capabilities because they reflect behaviors expected of effective driving policies, and use their presence or absence to assess policy quality across data-driven behavior cloning and simulation-driven reinforcement learning policies. To evaluate the presence of these capabilities, we investigate them as a function of scale, asking whether the closed-loop gains from larger datasets and longer simulation training reflect stronger prediction and planning or merely better behavioral heuristics. We use linear probing and targeted perturbations in both imitation learning and reinforcement learning models to track when these internal signals emerge, plateau, or fail. Despite good closed-loop performance, policies often fail to form timely surrounding-vehicle predictions during near-collision events, revealing a limitation in the predictive signals available for ego planning. Finally, causal intervention shows that correcting mistaken predictions improves ego planning toward safer trajectories.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.