당신의 주행 경로 이탈은 롱테일 환경에서 안전한가?
Is Your Trajectory Displacement Safe in Long-tail?
데이터셋이 기하급수적으로 증가함에도 불구하고, 롱테일 시나리오는 자율주행 시스템 평가에 여전히 큰 걸림돌입니다. 기존의 평가 파이프라인은 종종 인간의 판단과 일치하지 않거나, 안전성을 고려하지 않고, 검증 가능하거나 설명 가능하도록 설계되지 않습니다. 폐루프 메트릭은 우수한 계획 시스템에서 쉽게 포화되는 반면, 체계적인 프로토콜 없이는 비정형적인 인간 평가가 노이즈를 포함할 수 있습니다. 본 연구에서는 계획 평가를 추가적인 위험 감지 문제로 정의합니다: 특정 계획 시스템의 경로와 전문가 참조 경로를 비교했을 때, 해당 계획 시스템의 경로 이탈이 새로운 안전하지 않은 주행 행동을 유발하는가? 우리는 FluidTest라는 평가 파이프라인을 제안하며, 이는 세 가지 구성 요소로 이루어져 있습니다. 첫째, 신뢰할 수 있는 인간 어노테이션을 위한 쌍대 웹 인터페이스 프로토콜입니다. 둘째, 증거 기반 의사 결정 그래프를 사용하는 32가지 의미론적 위험 분류 체계입니다. 셋째, 정확성과 감사 가능성을 확보하기 위한 세 에이전트 검증 시스템과 피드백 루프 기능입니다. WOD-E2E 데이터셋에 대한 실험 결과, FluidTest는 교육받은 어노테이터 간에 일관된 레이블을 생성하며, Poutine 경로의 65%와 RAP 경로의 51%에서 추가적인 위험 요소를 식별합니다. 이러한 결과는 최첨단 계획 시스템도 높은 Rater Feedback Score (RFS) 및 낮은 Average Displacement Error (ADE)에도 불구하고 여전히 상당한 안전 관련 오류를 나타낼 수 있음을 보여줍니다. 추가 정보, 가이드라인 및 코드는 https://fluidtest.web.app 에서 확인할 수 있습니다.
Long-tail scenarios remain a major bottleneck for autonomous driving evaluation, even as datasets grow by orders of magnitude. Existing evaluation pipelines are rarely human-aligned, safety-aware, verifiable, and explainable at the same time: closed-loop metrics often saturate among strong planners, while unstructured human ratings can be noisy without a carefully designed protocol. We formulate planning evaluation as additional-threat detection: given a planner trajectory and an expert reference, does the planner's displacement introduce new unsafe driving behavior? We propose FluidTest, an evaluation pipeline with three components: a pairwise WebUI protocol for reliable human annotation; a taxonomy of 32 semantic threats with evidence-grounded decision graphs; and a three-agent verification system with reflection for precision and auditability. Experiments on the WOD-E2E dataset show that FluidTest produces consistent labels among trained annotators and identifies additional threats in 65% of Poutine trajectories and 51% of RAP trajectories. These results show that state-of-the-art planners can still exhibit substantial safety-relevant failures despite high Rater Feedback Scores (RFS) and low Average Displacement Error (ADE). Additional details, guidance, and code are available at https://fluidtest.web.app.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.