Radar4D-VLM: 제안 기반의 시간적 4차원 레이더 추론 모델 - 고정된 언어 모델 활용
Radar4D-VLM: Proposal-Grounded Temporal 4D Radar Reasoning Across Frozen Language Models
자율 주행을 위한 시각-언어 모델은 주로 카메라와 LiDAR에 의존하지만, 악천후 조건에서도 견고하고 방사 속도를 직접 측정할 수 있는 4D 레이더는 독립적인 인지 모달리티로서 큰 잠재력을 가지고 있습니다. 본 논문에서는 카메라나 LiDAR 입력 없이 연속된 10개의 4D 레이더 포인트 클라우드 데이터를 사용하여 시간적 추론을 수행하는 레이더 전용 시각-언어 모델인 Radar4D-VLM을 제안합니다. Radar4D-VLM은 기하학적으로 타당한 객체 제안을 추출하고, 레이더 데이터를 객체, 장면 및 운동 상태 토큰의 간결한 계층 구조로 구성합니다. 매개변수 효율적인 프로젝터는 이러한 토큰을 고정된 언어 모델 백본으로 매핑하며, 감사 가능한 예측 헤드는 객체의 개수, 공간 분포, 운동 상태, 충돌 위험, 의미 범주 및 방사 속도를 동시에 모델링합니다. Radar4D-VLM은 제안 기반의 시간적 객체 토큰화, 전역적인 장면 맥락 및 명시적인 운동 상태 토큰을 통합된 고정 백본 인터페이스 내에서 결합합니다. K-Radar 개발 검증 데이터셋에서, 시퀀스 분리 조건 하에 Radar4D-VLM의 Top-64 제안 재현율은 4m 거리에서 98.13%로, 기존의 격자 기반 및 균일 랜덤 방식보다 각각 6.40% 및 22.83% 더 높은 성능을 보였습니다. 또한, 동일한 적응 예산을 사용하여 Qwen, Phi, Mistral, Llama 및 Gemma 백본을 포함한 총 8개의 고정된 언어 모델에 대해 24개의 일관성 있는 실험을 수행했습니다. 레이더 토큰 인터페이스는 모든 다섯 가지 언어 모델 패밀리에서 호환성을 유지하며, 정렬된, 재배열된, 그리고 언어 정보를 사용하지 않은 제어 그룹과의 비교를 통해 센서 의존성은 확인되었지만, 정렬된 언어 지도 학습이 직접적인 성능 향상에 기여하는 안정적인 효과는 관찰되지 않았습니다. 이러한 결과는 레이더 전용의 다중 모드 장면 및 운동 추론을 위한 재현 가능한 기반을 제시하며, 인터페이스 호환성과 언어 지도 학습의 이점을 분리합니다.
Vision-language models for autonomous driving primarily rely on cameras and LiDAR, leaving 4D radar largely unexplored as a standalone perceptual modality despite its robustness to adverse visibility and direct measurement of radial velocity. We introduce Radar4D-VLM, a radar-only temporal vision-language model that reasons from ten consecutive 4D-radar point-cloud sweeps without camera or LiDAR input. Radar4D-VLM extracts geometrically grounded object proposals and organizes radar evidence into a compact hierarchy of object, scene, and kinematic tokens. A parameter-efficient projector maps these tokens into frozen language backbones, while auditable prediction heads jointly model object count, spatial distribution, motion state, collision risk, semantic category, and radial velocity. Radar4D-VLM combines proposal-grounded temporal object tokenization, global scene context, and explicit kinematic tokens within a unified frozen-backbone interface. On sequence-isolated K-Radar development validation, its Top-64 proposal recall reaches 98.13% at 4 m, exceeding fixed-lattice and uniform-random controls by 6.40 and 22.83 percentage points, respectively. We further evaluate 24 matched runs spanning eight frozen Qwen, Phi, Mistral, Llama, and Gemma backbones under an identical adaptation budget. The radar-token interface remains compatible across all five language-model families, while matched aligned, permuted, and no-language controls show sensor dependence but no stable direct-head gain from aligned language supervision. These results establish a reproducible foundation for radar-only multimodal scene and motion reasoning while separating interface compatibility from the benefit of language supervision.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.