Talk2Sensors: 센서 적응형 물리적 단서 매칭을 통한 자율 주행 시스템의 3차원 시각적 정렬
Talk2Sensors: 3D Visual Grounding in Autonomous Driving via Sensor-Adaptive Physical Cue Matching
체계적인 지능의 핵심 기능인 3차원 시각적 정렬(3DVG)은 주로 RGB-D 또는 포인트 클라우드 입력을 사용하는 실내 환경에서 연구되어 왔으며, 기존의 실외 환경으로 확장된 연구들은 대부분 단안 카메라 이미지에만 의존합니다. 이러한 방식 모두 실제 실외 환경에서의 인지 능력에는 한계가 있으며, 시각적 질감, 3차원 기하학 정보, 객체 운동 등 상호 보완적이면서도 고유한 물리적 특성을 가진 다양한 센서들을 활용하는 것이 유연하고 강력하며 상황에 적응적인 정렬을 위해 필수적이지만, 이러한 잠재력은 충분히 활용되지 못하고 있습니다. 이러한 격차를 해소하기 위해, 본 연구에서는 카메라, LiDAR 및 4D 레이더 데이터를 기반으로 구축된 최초의 다중 센서 3차원 시각적 정렬 데이터셋인 Talk2Sensors를 소개합니다. 이 데이터셋은 8,682개의 언어 지시문과 20,558개의 관련 객체로 구성되어 있으며, 다양한 프롬프트가 각 센서의 특정 물리적 단서와 명확하게 연결되어 있습니다. 또한, 본 연구에서는 자율 주행 시스템에서의 언어 기반 3차원 시각적 정렬을 위한 통합된 트랜스포머 기반 프레임워크인 TSFormer를 제안합니다. TSFormer는 거친 단계에서 세밀한 단계로 진행되는 속성 인지 융합 전략을 채택합니다. Language-Routed Property Sampler는 먼저 쿼리 수준의 언어적 단서를 사용하여 센서 샘플링 가중치를 조절함으로써, 텍스트 기반의 특징 추출을 수행합니다. 이후, Sparse-Preserving Modality Arbiter 모듈은 세밀한 단계에서 모달리티 간의 충돌을 조정하고, 텍스트 기반의 정제를 통해 정확한 참조 위치를 결정합니다. 이러한 설계는 각 프롬프트의 의미적 요구 사항에 따라 외형, 기하학 정보 및 운동 단서를 동적으로 라우팅하여, 풍부한 데이터가 희소하지만 중요한 센서 신호를 압도하는 것을 방지합니다. 광범위한 실험 결과는 TSFormer가 여러 벤치마크에서 최고 성능을 달성함을 보여줍니다. 특히 Talk2Sensors 데이터셋에서는 가장 강력한 기존 모델보다 8.05% 더 높은 mAP를 달성했으며, 단안 카메라 기반의 Mono3DRefer 벤치마크에서도 53.05%의 Acc@0.5를 기록하며 우수한 성능을 보였습니다.
As a key capability for embodied intelligence, 3D visual grounding (3DVG) has been predominantly studied in indoor scenes with RGB-D or point-cloud inputs, while existing outdoor extensions largely rely on monocular images alone. Both settings fall short of real-world outdoor perception, where heterogeneous sensors capture complementary yet distinct physical properties, such as visual texture, 3D geometry, and object kinematics, that are indispensable for flexible and robust query-adaptive grounding but remain under-exploited. To bridge this gap, we introduce Talk2Sensors, the first multi-sensor 3D visual grounding dataset built upon camera, LiDAR, and 4D radar. It contains 8,682 language instructions and 20,558 referred objects, with diverse prompts explicitly aligned with sensor-specific physical cues. Furthermore, we propose TSFormer, a unified Transformer-based framework for language-guided 3D visual grounding in autonomous driving. TSFormer adopts a coarse-to-fine property-aware fusion strategy: the Language-Routed Property Sampler first performs coarse text-conditioned feature retrieval by modulating sensor sampling weights with query-level linguistic cues, while the subsequent Sparse-Preserving Modality Arbiter module conducts fine-grained modality arbitration and text-guided refinement to determine the precise referred spatial location. This design enables dynamic routing of appearance, geometry, and motion cues according to the semantic requirements of each prompt, preventing dense modalities from overwhelming sparse but critical sensor signals. Extensive experiments demonstrate that TSFormer achieves state-of-the-art performance across multiple benchmarks: it improves over the strongest baseline by 8.05 mAP on Talk2Sensors, and transfers to the monocular Mono3DRefer benchmark with 53.05\% Acc@0.5.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.