FlowR2A: 다중 모드 자율 주행 계획을 위한 보상-액션 분포 학습
FlowR2A: Learning Reward-to-Action Distribution for Multimodal Driving Planning
다중 모드 자율 주행 계획은 오랫동안 두 가지 패러다임 사이의 긴장 관계에 직면해 왔습니다. 점수 기반 방법은 풍부한 보상 정보를 활용할 수 있지만, 고정된 액션 어휘로 제한되는 반면, 앵커 기반 방법은 동적으로 제안을 생성하지만 단일의 정답 경로에 의해 제한되는 희소한 정보에 의존합니다. 본 연구에서는 이러한 긴장을 해소하기 위해 시뮬레이션을 기반으로 한 보상을 판별적인 목표가 아닌 생성적 조건으로 재구성하는 FlowR2A를 제안합니다. FlowR2A는 트레일-보상 쌍의 밀집된 데이터를 사용하여 플로우 매칭 디코더로 보상에 따른 액션 분포를 학습함으로써, 점수 기반 방법의 밀집적인 정보 활용과 앵커 기반 방법의 제안 생성 방식을 하나의 생성 모델로 통합합니다. 이를 통해 모델은 안전, 진행, 편안함 및 규칙 준수를 포함한 액션과 그 결과 사이의 상관관계를 내재화하도록 유도합니다. 안전에 대한 엄격한 제약 조건과 진행 목표 간의 균형을 맞추기 위해, 우리는 세분화된 시간 단계별 보상 조건 설정과 보상 노이즈 증강 기술을 도입했습니다. 생성적 프레임워크는 자연스럽게 보상 기반 샘플링 및 앵커 기반 샘플링을 통해 테스트 시점에 제어 가능한 샘플링을 지원하여 고품질의 제안을 생성합니다. FlowR2A는 NAVSIM v1 및 v2 벤치마크에서 최첨단 결과를 달성했으며, 이전 방법보다 훨씬 높은 품질의 다중 모드 제안을 제공합니다.
Multimodal driving planning faces a long-standing tension between two paradigms: scoring-based methods benefit from dense reward supervision but are confined to a fixed action vocabulary, while anchor-based methods generate proposals dynamically yet suffer from sparse supervision constrained to a single ground-truth trajectory. In this work, we propose FlowR2A, which resolves this tension by reframing simulation-based rewards from discriminative targets into generative conditions. By learning the reward-conditioned action distribution from dense trajectory-reward pairs with a flow-matching decoder, FlowR2A unifies the dense supervision of scoring-based methods with the proposal generation of anchor-based methods in a single generative model, forcing the model to internalize the correlation between an action and its outcomes in safety, progress, comfort, and rule compliance. To balance hard safety constraints against soft progress objectives, we introduce fine-grained per-timestep reward conditioning and reward noise augmentation. The generative formulation naturally supports controllable test-time sampling via reward guidance and anchored sampling, producing high-quality proposals. FlowR2A achieves state-of-the-art results on the NAVSIM v1 and v2 benchmarks, with multimodal proposals of substantially higher quality than prior methods.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.