ReasonLight: 다중 모드 기반 파운데이션 모델을 활용한 강화 학습 프레임워크를 통한 제로샷 교통 신호 제어
ReasonLight: A Multimodal Foundation Model-Enhanced Reinforcement Learning Framework for Zero-Shot Traffic Signal Control
강화 학습(RL)은 교통 신호 제어(TSC) 분야에서 유망한 결과를 보여주었습니다. 그러나 RL은 미리 정의된 상태에 의존하기 때문에, 훈련 데이터에 없는 관찰 가능한 실제 상황에 대한 반응성이 제한됩니다. IoT 기술을 활용한 교차로는 도로 측 센서와 카메라로부터 다양한 정보를 제공하며, 이를 통해 RL의 실제 상황 적응성을 향상시킬 수 있는 기회를 제공합니다. 이에 따라, 우리는 제로샷 TSC를 위한 다중 모드 기반 파운데이션 모델을 활용한 강화 학습 프레임워크인 ReasonLight를 제안합니다. ReasonLight는 세 가지 정보 소스를 통합합니다: 구조화된 교통 측정 데이터, 다양한 시점의 카메라 관찰 결과, 그리고 사전 훈련된 RL 컨트롤러에서 생성된 후보 신호 변경 결정입니다. ReasonLight는 RL에서 제안된 신호 변경에 대해, 다중 시점 이미지로부터 시각적 의미를 추출하고 이를 센서 기반으로 얻은 간결한 장면 설명과 연결합니다. 이러한 연결을 통해, 의미론적으로 안내되는 정제 모듈이 교통 규칙 및 상황 의미를 고려하여 제안된 동작을 유지하거나 조정합니다. 시스템 운영의 신뢰성을 확보하기 위해, 정제된 동작은 가능한 신호 변경 옵션 범위 내에 제한됩니다. 유효하지 않은 결정은 거부되며, 시스템은 원래 RL 동작으로 되돌아갑니다. 우리는 ReasonLight를 RL 훈련 중에 관찰되지 않는 두 가지 유형의 드문 상황(긴급 차량 우선 처리 및 일시적인 교통 규제)에서 평가했습니다. 실험 결과는 ReasonLight가 재훈련 없이 제로샷 방식으로 적응할 수 있음을 보여줍니다. ReasonLight는 기존 RL 시스템에 비해 긴급 차량 대기 시간을 최대 88.7%까지 줄이는 동시에 일반적인 교통 흐름 성능을 유지합니다.
Reinforcement learning (RL) has shown promise in traffic signal control (TSC). However, its reliance on predefined states limits responsiveness to observable open-world events that are absent from training data. IoT-enabled intersections provide heterogeneous observations from roadside sensors and cameras, creating opportunities to improve RL adaptability to such events. To this end, we propose ReasonLight, a multimodal foundation model-enhanced RL framework for zero-shot TSC. ReasonLight integrates three sources of information: structured traffic measurements, multi-view camera observations, and candidate phase decisions from a pre-trained RL controller. Given an RL-proposed phase, ReasonLight extracts visual semantics from multi-view images and aligns them with compact sensor-derived scene descriptions. This alignment enables a semantic-guided refinement module to either preserve or adjust the proposed action according to traffic rules and event semantics. To ensure operational reliability, refined actions are constrained by the set of available phases. Any invalid decision is rejected, and the system falls back to the original RL action. We evaluate ReasonLight on two types of rare events not seen during RL training: emergency vehicle priority and temporary traffic regulation. Experimental results show that ReasonLight achieves zero-shot adaptation without retraining. It reduces emergency vehicle waiting time by up to 88.7% compared with the RL-only backbone while preserving comparable routine traffic performance.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.