InCarEmo: 차량 내부 감정 인식 및 운전자 상태 모니터링을 위한 다중 모드 데이터셋
InCarEmo: A Multimodal Dataset for In-Cabin Emotion Recognition and Driver State Monitoring
차량 내 지능형 시스템의 다음 세대는 안전을 보장하고 인간과 차량 간의 상호 작용을 향상시키는 데 매우 중요하며, 이를 위해서는 운전자의 감정과 상태를 이해하는 것이 필수적입니다. 그러나 현재 공개된 차량 내부 감정 컴퓨팅 데이터셋은 주로 시각 정보에 국한되어 있으며, 운전자 감정을 파악하는 데 중요한 언어적 및 상호 작용 단서를 포함하기 어렵습니다. 이러한 한계를 극복하기 위해, 본 연구에서는 차량 내부 감정 인식 및 운전자 상태 모니터링을 위한 다중 모드 데이터셋인 InCarEmo를 소개합니다. InCarEmo는 RGB 및 적외선 비디오, 차량 내부 오디오, 그리고 실제 운전 행동을 시뮬레이션하도록 설계된 시나리오에서 수집된 대화 텍스트를 통합하며, 다양한 조명 조건과 주행 환경을 포함합니다. 본 데이터셋은 세 가지 주요 작업을 지원합니다: 1) 다중 모드 감정 인식, 2) 피로 감지, 3) 주의 산만 감시. 원본 중국어 데이터 외에도, 초기 교차 언어 평가를 지원하기 위한 보조 영어 벤치마크를 구축했습니다. 우리는 단일 모드 및 다중 모드 방법 모두에 대한 광범위한 기본 결과와 함께 통일된 벤치마크를 제공하며, 모달리티 누락 및 노이즈 조건에서의 분석도 포함합니다. 실험 결과는 다중 모드 융합의 이점을 보여주며, 실제 환경에서 발생하는 노이즈와 저조도 조건 하에서의 여전히 존재하는 과제를 드러냅니다. InCarEmo 데이터셋을 공개함으로써, 우리는 강력하고 해석 가능하며 인간 중심적인 차량 내부 감정 이해를 위한 종합적인 기반을 구축하여, 더욱 안전하고 공감적인 운전자-차량 상호 작용을 촉진하는 것을 목표로 합니다.
Understanding driver emotion and state is critical for the next generation of intelligent in-cabin systems that ensure safety and enhance human-vehicle interaction. However, existing public datasets for in-cabin affective computing are largely limited to visual modalities and rarely include conversational information, making it difficult to capture the linguistic and interactive cues underlying driver emotion. To address these gaps, we introduce InCarEmo, a multimodal dataset for in-cabin emotion recognition and driver state monitoring. InCarEmo integrates RGB and infrared video, in-cabin audio, and dialogue text collected from scripted in-cabin scenarios designed to simulate realistic driver behaviors, covering diverse lighting conditions and driving contexts. The dataset supports three primary tasks: 1) multimodal emotion recognition, 2) fatigue detection, and 3) distraction monitoring. In addition to the original Chinese data, we construct an auxiliary English benchmark to support preliminary cross-lingual evaluation. We provide a unified benchmark with extensive baseline results across unimodal and multimodal methods, including analyses under modality-missing and noise conditions. Experimental results demonstrate the benefits of multimodal fusion and reveal remaining challenges under real-world noise and low-light conditions. By releasing InCarEmo, we aim to establish a comprehensive foundation for robust, interpretable, and human-centric in-cabin affective understanding, promoting safer and more empathetic driver-vehicle interaction.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.