스포츠에서 안전으로: 대규모 언어 모델(LLM)의 능동적 위험 추론 성능 평가
From Sports to Safety: Benchmarking Proactive Risk Inference in MLLMs
실제 환경에서의 안전을 위해서는 잠재적인 물리적 위험 요소를 사전에 예측하는 것이 중요하지만, 기존의 LLM 평가는 주로 유해 콘텐츠나 일반적인 위험에 초점을 맞추고 있어, 능동적인 물리적 위험 예측은 상대적으로 덜 연구되어 왔습니다. 스포츠는 이러한 연구를 위한 이상적인 실험 환경을 제공합니다. 왜냐하면 사고 원인은 다양한 부상 요인과 관련이 있으며, 사고 전의 시공간적 단서는 자율 주행 및 낙상 감지 등 광범위한 안전 영역에서 공유되는 추론 능력을 활용하기 때문입니다. 본 연구에서는 14개의 스포츠 종목과 3가지 환경 설정을 아우르는 2,888개의 실제 스포츠 동영상(사고 발생 영상 2,440개, 안전 영상 448개)으로 구성된 SPRINT (Sports Proactive Risk INference Testbed)라는 새로운 성능 평가 도구를 소개합니다. 사고 발생 영상에는 초기 위험 요소, 사고 발생 시점 및 계층적 원인에 대한 세부적인 주석이 포함되어 있으며, 안전 영상은 수동 검증을 통해 사고가 없음을 확인하고 프롬프트 유발로 인한 오탐 현상을 진단하는 데 사용됩니다. 다양한 프롬프트를 사용하여 최첨단 LLM을 평가한 결과, 위험 요소 감지 능력과 원인 이해 능력 사이에 큰 격차가 있음을 알 수 있었습니다. 최고 성능의 모델은 95% 이상의 정확도로 위험 요소를 감지했지만, 원인을 식별하는 데는 50% 미만의 낮은 성능을 보였습니다. 추가적인 진단 실험 결과, 명시적인 위험 관련 질문은 위험 요소가 없는 영상에서도 심각한 오탐 현상을 유발하는 것으로 나타났습니다. 이러한 결과는 현재 LLM이 표면적인 수준의 안전 기능만 제공하며, 안정적이고 원인 기반의 사전 경고 기능을 갖추지 못하고 있다는 점을 시사합니다. 또한, 동적인 물리적 환경에서 신뢰할 수 있는 능동적 안전 기능의 필요성을 강조합니다. 데이터와 코드는 논문 게재 확정 후 공개될 예정입니다.
Timely anticipation of physical hazards is essential for real-world safety, yet existing MLLM evaluations focus on harmful content or general risks, leaving proactive physical hazard prediction underexplored. Sports provide a well-suited testbed: accident causes span diverse injury dimensions and pre-accident spatiotemporal cues draw on reasoning capabilities shared with broader safety domains such as autonomous driving and fall detection. We introduce SPRINT (Sports Proactive Risk INference Testbed), a benchmark of 2,888 real-world sports videos (2,440 accident, 448 safe controls) spanning 14 sports and 3 environmental settings. Accident videos feature fine-grained annotations of early hazard cues, accident timing, and hierarchical causes; safe videos are manually verified as accident-free and serve to diagnose prompt-induced false alarms. Evaluating state-of-the-art MLLMs under diverse prompts and temporal windows reveals a sharp gap between hazard sensitivity and understanding: the best model exceeds 95% in signaling hazards yet falls below 50% in identifying their causes. Diagnostic experiments further show that explicit danger queries trigger severe false alarms even on hazard-free videos. These findings indicate that current MLLMs exhibit only superficial proactive safety, lacking stable, cause-grounded early warning, and underscore the need for reliable proactive safety in dynamic physical environments. Data and code will be open-sourced upon acceptance.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.