강화 학습 에이전트에 대한 퍼즈 테스트 평가
Evaluating Fuzz Testing for Reinforcement Learning Agents
강화 학습(RL) 에이전트는 로봇, 자율 주행 및 드론 제어와 같은 안전이 중요한 분야에서 점점 더 많이 사용되고 있으며, 예상치 못한 동작은 심각한 현실 세계의 결과를 초래할 수 있습니다. 퍼즈 테스트는 최근 RL 에이전트의 광대한 상태 공간을 탐색하고 오류를 발견하는 유망한 방법으로 부상했습니다. 다양한 RL 퍼징 방법이 제안되었지만, 기존 연구들은 평가 설정, 기준 및 지표가 서로 다르기 때문에 상대적인 효과와 실용성에 대한 신뢰할 수 있는 결론을 내리기가 어렵습니다. 이러한 격차를 해결하기 위해, 우리는 효과성, 다양성, 효율성 및 실용성을 중심으로 RL 퍼징 방법을 체계적으로 평가하는 최초의 종합적인 경험적 연구를 제시합니다. 우리는 세 가지 환경(MountainCar, BipedalWalker, CARLA)에서 통일된 구성 하에 최첨단 방법 5가지와 무작위 테스트를 비교하고, 탐지된 오류가 에이전트의 견고성 향상 및 안전 모니터링에 얼마나 유용한지를 평가합니다. 우리의 결과는 몇 가지 중요한 통찰력을 보여줍니다. 예를 들어, MDPFuzz와 같은 처리량 중심 방법은 오류 발견에서 뛰어난 효과성과 효율성을 보이는 반면, SeqDivFuzz와 같이 탐색을 장려하도록 설계된 방법은 다양한 유형의 오류를 발견하는 데 탁월합니다. 또한 퍼징으로 생성된 오류가 에이전트의 견고성을 의미 있게 향상시키고 다양한 방법을 통해 안전 모니터링을 정확하게 수행할 수 있음을 보여줍니다. 이러한 경험적 결과 외에도, 우리는 연구자와 실무자 모두에게 실행 가능한 지침을 제시하며, 상호 보완적인 퍼징 전략을 결합하고 다층 다양성 분석을 채택하여 보다 포괄적이고 실용적인 RL 테스트를 달성하는 이점을 강조합니다.
Reinforcement Learning (RL) agents are increasingly deployed in safety-critical domains such as robotics, autonomous driving, and drone control, where unexpected behaviors may lead to severe real-world consequences. Fuzz testing has recently emerged as a promising method for exploring the vast state spaces of RL agents and exposing crashes. Although numerous RL fuzzing methods have been proposed, existing studies often differ in evaluation settings, baselines, and metrics, making it difficult to draw reliable conclusions about their relative effectiveness and practical usefulness. To address this gap, we present the first comprehensive empirical study that systematically evaluates RL fuzzing methods from four complementary perspectives: effectiveness, diversity, efficiency, and practical utility. We benchmark five state-of-the-art methods alongside random testing under unified configurations across three environments of increasing complexity (MountainCar, BipedalWalker, and CARLA), and further assess the downstream usefulness of detected crashes for agent robustness improvement and safety monitoring. Our results reveal several key insights. For instance,throughput-oriented methods like MDPFuzz demonstrate superior effectiveness and efficiency in crash discovery, while methods explicitly designed to encourage exploration like SeqDivFuzz excel at uncovering diverse crash behaviors. We also show that fuzzing-generated crashes can meaningfully improve agent robustness and enable accurate safety monitoring with strong cross-method generalization. Beyond these empirical findings, we distill actionable guidance for both researchers and practitioners, highlighting the benefits of combining complementary fuzzing strategies and adopting multi-level diversity analysis to achieve more comprehensive and practical RL testing.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.