EXE-Bench: 실질적인 사용성을 고려한 AI 기반 Windows 악성코드 탐지기의 장단점 평가
EXE-Bench: Ranking the Tradeoffs of AI-based Windows Malware Detectors for Real-World Usability
체계적인 평가 부족으로 인해, 어떤 AI 기반 Windows 악성코드 탐지기를 실제 환경에 적용해야 할지 판단하기 어렵습니다. 기존의 평가는 (i) 학습 및 테스트에 사용되는 데이터가 서로 다르거나, (ii) 모델이 시간이 지남에 따라 성능 저하 없이 유지되는지를 보여주는 시계열 분석을 고려하지 않으며, (iii) 콘텐츠 주입 공격과 같은 적대적 공격을 통한 보안 평가를 회피하여 모델의 취약점을 드러내지 못하고, (iv) 배포에 필요한 컴퓨팅 자원을 간과하여 엔드포인트에서 느린 추론 속도를 유발할 위험이 있습니다. 이러한 이유로, 우리는 AI 기반 Windows 악성코드 탐지기를 종합적으로 평가하는 벤치마크인 EXE-Bench를 개발했습니다. EXE-Bench는 성능, 시계열 및 적대적 견고성, 컴퓨팅 오버헤드를 평가하고, 이를 하나의 점수로 통합하여 모델 간의 직접적이고 공정한 비교를 가능하게 합니다. EXE-Bench를 통해, 배포 후에만 수행되는 평가는 최적이 아니며 모델의 전체적인 성능을 완전히 파악할 수 없다는 것을 강조합니다. 특히, 우리의 분석 결과, 특징 엔지니어링을 통해 주입된 도메인 지식이 여전히 이 분야에서 매우 유용하며, 시간과 적대적 공격에 모두 강하다는 점을 보여줍니다. 이는 대부분의 심층 신경망이 배포 직후에는 뛰어난 성능을 보이지만, 시간이 지나면서 취약해지는 것과는 대조적인 결과입니다.
Due to the lack of systematic evaluations, we are not yet able to determine which AI-based Windows malware detector to deploy in production, since existing evaluations (i) differ in terms of data used for both training and testing; (ii) do not consider temporal analysis to showcase whether models withstand the passage of time; (iii) avoid security evaluations with adversarial attacks that could highlight their brittleness against content-injection attacks; and (iv) neglect the computational requirements for deployment, risking slow inference on endpoints. For these reasons, we develop EXE-Bench, a comprehensive benchmark of AI-based Windows malware detectors. EXE-Bench assesses performance, temporal and adversarial robustness, and computational overhead, aggregating them into a single score for direct and fair model comparison. Through EXE-Bench, we highlight how evaluations conducted only after deployment are suboptimal and unable to provide a complete picture of their performance. In particular, through our analysis, we remark how much domain knowledge instilled through feature engineering is still extremely useful in this domain, resisting both time and adversarial attacks, in stark contrast with most of the deep networks that only excel right after deployment.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.