SoK: 인공지능 기반 바이너리 역공학
SoK: AI-Augmented Binary Reversing
바이너리 역공학은 소프트웨어 이해, 취약점 발견, 악성코드 분석 및 펌웨어 감사에 필수적인 기술입니다. 그러나 컴파일 과정에서 의미 정보가 영구적으로 손실되기 때문에 본질적으로 어려운 과제입니다. 최근 머신러닝, 대규모 언어 모델(LLM) 및 에이전트 기반 AI 시스템의 발전으로 인해 인공지능을 활용한 바이너리 역공학이 빠르게 확산되고 있습니다. 하지만 관련 연구 결과는 역공학 영역, 분석 대상, 학습 방법 및 평가 방식에 따라 점점 더 세분화되고 파편화되는 경향을 보입니다. 본 논문은 인공지능 기반 바이너리 역공학 분야의 지식을 체계적으로 정리한 최초의 연구입니다. 2015년 이후 발표된 144개의 연구 논문을 분석하여 추론 작업에 따라 22가지 바이너리 역공학 영역으로 분류했습니다. 또한, 기존 방식과 인공지능 기반 방식을 포괄하는 통합적인 분류 체계를 제시합니다. 본 분류 체계는 전통적인 분석 기술, 바이너리에서 파생된 결과물, 표현 전략, 학습 패러다임 및 후속 추론 작업을 연결하며, LLM과 에이전트 기반 AI 시스템의 역할 변화를 명확히 설명합니다. 공통 용어와 구조화된 프레임워크를 확립함으로써, 지난 10년간 이 분야의 발전 과정을 종합적으로 보여줍니다. 본 연구는 표면적으로는 상반되는 접근 방식에도 숨겨진 공통점을 밝혀내고, 지속적인 기술적 과제 및 평가상의 격차를 강조하며, 향후 연구에 대한 유망한 기회를 제시합니다. 이러한 분석 결과는 현재 이 분야의 현황을 명확히 하고, 신뢰성 있고 확장 가능한 인공지능 기반 바이너리 역공학 시스템 개발을 위한 토대를 제공합니다.
Binary reversing is fundamental to software understanding, vulnerability discovery, malware investigation, and firmware auditing. However, it remains inherently challenging due to the irreversible loss of semantic information during compilation. Recent advances in machine learning, large language models (LLMs), and agentic AI systems have accelerated the adoption of AI-augmented binary reversing. Yet, the resulting body of work has become increasingly fragmented across reversing domains, artifact representations, learning approaches, and evaluation practices. This paper presents the first comprehensive systematization of knowledge on AI-augmented binary reversing. We analyze 144 research papers published since 2015, and organize them into 22 binary reversing domains according to the inference tasks. We further introduce a unified taxonomy spanning conventional and AI-augmented reversing pipelines. Our taxonomy connects traditional analysis techniques, binary-derived artifacts, representation strategies, learning paradigms, and downstream inference tasks, while clarifying the emerging roles of LLMs and agentic AI systems. By establishing a common vocabulary and structured framework, we provide a holistic view of the field's evolution over the past decade. Our study reveals common structures underlying seemingly disparate approaches, highlights persistent technical challenges and evaluation gaps, and identifies promising opportunities for future research. Collectively, these insights clarify the current state of the field and provide a foundation for the next generation of reliable and scalable AI-augmented binary reversing systems.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.