2601.22159v2 Jan 29, 2026 cs.CR

RedSage: 사이버 보안 분야의 범용 LLM

RedSage: A Cybersecurity Generalist LLM

Naufal Suryanto
Naufal Suryanto
Khalifa University
Citations: 238
h-index: 6
Muzammal Naseer
Muzammal Naseer
Citations: 8,514
h-index: 30
Pengfei Li
Pengfei Li
Citations: 45
h-index: 3
Syed Talal Wasim
Syed Talal Wasim
Citations: 739
h-index: 8
Jinhui Yi
Jinhui Yi
Citations: 20
h-index: 3
Juergen Gall
Juergen Gall
Citations: 47
h-index: 4
P. Ceravolo
P. Ceravolo
Citations: 2,879
h-index: 21
Ernesto Damiani
Ernesto Damiani
Citations: 35
h-index: 3

사이버 보안 운영은 민감한 데이터를 노출시키지 않으면서 다양한 워크플로우를 지원하는 어시스턴트 LLM이 필요합니다. 기존 솔루션은 개인 정보 보호 위험이 있는 독점 API에 의존하거나, 도메인 적응력이 부족한 오픈 모델에 의존하는 경우가 많습니다. 이러한 격차를 해소하기 위해, 우리는 대규모 웹 필터링 및 고품질 리소스의 수동 수집을 통해 28.6K개의 문서에 걸쳐 프레임워크, 공격 기술 및 보안 도구를 포함하는 11.8B 토큰의 사이버 보안 중심의 지속적인 사전 훈련 데이터를 구축했습니다. 이를 바탕으로, 우리는 전문가 워크플로우를 시뮬레이션하여 266K개의 다중 턴 사이버 보안 샘플을 생성하는 에이전트 기반 증강 파이프라인을 설계하여 지도 학습을 수행했습니다. 이러한 리소스와 일반적인 오픈 소스 LLM 데이터를 결합하여, RedSage는 도메인 인식 사전 훈련 및 사후 훈련을 갖춘 오픈 소스, 로컬 배포 가능한 사이버 보안 어시스턴트를 학습했습니다. 모델의 엄격한 평가를 위해, 우리는 사이버 보안 지식, 기술 및 도구 전문성을 다루는 30K개의 객관식 문제와 240개의 개방형 질의응답 항목으로 구성된 벤치마크인 RedSage-Bench를 도입했습니다. RedSage는 또한 기존의 사이버 보안 벤치마크(예: CTI-Bench, CyberMetric, SECURE) 및 일반적인 LLM 벤치마크에서 광범위한 일반화 능력을 평가했습니다. 8B 규모에서 RedSage는 일관되게 더 나은 결과를 보여주며, 사이버 보안 벤치마크에서 최대 +5.59점, Open LLM Leaderboard 작업에서 +5.05점의 향상을 보였습니다. 이러한 결과는 도메인 인식 에이전트 기반 증강 및 사전/사후 훈련이 사이버 보안 관련 전문성을 향상시킬 뿐만 아니라, 일반적인 추론 및 지시 따르기 능력 향상에도 도움이 될 수 있음을 보여줍니다. 모든 모델, 데이터 세트 및 코드는 공개적으로 이용 가능합니다.

Original Abstract

Cybersecurity operations demand assistant LLMs that support diverse workflows without exposing sensitive data. Existing solutions either rely on proprietary APIs with privacy risks or on open models lacking domain adaptation. To bridge this gap, we curate 11.8B tokens of cybersecurity-focused continual pretraining data via large-scale web filtering and manual collection of high-quality resources, spanning 28.6K documents across frameworks, offensive techniques, and security tools. Building on this, we design an agentic augmentation pipeline that simulates expert workflows to generate 266K multi-turn cybersecurity samples for supervised fine-tuning. Combined with general open-source LLM data, these resources enable the training of RedSage, an open-source, locally deployable cybersecurity assistant with domain-aware pretraining and post-training. To rigorously evaluate the models, we introduce RedSage-Bench, a benchmark with 30K multiple-choice and 240 open-ended Q&A items covering cybersecurity knowledge, skills, and tool expertise. RedSage is further evaluated on established cybersecurity benchmarks (e.g., CTI-Bench, CyberMetric, SECURE) and general LLM benchmarks to assess broader generalization. At the 8B scale, RedSage achieves consistently better results, surpassing the baseline models by up to +5.59 points on cybersecurity benchmarks and +5.05 points on Open LLM Leaderboard tasks. These findings demonstrate that domain-aware agentic augmentation and pre/post-training can not only enhance cybersecurity-specific expertise but also help to improve general reasoning and instruction-following. All models, datasets, and code are publicly available.

1 Citations
0 Influential
15 Altmetric
76.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!