AgentDoG 1.5: AI 에이전트의 안전 및 보안을 위한 경량화되고 확장 가능한 정렬 프레임워크
AgentDoG 1.5: A Lightweight and Scalable Alignment Framework for AI Agent Safety and Security
OpenClaw와 같은 현대적인 오픈 월드 에이전트는 강력한 크로스-환경 실행 능력을 제공하지만, 동시에 광범위하고 새로운 안전 위험 요소를 야기합니다. 동시에, 최첨단 AI 모델은 공격의 진입 장벽을 크게 낮추어, 현재 에이전트 정렬 프레임워크가 실제 환경에 적용되기에는 부적절하게 만듭니다. 이러한 새로운 위협에 대처하기 위해, 우리는 경량화되고 확장 가능한 에이전트 안전 정렬 프레임워크를 제안합니다. 구체적으로, Codex 및 OpenClaw 실행 시나리오에서 발생하는 새로운 위험을 수용하기 위해 에이전트 안전 분류 체계를 업데이트했습니다. 또한, 영향 함수 정제 기능을 갖춘 분류 체계 기반 데이터 엔진을 구축하여 약 1,000개의 샘플만 사용하여 경량화된 AgentDoG 1.5 변종(0.8B, 2B, 4B 및 8B 파라미터)을 학습시켰습니다. 이를 통해 GPT-5.4와 같은 선도적인 비공개 모델과 비교 가능한 성능을 달성했습니다. AgentDoG 1.5를 기반으로, 매우 효율적인 에이전트 안전 SFT(Supervised Fine-Tuning) 및 RL(Reinforcement Learning) 학습 환경을 구축하여 Docker 수준의 환경에서 배포 오버헤드를 두 배 이상 줄였습니다. 마지막으로, AgentDoG 1.5는 별도의 학습 없이 실시간 안전 관리를 위한 온라인 가이드레일로 배포됩니다. 광범위한 실험 결과는 AgentDoG 1.5가 다양한 복잡한 인터랙티브 에이전트 시나리오에서 최첨단 성능을 달성함을 보여줍니다. 모든 모델과 데이터셋은 공개적으로 제공됩니다.
Modern open-world agents such as OpenClaw exhibit powerful cross-environment execution capabilities yet introduce broad new safety risk sources. Meanwhile, advanced frontier AI models drastically lower attack barriers, rendering current agent alignment frameworks inadequate for real-world deployment. To tackle these emerging threats, we propose a lightweight and scalable agent safety alignment framework. Specifically, we update the agent safety taxonomy to accommodate emergent risks from Codex and OpenClaw execution scenarios. We further build a taxonomy-guided data engine with influence-function purification to train lightweight AgentDoG 1.5 variants (0.8B, 2B, 4B, and 8B parameters) using only around 1k samples, achieving comparable performance with leading closed-source models (e.g., GPT-5.4). Based on AgentDoG 1.5, we construct a highly efficient agentic safety SFT and RL training environment, which reduces deployment overhead in Docker-level environments by two orders of magnitude. Finally, we deploy AgentDoG 1.5 as a training-free online guardrail for real-time safety moderation. Extensive experimental results indicate that AgentDoG 1.5 achieves state-of-the-art performance in diverse and complex interactive agentic scenarios. All models and datasets are openly released.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.