ForesightSafety-SAGE: LLM 에이전트를 위한 완전 자동화된 시나리오 생성 및 안전 평가 프레임워크
ForesightSafety-SAGE:A Fully Automated Scenario Generation and Safety Evaluation Framework for LLM Agents
대규모 언어 모델(LLM)은 단순한 텍스트 기반 상호 작용 시스템에서 점차적으로 메모리 유지, 도구 사용, 외부 환경 접근 및 작업 실행이 가능한 LLM 에이전트로 발전하고 있습니다. 이러한 기능과 자율성이 확장됨에 따라, 에이전트가 직면하는 안전 위험 또한 더욱 다양해지고 있습니다. 기존의 평가는 종종 수동으로 작성된 시나리오, 정적인 프롬프트 또는 최종 결과에 대한 판단에 의존하여, 작업 실행 중에 에이전트가 겪을 수 있는 다양한 위험을 포착하기 어렵습니다. 본 논문에서는 LLM 에이전트를 위한 완전 자동화된 시나리오 생성 및 안전 평가 프레임워크인 ForesightSafety-SAGE를 소개합니다. 우리는 다섯 가지 위험 차원을 기반으로, 실제 작업 실행에서 추상적이고 다양한 안전 위험을 1,072개의 측정 가능한 평가 시나리오로 구체화했습니다. 자동화된 평가 파이프라인을 사용하여 12개의 LLM 에이전트를 두 가지 권한 환경에서 평가했습니다. 결과는 현재 에이전트가 작업 실행 중에 여전히 상당한 행동적 안전 위험에 노출되어 있으며, 평균 ASR(Adversarial Success Rate)은 47.1%이고, 여러 모델의 경우 70%를 초과한다는 것을 보여줍니다. 이러한 결과는 LLM 에이전트의 안전성을 이해하고 개선하기 위한 실행 가능하고 프로세스 수준의 평가의 중요성을 강조합니다.
Large language models (LLMs) are increasingly evolving from simple text-based interaction systems into LLM agents that can maintain memory, use tools, access external environments, and execute tasks. As their capabilities and autonomy expand, the safety risks they face also become more diverse. Existing evaluations often rely on manually written scenarios, static prompts, or final-output judgments, making it difficult to capture the diverse risks that agents may face during task execution. We introduce ForesightSafety-SAGE, a fully automated scenario generation and safety evaluation framework for LLM agents. Based on five risk dimensions,we instantiae abstract and diverse safety risks in real-world task execution into 1,072 measurable evaluation scenarios. Using the automated evaluation pipeline, 12 LLM agents are evaluated under two authority contexts. The results show that current agents still face substantial behavioral safety risks during task execution, with an average ASR of 47.1% and several models exceeding 70%. These findings demonstrate the importance of executable, process-level evaluation for understanding and improving LLM agent safety.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.