SafePro: 전문가 수준 AI 에이전트의 안전성 평가
SafePro: Evaluating the Safety of Professional-Level AI Agents
대규모 언어 모델 기반 에이전트는 단순한 대화형 어시스턴트에서 다양한 도메인의 복잡하고 전문적인 작업을 수행할 수 있는 자율 시스템으로 급격히 진화하고 있습니다. 이러한 발전은 상당한 생산성 향상을 예고하지만, 동시에 아직 충분히 탐구되지 않은 치명적인 안전 위험을 초래하기도 합니다. 기존의 안전성 평가는 주로 단순한 일상 보조 작업에 초점을 맞추고 있어, 전문적인 환경에서의 복잡한 의사결정 과정과 정렬되지 않은 행동이 초래할 수 있는 잠재적 결과를 제대로 포착하지 못하고 있습니다. 이러한 공백을 메우기 위해, 우리는 전문적인 활동을 수행하는 AI 에이전트의 안전 정렬을 평가하도록 설계된 포괄적인 벤치마크인 SafePro를 소개합니다. SafePro는 엄격한 반복적 생성 및 검토 과정을 통해 개발된, 안전 위험이 있는 다양한 전문 도메인에 걸친 고난도 작업 데이터셋을 포함합니다. 최첨단 AI 모델에 대한 우리의 평가는 심각한 안전 취약점을 드러내며, 전문적인 맥락에서 새로운 불안전한 행동들을 발견했습니다. 우리는 더 나아가 이러한 모델들이 복잡한 전문 작업을 수행할 때 불충분한 안전 판단력과 취약한 안전 정렬 상태를 모두 보인다는 것을 입증합니다. 또한, 우리는 이러한 시나리오에서 에이전트의 안전성을 향상시키기 위한 안전 완화 전략을 조사하고 고무적인 개선 사항을 관찰했습니다. 종합적으로, 우리의 연구 결과는 차세대 전문 AI 에이전트에 맞춘 견고한 안전 메커니즘의 시급한 필요성을 강조합니다.
Large language model-based agents are rapidly evolving from simple conversational assistants into autonomous systems capable of performing complex, professional-level tasks in various domains. While these advancements promise significant productivity gains, they also introduce critical safety risks that remain under-explored. Existing safety evaluations primarily focus on simple, daily assistance tasks, failing to capture the intricate decision-making processes and potential consequences of misaligned behaviors in professional settings. To address this gap, we introduce \textbf{SafePro}, a comprehensive benchmark designed to evaluate the safety alignment of AI agents performing professional activities. SafePro features a dataset of high-complexity tasks across diverse professional domains with safety risks, developed through a rigorous iterative creation and review process. Our evaluation of state-of-the-art AI models reveals significant safety vulnerabilities and uncovers new unsafe behaviors in professional contexts. We further show that these models exhibit both insufficient safety judgment and weak safety alignment when executing complex professional tasks. In addition, we investigate safety mitigation strategies for improving agent safety in these scenarios and observe encouraging improvements. Together, our findings highlight the urgent need for robust safety mechanisms tailored to the next generation of professional AI agents.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.