JANUS: 장기적인 안전을 위한 잠재적 위험 예측
JANUS: Foreseeing Latent Risk for Long-Horizon Agent Safety
에이전트 안전은 콘텐츠 검열에서 벗어나 도구를 사용하는 에이전트가 작동하기 전에 운영상의 실패를 예방하는 방향으로 발전하고 있습니다. 본 논문에서는 장기적인 에이전트 안전을 위한 예측 중심 프레임워크인 Janus를 제안합니다. Janus는 부분적인 경로로부터 발생하는 지연된 위험을 예측하도록 설계된 가드(guard) 모델을 학습시킵니다. Janus는 다중 에이전트 시뮬레이션을 통해 다양한 에이전트 경로를 합성하고, 안전과 관련된 미래를 예측하는 '예측 태스크'와 관찰된 초기 상태 및 예상되는 미래를 기반으로 안전 여부를 판단하는 '판단 태스크'라는 두 가지 연결된 태스크를 통해 공유 정책을 학습합니다. 이 두 태스크는 CoAA-RL을 사용하여 함께 최적화되며, 하위 단계의 안전 판단에 유용한 예측에 대해 보상을 제공합니다. 결과적으로 생성된 가드 모델인 Vanguard는 실행 전에 위험한 행동을 차단합니다. 네 가지 에이전트 안전 벤치마크에서 Vanguard는 기존 가드 모델보다 평균적으로 보호율을 15.9% 향상시키고, 동시에 정상적인 태스크 완료율을 5.1% 증가시킵니다.
Agent safety is moving from content moderation toward preventing operational failures before tool-using agents act. We propose Janus, a foresight-oriented framework for long-horizon agent safety that trains guards to anticipate delayed risks from partial trajectories. Janus synthesizes diverse agent trajectories via multi-agent simulation and learns a shared policy with two coupled tasks: an anticipation task that forecasts safety-relevant futures and an adjudication task that decides safety from both the observed prefix and anticipated future. The two tasks are jointly optimized with CoAA-RL, which rewards forecasts by their utility for downstream safety judgment. The resulting guard model, Vanguard, blocks unsafe actions before execution. Across four agent-safety benchmarks, Vanguard improves average protection by 15.9 percentage points over baseline guards while increasing benign task completion by 5.1 percentage points.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.