2602.01725v1 Feb 02, 2026 cs.CL

SafePred: 월드 모델 기반 컴퓨터 사용 에이전트를 위한 예측적 안전 장치

SafePred: A Predictive Guardrail for Computer-Using Agents via World Models

Yurun Chen
Yurun Chen
Citations: 185
h-index: 7
Zeyi Liao
Zeyi Liao
Citations: 925
h-index: 14
P. Yin
P. Yin
Citations: 28
h-index: 3
Taotao Xie
Taotao Xie
Citations: 7
h-index: 2
Keting Yin
Keting Yin
Citations: 140
h-index: 6
Shengyu Zhang
Shengyu Zhang
Citations: 85
h-index: 3

컴퓨터 사용 에이전트(CUA)가 복잡한 실제 환경에 광범위하게 배치됨에 따라, 흔히 발생하는 장기적인 위험은 심각하고 되돌릴 수 없는 결과를 초래할 수 있습니다. 대부분의 기존 CUA 안전 장치는 반응적인 접근 방식을 채택하여 에이전트의 행동을 현재 관찰 공간 내에서만 제한합니다. 이러한 안전 장치는 즉각적인 단기적인 위험(예: 피싱 링크 클릭)을 방지할 수 있지만, 장기적인 위험을 사전에 예방할 수 없습니다. 겉으로는 합리적인 행동이라도 지연되어 발생하는 고위험 결과를 초래할 수 있습니다(예: 로그 삭제는 향후 감사 추적을 불가능하게 함). 반응적인 안전 장치는 현재 관찰 공간 내에서 이러한 위험을 식별할 수 없습니다. 이러한 한계를 해결하기 위해, 우리는 예측적인 안전 장치 접근 방식을 제안하며, 이는 예측된 미래 위험을 현재 결정과 일치시키는 것을 핵심 아이디어로 합니다. 이러한 접근 방식을 바탕으로, 우리는 CUA를 위한 예측적인 안전 장치 프레임워크인 SafePred를 제시합니다. SafePred는 안전한 에이전트 행동을 보장하기 위해 위험-결정 루프를 구축합니다. SafePred는 다음 두 가지 핵심 기능을 제공합니다. (1) 단기 및 장기 위험 예측: SafePred는 안전 정책을 위험 예측의 기반으로 사용하여, 월드 모델의 예측 기능을 활용하여 단기 및 장기 위험에 대한 의미론적 표현을 생성하고, 고위험 상태로 이어지는 행동을 식별하고 제거합니다. (2) 의사 결정 최적화: 예측된 위험을 단계별 개입 및 작업 수준의 재계획을 통해 실행 가능한 안전한 의사 결정 지침으로 변환합니다. 광범위한 실험 결과, SafePred는 고위험 행동을 크게 줄이고, 반응적인 기준선과 비교하여 97.6% 이상의 안전 성능을 달성하고, 작업 유용성을 최대 21.4%까지 향상시키는 것으로 나타났습니다.

Original Abstract

With the widespread deployment of Computer-using Agents (CUAs) in complex real-world environments, prevalent long-term risks often lead to severe and irreversible consequences. Most existing guardrails for CUAs adopt a reactive approach, constraining agent behavior only within the current observation space. While these guardrails can prevent immediate short-term risks (e.g., clicking on a phishing link), they cannot proactively avoid long-term risks: seemingly reasonable actions can lead to high-risk consequences that emerge with a delay (e.g., cleaning logs leads to future audits being untraceable), which reactive guardrails cannot identify within the current observation space. To address these limitations, we propose a predictive guardrail approach, with the core idea of aligning predicted future risks with current decisions. Based on this approach, we present SafePred, a predictive guardrail framework for CUAs that establishes a risk-to-decision loop to ensure safe agent behavior. SafePred supports two key abilities: (1) Short- and long-term risk prediction: by using safety policies as the basis for risk prediction, SafePred leverages the prediction capability of the world model to generate semantic representations of both short-term and long-term risks, thereby identifying and pruning actions that lead to high-risk states; (2) Decision optimization: translating predicted risks into actionable safe decision guidances through step-level interventions and task-level re-planning. Extensive experiments show that SafePred significantly reduces high-risk behaviors, achieving over 97.6% safety performance and improving task utility by up to 21.4% compared with reactive baselines.

4 Citations
0 Influential
7 Altmetric
39.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!