2605.25707v1 May 25, 2026 cs.AI

AgentHijack: 일반적인 환경 오류에 대한 컴퓨터 사용 에이전트의 견고성 평가

AgentHijack: Benchmarking Computer Use Agent Robustness to Common Environment Corruptions

Xia Hu
Xia Hu
Citations: 31
h-index: 3
Jianing Zhu
Jianing Zhu
Citations: 1,112
h-index: 12
Jingwei Sun
Jingwei Sun
Citations: 100
h-index: 5
Yuanyi Li
Yuanyi Li
Citations: 1
h-index: 1
Tongliang Liu
Tongliang Liu
Citations: 1,930
h-index: 21
Bo Han
Bo Han
Citations: 252
h-index: 9

다중 모드 대규모 언어 모델(MLLM)로 구동되는 자율형 컴퓨터 사용 에이전트는 복잡한 디지털 워크플로우를 수행하는 데 유용한 도구로 부상하고 있습니다. 그러나 실제 실행 환경은 이상적이지 않으며, 팝업 창, 해상도 변경, 경쟁 애플리케이션 등이 에이전트의 인식 및 제어를 방해하는 경우가 많습니다. 본 연구에서는 일반적인 오류 상황에서 컴퓨터 사용 에이전트의 견고성을 평가하기 위한 벤치마크인 AgentHijack을 소개합니다. AgentHijack은 동적 환경으로 인한 불확실성이 의도적인 적대 행위 없이 실행 흐름을 방해하는 상황을 재현하기 위해 9가지 구성 가능한 일반적인 오류를 포함합니다. MLLM 기반 에이전트를 활용한 다양한 데스크톱 작업을 평가한 결과, 사소한 오류라도 상당한 성능 저하를 초래할 수 있다는 것을 발견했습니다. 이는 에이전트의 취약성을 강조하며 견고성 평가의 필요성을 뒷받침합니다. 이후, 액션 생성기, 향상된 근거 능력 및 행동 요약 및 환경 점검 기능을 갖춘 감시자를 통합한 프레임워크인 AgentHijack-Agent를 제안하고, 광범위한 실험을 통해 그 효과를 검증했습니다. 저희의 코드, 환경, 기준 모델 및 데이터는 다음 링크에서 공개적으로 이용할 수 있습니다: https://AgentHijack.github.io.

Original Abstract

Autonomous computer use agents that powered by multimodal large language models (MLLMs) are emerging as capable assistants for completing complex digital workflows. However, real-world execution environments are far from ideal: pop-ups, resolution changes, and competing applications frequently interfere with agent perception and control. We introduce AgentHijack, a benchmark designed to evaluate the robustness of computer-use agents under common corruptions, where the uncertainties in dynamic environment disrupt the execution flow without direct adversarial intent. Specifically, AgentHijack introduces 9 configurable common corruptions to replicate realistic imperfect scenarios. We evaluate a variety of desktop tasks that utilize MLLM-based agents and discover that even minor instances of corruption can result in substantial performance degradation, which emphasizes the fragility of agents and underscores the necessity of robustness evaluation. Afterward, we propose AgentHijack-Agent, a framework that integrates an action generator with enhanced grounding capabilities and an onlooker responsible for behavior summarization and environment checking. Extensive experiments validate its effectiveness. Our code, environment, baseline models and data are publicly available at: https://AgentHijack.github.io.

2 Citations
0 Influential
10.5 Altmetric
54.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!