2607.25904v1 Jul 28, 2026 cs.AI

인터랙티브 보상 에이전트: 환경-상태 검증을 통한 GUI 작업 평가

Interactive Reward Agent: GUI Task Evaluation via Environment-State Verification

Zhi Gao
Zhi Gao
Citations: 419
h-index: 8
Chenrui Shi
Chenrui Shi
Citations: 129
h-index: 7
Zirui Shang
Zirui Shang
Citations: 91
h-index: 5
Yang Liu
Yang Liu
Citations: 91
h-index: 1
Ruining Feng
Ruining Feng
Citations: 17
h-index: 3
Lifeng Fan
Lifeng Fan
Citations: 0
h-index: 0
Yuwei Wu
Yuwei Wu
Citations: 1,592
h-index: 20
Che Sun
Che Sun
Citations: 373
h-index: 8

GUI 작업 평가는 GUI 에이전트가 사용자 지시를 성공적으로 완료했는지 판단하는 것을 목표로 합니다. 자동화된 GUI 작업 평가는 테스트 시간 스케일링 및 사후 훈련에 대한 보상 신호로 활용될 수 있기 때문에 많은 관심을 받고 있습니다. 그러나 신뢰할 수 있는 GUI 작업 평가는 여전히 어려운 과제이며, 종종 시스템 구성, 파일 데이터, 애플리케이션 설정과 같은 환경 상태 정보에 접근해야 하지만, 실행 경로의 스크린샷만으로는 충분하지 않은 경우가 많습니다. 본 논문에서는 사후 실행 환경으로부터 증거를 획득하고 검증하기 위한 프레임워크인 제안-검증(propose-then-verify) 기반의 인터랙티브 보상 에이전트(IRA)를 제안합니다. IRA는 주어진 작업 지시와 GUI 에이전트 실행 후의 GUI 환경을 입력으로 받아, 먼저 작업 완료 조건을 제안하고 시스템 도구, 애플리케이션 도구 및 GUI 도구를 호출하여 이를 검증합니다. 이러한 설계는 가시적인 인터페이스뿐만 아니라 환경 상태로부터 얻은 증거를 활용하는 상호작용적인 프로세스를 통해 이루어집니다. 또한, 10개의 Ubuntu 데스크톱 애플리케이션 범주에 걸쳐 321개의 GUI 작업 경로를 포함하는 벤치마크인 GUI-RewardBench를 소개합니다. 실험 결과, IRA는 GUI-RewardBench에서 86.9%의 정확도를 달성하여 기존 평가 기준을 능가했습니다. 또한, IRA를 GUI 에이전트 강화 학습에 적용한 결과, OSWorld 성공률이 34.0%로 향상되었으며, 이는 IRA가 GUI 에이전트 훈련에 효과적인 보상 신호를 제공할 수 있음을 보여줍니다.

Original Abstract

Graphical user interface task evaluation aims to determine whether a GUI agent has successfully completed a user instruction. Automated GUI task evaluation has received increasing attention because the evaluation results can serve as reward signals for both test-time scaling and post-training. However, reliable GUI task evaluation remains challenging because the judgments often require access to environment states, such as system configurations, file data, and application settings, beyond the screenshots of execution trajectories. In this paper, we propose an interactive reward agent (IRA) based on a propose-then-verify framework to acquire and verify evidence from the post-execution environment. Given a task instruction and a GUI environment after the GUI agent execution, IRA first proposes the task completion conditions and then verifies them by invoking system tools, application tools, and GUI tools. This design combines evidence from both visible interfaces and the environment state in an interactive process. We further introduce GUI-RewardBench, a benchmark of 321 GUI task trajectories spanning 10 Ubuntu desktop application categories. Experiments show that IRA achieves 86.9% accuracy on GUI-RewardBench, outperforming existing evaluator baselines. We further apply IRA to reinforcement learning of GUI agents, achieving a 34.0% OSWorld success rate, which demonstrates that IRA can provide effective reward signals for training GUI agents.

0 Citations
0 Influential
10 Altmetric
50.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!