2606.11078v1 Jun 09, 2026 cs.AI

컴퓨터 사용 에이전트에 대한 시계열 정보 기반의 시각적 판단 비평기

A History-Aware Visually Grounded Critic for Computer Use Agents

Elias Stengel-Eskin
Elias Stengel-Eskin
Citations: 1,189
h-index: 19
Mohit Bansal
Mohit Bansal
Citations: 1,149
h-index: 20
Archiki Prasad
Archiki Prasad
UNC Chapel Hill
Citations: 941
h-index: 16
Sambit Sahu
Sambit Sahu
Citations: 87
h-index: 5
J. Chen
J. Chen
Citations: 988
h-index: 10
Supriyo Chakraborty
Supriyo Chakraborty
Citations: 11
h-index: 2
Zaid Khan
Zaid Khan
Citations: 243
h-index: 6
Kartik Balasubramaniam
Kartik Balasubramaniam
Citations: 2
h-index: 1
Jaewoo Lee
Jaewoo Lee
Citations: 3
h-index: 1
Hyunji Lee
Hyunji Lee
Citations: 21
h-index: 2

복잡한 그래픽 사용자 인터페이스(GUI) 환경에서 성능 향상을 위해, 실행 전 행동 평가를 통해 작동하는 다양한 형태의 컴퓨터 사용 에이전트(CUA)용 비평 모델들이 개발되어 왔습니다. 그러나 기존의 비평 모델들은 다음과 같은 두 가지 주요 한계점을 가지고 있습니다. (1) 주로 단기적인 의사 결정 루프에 집중하며, 이전 행동을 잊는 경향이 있고, (2) 잘못된 행동을 감지하는 데 필요한 시각적 정보를 활용하지 못합니다 (예: UI 요소를 잘못 클릭). 이러한 문제점을 해결하기 위해, 우리는 실제 GUI 환경의 데이터를 사용하여 과거 상호작용을 간결한 형태로 추상화하고, 시각적인 정보를 기반으로 행동을 평가하는 다중 모달 비평 모델인 HiViG라는 시계열 정보 기반의 시각적 판단 테스트 시간 프레임워크를 제안합니다. 테스트 시간에 HiViG는 정책 결정 루프에 비평 모델을 통합하여 정책이 달성한 결과를 요약하는 '매크로 액션 히스토리'와 현재 스크린샷과 실제 실행 좌표를 비교하여 오류를 사전에 감지하는 '시각적으로 판단된 비평'을 제공합니다. 웹, 모바일 및 데스크톱 벤치마크에서 HiViG는 기존의 숫자 기반 또는 언어 기반 비평 모델보다 일관되게 우수한 성능을 보이며, Qwen3-VL-32B 모델의 경우 평균 성공률을 가장 강력한 기준 모델보다 5.8% 향상시키고, Gemini-3-Flash 모델의 경우 9.0% 향상시키는 결과를 보여주었습니다. 또한 HiViG는 다양한 플랫폼에서 뛰어난 일반화 성능을 보입니다. 추가 분석 결과, 매크로 액션 히스토리는 단기적인 계획 문제를 완화하고, 시각적으로 판단된 비평은 실행 오류를 줄이는 데 중요한 역할을 하며, 이러한 요소들이 긴 시간 동안 지속되는 GUI 작업에서의 테스트 시간 성능 향상에 필수적임을 확인했습니다.

Original Abstract

Various test-time interventions for Computer Use Agents (CUAs), including critic models, have been developed to improve performance through pre-execution action evaluation in complex Graphical User Interface (GUI) environments. However, existing critics suffer from two key limitations: they (1) focus primarily on short-sighted decision loops (e.g., forgetting earlier actions) and (2) lack the visual grounding needed to detect flawed actions (e.g., clicking wrong UI elements). To address these, we introduce HiViG, a History-aware Visually Grounded test-time framework, built around a multimodal critic trained on real GUI trajectories to abstract past interactions into a compact record and to evaluate actions with visual grounding. At test time, HiViG integrates the critic into the policy decision loop to provide macro-action history, which summarizes the policy's completed achievements, and visually grounded critique, which verifies raw execution coordinates against the current screenshot to intercept errors before execution. Across web, mobile, and desktop benchmarks, HiViG consistently outperforms existing scalar and verbal critics, improving average success rates over the strongest baseline by 5.8% for Qwen3-VL-32B and 9.0% for Gemini-3-Flash, and demonstrates strong cross-platform generalization. Ablations show that macro-action history mitigates short-sighted planning and visually grounded critique reduces execution errors, with both components being critical for test-time scaling in long-horizon GUI tasks.

0 Citations
0 Influential
10 Altmetric
50.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!