2601.05787v1 Jan 09, 2026 cs.AI

오프라인 학습에서 온라인 학습으로: 양방향 전문가-정책 동화 기법을 활용한 GUI 에이전트 성능 향상

From Off-Policy to On-Policy: Enhancing GUI Agents via Bi-level Expert-to-Policy Assimilation

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Zezhou Wang
Zezhou Wang
Citations: 10
h-index: 1
Xiaoyi Zhang
Xiaoyi Zhang
Citations: 136
h-index: 6
Zhuzhong Qian
Zhuzhong Qian
Citations: 2
h-index: 1
Yan Lu
Yan Lu
Citations: 32
h-index: 2

최근 비전-언어 모델은 데스크톱 및 브라우저를 제어하는 컴퓨터 사용 에이전트(CUA)로 점점 더 많이 활용되고 있습니다. 뛰어난 성능을 보이는 CUA는 계획 및 실행을 분리하는 프레임워크 기반 시스템인 반면, 스크린샷에서 액션으로의 직접적인 정책은 배포가 용이하지만 OSWorld-Verified와 같은 벤치마크에서 성능이 떨어집니다. OSWorld와 같은 GUI 데이터셋은 몇 백 개의 상호 작용 가능한 검증된 작업 및 환경만을 제공하며, 이러한 환경과의 상호 작용을 통해 전문가 트레이저리를 수집해야 하므로 데이터 확장이 어렵습니다. 따라서 본 연구에서는 검증 가능한 보상을 활용한 강화 학습(RLVR)이 제한된 수의 기존 전문가 트레이저리를 활용하여 엔드-투-엔드 정책을 학습하는 데 어떻게 가장 효과적으로 활용될 수 있는지 탐구합니다. 전문가 트레이저리를 온라인 RLVR에 무작정 통합하는 방식은 불안정합니다. 형식 변환 후에도 전문가 트레이저리는 학습자와의 구조적 불일치와 분포 변화를 나타냅니다. 본 연구에서는 BEPA(Bi-Level Expert-to-Policy Assimilation)라는 기법을 제안합니다. BEPA는 기본 정책 하에서 생성된 도달 가능한 트레이저리를 활용하여 정적 전문가 트레이저리를 정책에 맞게 조정하고, RLVR에서 사용되는 각 작업별로 동적으로 업데이트되는 캐시를 활용합니다. OSWorld-Verified 데이터셋에서 BEPA는 UITARS1.5-7B의 성공률을 22.87%에서 32.13%로 향상시키고, 별도로 분리된 데이터셋에서 5.74%에서 10.30%로 향상시켰으며, MMBench-GUI 및 Online-Mind2Web 데이터셋에서도 일관된 성능 향상을 보였습니다. 본 연구의 코드 및 데이터는 다음 GitHub 저장소에서 확인할 수 있습니다: https://github.com/LEON-gittech/Verl_GUI.git

Original Abstract

Vision-language models are increasingly deployed as computer-use agents (CUAs) that operate desktops and browsers. Top-performing CUAs are framework-based systems that decompose planning and execution, while end-to-end screenshot-to-action policies are easier to deploy but lag behind on benchmarks such as OSWorld-Verified. GUI datasets like OSWorld pose two bottlenecks: they expose only a few hundred interactive, verifiable tasks and environments, and expert trajectories must be gathered by interacting with these environments, making such data hard to scale. We therefore ask how reinforcement learning from verifiable rewards (RLVR) can best exploit a small pool of exist expert trajectories to train end-to-end policies. Naively mixing these off-policy traces into on-policy RLVR is brittle: even after format conversion, expert trajectories exhibit structural mismatch and distribution shift from the learner. We propose BEPA (Bi-Level Expert-to-Policy Assimilation), which turns static expert traces into policy-aligned guidance via self-rolled reachable trajectories under the base policy (LEVEL-1) and a per-task, dynamically updated cache used in RLVR (LEVEL-2). On OSWorld-Verified, BEPA improves UITARS1.5-7B success from 22.87% to 32.13% and raises a held-out split from 5.74% to 10.30%, with consistent gains on MMBench-GUI and Online-Mind2Web. Our code and data are available at: https://github.com/LEON-gittech/Verl_GUI.git

2 Citations
0 Influential
23 Altmetric
11.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!