2608.06065v1 Aug 06, 2026 cs.CV

다음 스크린을 예측하는 방법: 모바일 GUI 에이전트를 위한 게이트된 후방지식 증류

The Next Screenshot Knows: Gated Hindsight Distillation for Mobile GUI Agents

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Weiwei Li
Weiwei Li
Citations: 28
h-index: 3
Tong Chu
Tong Chu
Citations: 106
h-index: 3
Hengfu Yu
Hengfu Yu
Citations: 5
h-index: 1
Wen Li
Wen Li
Citations: 3
h-index: 1

GUI 에이전트는 일반적으로 성공적인 상호 작용 경로를 기반으로 오프라인에서 학습됩니다. 표준 훈련 방식은 각 경로를 현재 스크린과 상호 작용 기록을 기반으로 행동을 예측하는 '접두사-행동' 쌍으로 분해합니다. 이후의 관측 정보는 버려지므로, 특정 행동이 올바른 이유에 대한 근거가 사라집니다. 예를 들어, 텍스트 줄 바꿈 기능을 활성화하려면 에이전트가 '편집' 또는 '보기' 메뉴를 클릭해야 하지만, 메뉴가 열리기 전에는 이를 알 수 있는 정보가 없습니다. 이러한 근거 없이 표준 모방 학습은 모델이 올바른 추론을 수행할 기회를 거의 제공하지 않습니다. 이 문제를 해결하기 위해, 우리는 다음 스크린 정보를 훈련 과정에서 활용하는 '게이트된 후방지식 증류(Gated Hindsight Distillation, GHD)'를 제안합니다. 학생 모델은 관측 가능한 경로 접두사를 기반으로 예측하고, 공유 파라미터를 가진 교사 모델은 추가적으로 다음 스크린 정보를 활용하여 학생 모델의 정책에 따른 응답을 재평가합니다. 학생 모델이 실패할 경우에만 증류를 수행하며, 후방지식 기반 교사 모델은 시연된 행동을 복구합니다. GHD는 AndroidWorld 및 AndroidLab 환경에서 두 가지 시각-언어 모델 모두에서 GRPO보다 작업 성공률을 향상시킵니다. 코드와 체크포인트는 공개될 예정입니다.

Original Abstract

GUI agents are commonly trained offline from successful interaction trajectories. Standard training decomposes each trajectory into prefix-action pairs: the agent predicts an action from the current screen and interaction history, while the subsequent observation is discarded. This removes the rationale of why an action is correct: the evidence often appears only on the subsequent screen. For example, to enable Soft Wrap, the agent should click Edit or View, but nothing reveals this until the menu opens. Without such evidence, standard imitation gives the model little chance of ever sampling and thus learning the correct reasoning. To address this issue, we propose Gated Hindsight Distillation (GHD), which uses the next screenshot as privileged information during training. A student predicts from the observable trajectory prefix, while a parameter-sharing teacher additionally observes the next screenshot and re-scores the student's on-policy responses. We apply distillation only when the student fails and the hindsight-conditioned teacher recovers the demonstrated action. GHD improves task success over GRPO on AndroidWorld and AndroidLab across two vision-language models. The code and checkpoints will be made available.

0 Citations
0 Influential
1.5 Altmetric
7.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!