Teach-and-Repeat: 모바일 화면 데모에서 운영 지식을 정확하게 추출하여 GUI 에이전트의 성능 향상
Teach-and-Repeat: Accurately Extracting Operational Knowledge from Mobile Screen Demonstrations to Empower GUI Agents
모바일 기기에서의 디지털 환경 이해는 정적인 UI 인식에서 동적인 동작 이해로 전환되고 있습니다. 이러한 능력은 모델이 시각적 상태 변화를 운영 지식으로 변환할 수 있도록 하며, 여기서 운영 지식은 동작 유형, 대상 UI 요소, 텍스트 인자 및 실행 순서를 설명하는 짧은 자연어 문장을 의미합니다. 그러나 애플리케이션 간의 매우 다양하고 이질적인 UI 디자인으로 인해 기존의 시각-언어 모델(VLM)은 이러한 기본 작업을 정확하게 추론하는 데 어려움을 겪습니다. 이러한 격차를 해소하기 위해, 우리는 모바일 화면 트래jectory를 단계별 운영 지식으로 변환하도록 설계된 핵심 모델인 Teach VLM을 소개합니다. Teach VLM은 데모 비디오에서 작업 관련 주요 프레임을 추출하고 분석하여 작동 방식을 파악합니다. 제한적인 학습 데이터 문제를 해결하기 위해, 확장 가능한 데이터 확보를 위한 체계적인 데이터 플라이휠 시스템을 개발했습니다. 또한, 세분화된 평가를 위한 새로운 중국 모바일 화면 Teach 벤치마크를 제시합니다. Teach VLM을 기반으로, 생성된 운영 지식을 해석 가능한 절차적 참조로 활용하여 다운스트림의 화면 기반 실행 에이전트를 안내하는 Teach-and-Repeat 패러다임을 제안합니다. 광범위한 실험 결과는 Teach VLM이 강력한 VLM 기준 모델보다 훨씬 뛰어난 성능을 보이며, 작업 의미 예측에서 최첨단 성능을 달성한다는 것을 보여줍니다. 또한, Android World 환경에서의 실험은 우리 패러다임이 다운스트림 에이전트의 작업 성공률을 지속적으로 향상시킨다는 것을 입증합니다. Teach VLM과 Teach-and-Repeat 패러다임은 원시 데모에서 재사용 가능한 작업 자동화로 이어지는 실용적인 방법을 제공합니다.
Understanding the digital world on mobile devices is shifting from static UI perception to dynamic action comprehension. This capability enables models to convert visual state transitions into operational knowledge, defined as short natural-language sentences that describe action types, target UI elements, textual arguments, and execution orders. However, due to the highly diverse and heterogeneous UI designs across applications, existing vision-language models (VLMs) struggle to accurately infer these underlying operations. To bridge this gap, we introduce Teach VLM, a core model designed to translate mobile screen trajectories into step-wise operational knowledge by extracting and analyzing operation-related keyframes from demonstration videos. To address the scarcity of aligned training data, we develop a systematic data flywheel for scalable data acquisition. We further introduce a novel Chinese Mobile Screen Teach Benchmark for fine-grained evaluation. Building upon Teach VLM, we propose the Teach-and-Repeat paradigm, where the generated operational knowledge serves as an interpretable procedural reference to guide downstream screen-based execution agents. Extensive evaluations demonstrate that Teach VLM significantly outperforms strong VLM baselines, achieving state-of-the-art performance in operation semantics prediction. Furthermore, experiments in Android World show that our paradigm yields consistent Task Success Rate improvements for downstream agents. Together, Teach VLM and the Teach-and-Repeat paradigm offer a practical pathway from raw demonstrations to reusable task automation.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.