StepReflect: 모바일 GUI 에이전트를 위한 구조화된 UI 전환 반영 기술
StepReflect: Structured UI Transition Reflection for Mobile GUI Agents
자율적인 모바일 GUI 에이전트는 안정적인 장기 실행을 위해 정확한 행동 반영 기능을 필요로 합니다. 기존 방법은 각 행동 후에 광범위하고 다중 양식의 추론에 의존하는데, 이는 비용이 많이 들고 GUI 상태 전환의 구조화된 특성과 잘 맞지 않습니다. 본 논문에서는 StepReflect를 제안합니다. StepReflect는 명시적인 전환 사양과 쌍을 이루는 시각적 증거에 기반하여 각 단계별 GUI 반영을 지도 학습 기반의 구조화된 예측 문제로 정의합니다. StepReflect는 지도 미세 조정, 교사-학생 증류, 그리고 선호도 및 보상 기반 개선을 결합한 단계별 파이프라인을 통해 훈련됩니다. 오프라인 실험에서, 결과적으로 생성된 80억 개의 매개변수를 가진 모델은 AndroidWorld 데이터셋에서 전환 수준의 정확도가 82.16%로 나타났으며, 동일한 구조화된 입력 조건 하에서 사전 학습된 GPT-5.2보다 11.83%p 더 높은 성능을 보였습니다. 온라인 실험에서는 M3A, Agent-SAMA, MAI-UI-8B 및 Seed-2.0-Pro 환경에서 StepReflect는 네 가지 에이전트 구성 중 세 가지에서 더 높은 작업 성공률을 달성했으며, 나머지 하나의 구성에서는 GPT-5.2 Reflection Agent와 1개의 성공적인 작업을 차이를 보였습니다. 또한 StepReflect는 모든 네 가지 구성에서 GPT 기반 반영 방식보다 유료 API 사용 비용을 절감했습니다. 이러한 결과는 StepReflect가 장기 실행이 필요한 모바일 GUI 에이전트를 위한 실용적이고 로컬 환경에 배포 가능한 반복적인 최첨단 모델 반영의 대안임을 입증합니다.
Autonomous mobile GUI agents require accurate action reflection for reliable long-horizon execution. Existing approaches rely on open-ended multimodal reasoning after each action, which is costly and poorly matched to the structured nature of GUI state transitions. We propose StepReflect, which formulates per-step GUI reflection as supervised structured prediction conditioned on explicit transition specifications and paired visual evidence. StepReflect is trained through a staged pipeline combining supervised fine-tuning, teacher-student distillation, and preference- and reward-based refinement. Offline, the resulting 8B model achieves 82.16% transition-level accuracy on AndroidWorld, exceeding zero-shot GPT-5.2 by 11.83 percentage points under the same structured input. Online, across M3A, Agent-SAMA, MAI-UI-8B, and Seed-2.0-Pro, StepReflect achieves higher task success in three of four agent configurations and remains within one successful task of the GPT-5.2 Reflection Agent in the fourth. It also reduces paid API charges relative to GPT-based reflection in all four configurations. These results establish StepReflect as a practical, locally deployable alternative to repeated frontier-model reflection for long-horizon mobile GUI agents.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.