2607.29320v1 Jul 31, 2026 cs.AI

MAGA: 구조화된 액션 증류를 통한 다중 플랫폼 GUI 에이전트의 자기 융합

MAGA: Multi-Platform Self-Fusion of GUI Agents via Structured Action Distillation

Changhua Meng
Changhua Meng
Citations: 140
h-index: 3
Runze Li
Runze Li
Citations: 12
h-index: 2
Shuheng Shen
Shuheng Shen
Citations: 140
h-index: 3
Beitong Zhou
Beitong Zhou
Citations: 80
h-index: 2
Zhangxuan Gu
Zhangxuan Gu
Citations: 841
h-index: 12
Jiaxuan Chen
Jiaxuan Chen
Citations: 11
h-index: 2
Hang Yan
Hang Yan
Citations: 0
h-index: 0
Yusong Hu
Yusong Hu
Citations: 45
h-index: 3

대규모 언어 모델 기반의 그래픽 사용자 인터페이스(GUI) 에이전트는 모바일, 웹 및 데스크톱 환경에서 점점 더 많이 사용되고 있습니다. 그러나 기존 에이전트는 일반적으로 특정 도메인에 특화되어 있어 활용 범위와 사용자 경험을 제한합니다. 이에 따라, 우리는 다양한 환경에서 작동하는 단일 정책으로 전문 모델들을 통합하고자 합니다. 가중치 병합은 도메인별 전문가들을 직접적으로 결합하지만, 전문가 간의 의견 불일치가 있을 경우 실행 가능한 동작이 손상될 수 있습니다. 반면, 온-정책 증류(OPD)는 충돌하는 교사 지시를 피하지만, 여전히 모든 응답 토큰을 동일하게 취급하여 증류 과정에서 에이전트와 환경 사이의 인터페이스인 액션 토큰의 중요성을 간과합니다. 이러한 문제를 해결하기 위해, 우리는 구조화된 액션을 기반으로 학습 신호를 재분배하는 MAGA를 제안합니다. MAGA는 생성된 액션의 정확성에 따라 불필요하거나 잘못된 증류 신호를 억제하고, 오류가 있는 액션에 대한 학습을 집중시킵니다. 또한, 학습 단계에서만 사용되는 힌트를 통해 학생 모델의 입력은 변경하지 않고 도메인별 교사가 제공하는 감독 신호를 최적화합니다. 두 가지 모델 크기에서 MAGA는 가장 높은 평균 성공률을 달성했으며, 8B 모델에서는 가장 강력한 기준 모델보다 2.0% 더 우수한 성능을 보였고, 전체적으로는 교사 모델과 거의 동일한 수준의 성능을 나타냈습니다.

Original Abstract

Graphical user interface (GUI) agents based on large language models are increasingly deployed across mobile, web, and desktop environments. However, existing agents are typically domain-specific, limiting the deployment and user experience. This motivates the consolidation of specialized models into a single cross-environment policy. Weight merging directly merges domain-specific experts but can corrupt executable actions under expert disagreement, while on-policy distillation (OPD) avoids conflicting teacher supervision yet still treats all response tokens equally during distillation, ignoring that action tokens are the only interface between the environment and the agent. To address this, We introduce MAGA that re-allocates training signal according to the structured action. Based on the correctness of the generated action, it suppresses unnecessary or invalid distillation signals and focuses learning on erroneous actions. Besides, a training-only hint optimizes the supervision signal provided by domain-specific teachers without changing the student input. Across two model scales, MAGA achieves the highest mean success rate, outperforming the strongest baseline by 2.0% at 8B and achieves almost the same average performance with teachers.

0 Citations
0 Influential
6 Altmetric
30.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!