2605.29532v1 May 28, 2026 cs.SE

GUITestScape: 탐색적 GUI 테스트를 위한 개방형 평가 연구

GUITestScape: Towards Open-set Evaluation on Exploratory GUI Testing

Xiaoyi Chen
Xiaoyi Chen
Citations: 11
h-index: 2
Jitao Sang
Jitao Sang
Citations: 37
h-index: 2
Y. Zhang
Y. Zhang
Citations: 171
h-index: 7
Xing Song
Xing Song
Citations: 2
h-index: 1
Yifei Gao
Yifei Gao
Citations: 7
h-index: 2
Yang Xu
Yang Xu
Citations: 3,263
h-index: 4

탐색적 GUI 테스트는 MLLM 에이전트에게 특히 어려운 과제입니다. 미리 정의된 테스트 스크립트 없이, 에이전트는 애플리케이션을 자율적으로 탐색하고 자체적인 상호 작용을 통해 결함을 발견해야 합니다. 그러나 현재 평가 방법은 두 가지 측면에서 한계점을 가지고 있습니다. 첫째, 기존 벤치마크는 거의 모든 경우 상호 작용 결함에만 초점을 맞추고 있으며, 화면 표시 관련 결함은 평가 대상에서 제외됩니다. 둘째, 평가 프로토콜은 미리 정의된 결함 주석에 의존하여 테스트 프로세스를 단일 종료 상태 판단으로 축소시키며, 질적으로 구별되는 실패 모드를 혼동시킵니다. 이러한 문제점을 해결하기 위해, 우리는 61개의 실제 Android 애플리케이션과 상호 작용 및 화면 표시 유형의 508개 미리 설정된 결함을 포괄하는 인터랙티브 벤치마크인 GUITestScape를 제시하고, 에이전트의 테스트 경로를 독립적으로 진단 가능한 기능으로 분해하는 개방형 평가 도구인 GUIJudge를 소개합니다. 실험 결과는 GUIJudge가 미리 정의된 주석을 넘어 프로세스 기반의 신뢰할 수 있는 평가를 수행하며, 모든 기본 모델보다 훨씬 뛰어난 성능을 보임을 보여줍니다. GUITestScape에 대한 벤치마킹은 또한 기존 모델이 두 가지 유형의 결함 모두에서 여전히 탐지 능력이 가장 중요한 병목 지점이며, GUIJudge의 검증기를 기존 에이전트에 통합하면 재학습 없이도 탐지 성능을 크게 향상시킬 수 있음을 보여줍니다.

Original Abstract

Exploratory GUI testing is a particularly demanding setting for MLLM agents: without predefined test scripts, an agent must autonomously navigate an application and discover defects through its own interaction. However, current evaluation falls short on two fronts. First, existing benchmarks focus almost exclusively on interaction defects, leaving display defects outside the evaluation frame. Second, evaluation protocols are bound to predefined defect annotations, collapsing the testing process into a single end-state judgment that conflates qualitatively distinct failure modes. To address these challenges, we present GUITestScape, an interactive benchmark covering 61 real-world Android applications and 508 preset defects spanning interaction and display types, and introduce GUIJudge, an open-set evaluator that decomposes an agent's testing trajectory into independently diagnosable capabilities. Experimental results demonstrate that GUIJudge achieves reliable process-aware evaluation beyond predefined annotations, substantially outperforming all baselines. Benchmarking on GUITestScape further reveals that detection remains the critical bottleneck for existing models across both defect types, and that integrating GUIJudge's verifiers into existing agents significantly boosts their detection performance without retraining.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!