GUI 에이전트: 사용자 민감 정보가 포함된 화면에 대한 안내 기반 탐색
GUI agent: Guided Exploration of User-Sensitive Screens
LLM(대규모 언어 모델) 에이전트는 점점 더 많은 분야에서 사용자를 위한 작업을 자동화하는 데 활용되고 있으며, 특히 개방형 GUI 환경에서 그 활용도가 높아지고 있습니다. 이러한 에이전트들은 필연적으로 사용자 민감 정보가 포함된 화면에 직면하게 되며, 이 경우 작업 실행의 제어를 사용자에게 넘겨주는 것이 매우 중요하거나 필수적입니다. 최첨단 LLM 기반 에이전트는 일반적으로 안전상의 문제를 고려하지 않고 작업을 완료하도록 설계되어 있어, 실제 환경에서의 배포가 어렵고 신뢰성에 부정적인 영향을 미칩니다. 따라서, 사용자 민감 상태를 식별하고 분류하며, 사용자 민감 쿼리를 정의하는 것이 매우 중요합니다. 본 데이터셋은 엔지니어들이 중요한 상황에서 사용자에게 작업 제어를 위임하도록 유도하기 위한 목적으로 사용됩니다. 본 논문에서는 탐색 에이전트를 개발하여, 제시된 하나의 작업을 기반으로 쿼리 공간을 체계적으로 탐색하면서 GUI 환경 내에서 실행될 경우 사용자 민감 상태로 이어질 수 있는 쿼리를 식별합니다.
LLM agents are increasingly being used to automate tasks for users within an open GUI environment. They inevitably encounter screens containing user-sensitive information, for which takeover of task execution by the user is highly desirable or even necessary. State-of-the-art LLM-driven agents are usually fine-tuned to complete tasks regardless of the safety implications of their actions. This makes their real-world deployment difficult and adversely affects the reliability. Therefore, it is crucial to identify and categorize user-sensitive states and define user-sensitive queries. This dataset would be to engineers to recognize and request handover to the user in critical scenarios. This short paper develops an explorer agent that systematically explores the query space starting from one demonstrated task to identify queries that, if executed, would lead to user-sensitive states in a GUI environment.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.