MIRAGE: 사용자 생성 콘텐츠를 이용한 모바일 GUI 에이전트에 대한 상황 인식 프롬프트 주입 공격
MIRAGE: Context-Aware Prompt Injection against Mobile GUI Agents via User-Generated Content
본 논문에서는 비전-언어 모델(VLM)에 의해 구동되는 모바일 그래픽 사용자 인터페이스(GUI) 에이전트가 화면을 렌더링된 픽셀로 인식하고, 보이는 것들을 기반으로 행동을 선택하기 때문에 신뢰할 수 있는 인터페이스 요소와 사용자 생성 콘텐츠를 명확하게 구분하지 못한다는 점에 주목합니다. 우리는 MIRAGE (Mobile Injection of Realistic Adversarial GUI Examples)라는 파이프라인을 제시합니다. 이 파이프라인은 무해한 모바일 스크린샷을, 에이전트, 애플리케이션 또는 운영 체제를 수정하지 않고도 공격자가 제어하는 텍스트를 일반적인 사용자 생성 콘텐츠 영역에 삽입하여 프롬프트 주입 샘플로 변환합니다. MIRAGE는 세 단계로 작동합니다. 첫째, Localizer가 스크린샷에서 사용자가 제어할 수 있는 영역을 식별합니다. 둘째, Generator는 상황 인식 페이로드를 합성하고 애플리케이션의 고유한 스타일로 렌더링합니다. 마지막으로, Curator는 현실성을 조정하고 샘플들을 애플리케이션, 영역 유형 및 공격 의도에 따라 균형 있게 구성합니다. 주요 과제는 주입된 스크린샷이 여전히 진정한 사용자 콘텐츠와 시각적으로 구별되지 않아야 한다는 것입니다. 동시에 에이전트를 오도해야 합니다. 우리는 이 문제를 해결하기 위해 영향 범위, 현실성 및 분포 균형을 제어하는 단계를 분리합니다. 10개의 애플리케이션과 11가지 공격 의도를 포괄하는 1,111개 샘플의 벤치마크에서, 평가된 모든 5개의 VLM 에이전트가 취약하며, 공격 성공률은 23%에서 30%에 이릅니다. 또한 MIRAGE는 기존의 가장 강력한 공격보다 인간의 현실성 평가에서 더 높은 점수를 받았습니다 (5점 만점에 3.02 대 2.52). 더욱이 샘플별 현실성과 공격 성공률은 상관관계가 없으므로, 시각적 품질 필터만으로는 이 위협을 효과적으로 방어할 수 없습니다.
Mobile graphical user interface (GUI) agents driven by vision-language models (VLMs) perceive the screen as rendered pixels and choose actions from what they see, so they cannot reliably separate trusted interface elements from user-generated content. We present MIRAGE (Mobile Injection of Realistic Adversarial GUI Examples), a pipeline that turns benign mobile screenshots into prompt-injection samples by placing attacker-controlled text into ordinary user-generated content regions, without modifying the agent, the application, or the operating system. MIRAGE operates in three stages: a Localizer identifies user-controllable regions on the screenshot, a Generator synthesises context-aware payloads and renders them in the application's native style, and a Curator moderates realism and balances the samples across applications, region types, and attack intents. A key challenge is that an injected screenshot must stay visually indistinguishable from genuine user content while still diverting the agent; we address this by separating the stages that control reach, realism, and distributional balance. On a 1,111-sample benchmark spanning ten applications and eleven attack intents, all five evaluated VLM agents are vulnerable, with attack success rates of 23%-30%, and MIRAGE scores higher on human realism ratings than the strongest prior attack (3.02 versus 2.52 out of 5). We further find that per-sample realism and attack success are uncorrelated, so visual-quality filtering alone cannot reliably defend against this threat.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.