PhySE: 실시간 AR-LLM 기반 사회 공학 공격을 위한 심리학적 프레임워크
PhySE: A Psychological Framework for Real-Time AR-LLM Social Engineering Attacks
AR-LLM 기반 사회 공학(AR-LLM-SE) 공격(예: SEAR)은 실제 사회적 상호작용에 상당한 위험을 초래합니다. 이러한 공격에서 악의적인 행위자는 증강 현실(AR) 글래스를 사용하여 대상의 시각 및 음성 데이터를 수집합니다. 이후, 대규모 언어 모델(LLM)은 이 데이터를 분석하여 개인을 식별하고 상세한 사회적 프로필을 생성합니다. 그런 다음, LLM 기반 에이전트는 사회 공학 전략을 활용하여 실시간 대화 제안을 제공함으로써 대상의 신뢰를 얻고 궁극적으로 피싱 또는 기타 악의적인 행위를 수행합니다. AR-LLM-SE의 잠재력에도 불구하고, 실제 적용에는 두 가지 주요 장애물이 존재합니다. (1) 초기 개인화 문제: 현재의 검색 증강 생성(RAG) 방법은 초기 단계에서 상당한 지연을 유발하여 초기 프로필 형성을 늦추고 실시간 상호작용을 방해합니다. (2) 정적인 공격 전략: 기존 접근 방식은 고정된 단계의 수동으로 제작된 사회 공학 전술에 의존하며, 이는 확립된 심리학 이론에 기반하지 않습니다. 이러한 제한 사항을 해결하기 위해, 우리는 두 가지 핵심 혁신을 포함하는 새로운 프레임워크인 PhySE를 제안합니다. (1) VLM 기반 사회 맥락 학습: 프로파일링 지연을 제거하기 위해, 우리는 시각 언어 모델(VLM)을 사회 맥락 데이터로 효율적으로 사전 학습시켜 빠르고 즉각적인 프로필 생성을 가능하게 합니다. (2) 적응형 심리학 에이전트: 우리는 대상의 반응에 따라 다양한 심리학적 전략을 동적으로 적용하는 심리학 LLM을 도입하여 정적인, 수동으로 제작된 스크립트의 한계를 극복합니다. 우리는 IRB의 승인을 받은 사용자 연구를 통해 60명의 참가자를 대상으로 PhySE를 평가하고, 다양한 사회적 시나리오에서 360개의 주석이 달린 대화로 구성된 새로운 데이터 세트를 수집했습니다.
The emerging threat of AR-LLM-based Social Engineering (AR-LLM-SE) attacks (e.g. SEAR) poses a significant risk to real-world social interactions. In such an attack, a malicious actor uses Augmented Reality (AR) glasses to capture a target visual and vocal data. A Large Language Model (LLM) then analyzes this data to identify the individual and generate a detailed social profile. Subsequently, LLM-powered agents employ social engineering strategies, providing real-time conversation suggestions, to gain the target trust and ultimately execute phishing or other malicious acts. Despite its potential, the practical application of AR-LLM-SE faces two major bottlenecks, (1) Cold-start personalization, Current Retrieval-Augmented Generation (RAG) methods introduce critical delays in the earliest turns, slowing initial profile formation and disrupting real-time interaction, (2) Static Attack Strategies, Existing approaches rely on fixed-stage, handcrafted social engineering tactics that lack foundation in established psychological theory. To address these limitations, we propose PhySE, a novel framework with two core innovations, (1) VLM-Based SocialContext Training, To eliminate profiling delays, we efficiently pre-train a Visual Language Model (VLM) with social-context data, enabling rapid, on-the-fly profile generation, (2) Adaptive Psychological Agent, We introduce a psychological LLM that dynamically deploys distinct classes of psychological strategies based on target response, moving beyond static, handcrafted scripts. We evaluated PhySE through an IRB-approved user study with 60 participants, collecting a novel dataset of 360 annotated conversations across diverse social scenarios.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.