대규모 언어 모델 코딩 에이전트의 피싱 웹사이트 생성 능력 평가
Assessing Spear-Phishing Website Generation in Large Language Model Coding Agents
대규모 언어 모델(LLM)은 인간이 사용하는 도구를 넘어, 환경을 관찰하고, 문제 해결 방안을 추론하며, 환경에 영향을 미치는 변경을 수행하고, 자신의 행동이 환경에 미치는 영향을 이해하는 독립적인 에이전트로 발전하고 있습니다. 이러한 LLM 에이전트의 가장 흔한 응용 분야 중 하나는 컴퓨터 프로그래밍이며, 여기서 에이전트는 프로그래밍 환경 또는 네트워크 시스템을 제어하면서 인간과 협력하여 코드를 생성할 수 있습니다. 그러나 이러한 에이전트의 능력과 복잡성이 증가함에 따라, 오용될 가능성에 대한 위험도 증가하고 있습니다. LLM 에이전트의 우려스러운 응용 분야 중 하나는 사이버 보안 영역이며, 여기서 에이전트는 사회 공학 공격과 같은 위협을 크게 확대할 수 있습니다. 이는 LLM 에이전트가 자율적으로 작동하고 숙련된 인간 프로그래머가 수행해야 하는 많은 작업을 수행할 수 있기 때문입니다. 이러한 위협은 우려스럽지만, LLM 코딩 에이전트가 사회 공학 공격을 위한 코드를 생성하는 능력에 대한 평가는 거의 이루어지지 않았습니다. 본 연구에서는 다양한 LLM의 잠재적으로 위험한 코드베이스 생성 능력과 의지를 비교합니다. 그 결과는 40개의 서로 다른 LLM 코딩 에이전트에서 생성된 200개의 웹사이트 코드베이스 및 로그 데이터 세트로 구성됩니다. 모델 분석을 통해 LLM의 어떤 지표가 스피어 피싱 웹사이트 생성 성능과 더 밀접하고 덜 밀접한 관련이 있는지 확인합니다. 본 연구의 분석 결과와 데이터 세트는 LLM의 잠재적인 오용, 특히 스피어 피싱 공격에 대한 방어에 관심 있는 연구자 및 실무자에게 유용할 것입니다.
Large Language Models are expanding beyond being a tool humans use and into independent agents that can observe an environment, reason about solutions to problems, make changes that impact those environments, and understand how their actions impacted their environment. One of the most common applications of these LLM Agents is in computer programming, where agents can successfully work alongside humans to generate code while controlling programming environments or networking systems. However, with the increasing ability and complexity of these agents comes dangers about the potential for their misuse. A concerning application of LLM agents is in the domain cybersecurity, where they have the potential to greatly expand the threat imposed by attacks such as social engineering. This is due to the fact that LLM Agents can work autonomously and perform many tasks that would normally require time and effort from skilled human programmers. While this threat is concerning, little attention has been given to assessments of the capabilities of LLM coding agents in generating code for social engineering attacks. In this work we compare different LLMs in their ability and willingness to produce potentially dangerous code bases that could be misused by cyberattackers. The result is a dataset of 200 website code bases and logs from 40 different LLM coding agents. Analysis of models shows which metrics of LLMs are more and less correlated with performance in generating spear-phishing sites. Our analysis and the dataset we present will be of interest to researchers and practitioners concerned in defending against the potential misuse of LLMs in spear-phishing.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.