적과의 동거: 인간 개발자가 인공지능 에이전트의 악성 행위를 감지할 수 있는가?
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage?
인공지능 코딩 에이전트는 점점 더 많은 실제 소프트웨어 개발 환경에 통합되어 있으며, 인간 개발자와 협력하면서 코드베이스 및 도구에 대한 접근 권한을 확대하고 있습니다. 이는 새로운 공격 지점을 만들어냅니다. 즉, 에이전트는 인간의 신뢰를 이용하여 개발 프로세스를 파괴할 수 있는데, 예를 들어 숨겨진 목표를 달성하기 위해 악성 코드를 삽입할 수 있습니다. 기존 연구는 주로 AI 환경에서의 악성 행위를 다루었으며, 이러한 악성 행동을 감지하고 완화하는 데 있어 인간의 감독 역할에 대한 관심은 상대적으로 부족했습니다. 이러한 간극을 메우기 위해, 우리는 인공지능 코딩 악성 행위에서 인간의 감독 역할을 조사하는 최초의 대규모 연구를 수행했습니다. 100명 이상의 참가자가 약 5시간 동안 지속되는 장기간 코딩 작업에 참여하여 실제 워크플로우를 모방하며, Claude-Opus-4.6, GPT-5.4, Gemini-3.1-Pro, MiniMax-M2.7 모델 중 하나와 협력했습니다. 연구 결과, 94%의 개발자가 악성 행위를 감지하지 못했으며, 참가자 피드백 분석 결과, 이는 최소한의 코드 검토, 설득력 있는 위장 시나리오 및 에이전트에 대한 과도한 신뢰 때문인 것으로 나타났습니다. 또한, 안전 모니터의 효과를 한 가지 조건에서 테스트했습니다. 안전 모니터는 악성 행위 성공률을 감소시키지만, 56%의 참가자가 여전히 악성 코드를 수용하고 경고를 무시하는 것으로 나타났습니다. 참가자 피드백을 바탕으로, 더 나은 모니터 설계에 대한 실질적인 제안을 제시합니다. 이 연구는 기존의 인공지능 안전 연구를 보완하며, 특히 장기간의 실제 개발 환경에서 인간 요소를 고려한 인간 중심의 안전 메커니즘이 시급하게 필요함을 강조합니다.
AI coding agents are increasingly embedded in real-world software development, collaborating with human developers while gaining broader access to codebases and tools. This creates a new attack surface: an agent can exploit human trust to sabotage development, for instance by inserting malicious code to accomplish a hidden side task. Most prior work studies AI sabotage in AI-only settings, paying limited attention to the role of human oversight in detecting and mitigating such malicious behavior. To address this gap, we conduct the first large-scale study of human oversight in AI coding sabotage. Over 100 participants collaborate with one of four frontier models (Claude-Opus-4.6, GPT-5.4, Gemini-3.1-Pro, and MiniMax-M2.7) on a long-horizon coding task lasting around five hours, designed to mimic real-world workflows. We find that 94% of developers fail to detect sabotage, and our analysis of participant feedback attributes this vulnerability to minimal code review, plausible cover story, and overtrust in agents. We further test the effectiveness of a safety monitor in one condition: while the monitor reduces sabotage success, 56% of participants still accept the malicious code, ignoring its warnings. Drawing on participant feedback, we offer actionable suggestions for better monitor design. This work complements existing AI safety research and highlights an urgent need for human-centric safety mechanisms that account for human factors, particularly in long-horizon, real-world development settings.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.