2607.26034v1 Jul 28, 2026 cs.AI

이상적인 인공지능 경쟁 실험: 뒤처지는 현상이 안전하지 않은 개발을 유발함

Falling Behind Drives Unsafe Development in an Idealised AI Race Experiment

Elias Fernández Domingos
Elias Fernández Domingos
Citations: 63
h-index: 4
T. Han
T. Han
Citations: 188
h-index: 7

기술 경쟁은 속도와 안전 사이의 긴장을 야기하며, 참가자들은 위험한 개발이 해로울 수 있음에도 경쟁자를 앞서나가기 위해 더 빠른 발전을 추구할 수 있습니다. 이는 인공지능(AI)에 대한 논쟁에서 두드러지는데, 경쟁적 압박은 종종 더 위험하고 안전에 덜 신경 쓰는 개발을 유도하는 요인으로 지목됩니다. 본 연구에서는 이상적인 AI 경쟁을 기반으로 한 행동 실험을 통해 이를 분석합니다. 참가자들은 불확실한 시간 제약 하에서 반복적으로 '안전' 개발과 '위험' 개발 중 하나를 선택했습니다. '위험' 개발은 더 빠른 발전과 높은 즉각적 보상을 제공했지만, 치료 조건에 따라 최대 10%, 60% 또는 90%까지 사적인 위험을 누적시켰습니다. 경쟁 구조는 일정하게 유지되었으며, 오직 최대 위험만이 변수였습니다. 사전 등록된 위험 수준 간의 비교 및 명시적으로 드러난 위험 선호도의 역할은 데이터에 의해 뒷받침되지 않았습니다. 대신, 과제 수행의 반복적인 구조에서 영감을 받은 탐색적 분석 결과, '위험' 행동은 위험 선호도보다는 경쟁 환경의 변화하는 전략적 상황에 더 큰 영향을 받는다는 것을 보여줍니다. 참가자들은 상대방이 '위험'을 선택한 후에 더 높은 확률로 '위험'을 선택하며, 앞서나가면 '위험' 선택 빈도가 감소하고 뒤처지면 증가합니다. 이러한 효과를 설명하기 위해 우리는 4가지 전략(항상 안전, 항상 위험, 조건부 안전, 그리고 조건부 반사회적 안전)을 포함하는 단순화된 진화 모델을 도입했습니다. 이 모델은 치료 효과를 재현하며, 경쟁적인 경쟁 환경에서 조건부 '위험' 행동이 어떻게 선호될 수 있는지 보여줍니다. 실험과 모델을 종합적으로 분석한 결과, '위험' 개발은 위험 선호도뿐만 아니라 초기 행동의 추세, 상대방의 행동 및 뒤처질 것에 대한 두려움으로부터 발생할 수 있습니다. 따라서 정책은 AI 개발에서 경쟁적 압박을 줄이고 협력을 장려하는 데 초점을 맞춰야 하며, 단순히 개인의 위험 관리에만 집중해서는 안 됩니다.

Original Abstract

Technological races create tension between speed and safety: actors may gain by moving faster than competitors, even when risky development is harmful. This is prominent in debates about artificial intelligence (AI), where competitive pressure is often argued to incentivise riskier, less safety-conscious development. We study this using a framed behavioural experiment based on an idealised AI race, in which paired participants repeatedly chose between Safe and Unsafe development under an uncertain time horizon. Unsafe development gave faster progress and higher immediate payoffs but accumulated private risk up to a treatment-specific maximum of 10\%, 60\%, or 90\%; the race's competitive structure was held constant, and only this maximum risk varied. Neither the pre-registered comparison between risk levels nor the role of elicited risk preferences was supported by the data. Instead, exploratory analyses motivated by the task's repeated structure show that Unsafe behaviour is shaped less by risk preferences than by the evolving strategic state of the race: participants are more likely to choose Unsafe after their opponent does so, being ahead reduces Unsafe play while falling behind increases it, and first-round choices predict later behaviour. To interpret these effects we introduce a reduced evolutionary model with four strategies -- Always Safe, Always Unsafe, Conditionally Safe, and Conditionally Antisocial Safe -- which reproduces the treatment effect and shows how conditional Unsafe behaviour can be favoured by competitive race dynamics. Together, the experiment and model show that unsafe development can emerge from early behavioural momentum, opponent behaviour, and fear of falling behind, rather than from risk preferences alone, suggesting policy should focus on reducing competitive pressure and promoting cooperation in AI development rather than only individual risk.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!