2607.19292v1 Jul 21, 2026 cs.CY

우리가 간과하는 안전 문제: 현대 인공지능 시스템의 숨겨진 핵심 안전 과제에 대한 고찰

The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems

Gjergji Kasneci
Gjergji Kasneci
Citations: 17,895
h-index: 40
Enkelejda Kasneci
Enkelejda Kasneci
Citations: 13,172
h-index: 45

현재 인공지능 안전 논의는 여전히 명백한 피해, 극단적인 오용, 그리고 가상적인 재앙 시나리오와 같은 눈에 띄는 실패 사례에 지나치게 집중하고 있습니다. 이러한 관점은 불완전합니다. 실제로 운영되는 시스템에서 가장 심각한 문제들은 종종 조용한 방식으로 나타납니다. 즉, 화려하기보다는 그럴듯하며, 단일 출력에 국한되지 않고 여러 구성 요소에 분산되어 있으며, 위험으로 인식되기 전에 워크플로우에 의해 정상화됩니다. 우리는 현대 인공지능 시스템의 핵심적인 안전 과제는 단순히 모델이 유해한 응답을 생성하는가 여부뿐만 아니라, 더 넓은 사회-기술 시스템이 오류가 여전히 가시적이고, 논쟁 가능하며, 통제 가능하고, 복구 가능한 상태로 유지될 수 있는 조건을 보존하는 것인지에 달려 있다고 주장합니다. 우리는 이러한 숨겨진 위험을 진단하기 위한 5가지 계층의 프레임워크를 제안합니다: (1) 인지적 무결성(증거와 불확실성이 신뢰할 만한 의사 결정을 지원할 수 있도록 충분히 정확하게 표현되는지 여부), (2) 통제적 무결성(권한, 허가 및 행동 경계가 공격 및 최적화 하에서도 안정적으로 유지되는지 여부), (3) 시간적 무결성(안전이 세션, 메모리 업데이트 및 배포 변경을 통해 지속적으로 유지되는지 여부), (4) 조직적 무결성(기관이 감사, 책임 부여 및 효과적인 개입 능력을 유지하는지 여부), 그리고 (5) 생태계적 무결성(인공지능 시스템이 미래의 감독에 의존하는 정보 환경을 파괴하지 않고 보존하는지 여부). 이러한 계층들을 통해 우리는 과도한 의존, 검색 과정에서의 불확실성과 정당성 조작, 프롬프트 주입 공격, 보상 해킹, 메모리 오염, 평가 속임수, 허구적인 인간 감독, 합성 증거 오염 및 모델 붕괴와 같은 간과된 위험 패턴을 식별합니다. 우리는 결론적으로 설계 및 거버넌스 권고 사항과 함께 인공지능 안전의 모델 중심 평가에서 사회-기술적 신뢰성으로 전환하기 위한 연구 과제를 제시합니다.

Original Abstract

Current AI safety discourse still focuses disproportionately on visible failures, including obvious harms, dramatic misuse, and hypothetical catastrophic scenarios. That focus is incomplete. In deployed systems, many of the most consequential failures are quieter: plausible rather than spectacular, distributed across components rather than localized in a single output, and normalized by workflows before they are recognized as hazards. We argue that a central safety challenge in modern AI systems is increasingly not only whether a model emits a harmful response, but whether the broader socio-technical system preserves the conditions under which errors remain visible, contestable, containable, and recoverable. We propose a five-layer framework for diagnosing these hidden risks: (1) epistemic integrity, concerning whether evidence and uncertainty are represented honestly enough to support calibrated reliance; (2) control integrity, concerning whether authority, permissions, and action boundaries remain robust under attack and optimization; (3) temporal integrity, concerning whether safety holds across sessions, memory updates, and deployment drift; (4) organizational integrity, concerning whether institutions retain the capacity to audit, assign responsibility, and intervene effectively; and (5) ecosystem integrity, concerning whether AI systems preserve rather than erode the information environment on which future oversight depends. Across these layers, we identify under-recognized risk patterns, including overreliance, uncertainty and legitimacy laundering in retrieval, prompt injection, reward hacking, memory poisoning, evaluation deception, fictional human oversight, synthetic evidence pollution, and model collapse. We conclude with design and governance recommendations and a research agenda for shifting AI safety from model-centric evaluation toward socio-technical reliability.

1 Citations
0 Influential
22.5 Altmetric
113.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!