2607.27951v1 Jul 30, 2026 cs.CR

복사 가능한 컨텍스트 기반의 안전장치는 LLM에 대한 신뢰할 수 있는 안전성을 제공하지 못한다

Safeguards Based on Copyable Context Cannot Provide Reliable Safety for LLMs

Neng H. Yu
Neng H. Yu
Citations: 2,242
h-index: 22
Pingyu Wu
Pingyu Wu
Citations: 211
h-index: 2
Lingyao Zhu
Lingyao Zhu
Citations: 0
h-index: 0
Weiming Zhang
Weiming Zhang
Citations: 941
h-index: 12

대규모 언어 모델(LLM)의 안전장치는 답변을 제공하기 전에 답변이 어떻게 사용될지 먼저 확인합니다. 이는 이중 용도 작업에서 근본적인 문제를 야기합니다. 동일한 답변은 권한이 있는 전문가에게 도움이 될 수도 있지만 공격자에게도 도움이 될 수 있으며, 공격자는 양성적인 요청 및 상호 작용 기록을 모방할 수 있습니다. 본 연구에서는 모델이 제공하는 기능과 하위 사용에 대한 이용 가능한 증거를 분리합니다. 해당 증거가 복사 가능할 때, 유용한 답변을 유지하면서 공격자가 얻을 수 있는 최악의 수준(worst-case floor)을 정확하게 파악했습니다. 그 결과는 '유용한 기능', '신뢰할 수 있는 안전성' 및 '개방적인 접근성'이라는 세 가지 요소가 공존할 수 없다는 삼자택일 문제(safety trilemma)를 제시합니다. 또한, 신뢰할 수 있는 자격 증명은 기존의 안전장치를 보완하여 실제 하위 사용을 예측하는 복제하기 어려운 정보를 추가함으로써 어떻게 작동하는지 보여줍니다. 더 나아가, 이러한 수준의 장벽을 제거하기 위해 필요한 강화된 조건을 식별했습니다. 이중 용도 평가, 적응형 공격 및 배포된 신뢰 기반 접근 프로그램에서 얻은 증거는 이러한 조건이 실제로 얼마나 중요한지를 뒷받침합니다.

Original Abstract

Large language model safeguards decide whether to answer before seeing how an answer will be used. This creates a basic problem for dual-use tasks: the same answer can help an authorized professional or an attacker, while an attacker can imitate a benign request and interaction history. We separate the capability released by the model from the evidence available about downstream use. When that evidence is copyable, we derive the exact worst-case floor on attacker assistance while preserving useful answers. The result yields a safety trilemma: Useful Capability, Reliable Safety, and Open Access cannot coexist. We then show how a trusted credential can complement existing safeguards by adding hard-to-copy information that predicts actual downstream use, and identify the stronger condition needed to eliminate the floor. Evidence from dual-use evaluations, adaptive attacks, and deployed trusted-access programs supports the practical relevance of these conditions.

0 Citations
0 Influential
11 Altmetric
55.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!