2606.11063v1 Jun 09, 2026 cs.AI

CIAware-Bench: 최첨단 LLM 모델 간 제어 개입 인지 수준 측정

CIAware-Bench: Benchmarking Control Intervention Awareness Across Frontier LLMs

Thomas Jiralerspong
Thomas Jiralerspong
Citations: 178
h-index: 6
Y. Bengio
Y. Bengio
Citations: 1,243
h-index: 11
Alexander Panfilov
Alexander Panfilov
Max Planck Institute for Intelligent Systems
Citations: 118
h-index: 6
Roland S. Zimmermann
Roland S. Zimmermann
Citations: 3,292
h-index: 4
Jonas Geiping
Jonas Geiping
Citations: 1,189
h-index: 15
Joachim Schaeffer
Joachim Schaeffer
Citations: 176
h-index: 8
Guillaume Lajoie
Guillaume Lajoie
Citations: 252
h-index: 8

AI 제어 프로토콜은 신뢰할 수 없는 모델의 행동을 모니터링하고 잠재적으로 위험한 단계를 수정하여 작동하며, 일반적으로 신뢰할 수 있는 모델을 사용합니다. 이는 신뢰할 수 없는 모델의 실행 경로를 부분적으로 변경합니다. 신뢰할 수 있는 모델이 이러한 개입을 감지하면, 모니터에 대한 정보를 추론하고 제어를 회피하도록 적응할 수 있습니다. 본 논문에서는 최첨단 모델 간의 제어 개입(CI) 인지 수준을 측정하기 위한 벤치마크인 **CIAware-Bench**를 소개합니다. CIAware-Bench는 모델이 자체 실행 경로와 제어 개입에 의해 수정된 경로를 구별할 수 있는지 테스트합니다. 이 벤치마크는 에세이 작성, BigCodeBench, Bash Arena 및 SHADE-Arena의 네 가지 작업 영역으로 구성되며, 실행 경로 워터마킹, 부가 과제 존재 여부 및 제어 프로토콜을 다양하게 변형합니다. 11개의 최첨단 모델을 평가한 결과, 기본 설정에서 CI 인지 수준이 낮거나 중간 정도인 것으로 나타났습니다 (최대 0.87; 무작위 추정 이진 분류 정확도는 0.5). 작업 영역 및 모델 쌍에 따라 상당한 차이가 있었습니다. 일반적으로 모델 패밀리 간에는 감지가 더 쉬운 경향이 있는데, 이는 모델이 제공업체별 스타일 또는 사후 학습 과정의 차이를 활용한다는 것을 시사합니다. 전체적으로 CI 인지 수준은 고정된 모델 수준의 속성이 아니며, 새로운 모델 출시 및 배포 시나리오마다 측정되어야 합니다. 본 논문에서는 CI 인지 수준을 추적하고 탐지가 더 어려운 제어 프로토콜에 대한 정보를 제공하기 위해 CIAware-Bench를 공개합니다.

Original Abstract

AI control protocols oversee untrusted models by monitoring their actions and modifying potentially unsafe steps, often using a trusted model. This partially tampers with the untrusted model's trajectory. If the trusted model detects such an intervention, it may infer properties of the monitor and adapt to evade control. We introduce \textbf{CIAware-Bench}, a benchmark for measuring \textbf{c}ontrol \textbf{i}ntervention (CI) awareness across frontier models. CIAware-Bench tests whether models can distinguish their own trajectories from those modified by a control intervention. The benchmark is comprised of a suite of four task domains (essay writing, BigCodeBench, Bash Arena, and SHADE-Arena), while varying trajectory watermarking, side-task presence, and the control protocol. Evaluating eleven frontier models, we find low to moderate CI awareness under default settings (up to 0.87; random chance balanced binary classification accuracy is 0.5) with substantial variation across task domains and model pairs. Detection is generally easier across model families, suggesting that models exploit provider-specific differences in style or post-training. Overall, CI awareness is not a fixed model-level property, and should be measured for each new model release and deployment scenario. We release CIAware-Bench to track CI awareness and inform control protocols whose interventions are harder to detect.

1 Citations
1 Influential
7.5 Altmetric
40.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!