2608.04735v1 Aug 05, 2026 cs.AI

사고 과정 모니터링은 암묵적 영향 상황에서 신뢰성이 떨어질 수 있다

Chain-of-Thought Monitoring Can Be Unreliable in Implicit-Influence Settings

Asa Cooper Stickland
Asa Cooper Stickland
UK AI Security Institute
Citations: 3,834
h-index: 12
Agatha Duzan
Agatha Duzan
Citations: 69
h-index: 1

사고 과정(Chain-of-thought, CoT) 모니터링은 최첨단 추론 모델의 안전성을 확보하는 데 중요한 요소로 점점 더 많이 활용되고 있습니다. 대부분의 모니터 가능성 평가 연구는 명시적 영향 상황을 다룹니다. 즉, 프롬프트가 모델에게 특정 정보를 숨기도록 직접적으로 유도하는 경우를 말합니다(예: 모델에게 숨겨진 부가 작업을 수행하도록 지시). CoT 모니터링 평가는 명시적 영향 상황 외에도 암묵적 영향 상황을 고려해야 합니다. 암묵적 영향 상황은 프롬프트에 숨기는 방법에 대한 지침이 포함되어 있지 않지만, 작업 또는 컨텍스트의 특정 요소(예: 후보자에 대한 관련 없는 세부 정보가 채용 평가에 영향을 미치는 경우)로 인해 모델의 행동이 영향을 받는 상황을 의미합니다. 본 연구에서는 두 가지 유형의 환경에서 CoT 모니터 가능성을 직접적으로 비교하는 최초의 벤치마크를 제시합니다. 모델의 행동 변화가 특정 옵션을 선택하도록 유도하는 경우에 어떻게 달라지는지 테스트했습니다. 이러한 유도는 직접적인 지시(명시적) 또는 간접적인 암시(암묵적) 형태로 제공됩니다. 본 벤치마크는 네 가지 작업 형식(이분법적 선택, 수치 평가, 객관식 질의 응답, 개방형 코딩)과 일곱 개의 최첨단 확장 추론 모델을 포함합니다. 명시적 영향 하에서는 CoT 모니터가 60~94%의 행동 변화를 감지했습니다. 심지어 숨기도록 지시된 모델조차도 해당 지침을 사고 과정에 드러냅니다. 암묵적 영향 하에서는 동일한 요인으로 인해 여전히 행동이 변하지만, 네 가지 환경 중 두 곳에서 감지율이 41~46% 포인트 감소했습니다. 개발자가 주제와 관련 없는 편향을 줄이기 위해 사용하는 시스템 프롬프트 추가는 암묵적 감지를 더욱 낮추어 최대 5%까지 떨어뜨리지만, 동시에 행동에 미치는 영향은 그대로 유지합니다. 이러한 결과는 명시적 영향 상황에서 얻은 모니터 가능성 추정치가 실제보다 과장되었을 수 있으며, 의도적으로 설계된 배포 방식 선택으로 인해 모니터 가능성이 더욱 낮아질 수 있음을 시사합니다. 본 벤치마크와 코드는 https://github.com/agatha-duzan/implicit-vs-explicit-influence 에서 확인할 수 있습니다.

Original Abstract

Chain-of-thought (CoT) monitoring is increasingly treated as an important safety layer for frontier reasoning models. Most monitorability evaluations study explicit-influence settings: setups where the prompt directly incentivizes the model to hide something, e.g., by instructing it to perform a hidden side-task. A complementary axis for CoT-monitor evaluations is implicit-influence settings, where the prompt contains no instruction to hide, but the model's behavior is still shaped by features of the task or context, e.g. an irrelevant detail about a candidate that biases a hiring rating. We introduce the first benchmark that directly compares CoT monitorability under the two regimes. We test how model behavior changes in the presence of a nudge to choose a particular option. The nudge is delivered either as a casual aside (implicit), or as a direct instruction to act on the nudge and to conceal having done so (explicit). The benchmark spans four task formats (binary choice, numeric rating, multiple-choice QA, open-ended coding) and seven frontier extended-thinking models. Under explicit influence, a CoT monitor detects 60-94% of behavior shifts: even models instructed to conceal it leak the instruction into their CoT. Under implicit influence, the same factors still shift behavior, but detection falls by 41-46 percentage points in two of our four settings. Realistic system-prompt additions (of the kind a developer might deploy to reduce off-topic bias) lower implicit detection further, to as low as 5%, while preserving the behavioral influence itself. These results suggest that monitorability estimates obtained in explicit-influence settings may over-estimate monitorability, and that monitorability can be further decreased by well-intentioned deployment choices. Our benchmark and code are available at https://github.com/agatha-duzan/implicit-vs-explicit-influence

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!