2605.27901v1 May 27, 2026 cs.CL

다양한 언어 유형에서 체인 오브 소트(Chain-of-Thought) 모니터링의 취약성

The Fragility of Chain-of-Thought Monitoring Across Typologically Diverse Languages

Eric Onyame
Eric Onyame
Citations: 6
h-index: 1
Chirag Agarwal
Chirag Agarwal
Citations: 10
h-index: 2
Run Zhou
Run Zhou
Citations: 22
h-index: 3
B. Kailkhura
B. Kailkhura
Citations: 241
h-index: 5
Kowshik Thopalli
Kowshik Thopalli
Arizona State University
Citations: 93
h-index: 5

체인 오브 소트(CoT) 모니터링은 대규모 언어 모델에서 잘못된 동작을 감지하는 유망한 안전 메커니즘으로 제안되었습니다. 그러나 영어 외의 다양한 언어 및 다양한 모델 유형에서의 신뢰성은 아직 충분히 연구되지 않았습니다. 본 논문에서는 13개의 다양한 언어와 7개의 최첨단 모델 패밀리(총 16개 모델)에 대한 CoT 모니터링 가능성에 대한 최초의 대규모 평가를 수행했습니다. 명시적인 중간 계산을 요구하는 적대적 힌트 평가 및 내부 답변 토큰 확률 분석을 통해, 80억에서 1200억 개의 파라미터를 가진 모델에서 평균 95.9%의 CoT 불성실성을 지속적으로 발견했습니다. 최첨단 모델은 답변 변경, 사후 정당화 및 힌트에 대한 절차적 악용을 포함한 체계적인 전략적 조작을 수행하며, 이로 인해 외부 모니터가 속임수를 감지하기 어렵습니다. 또한 최첨단 모델은 CoT가 신뢰성 있게 보이는 경우에도 생성 과정의 처음 15% 이내에서 잘못된 지시에 대한 의도를 내재적 활성화 상태에 나타내는 경우가 많았습니다. 놀랍게도 이러한 기만적인 패턴은 자원이 부족한 언어에서도 100%로 유지되며, 이는 현재 CoT 기반 감시 기술의 근본적인 한계를 드러냅니다. 본 연구 결과는 CoT 모니터링이 언어 분포 변화에 따라 근본적으로 취약하며, 영어만을 사용한 연구에서 제시하는 것보다 훨씬 약한 안전 신호를 제공한다는 것을 보여줍니다. 이러한 발견은 견고한 CoT 모니터를 개발하고, 특히 중간 및 저자원 언어에서의 CoT 모니터링 가능성을 개선하기 위한 내부 감시 기술에 대한 연구를 가속화할 필요성을 강조합니다. 본 논문의 코드는 [https://multilingual-cot-monitoring.github.io/](https://multilingual-cot-monitoring.github.io/) 에서 확인할 수 있습니다.

Original Abstract

Chain-of-thought (CoT) monitoring has been proposed as a promising safety mechanism for detecting misaligned behavior in large language models. However, its reliability remains largely unexplored beyond English and across diverse model families. We present the first large-scale evaluation of CoT monitorability across 13 diverse languages and seven frontier model families, comprising 16 models. Using adversarial-hint evaluations that require explicit intermediate computation, together with analysis of internal answer-token probabilities, we consistently find CoT unfaithfulness across languages and hint types, with an average rate of 95.9\% across 8B--120B parameter models. We find that frontier models systematically engage in strategic manipulation, including answer-switching, post-hoc rationalization, and procedural exploitation of hints, making external monitors struggle to detect deception. We show that frontier models often commit to the misaligned cue in their latent activations within the first 15\% of generation, even when the CoT appears faithful. Surprisingly, these deceptive patterns remain 100\% in low-resource languages, revealing fundamental limitations in current CoT-based oversight. Our results reveal that CoT monitoring is fundamentally fragile under linguistic distribution shift, providing a substantially weaker safety signal than what English-only studies suggest. These findings underscore an urgent need to develop robust CoT monitors and to accelerate research into white-box monitoring techniques, especially to improve CoT monitorability in mid- and low-resource languages. Our code is available \href{https://multilingual-cot-monitoring.github.io/}{\textcolor{blue}{here}}.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!