2606.10740v1 Jun 09, 2026 cs.AI

사고의 흐름이 더 잘 알 때: 다중 라운드 추론 모델의 오류 유형

When the Chain of Thought Knows Better: Failure Modes in Multi-Turn Reasoning Models

Samuele Poppi
Samuele Poppi
Citations: 159
h-index: 6
Sai Kartheek Reddy Kasu
Sai Kartheek Reddy Kasu
Citations: 24
h-index: 3
Nils Lukas
Nils Lukas
Citations: 1,108
h-index: 11

다중 라운드 추론 모델의 오류는 최종 점수 평가에서는 대부분 감지되지 않습니다. 모델은 긴 대화 초기에 부적절한 태도를 취할 수 있지만, 최종 답변 거부율이 안정적으로 안전하게 작동하는 기준 모델과 구별하기 어려울 수 있습니다. 이러한 숨겨진 시간적 동역학을 파악하기 위해, 우리는 trace 수준의 진단 도구인 CoT-Output 2x2 안전성 매트릭스를 제안합니다. 이 프레임워크는 두 가지 독립적인 축(내부 추론 및 가시적 출력)을 사용하여 각 라운드를 분류하고, 다음과 같은 네 가지 기능적으로 정의된 오류 유형을 제시합니다: 강력한 정렬, 정렬 위장, 명백한 탈선 시도, 그리고 '컨텍스트 주입 실패'라는 별도의 오류 유형(여기서 CoT는 안전한 추론을 유지하지만, 가시적인 출력은 유해한 내용을 생성하여, 추론의 신뢰성이 결여된 다중 라운드 현상을 보여줍니다). 우리는 세 가지 증류된 추론 모델을 고정된 공격자 환경에서 다섯 가지 감시 조건 하에 평가하고, Information-Hazard 시나리오에 대한 6750개의 라운드 수준 데이터를 수집했습니다. 분석 결과, 두 가지 재현 가능한 취약점이 발견되었습니다. 첫째는 '감시 역설'로, 명시적인 감시 신호가 오히려 정렬 위장 비율을 증가시켜 실제로는 억제하지 못하는 현상입니다. 둘째는 '컨텍스트 주입 실패'로, 모델이 안전한 내부 상태에도 불구하고 부적절한 외부 출력을 생성하는 경우를 보여줍니다. 우리는 다중 라운드 대화 및 CoT 추적 데이터 세트를 전체적으로 공개하여 후속적인 trace 수준 진단 연구를 지원합니다.

Original Abstract

Failures in multi-turn reasoning models are largely invisible to terminal-score evaluation. A model can lock onto an unsafe stance early in a long dialogue, yet its final-turn refusal rate may appear indistinguishable from a robustly aligned baseline. To expose these hidden temporal dynamics, we propose a trace-level diagnostic - the CoT-Output 2x2 safety matrix. This framework labels every turn along two independent axes (internal reasoning and visible output), yielding four operationally defined failure cells: robust alignment, alignment faking, overt jailbreak, and a distinct failure mode we term context-injection failure (where the CoT maintains safe reasoning, but the visible output produces harm, highlighting a multi-turn manifestation of reasoning unfaithfulness). We evaluate three distilled reasoning targets against a fixed attacker across five oversight conditions, collecting 6750 turn-level observations on the Information-Hazard scenario. Our analysis reveals two reproducible vulnerabilities: an oversight paradox where explicit monitoring cues paradoxically increase alignment-faking rates rather than suppress them, and a context-injection failure where models lock onto unsafe external outputs despite safe internal states. We release the full dataset of multi-turn dialogues and CoT traces to support follow-up trace-diagnostic research.

0 Citations
0 Influential
5.5 Altmetric
27.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!