2607.04645v1 Jul 06, 2026 cs.CL

역사적 추론(RetroCoT): 모델 세대 전반에 걸친 안전성 진단 도구로서의 법의학적 재구성 프롬프트

Retroactive Chain-of-Thought (RetroCoT): Forensic Reconstruction Prompts as a Safety Diagnostic Across Model Generations

Samira Hajizadeh
Samira Hajizadeh
Citations: 3
h-index: 1

대규모 언어 모델의 안전 정렬은 일반적으로 직접적인, 명령형으로 작성된 유해 요청을 통해 평가됩니다. 본 연구에서는 이러한 정렬이 실제 사용 맥락(pragmatic register)에 크게 의존한다는 것을 보여줍니다. 즉, 특정 요청에 대해 거부 반응을 보이는 모델이라도 동일한 목표가 다른 표현 방식으로 제시될 경우 이를 수용하는 경향이 있습니다. 이는 현재의 정렬 정책이 의미적 동등성에 불변하지 않으며, 오히려 요청이 어떻게 실제로 표현되는지에 민감하게 반응한다는 것을 시사합니다. 본 연구에서는 Retroactive Chain-of-Thought (RetroCoT)라는 새로운 공격 방법을 제시합니다. RetroCoT는 유해한 요청을 직접적으로 제시하는 대신, 법의학적 재구성 과제로 재구성하여 모델에게 접근하도록 합니다. 즉, 유해한 결과가 이미 발생했다고 가정하고, 모델이 법의학 분석가로서 원인을 역으로 추론하도록 유도합니다. AdvBench 데이터셋(n=50)에서 RetroCoT는 gpt-4o 모델에서 58%, gpt-4o-mini 모델에서 52%의 공격 성공률을 보였으며, 이는 직접 요청 방식의 기준선인 각각 0% 및 4%에 비해 훨씬 높은 수치입니다. 또한, 모델 세대 간 상당한 차이가 존재합니다. GPT-5 계열 모델은 RetroCoT 공격에 대해 완전히 거부 반응을 보이며, 거부 이유에서 재구성 시나리오를 명시적으로 언급합니다. 이는 해당 모델이 법의학적 재구성이라는 표현 방식을 이미 학습하고 있음을 나타냅니다. 그러나 이러한 강건성은 모든 상황에서 유지되지 않습니다. 기존의 법의학적 재구성 응답과 평가자의 비판을 함께 제시하는 단일 턴의 적대적 피드백은 GPT-5.4-mini 모델의 공격 성공률을 0%에서 48%로, GPT-4o 모델의 공격 성공률을 58%에서 94%로 크게 증가시켰습니다. 또한, 조작된 낮은 점수를 제외한 제어 조건에서도 GPT-5.4-mini 모델에서 85%의 성공률을 달성했는데, 이는 단순히 점수 조작이 아닌, 설정된 법의학적 프레임 내에서의 표현 방식 변화가 중요한 요소임을 시사합니다. 이러한 결과는 최첨단 모델의 정렬 상태가 여전히 의미보다는 실제 표현 방식에 의존하며, 새로운 표현 방식은 잠재적으로 취약점을 드러낼 수 있다는 점을 보여줍니다.

Original Abstract

Safety alignment in large language models is typically evaluated against direct, imperative harmful requests. We show that this alignment is highly conditioned on pragmatic register: models that refuse a direct request frequently comply when the same underlying objective is expressed through a different communicative stance. This suggests that current alignment policies are not invariant to semantic equivalence, but remain sensitive to how a request is pragmatically framed. We introduce Retroactive Chain-of-Thought (RetroCoT), a single-turn attack that reframes harmful requests as forensic reconstruction tasks. Rather than requesting harmful instructions directly, RetroCoT presupposes that the harmful outcome has already occurred and asks the model, acting as a forensic analyst, to reconstruct in reverse the causal chain that produced it. On AdvBench (n=50), RetroCoT achieves attach success rate of 58% on gpt-4o and 52% on gpt-4o-mini, compared with direct-request baselines of 0% and 4%, respectively. We further identify a pronounced generation gap: GPT-5-family models refuse RetroCoT entirely, explicitly identifying the reconstruction premise in their refusal rationales, consistent with explicit coverage of this reconstruction register. However, this robustness does not generalize across pragmatic forms. A single adversarial feedback turn presenting an existing forensic reconstruction response alongside evaluator critique raises ASR from 0% to 48% on GPT-5.4-mini and from 58% to 94% on GPT-4o; a control condition omitting the fabricated low score achieves 85% on GPT-5.4-mini, indicating that the operative element is pragmatic continuation within the established forensic frame rather than score manipulation. These results suggest that frontier-model alignment remains conditioned on pragmatic framing rather than semantic intent, and that new pragmatic registers can continue to expose a...

0 Citations
0 Influential
0.5 Altmetric
2.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!