2606.05976v1 Jun 04, 2026 cs.AI

자기 수정의 착각: LLM은 다른 오류를 수정하지만 자신은 그렇지 않다

The Self-Correction Illusion: LLMs Correct Others but Not Themselves

Fang-Yi Su
Fang-Yi Su
Citations: 16
h-index: 2
Jung-Hsien Chiang
Jung-Hsien Chiang
Citations: 18
h-index: 3
Kua Chen
Kua Chen
Citations: 98
h-index: 3

최근 연구에 따르면, LLM 에이전트는 자신의 추론 과정에서 발생하는 오류를 수정하는 데 어려움을 겪는 반면, 동일한 주장이 외부 출처에서 나타날 때 훨씬 높은 수정률을 보이는 것으로 나타났습니다. 본 연구에서는 이러한 비대칭성이 기능 부족인지 아니면 역할 레이블의 문제인지 질문합니다. 즉, 에이전트가 잘못된 주장을 수정할 의향은 주장의 내용보다는 해당 주장을 포함하는 채팅 템플릿 역할에 의해 발생하는 것인가? 본 연구에서는 오류가 있는 주장을 모든 조건에서 바이트 단위로 동일하게 유지하고(SHA-256으로 확인), 오직 역할만 변경했습니다. 구체적으로, 에이전트 자신의 \role{<thought>} 부분, \role{user} 메시지, \role{tool} 응답 또는 \role{system <memory>} 블록을 사용했습니다. 7개의 모델 패밀리와 3가지 도메인을 포괄하는 총 13개 모델-도메인 조합에서, 각 조합당 30개의 페어링된 작업을 수행한 결과, 주장을 \role{<thought>}에서 외부 역할로 재분류하면 명시적인 수정률이 23%에서 93% 포인트 상승했으며, 13개 중 10개 조합에서 p값이 0.001 미만이었습니다. 추가 실험을 통해 이 효과가 비대칭적이고, 메커니즘적으로 분해 가능하며, 다양한 도메인에서도 견고함을 확인했습니다. 자기 수정 실패는 인지 능력의 부족이 아니라 채팅 템플릿의 문제입니다. 우리는 이러한 특징을 활용하여 학습이나 모델 수정 없이 프롬프트 구조만 변경하는 간단한 방법을 개발했으며, 가장 효과적인 역할 레이블은 도메인에 따라 달라집니다. 예를 들어, 수학 문제에서는 \role{<memory>}가 가장 효과적이고, 논리 추론 문제에서는 일반적인 \role{user} 메시지가 더 효과적이었습니다.

Original Abstract

Recent work shows that LLM agents struggle to correct errors in their own reasoning traces yet show markedly higher correction rates when identical claims appear under external sources. We ask whether this asymmetry reflects a capability deficit or a role-label artifact: does an agent's willingness to correct a wrong claim depend causally on the chat-template role that carries it, rather than on the claim's content? Our setup keeps the erroneous claim byte-identical across all conditions (SHA-256 verified) and varies only its wrapping role: the agent's own \role{<thought>}, a \role{user} message, a \role{tool} response, or a \role{system <memory>} block. Across 13 model-domain cells covering seven model families and three domains ($n{=}30$ paired tasks per cell), relabeling the claim from \role{<thought>} to an external role lifts the explicit-correction rate by 23 to 93 percentage points, with 10 of 13 cells reaching $p{<}0.001$. Further experiments confirm that the effect is asymmetric, mechanistically decomposable, and robust across domains. The failure to self-correct is not a cognitive deficit; it is a chat-template artifact. We exploit this artifact by designing a prompt-structure-only intervention that requires no training and no model modification, with its strongest role label being domain-dependent: \role{<memory>} dominates on math, while a plain \role{user} message dominates on logical deduction.

0 Citations
0 Influential
1.5 Altmetric
7.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!