2607.28908v1 Jul 31, 2026 cs.LG

성찰인가, 재창조인가? 인간의 수정이 성공하는 이유와 LLM 수정 실패 원인

Reflection or Re-Generation? Why LLM Revision Fails Where Human Revision Succeeds

Luyang Kong
Luyang Kong
Citations: 148
h-index: 4
Yefan Tao
Yefan Tao
Citations: 17
h-index: 2
Gerald Friedland
Gerald Friedland
Citations: 13
h-index: 2
Madhusudhanan Chandrasekaran
Madhusudhanan Chandrasekaran
Citations: 0
h-index: 0

인간은 자신의 답변을 개선하기 위해 이전의 추론 과정을 되돌아보고 수정하는 '성찰' 능력을 활용합니다. 최근 대규모 언어 모델(LLM)은 '성찰'하도록 설계되고 있지만, 이것이 인간의 수정 과정과 얼마나 유사한지는 명확하지 않습니다. 본 연구에서는 인간과 LLM의 수정 과정을 동일한 조건에서 비교하는 통제된 두 단계 프로토콜인 Human-LLM Reflection Framework (HRF)를 소개합니다. 각 반복에서의 교차 엔트로피 감소를 기반으로 한 정보 이론적 분석을 통해, 우리는 LLM 성찰의 두 가지 실패 요인을 발견했습니다. 객관적인 작업(답변 공간이 유한한 작업)에서 성찰은 거의 0에 가까운 정보 이득(Delta I ≈ 0)을 가져오며, 이는 중립적인 재창조와 구별하기 어렵습니다. 주관적인 작업에서는 상당한 음의 정보 이득(Delta I < 0)을 가져와 예측이 목표와 멀어지는 현상이 나타납니다. 반면 인간의 수정은 두 가지 환경 모두에서 양의 정보 이득을 가져옵니다. 교차 에이전트 실험 결과, LLM의 실패는 입력 품질의 문제가 아니라 수정 단계 자체의 문제임을 확인했습니다. 진단 분석(첫 번째 단계에서의 정확성에 조건화된 수정 및 랜덤 재배열 기준선에 대한 오라클 가이드 수정) 결과, 하위 단계의 중요성은 작업과 모델에 따라 다르며 단일한 메커니즘으로 설명될 수 없습니다. 객관적인 다중 선택 문제에서는 자기 오류 감지가 존재하지만 주관적인 문제에서는 약하며, 오라클 오류 신호하에서의 회복은 일부 모델에서 기준선보다 우수하고 다른 모델에서는 미흡합니다. 통합된 관점은 구조적입니다. 외부 정보 없이는 자기 조건화된 수정이 목표에 대한 불확실성을 줄일 수 없으므로, LLM 성찰은 진정한 오류 기반 수정이라기보다는 조건부 재창조로 이해하는 것이 더 적절합니다.

Original Abstract

Reflection, the ability to revisit and revise prior reasoning, is central to how humans improve their answers. Large language models (LLMs) are increasingly prompted to "reflect," yet whether this resembles human revision remains unclear. We introduce the Human-LLM Reflection Framework (HRF), a controlled two-pass protocol comparing human and LLM revision under identical conditions across self-, peer-, and cross-agent settings. Using an information-theoretic analysis based on per-iteration cross-entropy reduction, we find two failure modes of LLM reflection. On objective tasks with finite answer spaces, reflection yields near-zero information gain (Delta I approx 0), behaving as neutral re-generation indistinguishable from re-sampling. On subjective tasks, it yields significant negative gain (Delta I < 0), moving predictions away from the target. Human revision, by contrast, yields positive gain in both settings. Cross-agent experiments localize the failure to the revision step, not input quality: LLMs degrade even high-quality human responses. Diagnostic analyses (revision conditioned on first-pass correctness, and oracle-guided revision against a random-reshuffle baseline) show that which sub-step dominates varies by task and by model rather than reducing to a single mechanism: self-error detection is present on objective multiple-choice tasks but weak on subjective ones, and recovery under an oracle error signal exceeds the baseline for some models and falls below it for others. The unifying account is structural: without external information, self-conditioned revision cannot reduce uncertainty about the target, so LLM reflection is better understood as conditioned re-generation than as genuine error-driven revision.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!