2605.29288v1 May 28, 2026 cs.AI

답변 정확도가 높은 긴 추론(CoT) 학습 데이터에서 유해한 추가 정보 진단

Diagnosing Harmful Continuation in Answer-Correct Long-CoT Training Traces

Wenxuan Zhang
Wenxuan Zhang
Citations: 2
h-index: 1
Lei Wang
Lei Wang
Citations: 248
h-index: 5
Yuhao Wu
Yuhao Wu
Citations: 1
h-index: 1
Chengyao He
Chengyao He
Citations: 0
h-index: 0
Fumin Shen
Fumin Shen
Citations: 268
h-index: 7

긴 추론(CoT) 체인은 LLM의 추론 능력 향상을 위한 지도 학습에 널리 사용되지만, 답변이 정확하더라도 여전히 다양한 미세 조정 결과를 초래할 수 있습니다. 본 연구에서는 답변이 충분히 뒷받침되는 것처럼 보이지만, 계속해서 추가적인 추론을 포함하는 '답변 정확도가 높은 긴 CoT 데이터'의 후속 설명 부분을 분석합니다. 이 후속 설명 부분이 학습에 미치는 영향을 평가하기 위해, 답변 내용을 유지하면서 불필요한 부분을 제거하는 편집기를 사용하여 원본 데이터와 처리된 데이터를 비교하고, CoT 기반의 지도 학습을 수행합니다. 실험 결과, 편집기가 식별한 후속 설명 부분을 제거했을 때 지도 학습 성능이 향상되는 것을 확인했습니다. 이는 해당 부분이 특정 환경에서 학습에 부정적인 영향을 미친다는 것을 시사하며, 이러한 현상을 '유해한 추가 정보'라고 명명합니다. 또한, 우리는 제거된 후속 설명 부분의 특징을 불확실성과 내부 상태 진행 측면에서 분석하여, 일관되지 않은 지역적 불확실성과 약화된 방향성 진행 간의 불일치를 확인했습니다. 마지막으로, 편집기가 식별한 후속 설명 부분의 경계를 근사하는 간단한 경계 추정 도구인 '유해한 추가 정보 제거(HCC)'를 제안합니다.

Original Abstract

Long chain-of-thought (CoT) traces are widely used as supervision for reasoning-oriented LLM SFT, yet answer-correct traces can still lead to markedly different fine-tuning outcomes. We study post-conclusion continuation in answer-correct long-CoT data: a continuation where the answer appears sufficiently supported, but the trace continues with additional reasoning that remains in the supervised target. To test its training effect, we use a delete-only editor to construct answer-preserving suffix removal and compare CoT-based SFT on the original and processed traces. We observe improved SFT outcomes after removing the editor-identified post-conclusion continuation, suggesting that this continuation is harmful to training in our setting. We therefore refer to this empirically supported phenomenon as harmful continuation. Beyond this intervention, we further characterize the removed post-conclusion continuation through uncertainty and hidden-state progress. We observe persistent local uncertainty together with weakened terminal-directional progress, forming an uncertainty--geometry mismatch. Finally, we instantiate Harmful Continuation Cut (HCC), a lightweight boundary proxy that approximates the editor-identified post-conclusion continuation boundary.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!