2605.28008v1 May 27, 2026 cs.AI

생각을 압축하다: LLM 후속 학습에서 압축된 추론 데이터가 언제, 어떻게 작동하는가

Zipping the Thought: When and How Compressed Reasoning Data Works in LLM Post-Training

Gouki Minegishi
Gouki Minegishi
Citations: 136
h-index: 6
Takeshi Kojima
Takeshi Kojima
Citations: 7,530
h-index: 7
Yusuke Iwasawa
Yusuke Iwasawa
Citations: 10,847
h-index: 23
Yutaka Matsuo
Yutaka Matsuo
Citations: 1,565
h-index: 13
Kohsei Matsutani
Kohsei Matsutani
Citations: 33
h-index: 3

대규모 언어 모델(LLM)은 이제 긴 연쇄적 사고(CoT) 방식을 통해 복잡한 문제를 해결할 수 있지만, 성능과 토큰 비용 간의 균형 문제는 여전히 중요한 과제입니다. 이러한 문제점을 해결하기 위해, 지도 학습(SFT)에서는 종종 CoT 과정을 짧고 간결한 형태로 압축한 데이터를 사용합니다. 그러나 이러한 압축된 추론 데이터가 후속 학습에 미치는 영향은 아직 제대로 이해되지 못하고 있습니다. 본 논문에서는 명시적 CoT(모든 연산을 집계 없이 출력), 합성 CoT(여러 연산을 하나의 단계로 결합), 그리고 암묵적 CoT(중간 연산을 생략)로 구성된 CoT 분류 체계를 제안합니다. 우리는 난이도, 압축 수준, 데이터 크기를 정밀하게 조절할 수 있는 인공적인 합성 추론 작업을 구축하고, 다양한 모델 계열 및 크기에 대한 광범위한 실험을 수행했습니다. 주목할 만한 결과는 다음과 같습니다 (i) 더 거친 CoT는 더 많은 SFT 데이터를 필요로 합니다 (ii) 명시적 CoT에 비해, 합성 CoT와 암묵적 CoT는 데이터 확장으로부터 더 큰 이점을 얻으며, 합성 CoT는 데이터 반복으로부터 이점을 얻고, 암묵적 CoT는 암기 현상을 유발하는 경향이 있습니다 (iii) SFT와 달리, 검증 가능한 보상(RLVR)을 사용한 강화 학습은 SFT 단계에서 학습된 압축된 단계를 분해합니다 (iv) 단방향 CoT 순서는 더 긴 순차적 작업에서 더 강력한 일반화 성능을 보여줍니다. 본 연구 결과는 데이터 자원 제약 조건 하에서의 CoT 설계에 대한 시사점을 제공하며, LLM 후속 학습에서의 SFT 및 강화 학습 메커니즘에 대한 중요한 통찰력을 제시합니다.

Original Abstract

Large language models (LLMs) can now solve complex problems through long chain-of-thought (CoT) reasoning, but the trade-off between performance and token cost remains a central challenge. To address this issue, supervised fine-tuning (SFT) often uses compressed reasoning data, where CoT traces are shortened into compact forms. However, the effect of such compressed reasoning data on post-training remains poorly understood. In this paper, we propose a taxonomy of CoT consisting of Explicit CoT, which outputs all operations without aggregation, Composed CoT, which combines multiple operations into a single step, and Implicit CoT, which omits intermediate operations. We construct a synthetic compositional reasoning task that allows controlled variation of difficulty, compression granularity, and data size, and conducted a comprehensive set of experiments across different model families and sizes. Notably, we find that (i) coarser CoT requires more SFT data, (ii) compared with Explicit CoT, Composed CoT and Implicit CoT benefit more from data scaling, while Composed CoT benefits from data repetition and Implicit CoT tends to lead to memorization, (iii) unlike SFT, subsequent reinforcement learning (RL) with verifiable rewards (RLVR) decomposes compressed steps learned during SFT, and (iv) unidirectional CoT ordering shows stronger generalization on longer sequential tasks. Our findings provide implications for CoT design under data resource constraints and offer important insights into the mechanisms of SFT and RL in LLM post-training.

1 Citations
0 Influential
11.5 Altmetric
58.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!