2603.07598v1 Mar 08, 2026 cs.AI

짧은 사고, 동일한 답변: 난이도 기반 세분화된 강화학습을 통한 CoT 압축

Shorter Thoughts, Same Answers: Difficulty-Scaled Segment-Wise RL for CoT Compression

Hongyu Lin
Hongyu Lin
Citations: 4,172
h-index: 29
Ye Tian
Ye Tian
Citations: 472
h-index: 10

체인 오브 씽킹(CoT)은 추론의 신뢰성을 향상시키지만 토큰 비용을 증가시키므로, 명시적인 추론 과정을 압축하는 것이 필요합니다. 하지만 최적의 추론 길이는 난이도, 모델 용량, 학습 상태에 따라 달라지므로, 고정된 길이의 목표는 유연성이 떨어집니다. 기존의 강화학습 기반 압축 방법은 사용자가 보는 답변을 부적절하게 짧게 만들 수 있는데, 이는 단일 완료 레벨의 학습 신호가 사고/답변 경계를 넘어 전달되기 때문입니다. 우리는 난이도 기반 세분화된 그룹 상대적 정책 최적화(DSS-GRPO)를 제안합니다. DSS-GRPO는 보상을 사고 및 답변 구성 요소로 분해하고, 각 세그먼트에 대한 그룹 상대적 이점을 계산하며, 하드 토큰 마스크를 사용하여 압축 업데이트가 사고 부분에만 영향을 미치도록 하고, 답변 정렬이 답변 부분에만 영향을 미치도록 합니다. DSS-GRPO는 프롬프트 단위의 그룹 내 성형과 난이도 인지적 스케일링을 사용하여 답변의 품질을 유지하면서 간결한 추론을 장려합니다.

Original Abstract

Chain-of-thought (CoT) improves reasoning reliability but increases token cost, motivating post-training compression of explicit reasoning traces. However, the shortest sufficient reasoning is not universal: it depends on difficulty, model capacity, and training state, making fixed length targets brittle. In practice, naive RL-based compression can also undesirably shorten the user-facing answer, because a single completion-level learning signal leaks across the think/answer boundary. We propose Difficulty-Scaled Segment-Wise GRPO (DSS-GRPO), which decomposes returns into think and answer components, computes group-relative advantages per segment, and routes them with hard token masks so compression updates act only on think while answer alignment acts only on answer. DSS-GRPO uses prompt-wise within-group shaping and difficulty-aware scaling to encourage concise reasoning without collapsing answer behavior.

0 Citations
0 Influential
14.5 Altmetric
72.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!