ETR: 효율적인 연쇄적 사고 추론을 위한 엔트로피 추세 보상
ETR: Entropy Trend Reward for Efficient Chain-of-Thought Reasoning
연쇄적 사고(Chain-of-Thought, CoT) 추론은 복잡한 작업에서 대규모 언어 모델의 성능을 향상시키지만, 종종 지나치게 길고 비효율적인 추론 과정을 생성합니다. 기존 방법들은 길이 페널티나 전역 엔트로피 감소를 사용하여 CoT를 단축하지만, 이는 추론 과정 전반에 걸쳐 낮은 불확실성이 바람직하다는 암묵적인 가정을 기반으로 합니다. 우리는 오히려 추론 효율성이 불확실성의 변화 경로에 의해 결정된다는 것을 보여줍니다. 뚜렷한 하향 엔트로피 추세를 보이는 CoT는 훨씬 짧은 경향이 있습니다. 이러한 통찰력을 바탕으로, 우리는 점진적인 불확실성 감소를 장려하면서 제한적인 지역적 탐색을 허용하는 경로 인지적인 목표 함수인 엔트로피 추세 보상(Entropy Trend Reward, ETR)을 제안합니다. 우리는 ETR을 Group Relative Policy Optimization (GRPO)에 통합하고 다양한 추론 모델과 어려운 벤치마크를 통해 평가했습니다. ETR은 지속적으로 우수한 정확도-효율성 균형을 달성하며, DeepSeek-R1-Distill-7B 모델의 정확도를 9.9% 향상시키면서 네 가지 벤치마크에서 CoT 길이를 67% 단축했습니다. 코드 및 관련 정보는 다음 링크에서 확인할 수 있습니다: https://github.com/Xuan1030/ETR
Chain-of-thought (CoT) reasoning improves large language model performance on complex tasks, but often produces excessively long and inefficient reasoning traces. Existing methods shorten CoTs using length penalties or global entropy reduction, implicitly assuming that low uncertainty is desirable throughout reasoning. We show instead that reasoning efficiency is governed by the trajectory of uncertainty. CoTs with dominant downward entropy trends are substantially shorter. Motivated by this insight, we propose Entropy Trend Reward (ETR), a trajectory-aware objective that encourages progressive uncertainty reduction while allowing limited local exploration. We integrate ETR into Group Relative Policy Optimization (GRPO) and evaluate it across multiple reasoning models and challenging benchmarks. ETR consistently achieves a superior accuracy-efficiency tradeoff, improving DeepSeek-R1-Distill-7B by 9.9% in accuracy while reducing CoT length by 67% across four benchmarks. Code is available at https://github.com/Xuan1030/ETR
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.