CAT: 효율적인 대규모 추론 모델을 위한 신뢰도 적응적 추론
CAT: Confidence-Adaptive Thinking for Efficient Reasoning of Large Reasoning Models
대규모 추론 모델(LRM)은 긴 연쇄적 사고(CoT) 경로를 활용하여 복잡한 작업에서 뛰어난 성과를 거두었지만, 종종 단순한 질문에 대해 과도하게 생각하는 경향이 있어 상당한 토큰 오버헤드와 감소된 추론 효율성을 초래합니다. 그러나 기존의 압축 방법은 주로 일관된 길이 축소 방식을 사용하거나 세분화되지 않은 난이도 추정 방식에 의존하여, 종종 어려운 문제에서 성능 저하를 유발합니다. 이러한 한계를 극복하기 위해, 본 논문에서는 모델 자체의 신뢰도를 활용하여 추론 길이를 자동으로 조절하는 프레임워크인 신뢰도 적응적 추론(CAT)을 제안합니다. 실험 결과는 CAT가 다양한 기본 모델에서 여러 벤치마크를 통해 추론 정확성 측면에서 최첨단 baseline 모델보다 일관되게 우수한 성능을 보임을 보여줍니다. 본 연구는 LRM이 확신하는 답변은 효과적으로 압축하고, 불확실한 답변에 대해서는 신중하게 고려함으로써 실제 산업 환경에서 정확성과 지연 시간의 균형을 맞추는 데 잠재적으로 강력한 솔루션을 제공합니다.
Large Reasoning Models (LRMs) have achieved remarkable success on complex tasks by leveraging long chain-of-thought (CoT) trajectories, yet they frequently exhibit overthinking on simple queries, resulting in significant token overhead and reduced inference efficiency. However, existing compression methods predominantly apply uniform length reduction or rely on coarse-grained difficulty estimation, often leading to performance degradation on difficult problems. To address this limitation, we propose Confidence-Adaptive Thinking (CAT), a framework that incorporates the model's intrinsic self-certainty signals as confidence into the preference optimization process, which autonomously modulates reasoning lengths based on problem difficulty. Experimental results show that CAT consistently outperforms state-of-the-art baselines on reasoning accuracy across multiple benchmarks on different base models. Our work enables LRMs to effectively compress confident responses while deliberating on uncertain ones, offering a potentially robust solution for balancing accuracy and latency in practical industrial scenarios.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.