2606.17687v1 Jun 16, 2026 cs.CL

SuCo: 충분성 기반의 연속적 적응 추론

SuCo: Sufficiency-guided Continuous Adaptive Reasoning

Xuebo Liu
Xuebo Liu
Citations: 477
h-index: 11
Min Zhang
Min Zhang
Citations: 24
h-index: 3
Jiahao Wang
Jiahao Wang
Citations: 8
h-index: 2
Bin Liang
Bin Liang
Citations: 19
h-index: 1
Chenhao Hu
Chenhao Hu
Citations: 15
h-index: 3
Longhui Zhang
Longhui Zhang
Citations: 55
h-index: 4
Jing Li
Jing Li
Citations: 8
h-index: 2
Xuelong Li
Xuelong Li
Citations: 135
h-index: 6

복잡한 작업에서 뛰어난 성능을 보이는 대규모 추론 모델(LRM)은 종종 지나치게 긴 사고 과정(Chain-of-Thoughts, CoT)을 생성하여 간단한 쿼리에도 계산 비용을 증가시키는 경향이 있습니다. 이러한 비효율성을 완화하기 위한 기존의 노력은 주로 이산적인 추론 모드 또는 고정된 예산 단계를 사용하며, 언제 추론이 충분한지에 대한 명확한 기준이 부족합니다. 본 연구에서는 정답을 도출하는 데 적절한 CoT 경로의 가장 짧은 부분을 '최소 충분 CoT (Minimal Sufficient CoT, MSC)'라고 정의합니다. 실험적으로 MSC는 추론에 필요한 토큰 수를 줄일 뿐만 아니라 난이도 수준 전반에 걸쳐 정확도를 향상시키는 것을 확인했습니다. MSC를 기반으로, 우리는 연속적인 스펙트럼에서 자율적인 추론 제어를 위한 두 단계의 훈련 프레임워크인 '충분성 기반의 연속적 적응 추론 (Sufficiency-guided Continuous Adaptive Reasoning, SuCo)'를 제안합니다. 1단계에서는 문제에 특화된 충분성 임계값을 사용하여 MSC 데이터를 구축하고, 이는 질문 난이도에 따라 자연스럽게 조정됩니다. 그런 다음 모델을 미세 조정하여 간결하면서도 충분한 추론 패턴을 내재화하도록 합니다. 2단계에서는 복잡성을 동적으로 추적하고 과도하거나 부족한 사고를 모두 처벌하는 충분성 기반의 보상을 사용하는 강화 학습을 통해 모델을 추가로 최적화합니다 (Sufficiency-Aware Policy Optimization, SAPO). 수학, 코드 및 과학 벤치마크에 대한 광범위한 실험 결과, SuCo는 정확도와 추론 효율성 측면에서 일관되게 성능 향상을 달성하는 것으로 나타났습니다.

Original Abstract

Despite remarkable performance on complex tasks, Large Reasoning Models (LRMs) often generate excessively long Chain-of-Thoughts (CoT), inflating computational costs even for simple queries. Existing efforts to mitigate this inefficiency typically rely on discrete reasoning modes or fixed budget tiers, lacking a principled criterion of when reasoning is sufficient. In this work, we introduce Minimal Sufficient CoT (MSC), defined as the shortest prefix of a CoT trajectory which is adequate for producing the correct answer. We empirically show that MSC not only reduces reasoning tokens, but also improves accuracy across difficulty levels. Building on MSC, we propose Sufficiency-guided Continuous Adaptive Reasoning (SuCo), a two-stage training framework for autonomous reasoning control along a continuous spectrum. In stage 1, MSC-Aligned Fine-Tuning (MFT) constructs MSC data using problem-adaptive sufficiency thresholds that naturally scale with question difficulty, then fine-tunes the model to internalize concise yet sufficient reasoning patterns. In stage 2, Sufficiency-Aware Policy Optimization (SAPO) further optimizes the model through reinforcement learning with dynamic complexity tracking and sufficiency-aware rewards that penalize both over- and under-thinking. Extensive experiments across mathematics, code, and science benchmarks show that SuCo consistently achieves improvements in both accuracy and reasoning efficiency.

0 Citations
0 Influential
5.5 Altmetric
27.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!