VibeThinker-3B: 소규모 언어 모델에서 검증 가능한 추론의 한계를 탐구하다
VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models
본 기술 보고서에서는 VibeThinker-3B를 소개합니다. 이는 엄격한 소규모 모델 환경 내에서 검증 가능한 추론이 얼마나 발전할 수 있는지 조사하기 위해 개발된 30억 개의 파라미터를 가진 컴팩트한 밀집 모델입니다. 스펙트럼-투-시그널(Spectrum-to-Signal) 사후 학습 패러다임을 기반으로, 우리는 커리큘럼 기반의 지도 미세 조정, 다중 도메인 강화 학습 및 오프라인 자기 증류를 포함하는 최적화된 파이프라인을 통해 모델을 체계적으로 향상시켰습니다. 실험 결과는 VibeThinker-3B가 매우 까다로운 검증 가능한 작업에서 최고 수준의 성능을 달성함을 보여줍니다. 특히, AIME26에서 94.3점(청구 수준 테스트 시간 스케일링 시 97.1점으로 향상), LiveCodeBench v6에서 80.2%의 Pass@1 점수를 기록했으며, 최근에 공개되지 않은 LeetCode 콘테스트에서 96.1%의 수용률을 보여주는 강력한 일반화 능력을 보입니다. 이는 VibeThinker-3B를 DeepSeek V3.2, GLM-5 및 Gemini 3 Pro와 같은 훨씬 더 큰 규모의 플래그십 모델과 동등하거나 뛰어넘는 성능 수준의 추론 시스템으로 자리매김합니다. 또한, IFEval에서 93.4점을 기록한 것은 극단적인 추론 향상이 엄격한 명령어 제어 기능을 저해하지 않음을 확인합니다. 이전 연구(1.5B)를 확장하여 얻은 이러한 결과는 검증 가능한 추론이 컴팩트한 추론 코어로 압축될 수 있다는 매개변수 압축-커버리지 가설을 뒷받침합니다. 반면, 개방형 도메인 지식 및 범용 능력에는 사실, 개념 및 장기적인 시나리오에 대한 광범위한 매개변수 커버리지가 필요합니다. 이러한 관점은 컴팩트 모델이 단순히 배포 효율성이 뛰어난 대체재가 아니라, 매개변수가 많은 성능 강도 영역에서 최첨단 성능을 달성하기 위한 보완적인 경로라는 것을 시사합니다.
This technical report introduces VibeThinker-3B, a compact dense model with 3B parameters developed to investigate how far verifiable reasoning can be pushed within a strictly small-model regime. Building upon the Spectrum-to-Signal post-training paradigm, we systematically enhance the model through an optimized pipeline that includes curriculum-based supervised fine-tuning, multi-domain reinforcement learning, and offline self-distillation. Experimental evaluations demonstrate that VibeThinker-3B achieves frontier-level performance on highly demanding verifiable tasks. Specifically, it attains a score of 94.3 on AIME26 (improving to 97.1 with claim-level test-time scaling), an 80.2 Pass@1 on LiveCodeBench v6, and exhibits strong out-of-distribution generalization with a 96.1\% acceptance rate on recent unseen LeetCode contests. This effectively places it in the performance band of first-tier reasoning systems, matching or exceeding flagship models that are orders of magnitude larger, such as DeepSeek V3.2, GLM-5, and Gemini 3 Pro. Furthermore, a score of 93.4 on IFEval confirms that this extreme reasoning enhancement does not compromise strict instruction controllability. Extending our previous 1.5B work, these findings motivate the Parametric Compression-Coverage Hypothesis, which views verifiable reasoning as compressible into compact reasoning cores, while open-domain knowledge and general-purpose competence require broad parameter coverage over facts, concepts, and long-tail scenarios. This perspective suggests that compact models are not merely deployment-efficient substitutes, but a complementary path toward frontier-level performance in parameter-dense capability regimes.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.