2608.05148v1 Aug 05, 2026 cs.CL

추론 코어: 완성 학습 기반 추론 훈련을 위한 광범위한 절차적 데이터 설계

Reasoning Core: Designing Broad Procedural Data for Completion-Supervised Reasoning Training

Damien Sileo
Damien Sileo
Citations: 7
h-index: 2
V. Lacombe
V. Lacombe
Citations: 5
h-index: 1
Dimitri Kachler
Dimitri Kachler
Citations: 0
h-index: 0

절차적 생성기는 대규모의 유용한 검증 가능한 추론 문제를 생성하지만, 완성 학습(completion-supervised) 미세 조정을 위한 데이터로 활용되는 경우는 상대적으로 적었습니다. 본 논문에서는 수학, 논리, 계획 수립, 상태 추적, 형식 언어, 구조화된 데이터, 게임, 인과 관계 및 코드를 포괄하는 50개의 생성기를 모아놓은 '추론 코어(Reasoning Core)'를 소개합니다. 이 컬렉션에는 의미론적 평가 도구, 난이도 조절 기능 및 작업 평가 도구가 포함되어 있습니다. 우리는 일관된 완성 학습 프로토콜을 사용하여 추론 코어를 Procedural Warmup, Reasoning Gym 및 SynLogic과 비교하고, 네 가지 기본 모델 설정과 다양한 훈련 기간 동안의 성능을 측정했습니다. 주요 실험에서 30억 개의 파라미터를 가진 모델을 사용했을 때, 추론 코어는 DROP, LogiQA 및 ARC-Challenge에서 가장 높은 평균 점수를 달성했으며, 절차적 데이터를 사용하지 않은 기준 모델 및 다른 세 가지 절차적 데이터 컬렉션보다 우수한 성능을 보였습니다. 작업 수준 분석 결과, 의미론적 유효성이 훈련에 도움이 된다고 보장하는 것은 아니며, 이는 간결한 목표 설정과 적절하게 조절된 난이도가 중요한 설계 요소임을 보여줍니다. 모델 지원 리뷰, 인간 검토 및 회귀 테스트를 결합한 감사 절차를 수행했습니다. 이러한 감사는 추론 코어 개발 과정 전반에 걸쳐 다른 컬렉션에도 적용되었으며, 생성, 렌더링, 목표 설정 및 평가 간의 미묘한 불일치를 드러냈습니다. 이는 절차적 생성이 정확성을 보장하지 않는다는 점을 상기시켜줍니다. 이 라이브러리, 생성된 데이터셋 및 감사 자료는 공개적으로 이용 가능합니다.

Original Abstract

Procedural generators produce useful verifiable reasoning problems at scale, but have received less attention as data for completion-supervised fine-tuning. We introduce Reasoning Core, a collection of 50 generators spanning mathematics, logic, planning, state tracking, formal languages, structured data, games, causality, and code, with semantic scorers, difficulty controls, and task evaluators. Under a matched completion-supervised protocol, we compare Reasoning Core with Procedural Warmup, Reasoning Gym, and SynLogic across four base-model settings and multiple training durations. In the primary 3B comparison, Reasoning Core achieves the highest mean scores on DROP, LogiQA, and ARC-Challenge, exceeding both the baseline without procedural data and all three alternative procedural collections. Task-level analyses show that semantic validity alone does not ensure training utility, highlighting compact targets and calibrated difficulty as important design factors. We ran audits combining model-assisted review, human adjudication, and regression testing. Applied throughout Reasoning Core development and to the other collections, they reveal subtle mismatches among generation, rendering, targets, and scoring, a reminder that procedural generation alone does not guarantee correctness. The library, generated datasets, and audit material are publicly available.

0 Citations
0 Influential
1 Altmetric
5.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!