2603.21140v1 Mar 22, 2026 cs.AI

ORACLE: 제약 조건 기반 합성 데이터 추출을 통한 대규모 언어 모델의 추론 능력 최적화

ORACLE: Optimizing Reasoning Abilities of Large Language Models via Constraint-Led Synthetic Data Elicitation

Zhuojie Yang
Zhuojie Yang
Citations: 43
h-index: 1
Wentao Wan
Wentao Wan
Citations: 20
h-index: 2
Keze Wang
Keze Wang
Citations: 873
h-index: 16

대규모 언어 모델(LLM)을 훈련할 때 합성 추론 데이터를 활용하여 추론 능력을 향상시키는 방법이 널리 사용되고 있습니다. 이 방법의 효과를 결정하는 핵심 요소는 생성된 다단계 추론 데이터의 품질입니다. 고품질의 추론 데이터를 생성하기 위해, 최근의 많은 방법들은 합성 추론 경로를 생성하고 최종 답변의 정확성을 기준으로 필터링하지만, 중간 추론 단계의 오류를 간과하는 경우가 많습니다. 중간 추론 단계의 검증을 강화하기 위해, 기존 연구에서는 주로 코드 실행 또는 심볼릭 추론 엔진을 사용합니다. 그러나 코드 기반 검증은 코딩 또는 수학 문제에만 적용 가능하며, 추론 엔진은 잘 구조화되고 완전한 맥락을 필요로 합니다. 결과적으로, 기존 방법들은 모호하거나 불완전한 맥락을 포함하는 자연어 추론 작업에서 효과적으로 작동하지 못합니다. 이러한 작업에서, 합성 데이터는 여전히 각 추론 단계를 검증하기 위한 신뢰할 수 있는 메커니즘이 부족합니다. 이러한 문제를 해결하기 위해, 우리는 삼단 논리에 영감을 받은 구조화된 데이터 생성 프레임워크인 ORACLE을 제안합니다. ORACLE은 LLM의 생성 능력을 심볼릭 감독과 통합합니다. LLM은 단계별 추론 맥락을 생성하고, 심볼릭 추론 엔진은 각 중간 단계의 유효성을 검증합니다. ORACLE은 통일된 프롬프트 템플릿을 사용하여 모듈화된 추론 체인을 생성하며, 이를 통해 세분화된 단계별 검증이 가능하여 고품질의 다단계 추론 데이터를 구축할 수 있습니다. 우리는 여섯 가지 논리적, 사실적, 상식적 추론 벤치마크에서 ORACLE이 다양한 모델에서 강력한 기준 모델들을 꾸준히 능가하는 것을 확인했습니다.

Original Abstract

Training large language models (LLMs) with synthetic reasoning data has become a popular approach to enhancing their reasoning capabilities, while a key factor influencing the effectiveness of this paradigm is the quality of the generated multi-step reasoning data. To generate high-quality reasoning data, many recent methods generate synthetic reasoning paths and filter them based on final answer correctness, often overlooking flaws in intermediate reasoning steps. To enhance the verification of intermediate reasoning steps, prior work primarily resorts to code execution or symbolic reasoning engines. However, code-based validation is restricted to code or mathematical tasks, and reasoning engines require a well-structured and complete context. As a result, existing methods fail to function effectively in natural language reasoning tasks that involve ambiguous or incomplete contexts. In these tasks, synthetic data still lack reliable checks for verifying each reasoning step. To address this challenge, we introduce ORACLE, a structured data generation framework inspired by syllogistic reasoning. ORACLE integrates the generative strengths of LLMs with symbolic supervision: the LLM produces step-wise reasoning contexts, while a symbolic reasoning engine verifies the validity of each intermediate step. By employing a unified prompting template to elicit modular reasoning chains, ORACLE enables fine-grained, step-level validation, facilitating the construction of high-quality multi-step reasoning data. Across six logical, factual, and commonsense reasoning benchmarks, our ORACLE consistently outperforms strong baselines on multiple models.

0 Citations
0 Influential
8 Altmetric
40.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!