2608.05254v1 Aug 05, 2026 cs.CL

제약 조건 우선 추론: 수학 문제 해결 시 답변 공간의 제약 조건을 활용하는 학습이 필요 없는 방법

Constraint-First Reasoning: A Training-Free Protocol for Exploiting Answer-Space Constraints in Mathematical Problem Solving

Jiajun Fan
Jiajun Fan
Citations: 78
h-index: 5
Ge Liu
Ge Liu
Citations: 32
h-index: 4
Bang Yang
Bang Yang
Citations: 17
h-index: 2
Hongbo Ma
Hongbo Ma
Citations: 1
h-index: 1
Hanwen Zhang
Hanwen Zhang
Citations: 0
h-index: 0
Yunqian Selina Cheng
Yunqian Selina Cheng
Citations: 0
h-index: 0

대규모 언어 모델은 타당한 수학적 결과를 도출할 수 있지만, 명시적인 요구 사항을 위반하는 경우가 있습니다. 예를 들어, 모듈러 환원을 생략하거나 정수가 아닌 값을 반환하거나 잘못된 인코딩 형식의 답을 사용하는 것입니다. 본 논문에서는 학습이 필요 없는 두 단계 프롬프팅 방식인 제약 조건 우선 추론(Constraint-First Reasoning, CFR)을 소개합니다. 1단계에서는 문제에서 요구하는 제약 조건을 추출하고 요약하며, 2단계에서는 그 요약을 바탕으로 중간 및 최종 결과를 검증하면서 문제를 해결합니다. Routed-CFR은 텍스트 기반 정규 표현식 라우터를 사용하여 제한적인 단서가 감지되면 두 단계 프로토콜을 활성화하고, 그렇지 않은 경우에는 직접적인 연쇄적 사고(Chain-of-Thought, CoT) 방식을 사용합니다. AIME, CMIMC, BRUMO 및 AIMO_AMC 데이터 세트에서 본 방법은 다양한 기반 모델에서 직접적인 CoT 성능을 향상시킵니다. 또한, 규칙 기반 라우팅 실험, 동일한 프롬프트를 사용하는 기준선 비교, 문제 수준의 페어링 테스트, 디코딩 안정성 검증, 제약 조건 품질 평가, 토큰 사용량 분석 및 OlympiadBench 평가 결과를 보고합니다. 이러한 분석을 통해 CFR은 일반적인 수학적 추론 방식을 대체하는 것이 아니라, 복구 가능한 제약 조건과 신뢰할 수 있는 1단계 추출에 의존하는 특정 상황에서 유용한 테스트 시간 개입 방법임을 알 수 있습니다.

Original Abstract

Large language models can derive a plausible mathematical object yet still violate explicit requirements--for example, by omitting a modular reduction, returning a non-integer, or using the wrong encoded answer form. We introduce Constraint-First Reasoning (CFR), a training-free two-stage prompting protocol: Stage 1 extracts and summarizes constraints entailed by the problem, and Stage 2 solves while checking intermediate and final results against that summary. Routed-CFR activates the two-stage protocol only when a text-only regex router detects restrictive cues; otherwise it uses direct chain-of-thought (CoT). Across AIME, CMIMC, BRUMO, and AIMO_AMC, the method improves direct CoT on multiple backbones. We further report convention-controlled routing experiments, matched prompting baselines, problem-level paired tests, decoding robustness, constraint-quality audits, total-token accounting, and an OlympiadBench evaluation. These analyses position CFR as a targeted test-time intervention whose benefit depends on recoverable constraints and reliable Stage 1 extraction, rather than as a general-purpose replacement for mathematical reasoning.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!