2606.16152v1 Jun 15, 2026 cs.AI

품질-유용성 역설: 왜 고보상 데이터가 소형 모델의 수학적 추론 능력을 저해하는가

The Quality-Utility Paradox: Why High-Reward Data Impairs Small Model Mathematical Reasoning

Lei Song
Lei Song
Citations: 32
h-index: 3
Jiang Bian
Jiang Bian
Citations: 173
h-index: 8
Chun Yuan
Chun Yuan
Citations: 208
h-index: 4
Lirong Che
Lirong Che
Citations: 0
h-index: 0
Haolong Qian
Haolong Qian
Citations: 26
h-index: 2
Xianliang Yang
Xianliang Yang
Citations: 95
h-index: 6
Yinuo Ma
Yinuo Ma
Citations: 17
h-index: 3
Feng Lu
Feng Lu
Citations: 13
h-index: 2
Ye Guo
Ye Guo
Citations: 28
h-index: 2

강력한 추론 모델로부터 지식을 전달하여 소형 언어 모델(SLM)의 수학적 추론 능력을 향상시키는 방법은 널리 사용되며, 일반적으로 더 높은 보상 점수를 가진 데이터가 더 유용한 지도 역할을 한다고 가정합니다. 본 연구에서는 수학적 추론 과정에서의 역설적인 **품질-유용성 역설**을 발견했습니다. 더 강력한 오라클(Oracle) 모델에 의해 정제되거나 생성된 데이터는 보상 모델에 의해 더 높은 품질로 평가되지만, Qwen2.5, LLaMA-3 및 DeepSeek 계열 모델에서 SLM 자체에 의해 생성되고 거부 샘플링을 통해 선택된 데이터보다 일관되게 성능이 떨어지는 것으로 나타났습니다. 분석 결과, 오라클 정제는 논리적 오류 수정과 함께 SLM의 원래 추론 분포로부터 벗어나는 분포 변화를 초래합니다. 이러한 변화는 학습 비용을 증가시키며, 향상된 추론 로직의 이점을 능가할 수 있습니다. 이 메커니즘을 검증하기 위해, 본 연구에서는 SLM의 원래 추세는 유지하면서 오라클 모델로부터 논리적 오류 수정 기능을 포함하는 **스타일 일관성 정제(Style-Aligned Refinement)** 방법을 도입했습니다. 이러한 접근 방식은 적응 비용을 줄이고 하위 작업에서의 유용성을 회복합니다. 본 연구 결과는 효과적인 수학적 추론 과정 지도가 단순히 보상 모델 점수에 의존하기보다는, 인지된 해결 품질과 학습자-데이터 호환성을 동시에 최적화해야 함을 시사합니다. 데이터셋 및 코드는 다음 링크에서 확인할 수 있습니다: https://github.com/Dracoqhl/Quality-Utility-Paradox.

Original Abstract

Knowledge distillation from powerful reasoning models is widely used to improve Small Language Models (SLMs) on mathematical reasoning, often assuming that traces with higher reward model scores provide more useful supervision. We identify a counterintuitive \textbf{Quality-Utility Paradox} in mathematical reasoning distillation. Data refined or synthesized by a stronger Oracle obtains higher perceived quality according to reward models, yet consistently underperforms traces generated by the SLM itself and selected through rejection sampling across Qwen2.5, LLaMA-3, and DeepSeek families. Our analysis shows that Oracle refinement couples logical repair with distributional drift away from the SLM's native reasoning distribution. This drift increases the learner's adaptation cost and can outweigh the benefit of improved reasoning logic. To test this mechanism, we introduce \textbf{Style-Aligned Refinement}, which preserves the native trajectory of the SLM while retaining logical repair from the Oracle. This intervention lowers adaptation cost and restores downstream utility. These findings suggest that effective mathematical reasoning distillation should jointly optimize perceived solution quality and learner-data compatibility, rather than relying solely on reward-model scores. The datasets and code are available at https://github.com/Dracoqhl/Quality-Utility-Paradox.

0 Citations
0 Influential
29.493061443341 Altmetric
0.0 Score
Original PDF
2

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!