2607.05199v1 Jul 06, 2026 cs.AI

이성, 보상, 개선: 구조화된 피드백을 활용한 단계별 오류 수정 - 소규모 언어 모델의 물리학 추론

Reason, Reward, Refine: Step-Level Errors Corrections with Structured Feedback for Physics Reasoning in Small Language Models

R. Shah
R. Shah
Citations: 6,839
h-index: 44
Tanuja Ganu
Tanuja Ganu
Citations: 982
h-index: 13
Raj Jaiswal
Raj Jaiswal
Citations: 67
h-index: 5
Dhruv Jain
Dhruv Jain
Citations: 28
h-index: 2
Rishabh Dhawan
Rishabh Dhawan
Citations: 0
h-index: 0
Sree Krishna Uppalapati
Sree Krishna Uppalapati
Citations: 0
h-index: 0
Shin’ichi Satoh
Shin’ichi Satoh
Citations: 90
h-index: 6

소규모 언어 모델에서 물리학 추론은 구조적으로 실패합니다. 어떤 단계에서의 오류든 전파되어 이후 모든 추론을 왜곡하게 됩니다. 제한적인 도메인 지식, 다단계 유도 과정에서의 환각 현상, 그리고 데이터 분포에 대한 민감성이 이러한 실패를 악화시킵니다. 우리는 단계별 보상 프레임워크를 제안합니다. 이 프레임워크는 첫 번째 추론 오류를 식별하고, 대상화된 구조화된 피드백을 생성하며, KL 정규화를 사용한 정책 경사법을 통해 모델이 정답을 직접 제공하지 않고 자체 솔루션을 수정하도록 훈련합니다. 어노테이션에 의존하는 단계별 방법과 달리, 선호도 데이터 구축이 필요 없으며, 외부 검증기는 오직 훈련 시간 동안만 작동합니다. 다섯 가지 물리학 벤치마크에서 우리의 프레임워크는 CoT 프롬프트보다 17-20%, 그리고 최상의 기준 모델보다 10-16%의 정확도 향상을 보여주었습니다. 계산 오류율은 56.9%에서 23.5%로, 오해 오류율은 22.3%에서 12.0%로 감소했습니다(가장 우수한 결과 기준). 개념적 오류는 89.7%에서 68.7%로 감소했지만, 모든 조건에서 가장 어려운 실패 요인으로 남아 있습니다.

Original Abstract

Physics reasoning fails structurally in small language models: an error at any step propagates forward, corrupting every inference that follows. Limited domain knowledge, hallucination under multi-step derivation, and distributional sensitivity compound this failure. We propose a step-level reward framework that identifies the first reasoning error, generates targeted structured feedback, and trains the model to revise its solution via policy gradient with KL regularization, without exposing it to ground truth solutions as generation targets. Unlike annotation-dependent step-level methods, no preference data construction is required and the external verifier operates exclusively at training time. Across five physics benchmarks, our framework delivers accuracy gains of 17-20% over CoT prompting and 10-16% over the strongest baseline, reduces calculation errors from 56.9% to 23.5%, and reduces miscomprehension errors from 22.3% to 12.0% in the best observed cases. Conceptual errors reduce from 89.7% to 68.7%, yet persist as the hardest failure mode across all conditions.

0 Citations
0 Influential
22 Altmetric
110.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!