MathShikkha: 방글라어 수학적 추론 능력을 갖춘 소규모 언어 모델에 대한 답변 기반 및 연쇄적 사고 지도 학습 연구
MathShikkha: A Controlled Study of Answer-Only and Chain-of-Thought Supervision for Bangla Mathematical Reasoning in Small Language Models
방글라어와 같은 저자원 언어에서 수학적 추론은 여전히 어려운 과제입니다. 본 연구에서는 교사가 생성한 방글라어 연쇄적 사고(CoT) 지도가 일반적인 지도 학습을 넘어 어떤 이점을 제공하는지 조사합니다. GPT-5.4가 생성한 근거 자료를 포함하는 방글라어 수학 추론 데이터셋인 extsc{MathShikkha}를 구축하고, 답변 기반과 CoT 조건을 동일하게 유지하면서 데이터 분할, 손실 마스킹, 디코딩 및 평가 방법을 통제하여 4B~7B 크기의 네 가지 모델을 학습했습니다. 실험 결과, 더 강력한 백본 모델의 경우, CoT는 답변 기반 학습보다 유의미한 개선 효과를 보이지 않았습니다(짝지어진 부트스트랩 95% 신뢰 구간에 0이 포함됨; 정확한 McNemar p-값 ≥ 0.17). 반면, 성능이 낮은 4B 모델에서는 CoT가 18.56 포인트의 상당한 개선을 가져왔습니다(p < 0.0001). 더 큰 규모의 BanglaMATH 벤치마크에서 이러한 경향은 역전되었으며, CoT는 네 가지 모델 모두에서 답변 기반 학습보다 유의미하게 높은 성능을 보였습니다(20.1~28.1 포인트, 모든 p < 0.0001). 답변 기반 학습은 세 모델에서 외부 데이터에 대한 정확도를 기본 모델 수준 이하로 낮추는 반면, CoT는 네 가지 모델 모두에서 이를 유지하거나 향상시켰습니다. 두 명의 공동 저자가 참여한 인간 평가에서는 외부 전문가의 판단과 Cohen's κ 값이 0.76~1.00인 것으로 나타났을 때, 추론 내용 측면에서 CoT가 기본 모델보다 유의미하게 개선되지 않았습니다. 대신, CoT의 측정 가능한 효과는 방글라어 준수 및 검토 가능한 추론 생성에 있었습니다. 전반적으로, 근거 기반 지도 학습의 가치는 백본 모델의 능력과 데이터 분포 변화에 따라 달라지며, 본 연구에서는 Bangla어 준수, 감사 가능성 있는 추론, 외부 데이터에 대한 강건성을 향상시키는 데 주로 기여합니다.
Mathematical reasoning remains challenging in low-resource languages such as Bangla. We study whether teacher-generated Bangla Chain-of-Thought (CoT) supervision provides benefits beyond ordinary supervised fine-tuning. We construct \textsc{MathShikkha}, a Bangla mathematical reasoning dataset with GPT-5.4-generated rationales, and fine-tune four 4B--7B student models under a matched protocol in which answer-only and CoT conditions share data splits, response-only loss masking, decoding, and scoring, differing only in the training target. In-domain, CoT provides no significant improvement over answer-only fine-tuning for three stronger backbones (paired bootstrap 95\% CIs include zero; exact McNemar $p \geq 0.17$), despite generating 15--52$\times$ more tokens, but significantly improves the weaker 4B model by 18.56 points ($p < 0.0001$). On the larger, contamination-audited BanglaMATH benchmark, this pattern reverses: CoT significantly outperforms answer-only supervision for all four models by 20.1--28.1 points (all $p < 0.0001$). Answer-only fine-tuning also reduces out-of-domain accuracy below the base model for three models, whereas CoT preserves or improves it for all four. A human study with two co-author annotators, external-expert adjudication, and Cohen's $κ= 0.76$--$1.00$ finds no significant CoT improvement over the base model on reasoning-content criteria; instead, its measurable effect is target-language adherence and producing inspectable reasoning. Overall, rationale supervision's value depends on backbone capability and distribution shift: in this setting, its main benefits are Bangla adherence, auditable reasoning, and out-of-domain robustness rather than improved in-domain reasoning validity.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.