2608.08503v1 Aug 09, 2026 cs.AI

MathShikkha: 방글라어 수학적 추론 능력을 갖춘 소규모 언어 모델에 대한 답변 기반 및 연쇄적 사고 지도 학습 연구

MathShikkha: A Controlled Study of Answer-Only and Chain-of-Thought Supervision for Bangla Mathematical Reasoning in Small Language Models

Jawad Hossain
Jawad Hossain
Citations: 107
h-index: 5
Rahmat Ali
Rahmat Ali
Citations: 0
h-index: 0

방글라어와 같은 저자원 언어에서 수학적 추론은 여전히 어려운 과제입니다. 본 연구에서는 교사가 생성한 방글라어 연쇄적 사고(CoT) 지도가 일반적인 지도 학습을 넘어 어떤 이점을 제공하는지 조사합니다. GPT-5.4가 생성한 근거 자료를 포함하는 방글라어 수학 추론 데이터셋인 extsc{MathShikkha}를 구축하고, 답변 기반과 CoT 조건을 동일하게 유지하면서 데이터 분할, 손실 마스킹, 디코딩 및 평가 방법을 통제하여 4B~7B 크기의 네 가지 모델을 학습했습니다. 실험 결과, 더 강력한 백본 모델의 경우, CoT는 답변 기반 학습보다 유의미한 개선 효과를 보이지 않았습니다(짝지어진 부트스트랩 95% 신뢰 구간에 0이 포함됨; 정확한 McNemar p-값 ≥ 0.17). 반면, 성능이 낮은 4B 모델에서는 CoT가 18.56 포인트의 상당한 개선을 가져왔습니다(p < 0.0001). 더 큰 규모의 BanglaMATH 벤치마크에서 이러한 경향은 역전되었으며, CoT는 네 가지 모델 모두에서 답변 기반 학습보다 유의미하게 높은 성능을 보였습니다(20.1~28.1 포인트, 모든 p < 0.0001). 답변 기반 학습은 세 모델에서 외부 데이터에 대한 정확도를 기본 모델 수준 이하로 낮추는 반면, CoT는 네 가지 모델 모두에서 이를 유지하거나 향상시켰습니다. 두 명의 공동 저자가 참여한 인간 평가에서는 외부 전문가의 판단과 Cohen's κ 값이 0.76~1.00인 것으로 나타났을 때, 추론 내용 측면에서 CoT가 기본 모델보다 유의미하게 개선되지 않았습니다. 대신, CoT의 측정 가능한 효과는 방글라어 준수 및 검토 가능한 추론 생성에 있었습니다. 전반적으로, 근거 기반 지도 학습의 가치는 백본 모델의 능력과 데이터 분포 변화에 따라 달라지며, 본 연구에서는 Bangla어 준수, 감사 가능성 있는 추론, 외부 데이터에 대한 강건성을 향상시키는 데 주로 기여합니다.

Original Abstract

Mathematical reasoning remains challenging in low-resource languages such as Bangla. We study whether teacher-generated Bangla Chain-of-Thought (CoT) supervision provides benefits beyond ordinary supervised fine-tuning. We construct \textsc{MathShikkha}, a Bangla mathematical reasoning dataset with GPT-5.4-generated rationales, and fine-tune four 4B--7B student models under a matched protocol in which answer-only and CoT conditions share data splits, response-only loss masking, decoding, and scoring, differing only in the training target. In-domain, CoT provides no significant improvement over answer-only fine-tuning for three stronger backbones (paired bootstrap 95\% CIs include zero; exact McNemar $p \geq 0.17$), despite generating 15--52$\times$ more tokens, but significantly improves the weaker 4B model by 18.56 points ($p < 0.0001$). On the larger, contamination-audited BanglaMATH benchmark, this pattern reverses: CoT significantly outperforms answer-only supervision for all four models by 20.1--28.1 points (all $p < 0.0001$). Answer-only fine-tuning also reduces out-of-domain accuracy below the base model for three models, whereas CoT preserves or improves it for all four. A human study with two co-author annotators, external-expert adjudication, and Cohen's $κ= 0.76$--$1.00$ finds no significant CoT improvement over the base model on reasoning-content criteria; instead, its measurable effect is target-language adherence and producing inspectable reasoning. Overall, rationale supervision's value depends on backbone capability and distribution shift: in this setting, its main benefits are Bangla adherence, auditable reasoning, and out-of-domain robustness rather than improved in-domain reasoning validity.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!