테스트 시간 스케일링을 활용한 소규모 언어 모델을 이용한 베트남어 추론 능력 격차 해소 연구
Bridging the Reasoning Gap in Vietnamese with Small Language Models via Test-Time Scaling
인공지능의 보편화는 제한된 자원을 가진 장치에서 정교한 추론 기능을 구현하는 데 달려 있습니다. 그러나 소규모 언어 모델(SLM)은 특히 베트남어와 같이 영어가 아닌 언어에서 일관된 사고 과정을 유지하는 데 어려움을 겪는 '추론 격차'를 보이는 경우가 많습니다. 본 연구에서는 베트남어 초등 수학 분야에서 Qwen3-1.7B 아키텍처에 대한 테스트 시간 스케일링 전략을 조사합니다. Gemini 2.5 Flash-Lite 기반 파이프라인을 통해 구축된 고품질 추론 데이터셋인 Vi-S1K와, 엄격한 평가를 위한 이중 자원 벤치마크인 Vi-Elementary-Bench를 소개합니다. LLM-as-a-Judge 프로토콜을 사용하여 분석한 결과, 기본 모델은 강력한 잠재적 지식을 보유하고 있지만(정확도: 5.00점 만점에 4.05점), 의사소통에서 심각한 '형식 격차'를 보입니다. 지도 미세 조정(SFT)은 중요한 '추론 능력 활성화' 역할을 하며, 설명 품질을 77% 향상시키고, 단순 계산과 교육적 일관성 사이의 격차를 해소합니다. 또한, 프롬프트 전략 분석 결과, ReAct와 같은 구조화된 프레임워크는 1.7B 파라미터 용량에 '인지적 부담'을 주어, 순수한 Chain-of-Thought(CoT)와 Self-Consistency를 결합한 방식에 비해 성능이 저하되는 것으로 나타났습니다. 이러한 연구 결과는 SLM의 배포 우선순위를 확립하며, 복잡한 에이전트 기반 워크플로우보다 단순화된 테스트 시간 스케일링과 결합된 SFT가 엣지 기반 추론에 더 우수함을 보여줍니다.
The democratization of ubiquitous AI hinges on deploying sophisticated reasoning capabilities on resource-constrained devices. However, Small Language Models (SLMs) often face a "reasoning gap", particularly in non-English languages like Vietnamese, where they struggle to maintain coherent chains of thought. This paper investigates Test-Time Scaling strategies for the Qwen3-1.7B architecture within the context of Vietnamese Elementary Mathematics. We introduce Vi-S1K, a high-fidelity reasoning dataset localized via a Gemini 2.5 Flash-Lite powered pipeline, and Vi-Elementary-Bench, a dual-resource benchmark for rigorous evaluation. Using an LLM-as-a-Judge protocol, we reveal that the base model possesses robust latent knowledge (Accuracy: 4.05/5.00) but suffers from a severe "formatting gap" in communication. Supervised Fine-Tuning (SFT) acts as a critical "reasoning unlocker", yielding a 77% improvement in Explanation Quality and bridging the gap between raw calculation and pedagogical coherence. Furthermore, our analysis of prompting strategies uncovers a significant trade-off: structured frameworks like ReAct impose a "cognitive tax" on the 1.7B parameter capacity, degrading performance relative to pure Chain-of-Thought (CoT) combined with Self-Consistency. These findings establish a deployment hierarchy for SLMs, demonstrating that SFT combined with simplified test-time scaling is superior to complex agentic workflows for edge-based reasoning.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.