2604.17794v1 Apr 20, 2026 cs.CL

테스트 시간 스케일링을 활용한 소규모 언어 모델을 이용한 베트남어 추론 능력 격차 해소 연구

Bridging the Reasoning Gap in Vietnamese with Small Language Models via Test-Time Scaling

Bui The Trung
Bui The Trung
Citations: 0
h-index: 0
Nguyen Van Vinh
Nguyen Van Vinh
Citations: 9
h-index: 2
B. Trinh
B. Trinh
Citations: 227
h-index: 9
Duc Do Minh
Duc Do Minh
Citations: 155
h-index: 4

인공지능의 보편화는 제한된 자원을 가진 장치에서 정교한 추론 기능을 구현하는 데 달려 있습니다. 그러나 소규모 언어 모델(SLM)은 특히 베트남어와 같이 영어가 아닌 언어에서 일관된 사고 과정을 유지하는 데 어려움을 겪는 '추론 격차'를 보이는 경우가 많습니다. 본 연구에서는 베트남어 초등 수학 분야에서 Qwen3-1.7B 아키텍처에 대한 테스트 시간 스케일링 전략을 조사합니다. Gemini 2.5 Flash-Lite 기반 파이프라인을 통해 구축된 고품질 추론 데이터셋인 Vi-S1K와, 엄격한 평가를 위한 이중 자원 벤치마크인 Vi-Elementary-Bench를 소개합니다. LLM-as-a-Judge 프로토콜을 사용하여 분석한 결과, 기본 모델은 강력한 잠재적 지식을 보유하고 있지만(정확도: 5.00점 만점에 4.05점), 의사소통에서 심각한 '형식 격차'를 보입니다. 지도 미세 조정(SFT)은 중요한 '추론 능력 활성화' 역할을 하며, 설명 품질을 77% 향상시키고, 단순 계산과 교육적 일관성 사이의 격차를 해소합니다. 또한, 프롬프트 전략 분석 결과, ReAct와 같은 구조화된 프레임워크는 1.7B 파라미터 용량에 '인지적 부담'을 주어, 순수한 Chain-of-Thought(CoT)와 Self-Consistency를 결합한 방식에 비해 성능이 저하되는 것으로 나타났습니다. 이러한 연구 결과는 SLM의 배포 우선순위를 확립하며, 복잡한 에이전트 기반 워크플로우보다 단순화된 테스트 시간 스케일링과 결합된 SFT가 엣지 기반 추론에 더 우수함을 보여줍니다.

Original Abstract

The democratization of ubiquitous AI hinges on deploying sophisticated reasoning capabilities on resource-constrained devices. However, Small Language Models (SLMs) often face a "reasoning gap", particularly in non-English languages like Vietnamese, where they struggle to maintain coherent chains of thought. This paper investigates Test-Time Scaling strategies for the Qwen3-1.7B architecture within the context of Vietnamese Elementary Mathematics. We introduce Vi-S1K, a high-fidelity reasoning dataset localized via a Gemini 2.5 Flash-Lite powered pipeline, and Vi-Elementary-Bench, a dual-resource benchmark for rigorous evaluation. Using an LLM-as-a-Judge protocol, we reveal that the base model possesses robust latent knowledge (Accuracy: 4.05/5.00) but suffers from a severe "formatting gap" in communication. Supervised Fine-Tuning (SFT) acts as a critical "reasoning unlocker", yielding a 77% improvement in Explanation Quality and bridging the gap between raw calculation and pedagogical coherence. Furthermore, our analysis of prompting strategies uncovers a significant trade-off: structured frameworks like ReAct impose a "cognitive tax" on the 1.7B parameter capacity, degrading performance relative to pure Chain-of-Thought (CoT) combined with Self-Consistency. These findings establish a deployment hierarchy for SLMs, demonstrating that SFT combined with simplified test-time scaling is superior to complex agentic workflows for edge-based reasoning.

0 Citations
0 Influential
4.5 Altmetric
22.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!