SLAI T-Rex: DeepSeek-V4 모델 패밀리에 대한 완전 파라미터 기반 후속 학습 연구 (Ascend SuperPOD 환경)
SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD
조규(trillion) 규모의 MoE 모델에 대한 완전 파라미터 기반 후속 학습은 대규모 분산 학습에서 심각한 메모리 압박, 통신 오버헤드 중첩 부재, 비효율적인 커널 실행 등 시스템 수준의 상당한 과제를 야기합니다. 대부분의 대규모 LLM 학습 시스템이 GPU 기반 클러스터를 중심으로 구축되는 반면, 본 연구에서는 Ascend NPU SuperPOD 환경에서 엔드투엔드 최적화 방안을 제시합니다. DeepSeek-V4 모델 패밀리를 대상으로 모델 수준 병렬 처리, 연산-통신 조율 및 저수준 커널 실행을 포괄하는 계층적 최적화 프레임워크를 개발했습니다. 결과적으로 개발된 시스템은 34.22%의 모델 FLOPs 활용률(MFU)을 달성했으며, 이는 오픈 소스 기준 레시피 대비 2.93배 향상된 성능이며, 동시에 학습 안정성을 유지합니다. 이 최적화된 인프라를 기반으로 복잡한 운영 연구(OR) 작업에 대한 지속적인 사전 학습(CPT) 및 지도 미세 조정(SFT) 워크플로우를 구축했습니다. 본 통합 프레임워크를 SLAI T-Rex라고 명명합니다. DeepSeek-V4-Flash 모델을 사용하여 수집된 도메인 리소스와 솔버 검증된 합성 최적화 문서를 결합한 OR 특화 CPT 및 SFT 데이터 파이프라인을 개발했습니다. 결과적으로 생성된 데이터셋은 네 가지 작업 범주와 세 가지 문제 표현 방식을 포괄하는 10,000개의 고품질 SFT 샘플로 구성되어 있습니다. 이 전문 모델은 평가 대상 모델 중 가장 높은 평균 제로샷 Pass@1 점수를 달성했으며, 이는 71.81%로 GPT-5.4-Mini 및 기본 DeepSeek-V4-Flash 모델보다 각각 3.98% 및 11.27% 더 높은 수치입니다. 전반적으로 본 연구는 Ascend 인프라를 활용한 조규 규모 모델의 효율적인 후속 학습부터 솔버 기반 수학적 모델링을 위한 도메인 특화 Flash 모델 개발까지, 복잡한 추론을 위한 최첨단 모델 시스템 구축에 기여합니다.
Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challenges for large-scale distributed training, including severe memory pressure, non-overlapped communication overhead, and inefficient kernel execution. While most large-scale LLM training systems are built around GPU-based clusters, this report presents an end-to-end optimization practice on the Ascend NPU SuperPOD. Using the DeepSeek-V4 model family as the target workload, we develop a hierarchical optimization framework spanning model-level parallelism, computation-communication orchestration, and low-level kernel execution. The resulting system achieves 34.22% Model FLOPs Utilization (MFU) with a 2.93x improvement over the open-source baseline recipe while maintaining training stability. Building on this optimized infrastructure, we further establish a CPT and SFT workflow for complex Operations Research (OR) tasks. We refer to the integrated framework as SLAI T-Rex. Using DeepSeek-V4-Flash, we develop OR-oriented CPT and SFT data pipelines that combine collected domain resources with solver-verified synthetic optimization documents. The resulting dataset contains 10K high-quality SFT samples spanning four task categories and three problem representations. The specialized model achieves the highest average zero-shot Pass@1 score among the evaluated models, reaching 71.81% and outperforming GPT-5.4-Mini and the base DeepSeek-V4-Flash model by 3.98 and 11.27 percentage points, respectively. Overall, this work demonstrates a full-stack pathway from efficient trillion-parameter model post-training on Ascend infra to domain-specialized Flash models for solver-grounded mathematical modeling, advancing frontier-model systems for complex reasoning.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.