2607.20145v1 Jul 22, 2026 cs.CL

SLAI T-Rex: DeepSeek-V4 모델 패밀리에 대한 완전 파라미터 기반 후속 학습 연구 (Ascend SuperPOD 환경)

SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Ziwei Zhu
Ziwei Zhu
Citations: 152
h-index: 5
Baotian Hu
Baotian Hu
Citations: 388
h-index: 10
Min Zhang
Min Zhang
Citations: 24
h-index: 3
Xuhui Chen
Xuhui Chen
Citations: 3
h-index: 1
Zhuohan Wang
Zhuohan Wang
Citations: 21
h-index: 2
Jie Yu
Jie Yu
Citations: 667
h-index: 9
Dong Zhang
Dong Zhang
Citations: 203
h-index: 4
Zhihang Lin
Zhihang Lin
Citations: 32
h-index: 3
Dongfang Li
Dongfang Li
Citations: 367
h-index: 10
Xiaodong Luo
Xiaodong Luo
Citations: 18
h-index: 3
Ruoyu Sun
Ruoyu Sun
Citations: 660
h-index: 11
Linyuan Qiu
Linyuan Qiu
Citations: 4
h-index: 1
Jian Meng
Jian Meng
Citations: 0
h-index: 0
Zhengxuan Lu
Zhengxuan Lu
Citations: 14
h-index: 2
Tao Guo
Tao Guo
Citations: 382
h-index: 6
Jing Li
Jing Li
Citations: 4
h-index: 1
Wei Dai
Wei Dai
Citations: 0
h-index: 0
Zirong Zeng
Zirong Zeng
Citations: 0
h-index: 0
I. Vasilyev
I. Vasilyev
Citations: 0
h-index: 0
Min Liu
Min Liu
Citations: 0
h-index: 0
Wei Sun
Wei Sun
Citations: 0
h-index: 0
Xin Chen
Xin Chen
Citations: 0
h-index: 0
Yi Gao
Yi Gao
Citations: 0
h-index: 0
Jinhua Zhou
Jinhua Zhou
Citations: 13
h-index: 3
Jinmin Xiang
Jinmin Xiang
Citations: 0
h-index: 0
Barkova Maria
Barkova Maria
Citations: 4
h-index: 1
U. Anton
U. Anton
Citations: 0
h-index: 0
Xia Jin
Xia Jin
Citations: 0
h-index: 0
Tian Ding
Tian Ding
Citations: 328
h-index: 7
Qian Chen
Qian Chen
Citations: 75
h-index: 3
Linxin Yang
Linxin Yang
Citations: 46
h-index: 4
Ming Yang
Ming Yang
Citations: 0
h-index: 0
Hong Yang
Hong Yang
Citations: 0
h-index: 0
Fan-jing Zhang
Fan-jing Zhang
Citations: 0
h-index: 0
Tolstykh Vasiliy
Tolstykh Vasiliy
Citations: 0
h-index: 0
Nosov Ivan
Nosov Ivan
Citations: 0
h-index: 0
Abdul Amir
Abdul Amir
Citations: 6
h-index: 2
Zhicheng Zhou
Zhicheng Zhou
Citations: 4
h-index: 1
Xin Zhang
Xin Zhang
Citations: 0
h-index: 0
Z. Ning
Z. Ning
Citations: 13
h-index: 2
Xutong Zhao
Xutong Zhao
Citations: 46
h-index: 4
Junjie Huang
Junjie Huang
Citations: 245
h-index: 8
Jiajun Liu
Jiajun Liu
Citations: 0
h-index: 0
Weiya Kong
Weiya Kong
Citations: 0
h-index: 0
Wenhan Luo
Wenhan Luo
Citations: 255
h-index: 5
Shihao Zeng
Shihao Zeng
Citations: 0
h-index: 0
Haizhou Li
Haizhou Li
Citations: 0
h-index: 0
Zhi-Ting Luo
Zhi-Ting Luo
Citations: 0
h-index: 0
Yuchen Xie
Yuchen Xie
Citations: 0
h-index: 0
Tianxiang Fang
Tianxiang Fang
Citations: 4
h-index: 1
Sihang Chen
Sihang Chen
Citations: 0
h-index: 0
Shihao Hong
Shihao Hong
Citations: 0
h-index: 0
Chang Liu
Chang Liu
Citations: 0
h-index: 0
Zhengjun Yue
Zhengjun Yue
Citations: 0
h-index: 0
Taolue Chen
Taolue Chen
Citations: 175
h-index: 6
Chenwei Wu
Chenwei Wu
Duke University
Citations: 361
h-index: 10
Wen-de Jin
Wen-de Jin
Citations: 0
h-index: 0
Bing Zhang
Bing Zhang
Citations: 0
h-index: 0
S. Qin
S. Qin
Citations: 0
h-index: 0
Cui Hu
Cui Hu
Citations: 0
h-index: 0
Zhengnian Zhang
Zhengnian Zhang
Citations: 0
h-index: 0
Linpeng Hu
Linpeng Hu
Citations: 0
h-index: 0
Yangbo Guo
Yangbo Guo
Citations: 0
h-index: 0
Linying Zeng
Linying Zeng
Citations: 3
h-index: 1

조규(trillion) 규모의 MoE 모델에 대한 완전 파라미터 기반 후속 학습은 대규모 분산 학습에서 심각한 메모리 압박, 통신 오버헤드 중첩 부재, 비효율적인 커널 실행 등 시스템 수준의 상당한 과제를 야기합니다. 대부분의 대규모 LLM 학습 시스템이 GPU 기반 클러스터를 중심으로 구축되는 반면, 본 연구에서는 Ascend NPU SuperPOD 환경에서 엔드투엔드 최적화 방안을 제시합니다. DeepSeek-V4 모델 패밀리를 대상으로 모델 수준 병렬 처리, 연산-통신 조율 및 저수준 커널 실행을 포괄하는 계층적 최적화 프레임워크를 개발했습니다. 결과적으로 개발된 시스템은 34.22%의 모델 FLOPs 활용률(MFU)을 달성했으며, 이는 오픈 소스 기준 레시피 대비 2.93배 향상된 성능이며, 동시에 학습 안정성을 유지합니다. 이 최적화된 인프라를 기반으로 복잡한 운영 연구(OR) 작업에 대한 지속적인 사전 학습(CPT) 및 지도 미세 조정(SFT) 워크플로우를 구축했습니다. 본 통합 프레임워크를 SLAI T-Rex라고 명명합니다. DeepSeek-V4-Flash 모델을 사용하여 수집된 도메인 리소스와 솔버 검증된 합성 최적화 문서를 결합한 OR 특화 CPT 및 SFT 데이터 파이프라인을 개발했습니다. 결과적으로 생성된 데이터셋은 네 가지 작업 범주와 세 가지 문제 표현 방식을 포괄하는 10,000개의 고품질 SFT 샘플로 구성되어 있습니다. 이 전문 모델은 평가 대상 모델 중 가장 높은 평균 제로샷 Pass@1 점수를 달성했으며, 이는 71.81%로 GPT-5.4-Mini 및 기본 DeepSeek-V4-Flash 모델보다 각각 3.98% 및 11.27% 더 높은 수치입니다. 전반적으로 본 연구는 Ascend 인프라를 활용한 조규 규모 모델의 효율적인 후속 학습부터 솔버 기반 수학적 모델링을 위한 도메인 특화 Flash 모델 개발까지, 복잡한 추론을 위한 최첨단 모델 시스템 구축에 기여합니다.

Original Abstract

Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challenges for large-scale distributed training, including severe memory pressure, non-overlapped communication overhead, and inefficient kernel execution. While most large-scale LLM training systems are built around GPU-based clusters, this report presents an end-to-end optimization practice on the Ascend NPU SuperPOD. Using the DeepSeek-V4 model family as the target workload, we develop a hierarchical optimization framework spanning model-level parallelism, computation-communication orchestration, and low-level kernel execution. The resulting system achieves 34.22% Model FLOPs Utilization (MFU) with a 2.93x improvement over the open-source baseline recipe while maintaining training stability. Building on this optimized infrastructure, we further establish a CPT and SFT workflow for complex Operations Research (OR) tasks. We refer to the integrated framework as SLAI T-Rex. Using DeepSeek-V4-Flash, we develop OR-oriented CPT and SFT data pipelines that combine collected domain resources with solver-verified synthetic optimization documents. The resulting dataset contains 10K high-quality SFT samples spanning four task categories and three problem representations. The specialized model achieves the highest average zero-shot Pass@1 score among the evaluated models, reaching 71.81% and outperforming GPT-5.4-Mini and the base DeepSeek-V4-Flash model by 3.98 and 11.27 percentage points, respectively. Overall, this work demonstrates a full-stack pathway from efficient trillion-parameter model post-training on Ascend infra to domain-specialized Flash models for solver-grounded mathematical modeling, advancing frontier-model systems for complex reasoning.

0 Citations
0 Influential
5.5 Altmetric
27.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!