본 논문에서는 2.8조 개의 파라미터를 가진 Mixture-of-Experts 모델인 Kimi K3를 소개합니다. Kimi K3는 1040억 개의 활성화된 파라미터, 내장된 이미지 처리 기능, 그리고 1백만 토큰의 컨텍스트 창을 갖습니다. Kimi K3는 시퀀스 길이와 모델 깊이 전반에 걸쳐 정보 흐름을 개선하는 Kimi Delta Attention 및 Attention Residuals 기술을 기반으로 합니다. 또한, 각 토큰당 896개의 라우팅된 전문가 중 약 16개를 효과적으로 활성화하는 Stable LatentMoE 기술과 개선된 학습 방법 및 데이터 처리 방식을 통해 Kimi K2보다 전체적인 확장 효율성이 약 2.5배 향상되었습니다. 학습 후 평가 결과, Kimi K3는 일반, 에이전트 기반, 코딩 영역 전반에 걸쳐 다양한 수준의 추론 능력을 갖추고 있으며, 이를 통해 복합적인 일반화 및 안정적인 장기 실행이 가능합니다. 2.8조 규모의 Kimi K3는 알고리즘과 시스템의 공동 설계(KDA), 효율적인 메모리 관리를 통한 균형 잡힌 전문가 병렬 학습, 지속적인 실행 및 샌드박스 상태를 갖춘 백만 토큰 에이전트 기반 강화 학습, 그리고 배포 혁신을 포함한 다양한 인프라 개선 덕분에 구현되었습니다. 광범위한 평가 결과, Kimi K3는 장기 코딩, 에이전트 기반 작업, 지식 처리, 추론 및 이미지 처리 작업에서 최첨단 수준의 성능을 달성합니다. 전반적인 성능은 가장 강력한 독점 모델인 Claude Fable 5 및 GPT-5.6 Sol에 미치지 못하지만, Kimi K3는 저희가 평가한 다른 개방형 및 독점 모델보다 꾸준히 우수한 성능을 보입니다. 향후 연구를 촉진하고 최첨단 지능의 광범위한 활용과 확산을 가속화하기 위해 Kimi K3 모델 전체의 파라미터를 공개합니다.
Original
Abstract
We introduce Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. Kimi K3 is built on Kimi Delta Attention and Attention Residuals, which improve information flow across sequence length and model depth. Together with Stable LatentMoE, which effectively activates 16 of 896 routed experts per token, and refined training and data recipes, these advances yield an approximately 2.5x improvement in overall scaling efficiency over Kimi K2. Post-training highlights reinforcement learning across general, agentic, and coding domains and multiple reasoning-effort levels, enabling compositional generalization and robust long-horizon execution. At 2.8T scale, Kimi K3 is supported by infrastructure advances in multiple areas: algorithm-system co-design for KDA, perfectly balanced expert-parallel training with efficient memory management, million-token agentic RL with persistent rollout and sandbox states, and deployment innovations. Extensive evaluations show that Kimi K3 achieves frontier-level performance across long-horizon coding, agentic, knowledge, reasoning, and vision tasks. While its overall performance still trails the most powerful proprietary models, namely Claude Fable 5 and GPT-5.6 Sol, Kimi K3 consistently outperforms other open and proprietary models evaluated in our suite. We release the full Kimi K3 model weights to facilitate future research and accelerate the broader deployment and adoption of frontier intelligence.