
Kimi K3: 2.8T Parameters, But Only 104B Active Per Token
Kimi K3 has 2. 8 trillion parameters-but only about 104 billion are activated for each token. This visual deep dive explains how 16 of 896 routed experts, 2 shared experts, Stable LatentMoE, and Kimi Delta Attention make that possible. Kimi K3 is a Mixture-of-Experts model with a 1M-token context window. We follow one token through expert routing, distributed communication, the 7168 → 3584 latent projection, Quantile Balancing, KDA, Gated MLA, and Block Attention Residuals.









