AI Mechanics
Kimi K3: 2.8T Parameters, But Only 104B Active Per Token
Kimi K3 has 2. 8 trillion parameters-but only about 104 billion are activated for each token. This visual deep dive explains how 16 of 896 routed experts, 2 shared experts, Stable LatentMoE, and Kimi Delta Attention make that possible. Kimi K3 is a Mixture-of-Experts model with a 1M-token context window. We follow one token through expert routing, distributed communication, the 7168 → 3584 latent projection, Quantile Balancing, KDA, Gated MLA, and Block Attention Residuals.
You’ll learn: • Why 2. 8T total parameters does not mean 2. 8T parameters run per token • The difference between selecting 16/896 experts and activating ~104B parameters • How Stable LatentMoE reduces routed width from 7168 to 3584 • Why Quantile Balancing matters for expert utilization • Why KV-cache growth and attention compute are different problems • How Kimi Delta Attention replaces token history with recurrent state • Why K3 combines 69 KDA layers with 24 Gated MLA layers • How Block Attention Residuals retrieve information across model depth Watch next - LLM Inference & Performance: Sources: Kimi K3 technical report: https://arxiv.
org/html/2607. 24653v1 Moonshot AI repository: https://github. com/MoonshotAI/Kimi-K3 Model card: https://huggingface. co/moonshotai/Kimi-K3 Kimi technical blog: https://www. kimi. com/blog/kimi-k3 vLLM Day-0 support: https://vllm. ai/blog/2026-07-27-k3
Watch the video for the full walkthrough. Use this page when you want the argument in writing without scrubbing the timeline.