Moonshot AI publishes the complete weights for Kimi K3 at 00:00 UTC on July 27 — seven o’clock this evening, Central time. It will be the largest open-weight model anyone has ever released: 2.8 trillion parameters in a sparse mixture-of-experts design, with a 1,048,576-token context window. DeepSeek’s V4 Pro runs 1.6 trillion. Zhipu’s GLM 5 series runs 744 billion. K3 is roughly 2.8 times the size of its own predecessor.
Then there is the file. In MXFP4 — a four-bit floating-point format with per-block scaling, supported natively on Nvidia’s Blackwell parts and AMD’s MI400 — the weights come to about 1.4 terabytes. At sixteen-bit precision the same weights run closer to 5.6 terabytes. And 1.4 TB is the floor, not the requirement: that is weights only, before microscaling metadata, KV cache, activations or a single token of context state.
Moonshot’s own guidance for serving the thing is a supernode of 64 or more accelerators. A top-end consumer graphics card holds 24 gigabytes — roughly 58 times too small to hold the weights alone, never mind run them. No consumer or prosumer hardware on the market fits this model, and none is close.
The quality is not the question
K3 matches or edges Anthropic’s Claude Fable 5 on several coding and agentic benchmarks, Terminal-Bench 2.1 and SWE-Marathon among them, while its API undercuts Fable’s token price by about two-thirds: $3 per million input tokens, $15 per million output, and $0.30 per million on a cache hit. There is no cheap non-thinking mode — K3 always reasons at maximum effort, and every reasoning token bills at the output rate. Moonshot still places it behind Fable 5 and GPT‑5.6 Sol on overall performance, ahead of everything else it tested.
So the model is real, the price is aggressive, and the license genuinely lets you download it, inspect it, fine-tune it, audit it. What it does not do is let you serve it. For all but a few dozen organizations on earth, running K3 in production means renting a cluster or calling somebody’s API — which is the same operational shape as a closed model, with a different invoice.
Our take: Open weights just crossed the line where the license stops being the binding constraint and the balance sheet starts. Every argument in Friday’s open-weights letter — inspection, sovereignty, no vendor lock-in — assumes someone can actually load the file. At 2.8 trillion parameters that someone is a hyperscaler, a national lab, or a government. Openness at this scale doesn’t democratize access to intelligence; it democratizes access to compute buyers, which is exactly why the companies that sell compute were the loudest names on that letter. For operators the practical read is unchanged and worth repeating: you are not going to self-host the frontier, so stop pricing your stack as if you might. What open weights buy you is leverage — a credible, near-frontier substitute that caps what any API can charge you. Build switchable, benchmark quarterly, and let the labs compete for the seat.
What to watch
- Whether the weights actually ship on time. Until the file is public, every VRAM figure, serving command and license reading in circulation is an estimate. Nobody has run this outside Moonshot.
- Who hosts it first. The inference providers that stand K3 up within days — and what they charge against Moonshot’s own $3/$15 — will set the real floor on frontier-class pricing.
- Hallucination and safety disclosure. Moonshot has published capability benchmarks freely and reliability numbers sparingly. Independent evals in week one matter more than the launch chart.
- Enterprise sign-off. Cheap and open is only useful if legal clears it. Watch whether U.S. and European buyers approve Chinese open weights for production — and whether Washington’s pending rules on downloadable models make the question moot.
