The Financial Times reported Friday that ByteDance is pre-training an AI model with as many as 10 trillion parameters, citing three people familiar with the project. There is no name, no benchmark, no release date and no official confirmation — ByteDance has said nothing publicly. What there is, is a number that reframes how much compute one Chinese company believes it can get its hands on.
The work sits inside Seed, the research division ByteDance stood up in early 2023 in the months after ChatGPT landed. The reported target is Anthropic’s Mythos — one of the frontier systems Chinese developers have so far struggled to match. People familiar with the project told the FT that neither the final scale nor the architecture has been settled. Ten trillion is the ceiling under consideration, not a spec sheet.
Pre-training at that size typically runs three to six months before fine-tuning even begins. Anything that ships from this run lands in 2027.
The comparison number is 2.8 trillion
Moonshot’s Kimi K3, at roughly 2.8 trillion parameters, is the largest model China has actually released. ByteDance’s reported ceiling is more than three times that. Moonshot is said to have used about 20,000 Nvidia accelerators to train K3 — and that is the part of this story that actually binds.
ByteDance has raised 2026 capital spending to roughly 200 billion yuan, about $29 billion and at least 25% above its earlier plan, with around half earmarked for semiconductors. A meaningful share goes to Nvidia hardware under Beijing’s capped H200 arrangement, which permits training use only. A growing share goes to domestic Chinese silicon — a hedge against the next round of export controls rather than a preference.
It also has somewhere to put the result. Doubao is China’s most-used consumer AI app, north of 330 million monthly users, with TikTok sitting behind it. Distribution is the one input ByteDance has never had to buy.
Our take: Parameter counts stopped measuring compute the moment sparse mixture-of-experts became the default. Alibaba’s Qwen3.8-Max carries 2.4 trillion total parameters and fires roughly 95 billion of them per token; a 10-trillion MoE model could be cheaper to serve than a far smaller dense one. So the headline isn’t “ByteDance builds the biggest model.” It’s that ByteDance is willing to commit a frontier-scale training budget in the same month DeepSeek beat its own flagship by getting smaller. Those two bets can’t both be right about where the next capability jump comes from.
What to watch
- Architecture disclosure. Dense or MoE, and what the active parameter count is. Without that, 10 trillion is a press number.
- The chip mix. If ByteDance can train a run this size predominantly on domestic accelerators, the export-control thesis needs rewriting. If it can’t, the cap is doing its job.
- Open or closed. Kimi K3 and Qwen3.8-Max both went open-weight. ByteDance has kept its strongest models closed and monetised them through Doubao.
- Confirmation. One FT story with three anonymous sources is a starting point, not a fact. ByteDance putting its name to the project is the actual event.
The thing worth holding onto: the constraint on Chinese frontier labs was never ambition or talent. It is how many accelerators they can legally buy, and how quickly domestic replacements close the gap. A 10-trillion-parameter pre-training run is a wager that the answer has become “enough.”
