AI Pulse by Inblix

ByteDance targets 10 trillion parameters — triple the size of China's largest public model

Ars Technica AI · Aug 7, 2026 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: ByteDance targets 10 trillion parameters — triple the size of China's largest public model

ByteDance is in the early stages of pre-training an AI model that could reach 10 trillion parameters, according to three people with direct knowledge of the project. That’s three times larger than Moonshot’s Kimi K3, the biggest Chinese model released publicly to date, and would put it in the same weight class as Anthropic’s most advanced system — widely estimated at around 8 trillion parameters for Mythos 5.

The model is currently in pre-training, a phase that typically runs three to six months before fine-tuning begins. The final parameter count isn’t locked yet; one source says it’ll only be determined at a later stage. But the sheer ambition of the project tells you something about how much the competitive landscape has shifted. A year ago, the conversation was about whether Chinese labs could close the gap with GPT-4-class systems. Now the question is whether they’ll surpass the frontier outright.

Parameter count alone isn’t destiny, of course. It sets the ceiling for what a model can store and recall, but actual capability hinges on data quality, training methodology, and post-training refinement. Anthropic doesn’t disclose its model sizes, but industry estimates peg Fable 5 at roughly 5 trillion parameters and Mythos 5 around 8 trillion. ByteDance’s target would leapfrog both — on paper, at least.

The timing is notable. In recent weeks, models from Moonshot and Alibaba have posted benchmark scores trailing only Fable 5 in certain categories. Mythos 5 itself remains restricted to approved organizations after a temporary ban in June over security concerns, which creates an odd dynamic: the most capable Western model is effectively gated, while multiple Chinese labs are reportedly training systems in the Fable 5 range. ByteDance is simply the most aggressive of the bunch. Whether a 10-trillion-parameter model translates into something genuinely more useful than what’s already out there — or just more expensive to run — is the question nobody can answer yet.

💡 Key Takeaways

  1. ByteDance's model would be three times larger than China's current largest public model, Moonshot's Kimi K3, signaling a dramatic escalation in scale ambitions.
  2. Multiple Chinese labs are now training models comparable to Anthropic's Fable 5, but ByteDance is the only one publicly pushing toward the 10 trillion parameter range.
  3. Parameter count sets memory capacity but doesn't guarantee better performance — data quality and training methods matter just as much, and ByteDance hasn't disclosed details on either.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles