AI Pulse by Inblix

Bytedance is secretly training a 10-trillion-parameter model to rival Anthropic

The Decoder · Aug 7, 2026 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: Bytedance is secretly training a 10-trillion-parameter model to rival Anthropic

Bytedance is quietly assembling a behemoth. The TikTok parent company is in the process of pretraining a new AI model that could scale up to a staggering 10 trillion parameters, according to the Financial Times. To put that in perspective, it’s roughly triple the size of Moonshot’s Kimi K3, currently the heavyweight champion of Chinese models, and it nudges Bytedance into a direct size comparison with Anthropic’s top-tier system, Mythos 5, which industry watchers peg at around eight trillion parameters. Anthropic has never officially confirmed that figure, but the ballpark alone signals serious intent from Beijing.

Three people familiar with the project told the FT that the model is currently in pretraining, a computationally brutal phase that usually eats up three to six months. Founder Zhang Yiming has reportedly directed the 2,000-strong Seed team to think in terms of decades, not quarters, aiming for world-leading capabilities. One particularly interesting operational detail emerged: a source claims Bytedance has sworn off distillation—the common practice of training on outputs from other companies’ models—for over a year. That’s a meaningful technical choice. It suggests they’re betting on genuinely native intelligence rather than a shortcut that often bakes another model’s flaws into your own.

This isn’t happening in a vacuum. Elon Musk recently disclosed that xAI is also spinning up Grok variants with six and ten trillion parameters on its Colossus 2 cluster. The raw parameter count is quickly becoming the new space race. But anyone who’s watched the AI industry mature knows that size is a vanity metric without the right data diet and training regimen. A bloated model trained on messy data is just an expensive way to be mediocre. Bytedance’s access to TikTok’s firehose of global, multimodal data could be the real differentiator here, not just the parameter count.

What’s unsaid in the FT report is just as interesting as what’s on the record. There’s no mention of a release timeline, no benchmarks, and no hint of what this model would actually power. Given the regulatory headwinds TikTok faces in the West, a world-class foundational model could be as much a geopolitical hedge as a product. If the training run succeeds, Bytedance won’t just have a bigger model—it will have a piece of critical infrastructure that operates outside the reach of Western export controls.

💡 Key Takeaways

  1. Bytedance's 10-trillion-parameter model would be the largest in China, dwarfing Moonshot's Kimi K3 by a factor of three.
  2. The company has deliberately avoided distilling from rival models for over a year, signaling a bet on fundamentally original AI capabilities.
  3. Founder Zhang Yiming has set a long-term vision for the 2,000-person Seed team, explicitly targeting world-leading model performance.
  4. The raw scale puts Bytedance in a direct race with both Anthropic's Mythos 5 and xAI's upcoming Grok variants on its Colossus 2 cluster.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles