ByteDance trains massive AI model in bid to rival Anthropic
Ars Technica Zijing Wu, Financial Times
ByteDance is training a new AI model with up to 10 trillion parameters, aiming to match Anthropic's top-tier system. It would dwarf China's current biggest model by 3x, showing just how fast the US-China AI gap is closing.
Based on reporting by Ars Technica, Zijing Wu, Financial Times — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
ByteDance isn't messing around. According to three people familiar with the effort, the company has quietly begun pre-training a model that could top out at 10 trillion parameters — a scale that would put it in the same conversation as Anthropic's most advanced system, known internally as Mythos. That's not a modest step up. It's roughly three times the size of Kimi K3, the model Moonshot released earlier this year that currently holds the title of China's largest.
Pre-training at this scale is a slow, expensive slog. Sources say the process typically runs three to six months before a model is even ready for fine-tuning, let alone a public release. And ByteDance is still early in that window, meaning the eventual parameter count isn't locked in — it could shrink, grow, or shift entirely depending on how training goes and what tradeoffs the company decides to make between size, cost, and performance.
What makes this notable isn't just the number attached to it. Parameter counts alone don't determine how good a model actually is — plenty of smaller, more efficiently trained systems have outperformed larger, clumsier ones. But the fact that a Chinese firm is even attempting something in this size class signals how quickly the perceived gap with US frontier labs has narrowed. A year ago, models of this scale were treated as almost exclusively an American ambition, bankrolled by companies like OpenAI, Google, and Anthropic with access to enormous compute clusters.
ByteDance has been investing heavily in AI infrastructure and talent, and this project looks like the clearest evidence yet that it intends to compete at the very top of the field rather than settle for a strong regional player status. Whether the finished model actually rivals Mythos in capability — not just in raw size — will be the real test once training wraps and benchmarks start circulating.
My take — AI-written commentary, not fact-checked reporting
Parameter count has always been a vanity metric dressed up as a technical one, and ByteDance chasing 10 trillion smells more like a headline strategy than a research breakthrough. The real story is compute access — if Chinese firms are quietly matching US labs on training scale despite export controls, that says a lot more about the effectiveness of chip restrictions than any benchmark will.
Read more about this at: Ars Technica