From TikTok to Frontier AI: Is ByteDance Sprinting towards a 10 Tn Model?
The TikTok owner’s biggest bet could be more than three times the size of Moonshot’s Kimi K3 and approach Anthropic’s Mythos in scale.
Opinions expressed by Entrepreneur contributors are their own.
You're reading Entrepreneur Asia Pacific, an international franchise of Entrepreneur Media.
ByteDance is training an artificial intelligence model that could contain as many as 10 Tn parameters, a leap in scale that would put the TikTok owner’s model broadly in the same size range as some of the most advanced US systems and sharply raise the stakes in China’s frontier AI race.
The Financial Times first reported the project on Friday (August 7), citing people with knowledge of the matter. Reuters, which subsequently reported the FT disclosure, said it could not independently verify the information.
ByteDance did not immediately respond to Reuters’ request for comment. The model is still in pre-training, and its final size has not been determined.
Parameters are the numerical settings a model learns during training, giving it the capacity to recognise patterns, generate responses and perform tasks. But size is only a rough indication of a model’s scale. Architecture, training data, post-training techniques and how efficiently those parameters are used can matter as much as the headline number when it comes to actual performance.
That qualification is particularly important when ByteDance is compared with Anthropic. The US developer does not disclose the parameter counts of its latest models. Industry estimates cited by the FT put its most advanced Mythos 5 at about 8 Tn parameters. If ByteDance reaches 10 Tn, the Chinese model would, therefore, move into roughly the same range of scale as Mythos.
It does not mean ByteDance will match Anthropic in capability. Anthropic describes Mythos 5 as its most capable model for cybersecurity and biology research and restricts access to vetted partners because of the sensitivity of those capabilities. ByteDance’s model has yet to complete training, let alone undergo independent benchmark testing, making a performance comparison premature.
The project, nevertheless, shows how quickly the scale of Chinese AI models is being reset. Moonshot AI’s Kimi K3 arrived only weeks ago with 2.8 Tn parameters, becoming the largest Chinese model reported at the time and pushing closer to leading US systems on several benchmarks. Before K3, Meituan’s LongCat-2.0 and DeepSeek’s V4-Pro were among China’s largest models at about 1.6 Tn parameters.
The AI race in China is accelerating, with companies pursuing different routes to stronger systems — from sheer scale and new architectures to multimodal reasoning and more efficient use of computing power. A potential jump from 2.8 Tn to as much as 10 Tn within weeks underlines how quickly one technical marker can be overtaken.
For ByteDance, the project is part of a transformation that has taken it well beyond the consumer internet business, best known globally for TikTok. The company established its Doubao Seed Team in 2023, an internal research organisation focussed on foundation models and areas such as large language models (LLMs), vision, speech, AI infrastructure, robotics and AI safety.
The Seed team now has about 2,000 people in China and overseas, according to the FT, and former Google DeepMind scientist Wu Yonghui leads it.
ByteDance has also built a sizeable consumer base for its AI products. Doubao, its AI assistant, has about 324 Mn monthly active users (MAU). Beyond chatbots, the company has moved into video generation with Seedance, multimodal and agentic models with Seed2.0 and enterprise AI services through its Volcano Engine cloud platform.
That distribution gives ByteDance an advantage over many pure-play AI laboratories. But its ambition is increasingly focussed on the technology underneath those products.
Founder Zhang Yiming recently told ByteDance’s Seed team not to rely on distillation — using outputs from a more powerful AI system to train or improve a smaller model — even if that meant sacrificing short-term performance, Reuters reported, citing China’s state-backed The Paper.
Zhang argued that relying on rival-model outputs could hamper longer-term technological breakthroughs, urging the team instead to pursue independent development.
But that approach requires resources.
ByteDance has expanded its data centre network, hired heavily for AI research, strengthened Volcano Engine and is exploring the development of custom AI chips. Training a model approaching 10 Tn parameters would add another test of that infrastructure, with larger systems placing heavier demands on computing power, memory, data and engineering.
The wider Chinese AI race is no longer moving in a single direction. Moonshot has pushed the scale of open-weight models. Alibaba and DeepSeek have also tested how far capability can be advanced while keeping inference efficient and costs down.
ByteDance is now testing another part of the frontier: how far one of China’s largest technology companies can push model scale through its own research and computing infrastructure.
Whether that produces a model capable of challenging Mythos or other leading US systems will depend on what emerges after training, fine-tuning and testing. But if ByteDance reaches the scale now under consideration, the gap in model size would have narrowed considerably. The harder question is whether it can turn that scale into comparable intelligence and performance.
ByteDance is training an artificial intelligence model that could contain as many as 10 Tn parameters, a leap in scale that would put the TikTok owner’s model broadly in the same size range as some of the most advanced US systems and sharply raise the stakes in China’s frontier AI race.
The Financial Times first reported the project on Friday (August 7), citing people with knowledge of the matter. Reuters, which subsequently reported the FT disclosure, said it could not independently verify the information.
ByteDance did not immediately respond to Reuters’ request for comment. The model is still in pre-training, and its final size has not been determined.