ByteDance is training a massive new artificial intelligence model with as many as 10 trillion parameters as the Chinese technology company steps up efforts to compete with leading US AI developers.
The model could approach or even exceed the size of Anthropic’s cutting-edge Mythos system, the Financial Times reported on Friday, citing people familiar with the project.
ByteDance’s new model is still at an early stage of training and could contain as many as 10 trillion parameters, according to the report.
Anthropic does not publicly disclose the size of its latest models, but industry estimates cited in the report put Mythos at around 8 trillion parameters. By comparison, Moonshot AI’s Kimi K3 has 2.8 trillion parameters, while Meituan’s LongCat-2.0 and DeepSeek’s V4-Pro are estimated at around 1.6 trillion parameters each.
However, a higher parameter count does not automatically mean a model will perform better. The figure mainly indicates the scale of the model, while performance also depends on areas such as training methods, data and architecture.
The ByteDance model is currently in the pre-training stage. This part of model development typically takes between three and six months before developers move to fine-tuning and other stages ahead of a possible release.
The project forms part of ByteDance’s wider push into advanced AI. Its Seed team focuses on areas including model pre-training, post-training, inference, memory, learning and interpretability. ByteDance also maintains dedicated infrastructure teams working on distributed training and high-performance inference for foundation models.
According to the Financial Times, ByteDance has spent heavily on AI over the past three years, expanding its data centres and hiring researchers. Its Seed model-development team has around 2,000 members in China and overseas and is led by former Google DeepMind scientist Wu Yonghui.
The company has also expanded its Volcano Engine cloud business, which provides AI services to enterprise customers, and has ambitions to develop its own AI chips.
ByteDance has reportedly followed a more independent approach to developing its models for more than a year instead of relying on model distillation from competing AI labs.
Model distillation generally involves training a smaller model to reproduce knowledge or outputs from a larger model.
ByteDance founder Zhang Yiming believes independent development is necessary if the company wants to eventually outperform its competitors, according to the report. During an internal meeting two weeks ago, Zhang reportedly told the Seed team to focus on achieving world-leading AI capabilities over the long term rather than worrying about temporarily falling behind competitors.
ByteDance did not respond to a request for comment on the report, while Reuters said it had not independently confirmed the information.
Get the latest tech news, telecom insights, and product launches wherever you prefer.
Add ProPakistani to Preferred Sources and see more of our stories in Google Search and Top Stories.
Technology and Automotive Specialist covering the latest cars, smartphones, AI breakthroughs, and...
Shares