Bytedance is training an AI model with up to ten trillion parameters, according to the Financial Times. This would make it one of the largest AI models in the world, surpassing the current largest Chinese model, Moonshot's Kimi K3, which has around three trillion parameters. The model is in the pretraining phase, a stage that typically takes three to six months. Parameters determine how much a model can store, but performance also depends on data quality and training methods. One of the sources said Bytedance has avoided distillation, meaning training on outputs from other companies' models, for over a year. Founder Zhang Yiming told the 2,000-person Seed team internally to aim for world-leading model capabilities over the long term. xAI is also training Grok variants with six and ten trillion parameters on its Colossus 2 cluster, according to Elon Musk.

The model's size is expected to give it a significant advantage in processing complex tasks and handling large volumes of data. However, the success of the model will also depend on the quality of the training data and the methods used during the training process. Bytedance's approach of avoiding distillation suggests a focus on developing the model from scratch, which could lead to more tailored and efficient performance. The company's long-term goal, as stated by Zhang Yiming, is to achieve world-leading capabilities in AI modeling.

According to the Financial Times, three insiders confirmed that Bytedance is working on an AI model with up to ten trillion parameters. The model is currently in the pretraining phase, which is a critical stage for developing large-scale AI systems. The Financial Times also noted that Anthropic's top system, Mythos 5, is estimated to have around eight trillion parameters, placing Bytedance's model in the same ballpark. The source added that Bytedance has avoided distillation for over a year, indicating a commitment to independent model development.

Source: thedecoder