China's AI development is facing a new bottleneck: a lack of training data in the Chinese language.

Improving the performance of advanced AI models requires a large amount of high-quality training data. While the US export restrictions on advanced AI chips have been highlighted as a major obstacle to China's AI development, Chinese AI experts warn that a lack of Chinese-language training data could become the next major bottleneck.
China faces new AI bottleneck as it runs out of Chinese-language training data | South China Morning Post
https://www.scmp.com/tech/tech-trends/article/3363318/china-faces-new-ai-bottleneck-it-runs-out-chinese-language-training-data

The lack of training data for AI is not just a problem for China; it has long been pointed out that the progress of AI could be slowed by a lack of high-quality data . Epoch AI, an AI research institute, estimates that the actual amount of human-created and published text, considering quality and the effectiveness of repeated learning from the same data, is approximately 300 trillion tokens. They note that if current trends continue, language models could exhaust this text data between 2026 and 2032.
One reason the problem appears particularly serious in China is the scarcity of websites using Chinese on the internet. According to web technology research service W3Techs , as of August 10, 2026, only 1.3% of websites whose language could be determined used Chinese. For comparison, English accounted for 49.5% and Japanese for 5.0%.
In an article contributed to the Guangming Daily , Professor Sun Maosong and Kong Cunliang of Tsinghua University point out that China's development of ' corpora '—large amounts of textual data used to train and evaluate AI models—is still in its early stages to cope with the era of large-scale language models. They note that Chinese language data in the large-scale web archive ' Common Crawl ' accounts for only about 4.8%, and that knowledge-density materials such as Chinese encyclopedias, scientific books, academic journals, and dictionaries are very scarce in existing Chinese language corpora.
As a way to increase the amount of Chinese language data, Sun and his colleagues propose digitizing publications, classical texts, local histories, and historical documents, and incorporating dictionaries, videos, dialects, and minority languages into the corpus. They also point out the need for a system to obtain permission from copyright holders, and in fact, China's Huaxia Publishing House has prohibited the use of books for AI training.

The Chinese government is also working to develop AI-related data in general, not just in Chinese. On June 8, 2026, the National Data Administration of China announced a plan to build high-quality datasets that have been validated through actual use in fields such as scientific research, industrial manufacturing, and medical hygiene by the end of 2028. The plan is to supplement data that is difficult to collect with simulation and synthesis technologies.
In an article published on the website of the National Data Administration, Yu Xiaohui, president of the China Academy of Information and Communications Technology, stated that 'competition in the age of AI is not only about models and computing power, but also about competition for high-quality data supply systems.' He argued that the country that first establishes an integrated system encompassing data generation, management, distribution, utilization, and value creation will take the lead in the next AI development race.
Related Posts:
in AI, Posted by log1b_ok







