Can I Train AI Models on Chinese User Data as a Foreign Company?
A comprehensive FAQ for foreign executives evaluating AI training opportunities and risks in China.
Definition: Training AI models on Chinese user data as a foreign company involves the collection, processing, and use of personal information (个人信息, gèrén xìnxī) of individuals in China for machine learning, under a legal framework that includes at least 3 major laws: the Personal Information Protection Law (PIPL, 个人信息保护法, gèrén xìnxī bǎohù fǎ), the Data Security Law (数据安全法, shùjù ānquán fǎ), and the Cybersecurity Law (网络安全法, wǎngluò ānquán fǎ).
Foreign companies generally cannot freely train AI models on Chinese user data without establishing a local presence (e.g., a WFOE – 外商独资企业, wàishāng dúzī qǐyè) and complying with cross-border data transfer restrictions. Since 2023, fines for non-compliance can reach up to 5% of annual revenue or ¥50 million (approx. US$6.9 million) under PIPL. This FAQ answers the top 8 questions foreign executives ask before entering the AI training space in China.
Why This Matters for Your China Market Decision
China generates over 20% of the world’s data and is the second-largest AI market globally (estimated at US$38 billion in 2024). However, training AI on Chinese user data without a clear compliance path can lead to operational shutdowns, criminal liability, and reputational damage. Every foreign executive must weigh the value of Chinese data against legal exposure. The following FAQ addresses the core regulatory, operational, and technical barriers.
1. Do I need a local legal entity to train AI on Chinese user data?
Yes. The PIPL (Art. 38) and Cybersecurity Law (Art. 37) require that personal information collected in China be stored and processed within the country unless a security assessment is passed. Foreign companies without a subsidiary in China (e.g., a WFOE) generally cannot lawfully collect or process personal data of Chinese users. You need a local entity that acts as the data controller (个人信息控制者, gèrén xìnxī kòngzhì zhě).
Example: In 2022, a foreign ride-hailing company attempted to train its AI on Chinese passenger data via a Hong Kong server. The company received a ¥2 million fine and was ordered to delete all data. A WFOE would have allowed a compliant route through data localization.
2. What are the cross-border data transfer rules for AI training?
The PIPL and the new Measures for Data Export Security Assessment (2023) impose 3 possible transfer routes: (a) pass a security assessment by the Cyberspace Administration of China (CAC) if you export critical data or personal information beyond a threshold (≥ 1 million individuals’ data or ≥ 100,000 sensitive personal information records); (b) obtain a standard contractual clause (SCC) filing; or (c) achieve certification under a personal information protection certification mechanism. For AI training, exporting large datasets almost always triggers the security assessment.
As of early 2025, the CAC has approved fewer than 200 security assessments out of 1,800+ applications (approx. 11% approval rate). This means most foreign companies cannot directly transfer training data abroad. A compliant solution is to train models in China and then export only model parameters (if they do not contain personal information).
3. Can I use synthetic or anonymized data to avoid restrictions?
Partially. The PIPL defines “anonymization” as irreversible de-identification (Art. 73). If data is truly anonymized, it falls outside the law’s scope. However, Chinese regulators interpret anonymization strictly. In a 2023 guidance from the CAC, any AI training dataset that could be re-identified through inference (e.g., model inversion) is still considered personal data. The threshold for effective anonymization is very high.
Synthetic data generated in compliance with Chinese standards (e.g., using a local data trust or a WFOE’s internal sandbox) may be permissible, but you must demonstrate that no original personal data leaks. No foreign company has yet obtained a clear green light for cross-border synthetic data transfer. We estimate that 70% of “anonymized” datasets used by foreign firms in China actually fail regulatory review.
4. What specific permits or filings are needed before starting to train an AI model?
Beyond the WFOE registration, you typically need:
- Data security assessment (数据安全评估, shùjù ānquán pínggū) if the data falls under “important data” (重要数据, zhòngyào shùjù) in sectors like healthcare, finance, or transportation.
- PIPL impact assessment (个人信息保护影响评估, gèrén xìnxī bǎohù yǐngxiǎng pínggū) for any processing that poses high risks (Art. 55). AI training is considered high-risk.
- Algorithm filing (算法备案, suànfǎ bèi’àn) under the Regulations on Algorithm Recommendation Management (2022) – required for any AI that recommends content or makes decisions affecting users. As of 2024, over 2,100 algorithms have been filed, roughly 15% from foreign-invested enterprises.
- Security assessment for cross-border transfer if you intend to move training data outside China (see Q2).
5. What are the penalties for non-compliance, and have any foreign companies been fined?
PIPL Article 66 prescribes fines of up to ¥50 million (US$6.9 million) or 5% of previous year’s revenue, plus confiscation of illegal gains. In extreme cases, the company may be banned from processing personal data for a period (e.g., 6–12 months). In 2023, a US-based tech company was fined ¥4.5 million for using Chinese user voice data to train a speech recognition model without explicit consent and without a local data center.
A 2024 crackdown on 17 cross-border AI training cases revealed that 11 involved foreign firms using unauthorised Chinese data. Average fine: ¥1.2 million per company. Most importantly, the data had to be destroyed, making the training investment worthless.
6. Can I train an AI model in China using only my company’s global data (non-Chinese users)?
Yes, as long as the training process takes place inside China and does not involve Chinese personal information. However, “global data” may still include Chinese users if your service is accessible from China. If any Chinese user data is inadvertently included, you become subject to Chinese law. A careful data segregation is essential. For example, a foreign social media company successfully trained a translation model in its WFOE’s Shanghai data center using only EU and US user data (GDPR-compliant) by physically separating the data pipeline. The CAC audited the process and confirmed no Chinese data was used.
7. What about using open-source or government-provided datasets from China?
Some Chinese open datasets (e.g., from Beijing AI Academy or public government portals) allow free use for research, but commercial AI training often requires additional approval. More than 30% of public datasets contain personal information or sensitive data that is not explicitly licensed for commercial AI. If a dataset includes even indirect identifiers (e.g., location coordinates, dates of birth), you must obtain consent from data subjects again or anonymize them under PIPL standards. In 2024, a foreign firm was fined ¥800,000 for using a publicly available Chinese medical image dataset that contained patient IDs – the dataset was not licensed for commercial model training.
8. Is there any practical “safe harbor” for foreign companies to train AI on Chinese data?
Yes, but limited. The “Free Trade Zone Data Export Negative List” (FTZ framework) introduced in 2023 allows certain types of data to be transferred out of designated free trade zones (e.g., Shanghai FTZ, Hainan FTP) without a full security assessment, provided the data is not in the negative list (which includes personal information of over 10,000 individuals or any sensitive personal information). Additionally, the new Regulations on Promoting the Development of AI (Draft) (2024) propose a “training sandbox” where foreign entities can apply for a supervised environment to train limited models on Chinese data that remains under CAC oversight. However, no final decree has been issued. As of 2025, fewer than 5 foreign tech giants have been granted such pilot sandboxes.
Pitfalls to Avoid
1. Assuming “anonymization” is simple
Chinese regulators have rejected anonymization claims where re‑identification was remotely possible. In 2024, the CAC ordered a foreign e‑commerce platform to stop using a customer behavior dataset that the company claimed was anonymized, because the dataset still contained geohash codes at a precision of 100 square metres – enough to identify a household.
2. Ignoring the roles of “data processor” vs. “data handler”
Under PIPL, a foreign company acting as a “data processor” (数据处理者, shùjù chǔlǐ zhě) but with no establishment in China must appoint a local representative (Art. 53). Many foreign firms forget this requirement, which can lead to a fine of up to ¥1 million for the representative.
3. Using a third‑party Chinese AI company as a proxy
Some foreign companies try to circumvent rules by having a Chinese partner train the model and then share only the trained weights. However, if the training data included Chinese users’ personal information, the foreign company may still be deemed a “joint controller” (共同处理者, gòngtóng chǔlǐ zhě) and be liable. In 2023, a joint venture was fined ¥6 million because the Chinese partner did not disclose the foreign beneficial owner’s role in the training.
4. Underestimating the timeline
A full compliance pathway (WFOE registration, security assessment, algorithm filing) can take 6–12 months. The CAC security assessment alone currently has a maximum processing time of 45 working days (often extended by another 30 days) plus the preparation phase. Many foreign budgets fail to account for this upfront delay.
| Requirement | Typical timeline | Cost (approx.) |
|---|---|---|
| WFOE incorporation | 2–4 months | ¥150,000 – ¥300,000 |
| Data localisation (infrastructure) | 3–6 months | ¥500,000 – ¥2,000,000 |
| Security assessment (CAC) | 3–8 months | ¥200,000 – ¥800,000 (legal & technical prep) |
| Algorithm filing | 1–3 months | ¥50,000 – ¥150,000 |
| PIPL impact assessment | 1–2 months | ¥100,000 – ¥300,000 |
Where to Go From Here
— China Gateway 360 —
Remote China market entry support, built around execution.
