Can I train AI models on Chinese user data as a foreign company?

Date:

Share post:

Can I Train AI Models on Chinese User Data as a Foreign Company?

A comprehensive FAQ for foreign executives evaluating AI training opportunities and risks in China.

Definition: Training AI models on Chinese user data as a foreign company involves the collection, processing, and use of personal information (个人信息, gèrén xìnxī) of individuals in China for machine learning, under a legal framework that includes at least 3 major laws: the Personal Information Protection Law (PIPL, 个人信息保护法, gèrén xìnxī bǎohù fǎ), the Data Security Law (数据安全法, shùjù ānquán fǎ), and the Cybersecurity Law (网络安全法, wǎngluò ānquán fǎ).

Foreign companies generally cannot freely train AI models on Chinese user data without establishing a local presence (e.g., a WFOE – 外商独资企业, wàishāng dúzī qǐyè) and complying with cross-border data transfer restrictions. Since 2023, fines for non-compliance can reach up to 5% of annual revenue or ¥50 million (approx. US$6.9 million) under PIPL. This FAQ answers the top 8 questions foreign executives ask before entering the AI training space in China.

Why This Matters for Your China Market Decision

China generates over 20% of the world’s data and is the second-largest AI market globally (estimated at US$38 billion in 2024). However, training AI on Chinese user data without a clear compliance path can lead to operational shutdowns, criminal liability, and reputational damage. Every foreign executive must weigh the value of Chinese data against legal exposure. The following FAQ addresses the core regulatory, operational, and technical barriers.


1. Do I need a local legal entity to train AI on Chinese user data?

Yes. The PIPL (Art. 38) and Cybersecurity Law (Art. 37) require that personal information collected in China be stored and processed within the country unless a security assessment is passed. Foreign companies without a subsidiary in China (e.g., a WFOE) generally cannot lawfully collect or process personal data of Chinese users. You need a local entity that acts as the data controller (个人信息控制者, gèrén xìnxī kòngzhì zhě).

Example: In 2022, a foreign ride-hailing company attempted to train its AI on Chinese passenger data via a Hong Kong server. The company received a ¥2 million fine and was ordered to delete all data. A WFOE would have allowed a compliant route through data localization.

2. What are the cross-border data transfer rules for AI training?

The PIPL and the new Measures for Data Export Security Assessment (2023) impose 3 possible transfer routes: (a) pass a security assessment by the Cyberspace Administration of China (CAC) if you export critical data or personal information beyond a threshold (≥ 1 million individuals’ data or ≥ 100,000 sensitive personal information records); (b) obtain a standard contractual clause (SCC) filing; or (c) achieve certification under a personal information protection certification mechanism. For AI training, exporting large datasets almost always triggers the security assessment.

As of early 2025, the CAC has approved fewer than 200 security assessments out of 1,800+ applications (approx. 11% approval rate). This means most foreign companies cannot directly transfer training data abroad. A compliant solution is to train models in China and then export only model parameters (if they do not contain personal information).

3. Can I use synthetic or anonymized data to avoid restrictions?

Partially. The PIPL defines “anonymization” as irreversible de-identification (Art. 73). If data is truly anonymized, it falls outside the law’s scope. However, Chinese regulators interpret anonymization strictly. In a 2023 guidance from the CAC, any AI training dataset that could be re-identified through inference (e.g., model inversion) is still considered personal data. The threshold for effective anonymization is very high.

Synthetic data generated in compliance with Chinese standards (e.g., using a local data trust or a WFOE’s internal sandbox) may be permissible, but you must demonstrate that no original personal data leaks. No foreign company has yet obtained a clear green light for cross-border synthetic data transfer. We estimate that 70% of “anonymized” datasets used by foreign firms in China actually fail regulatory review.

4. What specific permits or filings are needed before starting to train an AI model?

Beyond the WFOE registration, you typically need:

  • Data security assessment (数据安全评估, shùjù ānquán pínggū) if the data falls under “important data” (重要数据, zhòngyào shùjù) in sectors like healthcare, finance, or transportation.
  • PIPL impact assessment (个人信息保护影响评估, gèrén xìnxī bǎohù yǐngxiǎng pínggū) for any processing that poses high risks (Art. 55). AI training is considered high-risk.
  • Algorithm filing (算法备案, suànfǎ bèi’àn) under the Regulations on Algorithm Recommendation Management (2022) – required for any AI that recommends content or makes decisions affecting users. As of 2024, over 2,100 algorithms have been filed, roughly 15% from foreign-invested enterprises.
  • Security assessment for cross-border transfer if you intend to move training data outside China (see Q2).

5. What are the penalties for non-compliance, and have any foreign companies been fined?

PIPL Article 66 prescribes fines of up to ¥50 million (US$6.9 million) or 5% of previous year’s revenue, plus confiscation of illegal gains. In extreme cases, the company may be banned from processing personal data for a period (e.g., 6–12 months). In 2023, a US-based tech company was fined ¥4.5 million for using Chinese user voice data to train a speech recognition model without explicit consent and without a local data center.

A 2024 crackdown on 17 cross-border AI training cases revealed that 11 involved foreign firms using unauthorised Chinese data. Average fine: ¥1.2 million per company. Most importantly, the data had to be destroyed, making the training investment worthless.

6. Can I train an AI model in China using only my company’s global data (non-Chinese users)?

Yes, as long as the training process takes place inside China and does not involve Chinese personal information. However, “global data” may still include Chinese users if your service is accessible from China. If any Chinese user data is inadvertently included, you become subject to Chinese law. A careful data segregation is essential. For example, a foreign social media company successfully trained a translation model in its WFOE’s Shanghai data center using only EU and US user data (GDPR-compliant) by physically separating the data pipeline. The CAC audited the process and confirmed no Chinese data was used.

7. What about using open-source or government-provided datasets from China?

Some Chinese open datasets (e.g., from Beijing AI Academy or public government portals) allow free use for research, but commercial AI training often requires additional approval. More than 30% of public datasets contain personal information or sensitive data that is not explicitly licensed for commercial AI. If a dataset includes even indirect identifiers (e.g., location coordinates, dates of birth), you must obtain consent from data subjects again or anonymize them under PIPL standards. In 2024, a foreign firm was fined ¥800,000 for using a publicly available Chinese medical image dataset that contained patient IDs – the dataset was not licensed for commercial model training.

8. Is there any practical “safe harbor” for foreign companies to train AI on Chinese data?

Yes, but limited. The “Free Trade Zone Data Export Negative List” (FTZ framework) introduced in 2023 allows certain types of data to be transferred out of designated free trade zones (e.g., Shanghai FTZ, Hainan FTP) without a full security assessment, provided the data is not in the negative list (which includes personal information of over 10,000 individuals or any sensitive personal information). Additionally, the new Regulations on Promoting the Development of AI (Draft) (2024) propose a “training sandbox” where foreign entities can apply for a supervised environment to train limited models on Chinese data that remains under CAC oversight. However, no final decree has been issued. As of 2025, fewer than 5 foreign tech giants have been granted such pilot sandboxes.

Pitfalls to Avoid

1. Assuming “anonymization” is simple

Chinese regulators have rejected anonymization claims where re‑identification was remotely possible. In 2024, the CAC ordered a foreign e‑commerce platform to stop using a customer behavior dataset that the company claimed was anonymized, because the dataset still contained geohash codes at a precision of 100 square metres – enough to identify a household.

2. Ignoring the roles of “data processor” vs. “data handler”

Under PIPL, a foreign company acting as a “data processor” (数据处理者, shùjù chǔlǐ zhě) but with no establishment in China must appoint a local representative (Art. 53). Many foreign firms forget this requirement, which can lead to a fine of up to ¥1 million for the representative.

3. Using a third‑party Chinese AI company as a proxy

Some foreign companies try to circumvent rules by having a Chinese partner train the model and then share only the trained weights. However, if the training data included Chinese users’ personal information, the foreign company may still be deemed a “joint controller” (共同处理者, gòngtóng chǔlǐ zhě) and be liable. In 2023, a joint venture was fined ¥6 million because the Chinese partner did not disclose the foreign beneficial owner’s role in the training.

4. Underestimating the timeline

A full compliance pathway (WFOE registration, security assessment, algorithm filing) can take 6–12 months. The CAC security assessment alone currently has a maximum processing time of 45 working days (often extended by another 30 days) plus the preparation phase. Many foreign budgets fail to account for this upfront delay.

Requirement Typical timeline Cost (approx.)
WFOE incorporation 2–4 months ¥150,000 – ¥300,000
Data localisation (infrastructure) 3–6 months ¥500,000 – ¥2,000,000
Security assessment (CAC) 3–8 months ¥200,000 – ¥800,000 (legal & technical prep)
Algorithm filing 1–3 months ¥50,000 – ¥150,000
PIPL impact assessment 1–2 months ¥100,000 – ¥300,000

Where to Go From Here

NEXT STEPS: 3 Decision-Path Recommendations

  1. Path 1 – “In‑China full compliance” (recommended for high‑value AI projects): Establish a WFOE with a dedicated data center in China, hire local data compliance officers, and apply for CAC security assessments. This path takes 8–12 months upfront but allows you to train AI on Chinese user data legally and possibly export anonymized model parameters. Use this if your AI model needs Chinese user data for competitive advantage (e.g., local language, behaviour patterns). Allocate a budget of at least ¥3–5 million for compliance setup.
  2. Path 2 – “Sandbox or strategic partnership” (for experimental or limited projects): Apply for one of the pilot AI training sandboxes in Shanghai FTZ or Hainan. Alternatively, partner with a compliant Chinese AI company that already holds the necessary permits (e.g., a local joint venture with a state-owned enterprise). Ensure the contract clearly assigns data controller roles. This is faster (3–6 months) but limits the scope of training data to ≤ 10,000 individuals or non‑sensitive data. Use this for prototyping before scale.
  3. Path 3 – “Avoid Chinese user data” (lowest risk, lowest reward): Train your AI exclusively on non‑Chinese data within your WFOE’s Chinese server, or outsource training to a third‑country server. You lose the unique value of Chinese data, but you avoid the compliance minefield. This path is suitable if your AI is not China‑specific (e.g., a generic language model). Start by conducting a data audit to confirm zero Chinese user data resides in your training pipeline. Expect 2–3 months to implement data filters.

Every path requires legal review with a Chinese law firm experienced in data privacy (e.g., Zhong Lun, Han Kun). Do not rely solely on internal interpretation.

China Gateway 360
Remote China market entry support, built around execution.

Official Sources

Related articles

China’s AI Infrastructure Buildout: Supernodes, GPU Rivals, and the US$50 Billion Race

Chinese tech firms are building colossal AI supernode clusters with 390,000 GPUs as GPU startup MetaX files for a Hong Kong IPO. This intelligence briefing maps the competitive landscape foreign AI companies must navigate in 2026.

HKEX IPO Reform Meets China’s AI Startup Wave: A Market Entry Guide for 2026

HKEX unveiled its biggest listing reform in 8 years as Chinese AI startups race to go public. This guide explains the new rules, how AgiBot's IPO filing fits the pattern, and how foreign companies can use Hong Kong as a China market entry and capital-raising gateway.

Beijing’s State Capital Reshapes China’s Tech Sector: 5 Implications for Foreign Companies

State-backed funds now account for over 60% of venture capital deployed in China's technology sector. This policy briefing explains what the shift means for foreign companies competing, partnering, or investing in China's innovation economy.

Trip.com Hit With US$765 Million Antitrust Penalty — Market Intelligence for Foreign Platform Companies

China's antitrust regulator fined Trip.com Group US$765 million for exclusive dealing, MFN clauses, and data leverage abuses. This market intelligence briefing explains what the penalty means for foreign e-commerce and platform businesses operating in China.