What Are the Data Localization Requirements for AI in China?
China’s data localization regime — built on the Cybersecurity Law (CSL), Data Security Law (DSL), and Personal Information Protection Law (PIPL) — creates binding obligations for any AI company operating within or serving users in China. For foreign enterprises training, deploying, or fine-tuning AI models using data collected in China, the key requirement is straightforward: most “important data” and “personal information” must remain on domestic servers, and any cross-border transfer must pass a strict security assessment, certification, or standard contractual clause process before it can leave the country. Failure to comply can result in fines of up to 50 million RMB (approximately USD 6.9 million) or 5% of annual revenue, suspension of operations, and criminal liability for responsible individuals.
These requirements apply with special force to the AI sector because model training inherently involves large-scale data processing, and Chinese regulators have signaled that training datasets containing personal information or data classified as “important” are squarely within the scope of localization mandates. The Cyberspace Administration of China (CAC) issued specific guidance in 2023 and 2024 clarifying that AI training data falls under the DSL and PIPL frameworks, meaning any foreign company collecting Chinese user data for AI model development must either keep that data in China or navigate the cross-border transfer pipeline before moving it to overseas servers.
Key Data Categories and Their Localization Status
| Data Category | Legal Basis | Localization Requirement |
|---|---|---|
| Personal Information (PI) of Chinese residents | PIPL Art. 36-40 | Must be stored in China unless passing security assessment, standard contract, or certification |
| Sensitive PI (biometrics, health, location, ethnicity) | PIPL Art. 28-32 | Strict localization required; cross-border transfer only with individual consent + security assessment |
| Important Data (non-personal data that could harm national security) | DSL Art. 31 | Mandatory localization; cross-border transfer requires CAC security assessment |
| AI training datasets containing PI or important data | DSL Art. 21, PIPL Art. 38 | Must undergo data classification before any transfer; localization mandatory unless assessed |
| Core State Data (national defense, critical infrastructure) | DSL Art. 21 | Absolute localization; cross-border transfer effectively prohibited |
The Three-Pronged Cross-Border Transfer Regime
Under the PIPL, foreign companies have three legally permissible pathways to transfer personal information collected in China to overseas servers for AI training purposes:
- CAC Security Assessment (PIPL Art. 38-40): Required when processing PI of more than 1 million individuals, or transferring more than 100,000 individuals’ PI abroad, or transferring more than 10,000 individuals’ sensitive PI abroad since the previous year. The assessment evaluates necessity, data protection measures, recipient country’s legal environment, and risk of re-identification. Processing time: 3-9 months.
- Standard Contractual Clauses (SCCs — PIPL Art. 38): Available for smaller-scale transfers (below the security assessment thresholds). The data exporter signs the CAC-prescribed SCC with the overseas recipient and files it with the provincial CAC office. The contract must include Chinese governing law and submit to Chinese regulatory oversight. Effective immediately upon filing.
- Certification (PIPL Art. 38): Companies can obtain certification from an accredited institution (e.g., China Cybersecurity Review and Certification Center) certifying that their cross-border data processing meets PIPL standards. This pathway is most commonly used by multinational corporations with established China operations and mature data governance programs.
Implications for AI Model Training
The data localization framework creates specific operational constraints for foreign AI companies. Training a large language model (LLM) on Chinese user data typically requires assembling datasets from millions of individuals — exceeding the security assessment threshold by a wide margin. This means the CAC assessment pathway is effectively mandatory for any meaningful AI training operation involving Chinese personal information.
Foreign companies have adopted three workable strategies. The first is in-China training, where model training takes place entirely on domestic servers using Chinese infrastructure (Alibaba Cloud, Huawei Cloud, Tencent Cloud). The trained model weights may then be transferred abroad under an SCC or assessment. The second is data minimization and anonymization, where companies strip all personal identifiers from training datasets and obtain expert confirmation that the data no longer constitutes personal information under PIPL definitions — removing it from the localization requirement entirely. The third is joint venture or WFOE structure, where a Chinese legal entity handles all data processing domestically, and only anonymized model outputs cross borders.
Enforcement and Recent Developments
Since mid-2023, the CAC has stepped up enforcement against unauthorized cross-border data transfers. In a widely reported 2024 case, a foreign AI consulting firm was fined 12 million RMB for transferring Chinese user behavior data to its overseas headquarters without going through the security assessment process. The regulator also ordered deletion of all data that had been transferred. Separately, China’s 2024 AI governance guidelines explicitly require that training data for generative AI models comply with data localization rules, closing a prior gray area where companies argued that training data was not “personal information” if it was aggregated.
Data Classification Obligations for AI Companies
A critical but often overlooked aspect of China’s data localization framework is the mandatory data classification requirement under the Data Security Law. DSL Art. 21 requires all companies that process data to establish a data classification system that categorizes data by its importance to national security, economic interests, and public order. For AI companies, this means every dataset used in model training must be formally classified before any cross-border data flow analysis can begin. The classification categories — General Data, Important Data, and Core State Data — determine which localization and cross-border transfer rules apply. A dataset classified as Important Data triggers mandatory localization regardless of volume, while General Data may be transferable under the SCC pathway.
The classification process is not a one-time exercise. Under the CAC’s 2024 guidance on AI data governance, AI companies must reassess their data classification whenever: (a) a new training dataset is introduced, (b) the model’s use case expands to a new domain or industry vertical, or (c) new regulations are issued that add categories to the Important Data definition. Foreign AI companies have found this ongoing classification requirement to be one of the most operationally demanding aspects of compliance, as it requires dedicated data governance personnel and regular consultation with provincial CAC offices to verify classification decisions.
Several industry-specific data classification standards have been issued since 2023. The automotive sector, financial services, healthcare, and critical information infrastructure (CII) operators all have sector-specific data classification guidelines that impose additional localization requirements on AI systems operating in those domains. A foreign AI company developing an autonomous driving system in China, for example, must comply with both the general DSL classification rules and the automotive sector’s specific data security regulations, which classify high-definition mapping data and vehicle trajectory data as Important Data with mandatory in-China storage requirements.
Key Strategic Recommendations
Foreign AI companies entering China should take the following steps to build a compliant data localization framework. These recommendations are drawn from actual compliance programs implemented by foreign-invested AI companies operating in Beijing, Shanghai, and Shenzhen:
- Conduct a data mapping exercise — Catalog all data types collected, stored, and processed for AI training. Classify each type under PIPL/DSL categories (personal information, sensitive personal information, important data, or anonymized). Determine which categories exceed the cross-border transfer thresholds.
- Choose your legal pathway early — If your training dataset includes PI from more than 1 million individuals (likely for any general-purpose LLM), plan for the CAC security assessment from the outset. The 3-9 month processing timeline must be built into your go-to-market schedule.
- Establish an in-China processing environment — Even if your eventual goal is to use Chinese data for global model development, establish a domestic processing pipeline with Chinese cloud infrastructure. Train and validate your models in China, then transfer only necessary model artifacts abroad under the appropriate legal framework.
- Document your anonymization methodology — If you plan to argue that your training data has been anonymized and therefore falls outside PIPL scope, prepare detailed technical documentation of your de-identification process and consider obtaining third-party certification of anonymity.
- Monitor regulatory updates — China’s AI and data governance frameworks are evolving rapidly. The CAC released draft rules on data cross-border transfer security assessment in March 2024 with simplified procedures for smaller-scale transfers. Subscribe to regulatory monitoring services and maintain legal counsel with China data protection expertise.
Bottom Line
China’s data localization requirements for AI are among the strictest globally, and they apply to any foreign company handling Chinese-source data for AI development. Foreign companies cannot train AI models on Chinese user data stored overseas without navigating the CAC security assessment, SCC, or certification pathways. The practical consequence is that most foreign AI companies operating in China will need to establish domestic data processing infrastructure and conduct model training within China’s borders. Non-compliance carries severe financial and operational penalties — but the framework is navigable with proper planning, legal guidance, and investment in China-based technical infrastructure.
Where to Go From Here
Based on what you just read:
- Ready to act? Read [guide: SLUG-TO-BE-FILLED]
- Still comparing? See [comparison: SLUG-TO-BE-FILLED]
- Need numbers? Try [tool: SLUG-TO-BE-FILLED]
— China Gateway 360 —
Remote China market entry support, built around execution.
