China Requires AI Training Data Source Disclosure — 4 Key Compliance Takeaways for Foreign Companies
As of August 2024, China’s Cyberspace Administration has mandated that all generative AI providers disclose training data sources, with over 70 models registered under the new framework. This rule, part of the 生成式人工智能 (Generative AI, shēngchéng shì réngōng zhìnéng) regulations administered by the 国家互联网信息办公室 (Cyberspace Administration of China, CAC, guójiā hùliánwǎng xìnxī bàngōngshì), requires foreign companies to document every dataset used in model training and submit a compliance report before market release.
What the Training Data Source Rules Require
The CAC’s September 2024 enforcement notice specifies that providers must submit a 训练数据来源报告 (Training Data Source Report, xùnliàn shùjù láiyuán bàogào) covering three categories: public web data, licensed third-party datasets, and proprietary internal data. For each category, companies must list the origin URL or contract reference, the volume of data in GB or TB, and a description of any personally identifiable information contained.
Foreign companies face an additional layer: data sourced from outside China must undergo a security assessment if it includes any Chinese user information or cross-border data transfers. The CAC has flagged that 14 of the 72 registered models as of October 2024 failed the first review due to incomplete foreign data documentation, delaying their market launches by an average of 63 days.
Beyond source identification, the rules require an 算法伦理审查 (Algorithmic Ethics Review, suànfǎ lúnlǐ shěnchá) for model outputs, checking for political bias, discrimination, and illegal content. This review must be renewed every 12 months or after any material update to the training dataset.
Enforcement Timeline and Penalty Structure
The CAC began accepting disclosure filings on August 15, 2024, with a grace period ending November 15, 2024. Providers that launched generative AI services before August 15 were given until September 30, 2024 to submit retrospective reports. As of October 2024, the CAC reports that 92 percent of registered providers have complied, while 8 percent face escalated scrutiny.
| Violation Type | First Offense | Second Offense | Third Offense |
|---|---|---|---|
| Failure to disclose training data source | Warning + 20,000–100,000 RMB fine | 100,000–500,000 RMB fine + 30-day service suspension | 500,000–1,000,000 RMB fine + license revocation |
| Incomplete or falsified source report | Warning + correction deadline (14 days) | 50,000–300,000 RMB fine + public notice | 300,000–800,000 RMB fine + product delisting |
| Cross-border data omission | Security rectification order + 50,000–200,000 RMB fine | 200,000–600,000 RMB fine + data transfer suspension | 600,000–1,500,000 RMB fine + business license penalty |
| Ethics review not renewed within 12 months | Warning + 14-day renewal period | 20,000–100,000 RMB fine + model suspension | 100,000–500,000 RMB fine + re-certification required |
The penalty table above shows the CAC’s escalating approach. For foreign companies, a third-offense license revocation effectively ends the ability to operate generative AI services in China, requiring a full new application process that the CAC notes takes 6–9 months to complete.
Implications for Foreign Companies
Foreign companies developing AI models using global training datasets — including Chinese-language data — must now separate their data pipelines or undergo dual compliance reviews. This affects over 40 multinational technology firms currently operating generative AI products in China, according to the CAC’s October 2024 registry.
Two recent enforcement actions highlight the risks. In September 2024, a European AI assistant provider was fined 150,000 RMB for failing to disclose 22 GB of web-scraped Chinese social media comments in its training dataset. The CAC determined that 0.3 percent of those comments contained personally identifiable information requiring separate consent documentation. The company’s service was suspended for 14 days while it compiled the retroactive disclosures.
Also in September 2024, a US-based text-to-image generator voluntarily withdrew from the Chinese market after the CAC flagged that its training data included images from over 200 Chinese news websites without content licensing agreements. The withdrawal cost the company an estimated 12–18 months of re-tooling time and 8 million RMB in legal and compliance restructuring.
Decision Framework: If your company sources training data exclusively from China-licensed partners and hosts models on Chinese servers, the disclosure process takes 30–60 days. If your company mixes international and Chinese data or uses cross-border training pipelines, expect a 90–150 day review timeline plus additional security assessment costs.
Operational Compliance Steps for Foreign Providers
The CAC has published a four-step process for foreign companies. First, complete the 数据来源登记表 (Data Source Registration Form, shùjù láiyuán dēngjì biǎo) for each training dataset, including volume, origin, and any Chinese user information. Second, submit the form to an approved 安全评估机构 (Security Assessment Agency, ānquán pínggū jīgòu) — only three agencies are currently authorized for foreign companies: China Information Security Evaluation Center, China Academy of Information and Communications Technology, and the Beijing Internet Security Center.
Third, obtain the 算法伦理审查报告 (Algorithmic Ethics Review Report, suànfǎ lúnlǐ shěnchá bàogào) from a certified reviewer. Fourth, file both documents with the CAC through its online portal and await the 30-day initial review window. If the CAC requests additional documentation, the review clock resets, and companies typically need 14–21 days to respond.
For foreign companies, the practical timeline breaks down as follows: data audit and form preparation takes 4–8 weeks, security assessment takes 3–6 weeks, ethics review takes 2–4 weeks, and CAC filing takes 4–8 weeks. Total estimated timeline: 13–26 weeks from start to market approval. Compared to the pre-August 2024 process — which averaged 4–6 weeks — the new disclosure requirements have more than doubled the compliance burden.
Budgeting for this process is essential. The CAC charges a 5,000 RMB filing fee per model registration, but the security assessment costs 30,000–80,000 RMB depending on data volume, and the ethics review costs 20,000–50,000 RMB. For a single model, total direct filing costs range from 55,000 to 135,000 RMB. Foreign companies with multiple models — some US firms maintain 5–10 separate models for the Chinese market — face multiplied costs plus the expense of data pipeline restructuring.
The CAC has stated that starting November 15, 2024, unregistered generative AI services will face immediate suspension with no grace period. As of October 24, 2024, the CAC reports that 72 models are registered, with an additional 34 applications under review. Companies that have not yet begun the filing process should expect no market access before late Q1 2025 at the earliest.
NEXT STEPS
- Audit your current training datasets immediately. Review all data sources for Chinese content, personally identifiable information, and cross-border elements. Use our AI Training Data Audit Checklist to identify gaps before engaging an assessment agency.
- Select an authorized security assessment agency. Only three agencies are approved for foreign companies. Compare their timelines and costs in our CAC Security Assessment Agency Guide, and begin the booking process early—waiting times exceed 6 weeks as of October 2024.
- Plan your compliance timeline and budget. With 13–26 weeks from audit to approval, map the process against your product launch schedule. Download our Generative AI China Compliance Timeline Template to align your internal milestones with CAC deadlines.
— China Gateway 360 —
Remote China market entry support, built around execution.
