U.S. regulators launched a comprehensive investigation this week into the data used to train artificial intelligence models since 2018, officials said. The probe aims to assess the scope and legality of vast datasets—including web crawls, licensed content, and biometric information—employed by major AI developers to ensure compliance with emerging transparency and safety standards.
The investigation, initiated by U.S. regulatory agencies this week, targets the extensive datasets used by leading artificial intelligence developers to train foundation models since 2018, officials said. The probe seeks to determine whether the collection and use of such data comply with emerging transparency rules and safety standards.
These datasets reportedly include hundreds of billions to trillions of text tokens sourced from web crawls, licensed content, and biometric data, according to public documentation and corporate reports reviewed by regulators.
Records from major AI labs indicate that training data typically comprises a mix of publicly available internet content—including Common Crawl archives, digitized books, academic articles, and multilingual corpora spanning dozens of languages—as well as licensed proprietary datasets and synthetic data generated by earlier AI models. Parameter counts for state-of-the-art models often exceed hundreds of billions, implying data volumes far larger than those used in traditional machine learning, sources confirmed. Industry disclosures also highlight that a significant portion of training data remains “weakly curated” web content, supplemented by smaller, carefully filtered subsets for fine-tuning.
The sources of training data extend beyond text. Technical reports show that code-capable models are trained on open-source repositories such as public Git platforms and programming Q&A forums, while multimodal models incorporate large-scale image-text pairs from web scrapes, including photographs, diagrams, memes, and product images with captions. Privacy regulators have raised concerns that broad web scraping may inadvertently capture sensitive personal information, such as contact details, resumes, social media posts, and health-related discussions, prompting scrutiny of data handling practices.
Legal and regulatory developments over the past two years have intensified scrutiny of AI training data. Since 2023, multiple jurisdictions have examined whether large-scale scraping of copyrighted works constitutes infringement or qualifies for text-and-data mining exceptions, according to legal analyses. Data-protection authorities in regions with comprehensive privacy laws, including those modeled on the European Union’s General Data Protection Regulation (GDPR), have launched inquiries into whether the use of publicly available personal data in training meets standards for lawful basis, transparency, and data subject rights. Officials noted that court rulings increasingly distinguish between non-commercial research and commercial deployment, imposing stricter consent and licensing requirements on the latter.
Several countries have adopted or proposed AI-specific legislation mandating transparency in dataset composition and obligations to mitigate bias and privacy risks, regulatory sources said. International watchdogs have published lists of high-risk AI training practices requiring prior approval, underscoring the need for explicit legal and contractual frameworks governing web scraping and data extraction. Industry insiders report growing pressure to shift from implied consent toward formal agreements with rights holders.
Transparency and disclosure practices among major AI developers have evolved in response to regulatory and public pressure. Technical documentation now often includes high-level descriptions of dataset categories and filtering methods, though exhaustive lists of individual sources remain undisclosed, according to recent model cards and system reports. Some developers have published data governance policies stating that sensitive content categories—such as child sexual abuse material and explicit hate speech—are filtered out using automated tools and human review. Industry commitments also mention plans for impact assessments, privacy reviews, and mechanisms allowing rights holders to request removal of content from training datasets, sources confirmed.
Government agencies and international bodies have issued formal guidance emphasizing that AI training on personal data must comply with principles of purpose limitation, data minimization, accuracy, and storage limitation, with clear notices provided to individuals. Competition and consumer protection regulators have warned that opaque data sourcing could mislead users and customers, calling for truthful marketing and robust documentation. Nonbinding frameworks from digital governance organizations encourage states to adopt standards for dataset transparency, human rights impact assessments, and safeguards for vulnerable populations. National AI strategies in several countries promote the creation of trusted data spaces and public sector datasets with clear licensing to support lawful training. Regulatory sandboxes allow AI developers to test privacy-preserving techniques and disclosure mechanisms in collaboration with authorities, helping shape future rules.
Content creators and rights holders have responded with a mix of legal and commercial actions. Creative industry groups and authors’ organizations have publicly objected to unlicensed use of books, journalism, music lyrics, images, and video stills in AI training. Several media companies have negotiated licensing deals granting access to archives and real-time content in exchange for fees and attribution requirements. Individual creators have supported class-action lawsuits seeking judicial clarification of copyright violations related to AI training. Rights-holder coalitions have introduced technical opt-out mechanisms—such as metadata signals and robots-style directives—to instruct crawlers and AI datasets not to use specified content. Negotiations between AI developers and content owners increasingly explore hybrid models combining licensed premium content with restricted use of general web data and stricter compliance with takedown requests.
Technical and policy innovations aim to improve the safety and accountability of training data. Research teams are developing privacy-preserving training techniques, including differential privacy, federated learning, and advanced de-identification pipelines, to minimize risks of memorizing or revealing personal data. Dataset curation efforts now often combine automated filters for illegal material, hate speech, explicit imagery, and copyrighted works with targeted human quality control. Proposals for data lineage and provenance tracking seek to attach machine-readable metadata to content indicating licensing terms and usage restrictions, enabling more granular compliance. Policy experts are exploring collective licensing schemes and remuneration models to allow scaled training on creative works while providing financial or reputational benefits to rights holders. Standards bodies and multi-stakeholder initiatives are working on benchmark norms for dataset documentation, risk assessment, bias evaluation, and transparency reports to enhance auditability.
The investigation will likely consider these complex legal, technical, and policy dimensions as regulators assess compliance with emerging AI transparency and safety standards. Officials have not disclosed a timeline for the inquiry or specific targets but indicated ongoing coordination with international counterparts to develop joint frameworks for foundation model safety audits.
