France's CNIL Sets Rules for Scraping Personal Data in AI Training

Two guidance sheets published June 19 clarify when scraping personal data is lawful.
On June 19, 2025, France's data protection authority CNIL published two how-to sheets targeting AI developers. One addresses legitimate interest as a legal basis for building training datasets. The second sets conditions for lawful web scraping of personal data.
The guidance, reported by www.hlc.com, acknowledges that scraping has fundamentally changed how internet data is accessed and reused. CNIL identifies significant risks to data subjects from mass collection. The authority stops short of prohibition but requires developers to assess each scraping use case individually and implement appropriate safeguards.
CNIL also recommends that lawmakers introduce dedicated legislation regulating public-authority scraping practices. Labs and operators building training pipelines in France now have clearer compliance benchmarks, but no safe harbor. The call for new legislation signals that current GDPR frameworks are straining under the weight of AI-scale data collection. Watch whether other EU member-state regulators follow with parallel guidance.
Analysis
Capability is not the constraint here, compliance is. Labs that scrape at scale now face a documented standard, and the regulator has flagged that existing law is not enough.
Research this with your AI
Copy the research prompt into your AI assistant to see how this story affects you.
Show the prompt
I just read this AI news story and want to understand it in my own context. Title: France's CNIL Sets Rules for Scraping Personal Data in AI Training Summary: CNIL published two how-to sheets on June 19, 2025, covering legitimate interest and web scraping for AI training datasets. The regulator stops short of banning scraping but demands case-by-case assessment and calls for specific legislation. Category: Policy Source: Hogan Lovells, https://www.hlc.com/en/publications/development-of-an-ai-system-cnil-issues-guidelines-regarding-collection-of-data-via-web-scraping Using my own history and context, help me understand: 1. What is the core development and why does it matter? 2. Who are the major players involved and what are their motivations? 3. How does this fit into the broader AI landscape right now? 4. How does this apply to my own work, and what should I do or watch next? Be specific and plain spoken.
Newsletter
The day's AI stories, with the editor's take, in one email.
Free. Unsubscribe in one click.