Skip to content

TheLLM Brief

← All stories

Policy

France's CNIL Sets Rules for Scraping Personal Data in AI Training

France's CNIL Sets Rules for Scraping Personal Data in AI Training
Image: www.hlc.com

Two guidance sheets published June 19 clarify when scraping personal data is lawful.

Sourced from Hogan Lovells

On June 19, 2025, France's data protection authority CNIL published two how-to sheets targeting AI developers. One addresses legitimate interest as a legal basis for building training datasets. The second sets conditions for lawful web scraping of personal data.

The guidance, reported by www.hlc.com, acknowledges that scraping has fundamentally changed how internet data is accessed and reused. CNIL identifies significant risks to data subjects from mass collection. The authority stops short of prohibition but requires developers to assess each scraping use case individually and implement appropriate safeguards.

CNIL also recommends that lawmakers introduce dedicated legislation regulating public-authority scraping practices. Labs and operators building training pipelines in France now have clearer compliance benchmarks, but no safe harbor. The call for new legislation signals that current GDPR frameworks are straining under the weight of AI-scale data collection. Watch whether other EU member-state regulators follow with parallel guidance.

Analysis

Capability is not the constraint here, compliance is. Labs that scrape at scale now face a documented standard, and the regulator has flagged that existing law is not enough.

Research this with your AI

Copy the research prompt into your AI assistant to see how this story affects you.

Then paste it into ChatGPT, Claude, Gemini, Grok and others.
Runs in your own assistant with your own context. Nothing is sent to us.
Show the prompt
I just read this AI news story and want to understand it in my own context.

Title: France's CNIL Sets Rules for Scraping Personal Data in AI Training
Summary: CNIL published two how-to sheets on June 19, 2025, covering legitimate interest and web scraping for AI training datasets. The regulator stops short of banning scraping but demands case-by-case assessment and calls for specific legislation.
Category: Policy
Source: Hogan Lovells, https://www.hlc.com/en/publications/development-of-an-ai-system-cnil-issues-guidelines-regarding-collection-of-data-via-web-scraping

Using my own history and context, help me understand:
1. What is the core development and why does it matter?
2. Who are the major players involved and what are their motivations?
3. How does this fit into the broader AI landscape right now?
4. How does this apply to my own work, and what should I do or watch next?

Be specific and plain spoken.

Newsletter

The day's AI stories, with the editor's take, in one email.

Free. Unsubscribe in one click.