Skip to content

TheLLM Brief

← All stories

Policy

CNIL Clears Legitimate Interest for AI Training Data

CNIL Clears Legitimate Interest for AI Training Data
Image: Skadden

Public-source scraping can be lawful under GDPR, with conditions on balancing, safeguards, and documentation.

Sourced from Skadden

France's data regulator, the CNIL, has issued guidance confirming that legitimate interest is a viable legal basis under GDPR for training AI models on personal data scraped from public sources. Three conditions apply: a credible balancing of interests, demonstrable safeguards, and clear documentation. Skadden published the analysis.

This is one layer cleared, not the whole stack. The guidance addresses training-phase GDPR exposure only. Copyright, database rights, commercialisation constraints, and post-training litigation risk sit outside its scope entirely. Labs operating in Europe now have a cleaner path on one specific question, but the adjacent liabilities remain open.

Watch how enterprise buyers and cloud providers respond. A GDPR green light on training data sourcing makes European AI development marginally more viable, but operators deploying models still face deployment-phase constraints the CNIL did not touch. The compliance picture is narrower than the headline suggests.

Analysis

GDPR clarity at the training stage is real, but it is plumbing, not a palace. Deployment-phase liability, copyright, and database rights are where the next legal battles land.

Research this with your AI

Copy the research prompt into your AI assistant to see how this story affects you.

Then paste it into ChatGPT, Claude, Gemini, Grok and others.
Runs in your own assistant with your own context. Nothing is sent to us.
Show the prompt
I just read this AI news story and want to understand it in my own context.

Title: CNIL Clears Legitimate Interest for AI Training Data
Summary: France's CNIL confirmed AI model training on publicly scraped personal data can meet GDPR's legitimate interest standard. Copyright, database rights, and deployment-phase liability remain unresolved.
Category: Policy
Source: Skadden, https://www.skadden.com/insights/publications/2025/06/cnil-clarifies-gdpr-basis-for-ai-training

Using my own history and context, help me understand:
1. What is the core development and why does it matter?
2. Who are the major players involved and what are their motivations?
3. How does this fit into the broader AI landscape right now?
4. How does this apply to my own work, and what should I do or watch next?

Be specific and plain spoken.

Newsletter

The day's AI stories, with the editor's take, in one email.

Free. Unsubscribe in one click.