TLDRocket
Sign in

Why You Will Probably Never Know Which of Your Data Went Into AI Training

Trending Topics Jakob Steinschaden

The EU’s AI Act transparency requirement forces general-purpose AI providers to publish training-data summaries, but the EU Commission’s FAQ limits what can be identified about any specific user’s content. Providers must list only the top 10 percent of domains by training-data volume (or the top 5 percent / 1,000 domains for smaller firms), leaving most scraped sources unreported. As a result, summaries become category-level disclosure rather than a way to determine whether your specific data or URLs were used, pushing people toward indirect tests, legal discovery, or future opt-outs like rights reservation and platform objections.

Why it matters

Since last year, the EU has had a rule on the books that sounds, at first glance, like the end of the secrecy: anyone placing a general purpose AI model on the European market has to publish a summary of the content it was trained on. The obligation sits in Article 53 of the AI […] Der Beitrag Why You Will Probably Never Know Which of Your Data Went Into AI Training erschien zuerst auf Trending Topics.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.