Why You Will Probably Never Know Which of Your Data Went Into AI Training
Trending Topics Jakob Steinschaden
The EU’s AI Act transparency requirement forces general-purpose AI providers to publish training-data summaries, but the EU Commission’s FAQ limits what can be identified about any specific user’s content. Providers must list only the top 10 percent of domains by training-data volume (or the top 5 percent / 1,000 domains for smaller firms), leaving most scraped sources unreported. As a result, summaries become category-level disclosure rather than a way to determine whether your specific data or URLs were used, pushing people toward indirect tests, legal discovery, or future opt-outs like rights reservation and platform objections.
Why it matters
Since last year, the EU has had a rule on the books that sounds, at first glance, like the end of the secrecy: anyone placing a general purpose AI model on the European market has to publish a summary of the content it was trained on. The obligation sits in Article 53 of the AI […] Der Beitrag Why You Will Probably Never Know Which of Your Data Went Into AI Training erschien zuerst auf Trending Topics.