TLDRocket
Sign in

Fed up with Big Tech, communities turn to data collectives for control

Rest of World Rina Chandran

Communities are forming data collectives to control how their data trains AI, instead of letting Big Tech scrape it for free. Small or endangered languages finally get a shot at chatbots and speech tools built for them.

Based on reporting by Rest of World, Rina Chandran — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

There's a quiet rebellion brewing against the way AI companies have treated the internet as an all-you-can-eat buffet. Meta, OpenAI, Google, Anthropic — they've hoovered up nearly everything scrapeable, often without asking, and defended it as fair use. Now communities that hold smaller, unusual, or culturally specific data sets are deciding they'd rather set their own terms than get scraped again.

Mozilla Foundation is one organization trying to formalize that shift. Its CTO, Raffi Krikorian, says people are increasingly aware they're sitting on data that doesn't exist anywhere else online, and they want a say in how it's governed. Mozilla launched the Mozilla Data Collective last year specifically to give those data sets a home. Krikorian is blunt about the motivation: the big AI companies built their products on the work of people who are now realizing they never agreed to the deal.

The language angle is where this gets genuinely moving. Meesum Alam, a computational linguistics PhD candidate at Indiana University, grew up Baloch in Pakistan without speaking Balochi, his own ancestral language — what he calls a linguistic identity crisis. That pushed him to start documenting dying languages, starting with Dawoodi, spoken by roughly 300 people, and Kalasha, with about 3,000 speakers. He's since gathered around 700 hours of voice data across 39 languages, uploaded to Mozilla's platform, where companies including Meta have used it to build speech recognition tools. The communities involved set the rules themselves — research and non-commercial use only — and companies have to negotiate directly with them.

This isn't really about money, Alam says. It's about not getting exploited again by companies that could otherwise take the data and monetize it without consent. In Africa, a similar framework called the Nwulite Obodo Open Data License, launched in 2024, lets creators share data without forfeiting rights to benefit from it — about 70 African data sets, covering more than 20 languages plus music and oral traditions, now sit inside Mozilla's collective through that license.

None of this is friction-free. Astha Kapoor of India's Aapti Institute points out that data cooperatives built purely to steward information without a sustainability plan often end up forced into monetizing the very thing they were meant to protect — the opposite of the point. Mozilla is already testing paid commercial licensing for a handful of community data sets, with plans to widen it. Whether that solves the sustainability problem or just recreates a smaller version of the extraction it was built to avoid is still an open question. But for people like Alam, who says he now regularly gets voice notes from someone using a chatbot in their own language for the first time, the emotional payoff is already real, even if the governance model is still being figured out.

My take — AI-written commentary, not fact-checked reporting

Big Tech didn't ask permission the first time around, so it's hard to feel sorry that communities are now making them negotiate. The real test isn't whether collectives can attract a Meta or an Anthropic to the table — it's whether they can pay their own bills without turning into the same monetization machine they were built to resist. Watch the funding model, not the mission statement.

Read more about this at: Rest of World

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.