TLDRocket
Sign in

Data is better together: Enabling communities to collectively build better datasets together using Argilla and Hugging Face Spaces

Hugging Face

Hugging Face and Argilla got 350 volunteers to rank 11,000 AI prompts in just days. Now they're inviting more communities to build datasets together, no coding required.

Based on reporting by Hugging Face — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Argilla and Hugging Face just proved something that a lot of AI labs would rather you didn't think about too hard: you don't need a giant budget or an army of contractors to build a decent dataset. You need a low-friction tool and a community willing to show up. Over a few days, 350 people logged in and rated more than 11,000 prompts, producing a new dataset called 10k_prompts_ranked — 10,000 prompts scored for quality by actual humans, not scraped and hoped-for.

The trick wasn't some clever labeling algorithm. It was removing the annoying parts. Argilla recently added Hugging Face account login to its annotation tool when hosted on Spaces, which means someone can go from clicking a link to rating prompts in the time it takes to read this sentence. No account creation, no onboarding friction, no waiting for IT to approve access. That sounds like a small technical tweak, but it's the kind of small tweak that determines whether 350 people show up or 15 do.

The bigger point here is about where good training data actually comes from. Plenty of languages, specialized domains, and niche tasks still don't have solid datasets, even as the Hugging Face Hub fills up with new models and demos every day. Big labs solve this with money — hire annotators, buy data, move on. Open communities solve it differently: they distribute the work across people who care about a language or a subject and don't need to be paid to spend twenty minutes on it.

So now Argilla and Hugging Face are trying to turn a one-off experiment into a repeatable pattern. They're opening applications for a first cohort of community groups who want to build their own datasets — chat data for underrepresented languages, benchmarks for specific domains, preference data from more diverse participants, whatever a community actually needs. Accepted groups get a properly configured Argilla Space with free persistent storage and better CPU resources, help with setup, and promotional support so people actually find the project. The catch, if you can call it that, is that this round is text-only; multimodal ideas are welcome but not guaranteed a spot.

Applications aren't a form to fill out — you just show up in the Data is Better Together channel on the Hugging Face Discord and pitch what you want to build. It's a low bar on purpose. Given how fast the first experiment moved, betting against a second cohort producing something useful within weeks seems unwise.

My take — AI-written commentary, not fact-checked reporting

This is the unglamorous, unsexy work that actually keeps open-source AI competitive with the closed labs sitting on mountains of proprietary data, and I'd rather see Hugging Face's energy go here than into another leaderboard flex. The lesson people keep missing is that the bottleneck was never enthusiasm, it was friction — fix the login screen and 350 volunteers show up overnight. Every underrepresented language and niche domain in AI has a community that would do this for free if someone just built them the on-ramp.

Read more about this at: Hugging Face

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.