OpenWALDO aims to blow the doors off proprietary AI training models
The Register ● Covered by 2 sources
Gregory Kurtzer’s new OpenWALDO project wants a shared, open AI training set anyone can help build. It’s meant to make model training auditable, instead of hiding the data behind closed doors.
Based on reporting by The Register — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Gregory Kurtzer, the CentOS and Rocky Linux founder, is taking a run at one of AI’s messier problems: training data nobody can inspect. His new project, Open Weights, Artifacts, Licenses, Data, Origins — OpenWALDO for short — is trying to build a shared dataset that people can contribute to like an open source software project.
The idea is simple enough. If model builders can see what went into the training set, and under what license, then the whole process gets less mysterious. CIQ, Kurtzer’s AI infrastructure company and the project’s backer, says today’s open-weight models still rely on training data that is effectively closed off. That can include copyrighted material, distilled outputs from other models, and user-generated content that may never have been offered under truly open terms.
CIQ’s pitch goes beyond philosophy. It argues that hidden training data can create legal and security headaches, and that the AI industry is duplicating a lot of expensive work by training behind closed doors. A shared public corpus, the company says, would let different teams start from the same verified baseline, then add their own proprietary data and keep a clear record back to the source material.
Kurtzer is framing that as the next move for open source itself. He points to the old arguments against open code — that it was unsafe, untrustworthy, impossible to manage — and says Linux won because it could be inspected, forked, and validated by a community. OpenWALDO is meant to give AI that same property. Whether the industry buys it is another matter.
For now, the dataset is still tiny next to the training piles used by frontier models. OpenWALDO lists 167.3 billion reference tokens drawn from government records, open-source academic papers, mailing lists, and public domain literature, while the big AI systems are trained on sets measured in the tens of trillions of tokens. CIQ didn’t say whether anyone has already trained a model on the dataset, and that silence says plenty about how early this is.
My take — AI-written commentary, not fact-checked reporting
This is the right fight, and it’s overdue. AI vendors love to act like secrecy is a feature until somebody asks what exactly they fed the machine. Open source won this argument in software by being inspectable; AI won’t get a free pass just because the models are bigger and the invoices are uglier.
Read more about this at: The Register