How Domyn and AISquared built on Ai2's open releases
Allen Institute (AI2)
Two startups, Domyn and AISquared, built enterprise AI on Ai2's fully open Olmo, Dolma and Dolci releases. Why it matters: in regulated industries, showing your training data is becoming the price of admission.
Procurement teams at banks and federal agencies don't care how clever your model is if they can't trace what went into it. That's the wall a lot of vendors hit, and it's exactly the wall Domyn and AISquared decided to build around instead of through, by starting from Ai2's open-source Olmo family rather than a black-box foundation model.
AISquared, out of Washington, D.C., used Olmo 2 and Olmo 3 as the base for Bolt Instruct, a set of small models at 1B, 7B, and 32B parameters that now do guardrail filtering and request routing inside the company's UNIFI platform. Co-founder Jacob Renn says other open-weight options he tested came with murkier architectures that made fine-tuning painful and expensive, while Olmo's full transparency let his team trust it enough to customize aggressively. The payoff, he claims, was blunt and financial: roughly a 50% cut in hosting costs, passed along to customers too.
Domyn, based in Milan, took a longer route. It trained its own 10-billion-parameter foundation model, Italia 10B, then needed to stretch its context window and sharpen its reasoning for client work involving long documents. Rather than scrape the open web and hope for the best, engineering manager Martin Cimmino turned to Ai2's Dolma dataset, which comes with documented sourcing and filtering, to extend training in a way Domyn could defend to auditors. Layering in Dolci, Ai2's 260,000-pair instruction dataset released alongside Olmo 3 last year, then produced a 10.1-point jump on GPQA-Diamond — the single biggest gain anywhere in Domyn's post-training pipeline.
What both companies are really buying isn't performance, it's paperwork that holds up. The EU AI Act now requires general-purpose model providers to disclose training data summaries, and U.S. federal contracts carry their own provenance demands that most commercial labs would rather not answer. Ai2 publishing the entire stack — weights, code, data, and the blueprints connecting them — means Domyn and AISquared can hand over documentation instead of vague promises when a compliance officer starts asking questions.
None of this makes Olmo the most powerful model on the market, and nobody involved is claiming that. But for labs selling into finance, government, and healthcare, capability was never going to be the whole pitch. Being able to show your work turns out to matter just as much, and right now Ai2 is one of the few outfits building models where that's actually possible.
My take
This is the quiet argument for open-weight models that hype cycles keep burying under benchmark chest-thumping: transparency isn't a nice-to-have, it's the actual product for anyone selling into regulated markets. Watching a Milan startup and a D.C. federal contractor both land on the same open dataset for compliance reasons should worry the closed-model labs more than another leaderboard loss ever will, because procurement officers, not Twitter, decide who gets paid.
Read more about this at: Allen Institute (AI2)