TLDRocket
Sign in

Exclusive: Synthefy raises $6.5M for its number-crunching models trained on numerical data instead of words

SiliconANGLE Mike Wheatley

Synthefy just raised $6.5M to build AI models trained on numbers instead of words. Its tiny open-source model reportedly beat Google's much bigger one at crunching tables and time-series data.

Based on reporting by SiliconANGLE, Mike Wheatley — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

There's a version of the AI boom that everyone already knows: massive language models trained on oceans of text, learning to predict the next word. Synthefy Inc. thinks there's a parallel opportunity nobody's built properly yet — the same trick, but for numbers. The startup announced a $6.5 million seed round today, led by Wing Venture Capital with Haystack, Samsung Next, Canonical Crypto and Lightscape joining in, plus angel money from people at OpenAI, Microsoft and Meta.

The pitch is what Synthefy calls Structured Data Foundation Models, or SDFMs. Instead of ingesting text, these models train on tables and time-series data, learning the relationships buried in that structure so they can generalize to new numerical problems rather than starting from scratch each time. The company's first open release, a lightweight model called Nori, quietly went live a few weeks back. A 30-million parameter version reportedly outperformed Google's TabFM model, which has 1.6 billion parameters — and with its "Thinking" mode switched on, Nori apparently does even better, despite running at roughly 2% of TabFM's size.

What makes this interesting isn't just the benchmark bragging rights. It's the workflow problem Synthefy claims to be solving. Fraud detection and dynamic pricing already run on machine learning frameworks like LightGBM and XGBoost, but CEO Somi Agarwal says most enterprises burn weeks preparing data and tuning models for each new problem, and none of that effort carries forward. Point Nori at a fresh table, he argues, and it can produce a solid prediction without a fresh training cycle — turning a weeks-long evaluation into something that takes minutes, because the model already learned from millions of synthetic datasets before it ever saw the customer's data.

The business model looks familiar from the LLM playbook: give away the open foundation layer, then sell the enterprise trimmings. Agarwal is eyeing managed API access, private deployments, security and governance tooling, and large-scale production infrastructure as the actual revenue engine. Early traction backs up the interest — Nori has racked up more than 600,000 downloads in just a few weeks. Wing's Gaurav Garg is betting that structured data models become the next real expansion of the AI market, calling out Synthefy's mix of performance, efficiency and open models as a shot at defining the category before anyone else claims it.

What happens next depends on whether enterprises actually swap out battle-tested tools like XGBoost for a foundation-model approach they haven't lived with for years. Synthefy says the funding goes toward research, hiring engineers and building the next Nori, along with chasing industry partnerships. The download numbers suggest curiosity is already there — the harder test is whether that curiosity turns into production deployments handling real fraud and pricing decisions.

My take — AI-written commentary, not fact-checked reporting

A 30-million parameter model beating a 1.6-billion parameter one is the kind of headline number that begs for independent scrutiny before anyone crowns a new category. Enterprises don't switch away from XGBoost because a startup's benchmark looks good in a press release; they switch when the thing survives a messy production dataset with missing values and drifting distributions. The open-source download count is a genuinely good signal of curiosity, though — that part's real, and worth watching.

Read more about this at: SiliconANGLE

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.