TLDRocket
Sign in

a16z leads $40M Vals AI round at $400M valuation to test AI on real-world tasks

Tech Funding News Abhinaya Prabhu

Vals AI raised $40M at a $400M valuation to judge AI on real jobs, not toy tests. That matters because the usual benchmarks are getting gamed, leaked, and beaten for show.

Based on reporting by Tech Funding News, Abhinaya Prabhu — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Vals AI has closed a $40 million Series A at a $400 million valuation, with Andreessen Horowitz leading the deal. Existing investors 8VC, Pear VC, and Bloomberg Beta came back in, while HRT Ventures and Next Ladder Ventures joined the round.

The pitch is simple and a little brutal: frontier models keep claiming to be the best, and the public tests used to rank them are running out of credibility. Academic benchmarks worked for a while, but once everyone started optimizing for the same lists, the lists stopped meaning much. Models can ace a leaderboard and still stumble when the task is messy, multi-step, and closer to what a customer actually pays for.

Vals was built to measure that gap. Rayan Krishnan and Langston Nashold, who studied computer science together at Stanford, pair domain experts in law, finance, healthcare, and coding with automated grading systems that score model output to professional standards. The company keeps its test sets private, limits how often they’re run, and says it can turn around results within hours of getting a new model.

That last part matters because the tests themselves are now part of the race. Vals says it retires benchmarks once they stop separating strong models from weak ones; in May it replaced a corporate-finance benchmark called CorpFin with a new Excel-modeling test. Its evaluations have already shown up in model cards from OpenAI, Anthropic, Google, Meta, and xAI, and enterprises are using the scores to decide what gets deployed.

The company says revenue is up eightfold across 2025, the customer base has doubled, and the team has tripled in six months. With the new funding, Vals is expanding its infrastructure and rolling out Vals Smith for coding benchmarks built from customers’ own GitHub repositories, along with frontier-risk benchmarks for cybersecurity, mental health, and AI safety. It also launched Vals Index 2.0, which pushes its measurement system beyond model tests and toward the broader economy.

My take — AI-written commentary, not fact-checked reporting

This is the less sexy but more important AI story: whoever controls the scorecard quietly controls the market. Public benchmarks were always going to break the moment the incentives got serious, and now the adults are being paid to rebuild the referee instead of pretending the old scoreboard still works. Good. The industry needs fewer victory laps and more auditors.

Read more about this at: Tech Funding News

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.