a16z leads $40M Vals AI round at $400M valuation to test AI on real-world tasks
Tech Funding News Abhinaya Prabhu
Vals AI raised $40M at a $400M valuation to judge AI on real jobs, not toy tests. That matters because the usual benchmarks are getting gamed, leaked, and beaten for show.
Based on reporting by Tech Funding News, Abhinaya Prabhu — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Vals AI has closed a $40 million Series A at a $400 million valuation, with Andreessen Horowitz leading the deal. Existing investors 8VC, Pear VC, and Bloomberg Beta came back in, while HRT Ventures and Next Ladder Ventures joined the round.
The pitch is simple and a little brutal: frontier models keep claiming to be the best, and the public tests used to rank them are running out of credibility. Academic benchmarks worked for a while, but once everyone started optimizing for the same lists, the lists stopped meaning much. Models can ace a leaderboard and still stumble when the task is messy, multi-step, and closer to what a customer actually pays for.
Vals was built to measure that gap. Rayan Krishnan and Langston Nashold, who studied computer science together at Stanford, pair domain experts in law, finance, healthcare, and coding with automated grading systems that score model output to professional standards. The company keeps its test sets private, limits how often they’re run, and says it can turn around results within hours of getting a new model.
That last part matters because the tests themselves are now part of the race. Vals says it retires benchmarks once they stop separating strong models from weak ones; in May it replaced a corporate-finance benchmark called CorpFin with a new Excel-modeling test. Its evaluations have already shown up in model cards from OpenAI, Anthropic, Google, Meta, and xAI, and enterprises are using the scores to decide what gets deployed.
The company says revenue is up eightfold across 2025, the customer base has doubled, and the team has tripled in six months. With the new funding, Vals is expanding its infrastructure and rolling out Vals Smith for coding benchmarks built from customers’ own GitHub repositories, along with frontier-risk benchmarks for cybersecurity, mental health, and AI safety. It also launched Vals Index 2.0, which pushes its measurement system beyond model tests and toward the broader economy.
My take — AI-written commentary, not fact-checked reporting
This is the less sexy but more important AI story: whoever controls the scorecard quietly controls the market. Public benchmarks were always going to break the moment the incentives got serious, and now the adults are being paid to rebuild the referee instead of pretending the old scoreboard still works. Good. The industry needs fewer victory laps and more auditors.
Read more about this at: Tech Funding News