AI benchmarking is, at this point, a marketing exercise as much as a technical one. Companies know how the tests work, they train toward them, and the numbers come out looking great. So what does a benchmark actually tell you? Vals, a San Francisco startup, thinks the answer is: not much. And according to TechCrunch, the company just raised $40 million in a Series A led by Andreessen Horowitz to prove there’s a better way.
Founded in 2024 by 25-year-old Rayan Krishnan, a Stanford grad with stints at Palantir and Microsoft on his resume, Vals was built around a specific frustration. Academic benchmarks, many of them years old, were not keeping pace with the speed at which new models were reaching the market. The tests were measuring abstract intelligence, things like whether a model could pass a bar exam, rather than whether a model could actually do useful work in the real world. That gap is what Vals is trying to close.
The core difference is in how Vals structures its evaluations. Most public benchmarks publish their test materials openly, which creates an obvious problem: labs can train models directly against those tests. Vals keeps its specific test content private. And instead of general knowledge checks, it measures how well models complete complex, domain-specific tasks in areas like law, finance, and software development. The goal is closer to a real performance review than a standardized test.
Vals also goes a step further. Krishnan says the company is actively working on evaluations in areas like recursive self-improvement, cybersecurity, biosecurity, mental health, and even how models interpret the Geneva Convention in conflict scenarios. That breadth puts Vals in a different category from players like Scale AI’s evaluation tools or internal red-teaming efforts at labs like Anthropic and OpenAI. This is systematic, third-party testing with a commercial model behind it.
The commercial angle is worth understanding. Companies pay Vals to evaluate their models, similar to how students pay to sit the SAT. The value isn’t just knowing where you rank. It’s knowing what to fix. Revenue is currently eight times what it was a year ago, the team has grown from eight to 25 people, and a federal agency program has recently launched.
Krishnan sees the timing as significant. As AI companies like Anthropic and OpenAI move toward public markets, third-party evaluation data becomes something regulators, investors, and enterprise buyers will increasingly demand. Independent benchmarking may shift from optional to expected. That’s a large surface area for Vals to grow into, if the evaluations hold up under scrutiny.



