Skip to content

TheLLM Brief

← All stories

Industry

Vals Raises $40M to Fix Broken AI Benchmarking

Vals Raises $40M to Fix Broken AI Benchmarking
Image: TechCrunch

Andreessen Horowitz leads the Series A for the 2024-founded startup.

Sourced from TechCrunchBy Lucas Ropek

Vals AI closed a $40 million Series A led by Andreessen Horowitz, TechCrunch reports. The 2024-founded startup targets a specific failure in the AI stack: benchmarks that no longer measure what modern models actually do.

The core problem is structural. Legacy benchmarking systems were not built for current model capabilities, and companies have learned to optimize for the metrics rather than the outcomes. Good benchmark scores became a PR tool, not a signal of real performance. Vals was founded to close that gap.

Co-founder Rayan Krishnan, 25, came out of Palantir, Microsoft, and Stanford's AI lab. The company previously raised a seed round led by 8VC and Bloomberg Beta. Watch whether enterprise buyers start requiring Vals scores the way they once required SOC 2 reports. The signal is not the funding. The signal is whether procurement teams make independent benchmarking a contract condition.

Analysis

Benchmarks are where capability claims get tested by someone other than the lab. If Vals captures that role, the buyer gains leverage and the model vendor loses the ability to grade its own homework.

Research this with your AI

Copy the research prompt into your AI assistant to see how this story affects you.

Then paste it into ChatGPT, Claude, Gemini, Grok and others.
Runs in your own assistant with your own context. Nothing is sent to us.
Show the prompt
I just read this AI news story and want to understand it in my own context.

Title: Vals Raises $40M to Fix Broken AI Benchmarking
Summary: Vals AI raised $40 million in a Series A led by Andreessen Horowitz. The startup, founded in 2024, aims to replace legacy benchmarks that AI companies have learned to game.
Category: Industry
Source: TechCrunch, https://techcrunch.com/2026/09/19/vals-backed-by-andreessen-horowitz-is-looking-to-become-the-gold-standard-for-ai-benchmarking/

Using my own history and context, help me understand:
1. What is the core development and why does it matter?
2. Who are the major players involved and what are their motivations?
3. How does this fit into the broader AI landscape right now?
4. How does this apply to my own work, and what should I do or watch next?

Be specific and plain spoken.

Newsletter

The day's AI stories, with the editor's take, in one email.

Free. Unsubscribe in one click.