Vals AI: The $40M Bet on AI Evaluation as Blockchain’s Next Infrastructure Layer
CryptoBen
The ledger was clean, but the vision was fragile. When a16z led a $40 million Series A for Vals AI—a tool promising to audit AI models—the crypto-native crowd barely blinked. Yet beneath the surface, this funding signals a structural shift: AI evaluation is becoming the new smart contract audit. In a bull market where every DeFi protocol rushes to integrate LLM agents, the ability to verify model behavior programmatically is no longer optional. It is the gatekeeper of trust.
Context: The AI Evaluation Land Grab
Vals AI operates in the crowded AI evaluation space, a sector that has seen explosive growth since 2024. The company’s core thesis is simple: as enterprises deploy large language models into production, they need a standardized, auditable way to measure performance, safety, and alignment. This mirrors the evolution of smart contract auditing in 2018—when the first DeFi hacks exposed the fragility of unaudited code. Today, AI agents execute transactions, manage loans, and interact with on-chain protocols. A single hallucination can drain a liquidity pool. Vals AI positions itself as the auditor for this new frontier.
But the market is already saturated. Competitors like LangSmith, Galileo, and Arthur AI offer similar tooling. The difference? Vals AI’s $40 million raise—led by a16z, a firm with deep crypto roots—suggests a strategic bet on the intersection of AI and blockchain. In my 2020 DeFi arbitrage days, I learned that infrastructure investments often precede the next wave of innovation. The question is whether Vals AI can deliver what the market demands: a battle-tested evaluation framework that works across both centralized and decentralized environments.
Core: The Technical Architecture of Trust
Based on my experience auditing Power Ledger’s smart contracts in 2018, I know that trust is built on verifiable logic, not promises. Vals AI’s approach follows a similar pattern. Their evaluation tool likely relies on a “LLM-as-Judge” architecture, where a frontier model (e.g., GPT-4o, Claude) grades outputs of other models. This creates a feedback loop—but who audits the judge? This is the same problem that plagued early DeFi oracles. Without a transparent, decentralized verification mechanism, the evaluation layer itself becomes a single point of failure.
On-chain, the challenge is more acute. Consider a lending protocol that uses an AI agent to assess collateral risk. If the agent’s evaluation is flawed, the entire pool is at risk. Vals AI must provide not just off-chain metrics, but on-chain attestations that can be verified by smart contracts. This requires a hybrid approach: off-chain computation for heavy inference, on-chain settlement for immutability. The team’s technical background—likely heavy in MLOps and zk-proofs—will determine if they can bridge this gap.
I recall a pattern from 2021: every NFT project claimed to be “alpha” until the wash-trading bots exposed the illusion. Similarly, many AI evaluation tools today claim to catch hallucinations, but their methodologies are opaque. Code does not lie, but people certainly do. Vals AI’s real innovation might be in making evaluation results auditable via cryptographic proofs—a move that would resonate with the Web3 community.
Contrarian: The Simulation Trap
The conventional wisdom is that better evaluation tools will lead to safer AI. But I see a darker parallel to the Terra/Luna collapse. In 2022, algorithmic stablecoins were touted as “self-auditing” until the death spiral proved otherwise. Evaluation tools can create a false sense of security. If a model passes Vals AI’s benchmarks, developers might assume it is safe for production, ignoring edge cases that the tests didn’t cover. This is the “audit theater” problem—the same reason why many smart contract audits miss critical vulnerabilities.
Moreover, the competitive landscape is brutal. OpenAI and Anthropic are building their own evaluation frameworks, which will be deeply integrated into their APIs. Why would a developer pay for a third-party evaluator when the model provider offers a free one? The answer lies in independence. In crypto, we learned that trusting a single oracle is dangerous. Decentralized evaluation networks, where multiple assessors vote on model quality, could be the solution. But Vals AI is centralized—a single point of failure. The contrarian bet is that this centralization will be its undoing in a market that values trust minimization.
I also question the timing. Bull market euphoria often masks technical flaws. During the 2024 ETF approval frenzy, I saw institutional investors pour money into projects with weak fundamentals. Vals AI’s $40 million raise could be a similar case: a16z is betting on a narrative, not necessarily on a proven product. The company needs to deliver real revenue and customer traction before the next bear market reveals the fragility.
Takeaway: The Alpha Hides in the Noise
We bet on the pattern, not the hype. The pattern here is clear: AI evaluation is becoming a critical infrastructure layer, much like smart contract auditing was for DeFi. Vals AI has the funding and the pedigree to capture a slice of this market. But the edge lies in technical depth—specifically, the ability to produce on-chain verifiable evaluations that resist manipulation. If they can do that, they will be the Trail of Bits of the AI era. If not, they will be another footnote in the a16z portfolio. The chart doesn’t lie, but the narrative does. Watch the code, not the press release.