A $40 million Series A. Led by a16z. For an AI evaluation tool.
At first glance, this is just another infrastructure investment in the AI hype cycle. But as a quant trader who has survived the 2017 ICO carnage, the 2020 DeFi yield farming wars, and the 2022 Terra-Luna collapse, I see a different story.
This is the birth of the smart contract audit—for AI agents.
History is just data waiting to be backtested.
Context: The Vals AI Play
Vals AI builds evaluation tools for AI models. The company just raised $40M from a16z, one of the most influential VCs in crypto and AI. The narrative: reliable AI evaluation is critical for enterprise adoption.
But the crypto market is already integrating AI agents into DeFi, governance, and trading. Uniswap V4 hooks are programmable. Layer2s are fragmented. Liquidity is already thin. Now, imagine deploying an AI model that makes trading decisions without rigorous testing.
That's a disaster waiting to happen.
From my experience auditing ICO smart contracts in 2017, I know the cost of skipping verification. An integer overflow wiped out millions. The same applies to AI models. The evaluation tool is the new audit.
Core: The Evaluation Infrastructure
The technology behind Vals AI is not about training models. It's about building a repeatable, quantifiable framework to test model behavior. This is the same discipline I use in my trading bots: backtest, stress test, then deploy.
But here's the catch—most AI evaluation tools today rely on static benchmarks like MMLU. These are the equivalent of backtesting on historical data without considering regime change. In 2020, I learned that theoretical yields from DeFi farming were offset by hidden transaction costs and impermanent loss. The same applies to AI evaluation. The tool must test for adversarial inputs, drift, and edge cases.
Based on my experience building quantitative models, I know that the real value is in the scenario design, not the algorithm. Vals AI's product likely includes dataset construction, evaluation orchestration, and automated judgment. But without transparency, it's just another black box.
History is just data waiting to be backtested—but only if the backtest is honest.
Contrarian: The False Sense of Security
The market sees a16z's investment as a validation of the evaluation tool category. I see a trap.
Evaluation tools are double-edged. They can create a false sense of security. In 2022, Terra's algorithmic stablecoin was evaluated as safe by many metrics—until it wasn't. The death spiral mechanism was an inherent flaw, but the evaluation framework didn't capture it.
The same risk applies to AI evaluation. If the tool only tests for known failure modes, it will miss the unknown unknowns. The industry needs "red teaming" that goes beyond surface-level checks. But Vals AI's product details are not public.
From my perspective, the smart money is not betting on the tool itself. It's betting on the narrative: that every AI agent in crypto will need an audit. But the auditors themselves need to be audited.
Who evaluates the evaluator?
During the 2020 DeFi summer, I saw protocols with flashy dashboards and high APYs attract liquidity, only to collapse when the code broke. The same pattern will repeat with AI agents. The evaluation tool is a necessary condition for safety, but not sufficient.
Takeaway: The Verdict for Crypto Traders
Vals AI's $40M raise is a signal that the market is waking up to the need for AI safety. But as a trader, I'm not buying the hype. I'm watching the data.
If evaluation tools become standard for crypto AI agents, they will reduce tail risk. But the real alpha is in understanding the limitations of these tools. Treat them as a starting point, not a guarantee.
History is just data waiting to be backtested. The question is: who designs the test?
Until Vals AI publishes its methodology, I'll keep my capital in cold storage and my models in a sandbox. The market will reward those who survive the next black swan—not those who trust an unverified evaluation tool.