Passing Your Evals Doesn’t Mean You’re Safe
Evals are part of every serious conversation about putting AI into production. Teams define benchmarks, set thresholds, and increasingly run red teams to see how the system holds up against someone actively trying to break it. That combination is reasonably good at telling you whether a model is accurate, reliable, fast enough for production, and resistant to an adversarial attack. It says almost nothing about what actually gets a company into trouble once the system is live.
If this newsletter is worth your time, consider becoming a paid supporter 🙏
Look at the AI lawsuits and regulatory findings piling up this year and the gap becomes obvious. A man in California is suing OpenAI, arguing that ChatGPT used his own disclosed bipolar diagnosis to keep him engaged in conversation rather than steering him toward help, a product liability claim built entirely around a session memory feature. A federal court let a discrimination suit against Workday proceed on the theory that its hiring software leaned on medical leave and other proxy signals to screen out an older applicant with a disability, exposing the vendor itself, not just the employer who bought the tool. A German court ruled that a company is responsible for its own chatbot inventing a doctor’s medical credentials, reasoning that a chatbot isn’t a third party, it’s simply the business talking. None of these are the kind of failure a red team session or a bias checklist is built to catch. They’re everyday production risk, legal, reputational, and regulatory, sitting inside systems that had already passed whatever evals their teams ran.

The High-Dimensional Fix
I recently sat down with Andrew Burt, CEO of Luminos, about the whitepaper his team just published on agentic and generative AI risk evals. The report’s core argument is that the standard setup, one broad prompt, one model, a pass or fail score, is too coarse to do its job. Ask a model whether an output is “biased” or “manipulative” and you get a shallow answer back. Real problems slip through, harmless outputs get flagged as risky, and either way you’re left without enough detail to know what to actually fix.
Luminos calls its answer high dimensionality, and the idea is more concrete than the name suggests. Take an ad that shouldn’t manipulate the people looking at it. A typical eval asks one broad question, is this manipulative? A high-dimensional eval tests separately for a fake countdown timer, a hidden renewal fee, a bait-and-switch price, and a handful of other specific tactics, then comes back with something like: false urgency (yes), hidden fee (yes), bait and switch (no). That’s a finding a product team can actually act on. The same decomposition has to run across more than one model, because models have real personalities when it comes to flagging risk. Claude tends to run conservative and overflag, other models underflag, and a prompt that works cleanly on one model can fall apart entirely on another. Underneath all of it, the standards need to come from actual legal, privacy, and compliance expertise, not just engineers guessing at what a regulator will care about.

None of these three ideas is new by itself. Plenty of teams already run more than one model, and some already pull in legal or compliance reviewers. What’s rare is doing all three together, and treating the result as an ongoing job rather than a pre-launch checklist, since models get updated, laws vary by jurisdiction, and a test suite that felt thorough in January can be stale by summer.
It’s worth being clear about what the whitepaper is and isn’t: it doesn’t present a controlled study proving this approach beats simple evals by some specific margin. The case is structural, built from Luminos’s own experience running this at scale, but it’s a pretty convincing one. Testing more of the actual risk surface should catch more real problems, calibrated sub-risks should cut down on false alarms, and requiring an explanation for every flag should make results something a team can actually audit and fix.

Treat AI Risk Like Reliability
The pattern across this year’s AI incidents isn’t that companies skipped evals entirely, most had something in place. This isn’t an argument for ditching red teaming, guardrails, or standard performance evals either, they answer different questions and are still worth running. The mistake is assuming that together they add up to a full picture of production risk.

Once an AI system can influence a customer, an employee, an account, or a regulated decision, its risk testing deserves the same rigor teams already give to reliability and uptime. That means specific tests, evidence you can point to, clear ownership, and monitoring that continues after launch, not a box checked once before shipping. That bar covers more systems than it sounds like, a support chatbot, a hiring tool, an internal scoring system, an agent with write access to anything, almost any AI system doing real work in a business. There’s no comfortable middle category of “important enough to ship, not important enough to evaluate properly.”
The Luminos whitepaper is worth reading in full, not just skimming, whether or not you end up building this exact approach yourself.
AMD’s Long Game in AI

Two AI Labs Carry More of the Cloud Boom Than You Think

