A Clean Failure Record Is a Red Flag, Not a Green One

Companies that build a robust context layer are more likely to report AI agent failures than firms that don’t.

I have to say I was surprised to see this, but I shouldn’t have been.

AI agents confidently delivering incorrect answers is a known problem. My experience in practice, and with the Kendall Framework, is that rigorously curating appropriate context and feeding it into AI models greatly decreases hallucination and other errors.

So why are the reported rates of errors climbing?

It’s simple really. Those companies that have first defined the appropriate context, and are managing it, are surfacing the errors better. It turns out that a low rate of agent errors doesn’t mean those errors aren’t happening. It means they’re going unreported. Suddenly, it all clicks into place.

That’s the finding from VentureBeat’s July 2026 VB Pulse survey of enterprise AI teams. The full numbers are worth a look. What they point to matters more than any single stat. Governance has risen to the top of buyer selection criteria, but what firms are selecting for and what they’re actually measuring are different things.

Every AI agent needs some way to know what the business actually means. A robust context layer gives every agent, and every person, a shared, agreed-on definition of what the business’s data means: what counts as a sale, which customer record is current, which number is the real one. That shared reference point is what lets someone catch a wrong answer and trace it back to where it broke.

Take that reference point away, and the wrong answer is still sitting there. Nobody’s tracking it, so nobody’s counting it.

A client story that shows the gap

I saw a version of this with a client recently. The company had two systems tracking sales, each on a different platform. The trouble was each system defined “sale” its own way, different from the other. Most importantly, both differed from how the company itself defined sales. Not that the discrepancy was hard to track down. It did take real effort to enforce one common, companywide definition across two different reporting platforms.

So when a company tells you their AI systems haven’t produced a wrong answer, don’t take too much comfort in it. It’s just as likely nobody’s checking.

We’ve seen this before

Kyle Nesbit, founder of the semantic layer startup Credible Data, has watched this play out for years, long before AI entered the picture. “It’s the same pain point people have had for 30 years, the lack of governed data analysis,” he told VentureBeat. “Now with AI, it’s the same problem, but orders of magnitude more chaos and pain.” I’ve seen the same issue play out in BI systems as well. The pace has changed, though.

Part of why these errors go unnoticed is that they don’t look wrong. Srijith Rajamohan, an AI research leader at Redis, put it this way in an interview with VentureBeat: embedding retrieval can’t tell “Rome is closer than Paris” from “Paris is closer than Rome,” because the words match even when the meaning is reversed. A confidently wrong answer usually sounds exactly like a right one. Someone has to be looking for the difference on purpose.

One of the things I appreciate about the Kendall Framework is that it provides a structure that makes it easier to surface and identify these kinds of context errors. It won’t solve them all for you. That’s for your team to work out. But knowing there’s a context discrepancy is better than sweeping it under the rug.