← Back to blog

AI Hallucination Gets All the Blame. Human Error and Selective Evidence Get a Pass.

We asked an auditor directly: do you trust AI-generated evidence more or less than human-generated evidence? The answer was less. They dig deeper when AI is involved, probe harder, and approach it with a level of skepticism they do not apply to evidence a human produced.

That answer is not grounded in data. It is grounded in instinct. And that gap, between how the industry perceives AI risk and where the actual risk tends to live, is worth talking about.

The Risk Nobody Keeps Score On

Non-conformities exist because we expect humans to make mistakes. That is not a criticism; it is just how compliance works. You build procedures around human fallibility, you log the failures, you correct them, and you move on. The whole framework assumes non-conformity is expected.

Human error's silent risk, AI's visible flaws.
Human error's silent risk, AI's visible flaws.

So why does AI entering the picture suddenly shift the standard?

If a human produces a compliance record with an error in it, that is a non-conformity. It gets documented, a corrective action gets opened, and the process continues. If AI produces something with an error, the conversation changes entirely. Suddenly it is a hallucination, a systemic trust problem, a reason to probe the entire body of evidence more deeply.

The double standard is real and observable. We have built an entire procedural apparatus around the assumption that humans get things wrong, and then applied a completely different bar to AI without asking whether that bar is justified.

Hallucination is not the only failure mode worth tracking. Selective evidence, where a compliance team surfaces what looks good and buries what does not, and plain human error, where someone makes a mistake that never gets flagged, are both more likely to slip through undetected. Hallucinations are at least visible. You can catch them. A human who selectively presents evidence to an auditor is much harder to surface, and the auditor may never know what they did not see.

Misapplication Is the Real Problem

A lot of what gets called an AI hallucination problem is actually a deployment problem. Cramming too much into a single context window. Choosing the wrong model for the task. Applying AI to things that genuinely require careful, considered human reasoning, and then being surprised when the output does not hold up.

Misapplication, not AI, is the real problem.
Misapplication, not AI, is the real problem.

AI does some things well. It does other things badly. The industry is still figuring out where that line sits, and that is fine. But conflating misapplication with fundamental untrustworthiness is a different argument, and it is one that is being made without much evidence behind it.

The auditor who told us they trust AI evidence less did so based on their experience and intuition. That is a legitimate starting point. It is not a conclusion. Intuition shaped by anecdote is exactly the kind of evidence that compliance frameworks are supposed to interrogate, not rely on.

There is something worth sitting with in that irony. The industry whose entire function is evidence-based judgment is applying a lower evidentiary standard to its own skepticism of AI than it applies to the systems it audits.

Show Us the Numbers

Here is the challenge we want to put to the compliance and audit community: produce the data.

Show us the data: AI vs. human compliance evidence.
Show us the data: AI vs. human compliance evidence.

If AI-generated evidence is meaningfully less reliable than human-generated evidence, that should be measurable. What is the error rate of AI-assisted compliance documentation compared to human-produced documentation across the same controls? How do those numbers move when you account for model selection, context design, and review processes? How does selective evidence, a human choice, not a model failure, factor into audit outcomes?

Has anyone published that research in a rigorous way? The skepticism exists. The data does not.

We are not arguing that AI should get a free pass. We are arguing that the current standard is being applied unevenly, based on anecdote rather than evidence, and that the risks getting the least attention, selective human presentation and plain human error, are the ones most likely to cause an actual audit failure.

If you work in compliance or audit and you have the data, we want to see it. Not the intuition. The numbers.