Sep 16, 2026 Sina Schäfer
ShareGenerative AI can be remarkably convincing. Its language is polished, its answers sound confident, and its knowledge can seem almost limitless. That confidence, however, can make mistakes especially difficult to spot. AI hallucinations may sound plausible and technically sound even when they are factually wrong. Humans, meanwhile, are just as vulnerable to errors in judgment shaped by experience, expectations, and intuition. Problems arise when these two sources of error reinforce each other and influence business decisions without being challenged. For companies, the key question is how to design AI-supported decision-making processes so that human and machine errors counterbalance rather than amplify one another.
75 percent. That was the hallucination rate of OpenAI’s o4-mini in the SimpleQA benchmark when the model was introduced in April 2025. Although it was designed as a reasoning model for more complex tasks and step-by-step problem solving, it still provided an answer in 99 percent of cases. It declined to answer in only one percent, even when the information it produced was not reliable. By August 2025, the picture had changed. In the same benchmark, GPT-5 Thinking Mini reduced its hallucination rate to 26 percent and declined to answer in 52 percent of cases. The improvement was not only about getting more answers right. It was also about becoming better at recognizing when a reliable answer was not possible.
When Plausibility Becomes a Trap
People tend to question information less when it sounds plausible and fits what they already expect to be true. The same applies to AI-generated responses. If an answer confirms an assumption a user already holds, it can quickly be interpreted as validation rather than something that still needs to be checked. The AI provides a plausible but incorrect answer, while the human sees little reason to question it because it matches their expectations.
Machine error and human misjudgment can therefore quickly reinforce one another.
This risk is closely tied to how large language models work. LLMs are not databases of verified knowledge. They are statistical models trained to generate language. Unless they are connected to trusted external sources, they do not retrieve facts in the same way a database does. Instead, they generate answers based on patterns and probabilities. When information is missing, ambiguous, or contradictory, a model may fill in the gap with the most statistically likely continuation. The result can sound perfectly coherent while still being factually wrong.
A well-known example came in 2023, when Google’s chatbot Bard incorrectly claimed that the James Webb Space Telescope had taken the first images of an exoplanet, even though such images had already been captured in 2004. The statement sounded convincing, but it was wrong. In the aftermath, the market value of Google parent company Alphabet temporarily dropped by around $100 billion.
How Experience and Expectations Shape Judgment
AI may hallucinate for statistical reasons, but human errors follow their own patterns. Human judgment is shaped by experience, emotions, expectations, and intuition. We are naturally wired to recognize patterns and make sense of information quickly. That ability is essential in social situations and intuitive decision-making, but it also makes us vulnerable to systematic biases.
One of the best-known examples is confirmation bias: we are more likely to accept information when it supports what we already believe. If an AI confirms an existing assumption, its response may receive less scrutiny than it otherwise would.
Another related phenomenon is confabulation. Humans, too, can unconsciously fill gaps in their knowledge with details that seem logical or consistent. We complete the picture in a way that feels coherent, even when parts of it may be inaccurate.
It becomes especially problematic when human and machine errors validate each other. An AI produces a plausible but false statement, and an employee accepts it because it fits their expectations. The reverse can happen as well: a correct AI-generated answer may be rejected simply because it conflicts with the human decision-maker’s existing view. Once such errors feed into a forecast, a procurement decision, or a management briefing, a small mistake can quickly turn into a business problem.
Four Ways to Make AI Decisions More Reliable
So how can companies design processes in which human and machine errors challenge rather than reinforce each other? Four principles can help keep these risks in check:
- Ground answers in reliable sources: In fact-sensitive applications, users should be able to understand what an answer is based on and where uncertainty remains.
- Match the level of review to the potential impact: The greater the possible business impact, the more rigorous the review and approval process should be.
- Build in challenge and alternative perspectives: Important decisions should deliberately allow for competing explanations, second opinions, or additional expert review.
- Define clear accountability: It should always be clear who is responsible for reviewing the output, questioning assumptions, and ultimately making the decision.
Less Room for Plausible Mistakes
These principles will not make AI error-free. But they can reduce the likelihood that errors go unnoticed and influence decisions. Reliable sources leave less room for unsupported but plausible additions. Structured review and deliberate challenge increase the chances that errors are caught before they have a real business impact.
Neither humans nor generative AI can realistically be expected to operate without error.
What matters is not which side makes fewer mistakes, but how the decision-making process deals with the fact that mistakes will happen.
In this context, human-in-the-loop means designing the interaction between people and AI in a way that encourages questionable judgments to be challenged and corrected as early as possible. That requires reliable guardrails, transparent processes, and clearly defined responsibilities. The next phase of AI adoption will therefore be less about generating the most impressive answers and more about producing dependable ones, while maintaining a clear understanding of the limitations on both sides.
About our Expert

Sina Schäfer
Corporate Communications Manager
Sina Schäfer has been working as Corporate Communications Manager in Corporate Marketing at INFORM since 2021. Her focus is on external communication in the areas of inventory & supply chain, production and industrial logistics.
