AI doesn't need to hallucinate to mislead us
The more subtle risk of generative AI: a plausible, often repeated but scientifically shaky conclusion, convincingly reproduced.

The more subtle risk of generative AI: a plausible, often repeated but scientifically shaky conclusion, convincingly reproduced.
When we talk about the risks of generative AI, the conversation quickly turns to hallucinations: answers in which an AI model invents facts, sources, or events. But there is a more subtle risk. AI doesn't have to invent anything at all to still mislead us. It can also reproduce a plausible, frequently repeated, but scientifically shaky conclusion in a particularly convincing way. That problem becomes more significant as AI gets better.
Generational Conflict
This holiday, I read Generatiestrijd by Paul de Beer. In the book, he critically examines how robustly many popular statements about generational differences are actually supported statistically. We all know the statements. Generation Z values freedom more than older generations. Millennials view work and career differently. Baby boomers are more loyal to their employers. Younger generations value meaning over salary. They sound plausible, and sometimes studies even show clear differences between age groups. But that doesn't tell us what causes those differences.
Take, for example, the observation that 25-year-olds value freedom more than 60-year-olds. The temptation is strong to immediately turn this into a generational difference. After all, the 25-year-old belongs to a different generation than the 60-year-old. But that conclusion does not automatically follow from the observation. It could be an age effect, because people start thinking and acting differently as they get older. It could be a period effect, where economic, social, or technological circumstances influence everyone at a certain point. Or there could genuinely be a cohort effect, where people born in a particular period systematically differ from people from other birth periods.
A well-known statistical problem
This is precisely where a well-known statistical problem arises. Age, measurement year, and birth year are inextricably linked. Age is measurement year minus birth year. If you know someone's age and the year of the study, you automatically know their birth year. As a result, you cannot simply separate age effects, period effects, and cohort effects with more data. Even a more advanced model does not automatically solve this fundamental problem. Additional assumptions are needed to draw conclusions.
This highlights something fundamental about how we approach research. Data does not speak for itself. A statistical correlation is an observation, and the explanation of that correlation is an inference. There is often much more uncertainty between the two than the final conclusion suggests. For example, we can measure that two groups differ from each other. But as soon as we say why they differ, we take a further step. This requires theory, research design, and assumptions.
Why this affects AI
This is precisely where generative AI becomes interesting. A language model is exceptionally good at producing plausible language. Ask it about the differences between Generation Z and millennials, and it can produce a convincing overview in seconds, including characteristics, explanations, and practical advice for employers. The model doesn't need to hallucinate to do this. It can simply reproduce what has been written thousands of times on websites, in articles, presentations, and management books. Yet the conclusion can be far less scientifically robust than its phrasing suggests.
Persuasiveness and evidential power are two different things.
This makes this risk more insidious than a classic hallucination. An invented source can be checked, and a wrong date can be corrected. But a frequently repeated line of reasoning that sounds logical and is partly supported by research demands something different from the user. We must not only ask if something is correct, but especially what conclusion the evidence actually permits. Is something an observation or already an explanation? What alternative explanations are possible? What assumptions are needed to get from the data to the conclusion? And has the conclusion truly been proven, or is it mainly just frequently repeated?
Paradoxically, this problem grows as AI gets better. A poor language model is more easily exposed because it makes strange errors or gives illogical answers. A good language model presents information clearly, structured, and convincingly. But persuasiveness and evidential power are two different things.
With agentic AI, it becomes more urgent
With the rise of agentic AI, this distinction becomes even more important. AI systems will then not only produce text but also gather information, analyse data, formulate conclusions, and take action based on them. An incorrect factual observation is one risk. A correct observation with an unduly firm explanation can be at least as problematic.
Methodological caution is therefore increasingly becoming an important AI skill. We don't just need to learn how to write better prompts, but also how to better assess what an AI model returns. This means distinguishing between facts, correlations, explanations, and assumptions. Sometimes it also means accepting that the most scientifically sound conclusion is less spectacular than the story we would like to tell.
Generative AI makes knowledge more accessible than ever, but accessibility does not automatically make knowledge more reliable. The most important question, therefore, may not be whether AI hallucinates. The more important question is how we know that the convincing story AI tells us is actually the conclusion that the available evidence justifies. That's where critical AI use begins.
Job van den Berg is an AI keynote speaker, tech entrepreneur, and author of five books on AI. He puts agents into production himself weekly and gives 150+ keynotes per year on AI agents and agentic commerce.
Counting Cybertrucks in the back seat is not representative research. But it is a perfect example of how statistics, logical thinking, and AI come together: observe, formulate hypotheses, and above all, try to disprove them.
After half a year back in Silicon Valley, one thing stands out: the conversation is shifting from 'what can AI do?' to 'how do we truly build on it?'. Billboards about safe AI, control, and infrastructure indicate a new theme.
The promise of AI agents is friction-free, ease first. But it was precisely in that friction that the contact which forges our social capital always arose. An essay on connection, technology, and the default setting.
Answers are more accessible than ever. The real skill will not be formulating the question, but opening the hatch under the bonnet.

Job van den Berg is an AI keynote speaker, tech entrepreneur and author of five books on AI. He ships AI agents into production every week and delivers 150+ keynotes a year on AI agents and agentic commerce.
On EditieNL I discussed whether AI could threaten humanity. About Anthropic researcher Evan Hubinger, agentic AI, the black box, and why we are building faster than we understand.
A simple AI video already costs about 4 litres of water and as much electricity as a 10-watt LED lamp burning for 42 hours. What happens with full films and commercials, and why digital is not automatically sustainable.
AI keeps getting better, yet workplace sentiment about AI is deteriorating. Research shows why adoption is as much a social challenge as a technological one: from the Matthew effect to psychological safety.
Human in the loop sounds reassuring. But researchers warn that prolonged use of autonomous AI systems can undermine the cognitive capacities of the very supervisors we depend on. The question is not whether a human is formally present, but whether that human can still meaningfully intervene.










































