Why AI makes things up (and how to prevent it)
AI makes things up because it does not know what is true: a language model such as ChatGPT does not look up the answer anywhere, it builds it word by word, choosing the most likely word each time. When it does not have the information, it does not stay quiet: it writes whatever sounds most likely, with the same confidence it uses for a true fact. This is called a hallucination. And you prevent it by giving the model the source and asking it to cite where each point comes from. If it cannot cite it, it does not know it.
We explain it in just over half a minute in this video (in Spanish), and in more detail below.
How an AI writes
A language model has learned, from enormous amounts of text, which words tend to follow others. When you ask it something, it does not check a database or open a document: it calculates which word fits best after the previous one, writes it and repeats.
In the video we show it with a simple example. If the sentence is "The capital of France is", the most likely word is "Paris", by a huge margin over any other. There, probability and truth coincide, because that fact appears again and again in everything the model has read.
What happens when it does not have the information
Now change the question to something the model cannot know: what a specific article of a regulation says, the price of one of your products, your store's opening hours. The mechanism is the same. The model keeps looking for the most likely word, and the most likely one is almost never "I do not know". It is something that sounds right, such as "the deadline is 30 days".
The answer has the same confidence, tone and format as a true fact. That is why it misleads: nothing in the text warns you that this part was made up.
Why asking it not to make anything up does not work
The natural reaction is to add "do not make anything up" to the request. It helps very little. The model has no internal list of what it knows. It cannot check whether one of its sentences is true, because it has nothing to compare it with. You are asking it to tell apart something it cannot see.
How to spot a made-up fact
There is no foolproof sign, but there are clues worth a second look:
- Figures that are too round or too precise with no indication of where they come from.
- Very specific references: the number of an article, the title of a study, the name of a regulation. They are easy to invent and easy to check.
- Information that changes over time, such as prices, opening hours or dates, given as if it were fixed.
If you see any of these, look the information up in its source before using it.
How to prevent it: the source and the citation
What does work is changing the job you give it. Instead of "answer me", ask it to "answer me using this".
- Give it the source. Paste the document, the policy, the product sheet or the text of the regulation. That way the answer comes from something real that you control, not from what the model remembers.
- Ask it to cite where each point comes from: the section, the sentence or the page. A citation can be checked in a moment.
- Let it say "I do not know". Tell it explicitly that, if the answer is not in the text you give it, it should say so. If you do not give it that way out, it will fill the gap.
- Check whatever changes. Prices, opening hours, dates and rules may have changed since the model was trained: always look them up in the original source.
The short rule is the one from the video: if it cannot cite it, it does not know it.
Answer only with the information in the text I paste below. For every point you give me, state which part of the text it comes from. If the answer is not in the text, tell me "I do not know" and do not fill it in with what you think.
Text:
[paste the document here]
Question:
[write your question here]
What it means for a small business
In a private chat, a made-up fact is a nuisance. In front of a customer, it is not. If you use AI to answer messages, write product descriptions, prepare quotes or help visitors on your website, a made-up answer goes out under your business's name.
That is why, in a chatbot or an assistant for a website, the first thing we take care of is where the answers come from. The assistant on our own website answers only with what the website says and, if something is not there, it says so instead of making it up. It is the same idea of the source and the citation, applied to a system that works on its own.
The same goes for AI automations. When a process uses a model to read an email or summarize a document, you have to decide what information it can use, what it does when something is missing and what goes through a person before it goes out.
Frequently asked questions
What is an AI hallucination?
It is an answer that the model presents as true but has made up. It happens because the model writes the most likely continuation of the text, not the true one, and when it does not have the information it fills the gap with something that sounds right.
Can you make AI never make anything up?
Not completely. What you can do is reduce it a lot and make it visible: give it the source, ask it to cite where each point comes from and allow it to say "I do not know". That way, anything without a source is spotted right away.
Does this happen with ChatGPT and the other AI tools?
Yes. It is a consequence of how language models work, not of a particular brand. Some fail more and others less, but all of them can invent a fact you do not give them.
How do I know if a chatbot makes up its answers?
Ask it something that is not in your business's information and see what it replies. If it gives a specific answer instead of saying it does not know, it is making it up.