Conversational AI chats, such as ChatGPT or Gemini, have been a blessing for many people, both professionally and personally. However, these AIs conceal a dangerous double-edged sword: hallucinations. This phenomenon occurs when the AI generates information that appears plausible and convincing but is actually as false as a three-dollar bill.
Although this error may be harmless in informal conversations, it becomes an unacceptable risk in high-stakes fields such as medicine or law. A study by OpenAI and the Georgia Institute of Technology has revealed the truth behind this behavior: these models hallucinate not because they have been designed to do so, but because their training and evaluation systems reward them for it.
This occurs more frequently in advanced models (in other words, if you pay, you are likely to receive even more deceptive answers).
The main thesis of the study is that models hallucinate not because they are inherently prone to fantasy, but because their training and evaluation procedures reward guessing instead of admitting uncertainty.
The analysis of the causes was conducted in two phases of the model’s lifecycle:
At a very basic level, AI models learn to predict the next word in a sequence, a task that can be seen as a series of binary decisions. The model is constantly attempting to classify whether a statement is valid or not based on the billions of data points on which it has been trained.
However, when the model encounters very specific information, rare, or lacking a clear pattern (for example, the exact year a little-known historical figure was born, mentioned only once in a text), the model faces a dilemma. It does not have sufficient data to determine whether this information is an isolated fact or simply an error. In this situation, statistical pressures compel it to generate an answer that sounds correct, even if it is incorrect.
According to the study, the AI is not “making things up” in a malicious way, but rather is committing a classification error that is a natural byproduct of the learning process. It is similar to a student who, lacking a clear rule, is forced to guess an answer on an exam.
This is where the study becomes more critical and refers to it as a “socio-technical” cause. The manner in which language models are measured in the industry, through benchmarks (standardized tests), is the main driver behind hallucinations. These tests are designed to reward binary accuracy: an answer is correct (1 point) or incorrect (0 points). There is no in-between.
If a model is confronted with a difficult question and answers “I do not know,” it receives a score of zero. If, instead, it takes a risk, even with an answer that only seems plausible, it has a chance of being correct and earning a point. The model, being optimized to maximize its score, “learns” that it is better to guess than to abstain. This incentive creates a vicious cycle in which unreliability is indirectly rewarded. The study refers to this as an “epidemic of penalizing uncertainty” that has shaped AI behavior.
According to the study, the more advanced models hallucinate even more. For example, in the PersonQA dataset (a set of questions and answers about people), an advanced OpenAI model fabricated answers 33% of the time, doubling the rate of its previous version. The smallest model, o4-mini, performed even worse, with a rate of 48%. And this is despite both models demonstrating superior mathematical skills.
This suggests that as models become more proficient at certain tasks, they are also incentivized to take greater risks and exercise less caution, which leads them to generate more false but convincing statements. The main conclusion of the study is that hallucinations are not an unfathomable flaw, but rather the logical consequence of a system that penalizes honesty.
We can confirm this assertion from Marketing4ecommerce, as we often utilize paid ChatGPT for various tasks. In fact, we have created several GPTs for specific purposes, and we have observed that, over time, instead of improving through learning, these models become increasingly imprecise and deviate further from the initial directives. To such an extent that the best solution has been to delete those GPTs and create new ones.
The study proposes a solution that goes beyond building bigger and more complex models. The answer does not lie in technology, but rather in the philosophy of evaluation. The researchers argue that, rather than developing new specialized tests to detect hallucinations, the industry should change the criteria of existing tests.
The proposed approach is to modify the way models are scored so that honesty is properly rewarded. The new evaluation metrics should follow a more sophisticated system than the simple “all or nothing” approach:
This logic is the same as that used in some multiple-choice examinations where incorrect answers subtract points, encouraging students to leave the question blank if they are unsure. By making this change, the industry can be guided to develop models that are not only more accurate, but also more reliable and honest regarding their own limitations.
Our friendly advice: you should never trust an AI.
Photo: Nano Banana
Your email address will not be published. Required fields are marked *
Δ