OpenAI admits it: ChatGPT hallucinates quite frequently, especially if it is a paid model

OpenAI conducted a study where they discovered that an advanced OpenAI model fabricated 33% of its responses.
September 10, 2025

Conversational AI chats, such as ChatGPT or Gemini, have been a blessing for many people, both professionally and personally. However, these AIs conceal a dangerous double-edged sword: hallucinations. This phenomenon occurs when the AI generates information that appears plausible and convincing but is actually as false as a three-dollar bill.

Although this error may be harmless in informal conversations, it becomes an unacceptable risk in high-stakes fields such as medicine or law. A study by OpenAI and the Georgia Institute of Technology has revealed the truth behind this behavior: these models hallucinate not because they have been designed to do so, but because their training and evaluation systems reward them for it.

This occurs more frequently in advanced models (in other words, if you pay, you are likely to receive even more deceptive answers).

Why AI is a liar

The main thesis of the study is that models hallucinate not because they are inherently prone to fantasy, but because their training and evaluation procedures reward guessing instead of admitting uncertainty.

The analysis of the causes was conducted in two phases of the model’s lifecycle:

  1. Training phase: errors originate from statistical pressures.
  2. Evaluation phase: errors persist due to poor test design.

Cause 1: Statistical pressures during training

At a very basic level, AI models learn to predict the next word in a sequence, a task that can be seen as a series of binary decisions. The model is constantly attempting to classify whether a statement is valid or not based on the billions of data points on which it has been trained.

However, when the model encounters very specific information, rare, or lacking a clear pattern (for example, the exact year a little-known historical figure was born, mentioned only once in a text), the model faces a dilemma. It does not have sufficient data to determine whether this information is an isolated fact or simply an error. In this situation, statistical pressures compel it to generate an answer that sounds correct, even if it is incorrect.

According to the study, the AI is not “making things up” in a malicious way, but rather is committing a classification error that is a natural byproduct of the learning process. It is similar to a student who, lacking a clear rule, is forced to guess an answer on an exam.

Cause 2: The evaluation epidemic

This is where the study becomes more critical and refers to it as a “socio-technical” cause. The manner in which language models are measured in the industry, through benchmarks (standardized tests), is the main driver behind hallucinations. These tests are designed to reward binary accuracy: an answer is correct (1 point) or incorrect (0 points). There is no in-between.

If a model is confronted with a difficult question and answers “I do not know,” it receives a score of zero. If, instead, it takes a risk, even with an answer that only seems plausible, it has a chance of being correct and earning a point. The model, being optimized to maximize its score, “learns” that it is better to guess than to abstain. This incentive creates a vicious cycle in which unreliability is indirectly rewarded. The study refers to this as an “epidemic of penalizing uncertainty” that has shaped AI behavior.

The most advanced AIs are the biggest liars

According to the study, the more advanced models hallucinate even more. For example, in the PersonQA dataset (a set of questions and answers about people), an advanced OpenAI model fabricated answers 33% of the time, doubling the rate of its previous version. The smallest model, o4-mini, performed even worse, with a rate of 48%. And this is despite both models demonstrating superior mathematical skills.

This suggests that as models become more proficient at certain tasks, they are also incentivized to take greater risks and exercise less caution, which leads them to generate more false but convincing statements. The main conclusion of the study is that hallucinations are not an unfathomable flaw, but rather the logical consequence of a system that penalizes honesty.

We can confirm this assertion from Marketing4ecommerce, as we often utilize paid ChatGPT for various tasks. In fact, we have created several GPTs for specific purposes, and we have observed that, over time, instead of improving through learning, these models become increasingly imprecise and deviate further from the initial directives. To such an extent that the best solution has been to delete those GPTs and create new ones.

Rules of the game must change

The study proposes a solution that goes beyond building bigger and more complex models. The answer does not lie in technology, but rather in the philosophy of evaluation. The researchers argue that, rather than developing new specialized tests to detect hallucinations, the industry should change the criteria of existing tests.

The proposed approach is to modify the way models are scored so that honesty is properly rewarded. The new evaluation metrics should follow a more sophisticated system than the simple “all or nothing” approach:

  • Penalize confident errors: a model that provides an incorrect answer with high confidence should be penalized by deducting points.
  • Reward uncertainty: a model that states “I do not know” when the answer is uncertain should receive credit or, at the very least, not be penalized with a zero.

This logic is the same as that used in some multiple-choice examinations where incorrect answers subtract points, encouraging students to leave the question blank if they are unsure. By making this change, the industry can be guided to develop models that are not only more accurate, but also more reliable and honest regarding their own limitations.

Our friendly advice: you should never trust an AI.

Photo: Nano Banana

Other articles related to

Published by

Content Manager in Marketing4eCommerce

Stay up to date!

Únete a nuestro canal de Telegram

All you need to know!

Sign up for our newsletter and receive our best articles on eCommerce and digital marketing in your email for free.