OpenAI launches ChatGPT image generation and the results will blow your mind

The 4o image generator creates hyper-realistic images, allows for the generation of consistent variations, inserts quality text, and so forth.
March 26, 2025

OpenAI has just made a significant leap forward in AI image creation by integrating its new model into GPT-4o. The image generator 4o surpasses the capabilities of the DALL-E model family, allowing for more realistic works, higher levels of detail, coherent text inclusion, and maintaining consistency for generating variations, among other useful capabilities.

The developer has already activated the 4o image generation in the GPT-40 model for the Plus, Pro, Teams, and free plans, and also within Sora, its video-generating AI. OpenAI has announced that it will soon be rolled out in the API and for Enterprise and Edu plans.

Those who wish to continue using DALL-E to create their AI works can access this model through a dedicated DALL-E GPT account.

Capabilities of the 4o Image Generator

OpenAI developed the 4o image generator focusing on enhancing its utility. Beyond the appeal of the images, they should also serve to communicate, explain, or persuade, and for this, they must be coherent, high-quality, and clearly organize the information.

We train our models on the joint distribution of images and text found online, learning not only how images relate to language but also how they relate to each other. Combined with intensive post-training, the resulting model possesses surprising visual fluency, capable of generating useful, consistent, and contextual images,” explains from the developer.

These premises have materialized in a vast array of new capabilities or enhanced functions. Something that has not changed is the ease of use of the tool, as this model continues to function conversationally. You only need to enter the description of the image you want to create or the instructions to generate variations, and ChatGPT-4o will handle the rest.

However, you may notice that the AI takes a bit longer to generate the images (something that can take up to a minute). This is because the model has a longer “thought” process to create more precise and detailed works.

Enhanced Photorealism and More Styles

The 4o image generator model is capable of creating photorealistic works with great quality and detail. This contrasts with the capabilities of DALL-E, which in this regard were not as advanced as those of other image-generating AIs.

Similarly, OpenAI explains that “training with images that reflect a wide variety of image styles allows the model to create or transform images convincingly.”

I asked the AI to create this image: “a female climber resting seated on a rock ledge on a mountain. Her companion is about to reach her, climbing from below.” Then, I asked it to do so in a steampunk style, which it recreated perfectly.

Left: photorealistic image of two mountaineers created with 4o / Right: same steampunk style image created with 4o

Improved Text Inclusion

While DALLE-3 could insert text into the images it generated, this capability was not by any means infallible. Often, it invented a new language or did not write the letters correctly. The 4o image generator is much more precise in this regard, ensuring effective inclusion of text.

Birthday invitation created with 4o

Consistent Image Variation

The fact that image generation is now native to GPT-4o allows the model to leverage images and text in chat context, promoting consistency. In this manner, you will be able to refine your works through natural conversations with the tool and create variations while maintaining character and context consistency of the work.

To test this capability, I asked 4o to generate the image of “a duckling with a flower on its head.” After this, I requested “change the style of the image you created to a 3D style,” then change it to origami style, and finally to make the duckling glass and iridescent.

Artistic style variations of the image of a duckling with a flower on its head

Likewise, besides varying the artistic style, the 4o image generator is also capable of maintaining consistency, adding new elements to the work, applying color changes, and even placing our character in new scenarios.

Variations in style, color, elements, and background of the image of a duckling with a flower on its head

And yes, if you were wondering, it also allows generating hyper-realistic variations of images you upload from your device. In this case, I uploaded a Depositphotos photograph and asked it to “Create an image that recreates the attached image, but the woman is using a video game console and is seated at the North Pole.”

Variation of a real photograph

Management of 10 to 20 Objects

OpenAI highlights that while other image-generating AIs struggle to follow instructions involving 5 to 8 objectives, the 4o model can manage up to 10-20 different objects. This capability allows it to follow detailed instructions with attention to detail. “The greater linking of objects to their characteristics and relationships allows for better control.”

For example, I tried asking the AI to create “a square image with a grid of 4 rows by 4 columns and 16 sticker-style objects in png. From left to right and from top to bottom. Here is the list: 1. A raccoon, 2. A yellow lightning bolt, 3. A lilac spiral, 4. A green circle, 5. A red tulip, 6. An hourglass, 7. A gray cloud with a sad face, 8. A 26 with eyes, 9. A ginger cat with a black bowtie, 10. A globe, 11. A magnifying glass, 12. A lovestruck emoji, 13. A hot air balloon, 14. A blue and white walrus, 15. The word “Marketing4eCommerce” written in light blue and 16. A rainbow lightning bolt.”

Image composed of 16 creations

As you can see, it nailed the image. This represents a huge qualitative leap, and I tell you this firsthand, having had multiple struggles with ChatGPT to include a specific number of people, phones, screens, etc., in an image, and most of the time, I ended up losing.

Additionally, as an extra bonus, we inform you that you can create pngs (like the case of these stickers) and then use them to complete other images. I have used them to decorate this photograph of two mugs.

Photograph of two mugs over which stickers created with 4o have been inserted

But… if you want to save yourself the work, 4o does it for you! I asked the tool to use the stickers it had just created and paste them on the image of a corporate mug. It replied, “I can help you create an image that includes the stickers over a corporate mug. However, I will need you to send me the image of the mug or give me details about the design. Could you provide me with more information or upload the mug’s image?” I uploaded an image of a mug… and voilà!

Insertion of stickers created by 4o into an image subsequently generated by the AI

Focusing on the world of eCommerce, this could become a simple and accessible way to optimize product images.

Contextual Learning

The 4o model is capable of analyzing and learning from images uploaded by the user, thanks to which it can skillfully integrate their details and context to elevate image generation. OpenAI provides the following example:

Image creation from sample works

General Knowledge

Another advantage of the native image generation of 4o is that the AI can directly link its knowledge, resulting in a more intelligent and efficient model. For example, if you ask it to create a herbal-style poster with 4 typical spring flowers from the US, it will use its knowledge to identify which flowers to create.

Image creation using 4o’s general knowledge

Limitations

As we have just observed in this last example, while OpenAI’s new image-generating AI showcases astounding progression, it is still not perfect. The company itself states: “we are aware of multiple current limitations that we will address through improvements to the model following the initial launch.”

Some of the flaws the tool may present include creating cropped images, experiencing hallucinations, binding issues, generating inaccurate graphics, difficulties including text in non-Latin languages, and alterations in certain aspects of the image when requested edits.

Nevertheless, we strongly encourage you to try it, as it is by far the most comprehensive image-generating AI we have tested to date.

Photo: GPT-4o

Other articles related to

Published by

Content Manager in Marketing4eCommerce
Content Manager in Marketing4eCommerce, which translates to: writer, editor, and absolute fan of generating images with AI.

Stay up to date!

Únete a nuestro canal de Telegram

All you need to know!

Sign up for our newsletter and receive our best articles on eCommerce and digital marketing in your email for free.