Meta has just made a major statement with the launch of Muse Image and the preview release of Muse Video, its first AI-powered multimedia content generation models developed by its Meta Superintelligence Labs division.
Introducing Muse Image and Muse Video, the first media generation models developed by Meta Superintelligence Labs. Muse Image is our most advanced image generation model yet. It follows instructions faithfully, edits with precision, composes from multiple references, and draws… pic.twitter.com/byNpQZO1RW — AI at Meta (@AIatMeta) July 7, 2026
Introducing Muse Image and Muse Video, the first media generation models developed by Meta Superintelligence Labs.
Muse Image is our most advanced image generation model yet. It follows instructions faithfully, edits with precision, composes from multiple references, and draws… pic.twitter.com/byNpQZO1RW
— AI at Meta (@AIatMeta) July 7, 2026
Muse Image is Meta’s most advanced image generation model to date. According to what the company explained: “it follows instructions faithfully, edits with precision, composes from multiple references, and uses Instagram to create social context.”
One of Muse Image’s main attributes is that it operates as an agent, since it can use search and coding tools to create more accurate images, has self-improvement capabilities, and is able to optimize its own performance.
In addition, Muse Image also integrates with Muse Spark, allowing them to share tools and plan jointly in order to create interactive multimedia content.
Now that we know the general context behind Meta’s new AI image generation model, let us take a closer look at its capabilities:
This AI can search the web to access real-time information and images, supporting the generation of truthful and realistic content.
Meta also notes that “the search feature improves factual accuracy on questions that require a high level of knowledge, especially those related to current events and real-world facts.”
Meta Image relies on reinforcement learning to write and execute code to create charts, accurate QR codes, and optimized images from rendered figures.
It also integrates with Muse Spark to combine code generation and multimedia content. This combination makes it possible to automatically and interactively develop animated GIFs, websites with embedded images, and interactive visual games.
Muse Image can follow the user’s instructions precisely to edit an existing image, modifying only the specified elements.
It also makes it possible to carry out multiple editing stages on the same image, since it preserves the coherence and consistency of everything that the user did not ask to change.
This AI model can process several reference images provided by the user and generate works that include the specified elements from each one (characters, objects, clothing, settings, design styles, etc.). While writing the prompt, you can alternate between text and attached images for greater precision in the instructions.
Muse Image reflects on and refines its work autonomously within its chain of thought. This self-refinement adapts depending on the error: it performs local edits to correct small details, generates an image from scratch if the mistakes are more significant, or uses tools to improve accuracy.
It is worth noting that this behavior was not deliberately designed by Meta. It emerged spontaneously during reinforcement learning training, as the system discovered that self-correction produced higher-quality images and, therefore, a higher reward.
Like language models, Muse Image improves when it processes more information during inference. Greater test-time compute allows it to reason more, use tools, and self-refine. This increase in reasoning capacity improves human preference Elo scores following a log-linear relationship, where final quality depends on the total combined processing of text and visual tokens.
Optimizing this token budget is crucial. While the Best-of-N method (generating multiple options and selecting the best one) saturates quickly, investing that compute in deliberate reasoning offers greater scalability. In addition, reasoning and tools reinforce one another: they allow the model to search for external references or write code, filling logical gaps with precision.
For now, Muse Image has begun rolling out in the Meta AI app and on the meta.ai website, in Instagram Stories in the United States, and on WhatsApp in several countries (although the company has not specified which ones).
However, the tech giant’s ambitions do not stop there, as it plans to expand Muse Image soon across the rest of its ecosystem (Facebook and Messenger), as well as make it available to advertisers through Meta Advantage+ Creative.
In addition to officially launching Muse Image, Meta has also taken the opportunity to share a preview of Muse Video, its video AI, which also includes native audio support.
To develop Muse Video, the tech company started from the same pretraining foundation used to build Muse Image. The model stands out for its accuracy, visual fidelity, and temporal consistency, and Meta is investing to further optimize audio-video synchronization and the physically accurate rendering of fast motion.
The company has not given a specific date for Muse Video’s official launch. For now, the only thing we know is that “it will be available soon for creators and Meta AI.”
To allow users to verify whether an image has been generated by AI, Muse Image incorporates Content Seal, an invisible watermarking system similar to Google’s SynthID. Images created in the Meta AI app and on meta.ai include this hidden provenance signal, which remains intact despite cropping, compression, resizing, or screenshots. Soon, this technology will also be extended to videos.
In addition, a detection tool is being introduced to check whether a file carries this watermark, making it easier to identify content created with Meta AI.
Photo: Muse Image
Your email address will not be published. Required fields are marked *
Δ