Sol, Terra and Luna: the solar system of OpenAI’s new GPT-5.6 models

GPT-5.6 introduces a new approach to organizing its capabilities. Discover what each model offers and the enhancements they bring.

Sol, Terra, and Luna are not friendly nicknames meant to make a technical launch easier to digest; they are, according to OpenAI itself, a shift in logic: from now on, the number (5.6) identifies the model generation, while Sol, Terra, and Luna identify lasting capability tiers, which the company says will be able to advance at their own pace, without waiting for the next full version.

It is a shift worth understanding well, because it changes how we are going to talk about “AI updates” from here on out: it will no longer be just “GPT-6 is out”, but “Terra moved up a level” or “Luna now reaches what Sol was doing six months ago.”

Applied to specific products, this is how the family now looks: Sol is the model for demanding reasoning, long-range programming, cybersecurity, and agentic work, at a cost of $5 per million input tokens and $30 per million output tokens. Terra is the model designed for balanced day-to-day work, with competitive performance against GPT-5.5 but at half the price, at $2.50 for input and $15 for output per million tokens. And Luna is the option for fast, low-cost tasks, with surprisingly solid capabilities for its price: just $1 for input and $6 for output per million tokens.

What really improved (and it is not just that “it is smarter”)

Here is a summary of what, in my reading, are the four improvements with the greatest practical impact:

  1. A much larger context window. Sol works with a window of 1.05 million input tokens and up to 128,000 output tokens. In practical terms: it can sustain conversations or agentic work projects lasting several hours without “losing” the thread, something that previously forced long tasks to be broken into smaller pieces.
  2. An “ultra” mode that distributes work across multiple agents. Sol introduces a maximum reasoning effort mode (max reasoning) and an ultra mode that coordinates multiple sub-agents working in parallel on different parts of the same complex task. Sam Altman summarized it to CNBC by noting that the model is 54% more token-efficient for agentic programming than its predecessor; in other words, it does the same thing, or more, while using significantly fewer tokens.
  3. Much cheaper per unit of intelligence. According to OpenAI’s own tests, Terra and Luna outperform Anthropic’s Claude Fable 5 at roughly one-sixteenth of the estimated cost, while Sol, in its maximum reasoning configuration, nearly matches the performance of Fable 5 while completing tasks in 61% less time and at half the estimated cost. In addition, the prompt caching system has been improved: it now includes explicit cache breakpoints and a minimum cache lifespan of 30 minutes, which lowers costs when working with documents or contexts that are constantly reused.
  4. Better for visual and product tasks, not just code. In OpenAI’s own frontend tests, GPT-5.6 received a score of 4.4 out of 5 on interface design tasks (turning briefs, dashboards, and products into full, responsive interfaces), compared with 4.0 for GPT-5.5 and 3.5 for Claude 4.8 in the same test. In presentations, OpenAI reports that the model is approximately 1.6 times more token-efficient than competing models for slide generation at the scale of tools like Canva.

The key takeaway from its safety testing

Here is the detail I am not going to gloss over just because it does not appear in the official announcement. METR, an independent AI safety evaluator, found that Sol “gamed” its own agentic behavior benchmark at the highest rate ever recorded in an OpenAI model, a finding that (according to that same report) makes those specific scores less reliable. Put simply: the model learned, to some extent, to optimize for looking good on the test, not just for actually solving the task well.

This does not invalidate the real improvements documented above: the frontend, token efficiency, and context benchmarks come from different methodologies. But it is a reminder that it is worth viewing any “it beats X model by Y%” figure through the healthy filter of who ran the test and under what conditions. It is, quite literally, why I always prefer to cross-check the manufacturer’s announcement with at least one independent source before repeating a number.

Where it is available

GPT-5.6 was immediately integrated into the OpenAI ecosystem and that of its partners: it is available through ChatGPT, Codex, and the official API as of July 9, it has already reached GitHub Copilot in its three variants, and OpenAI announced that Sol will run on Cerebras infrastructure at speeds of up to 750 tokens per second, a figure aimed at use cases where latency matters just as much as response quality, such as live customer support or real-time content generation during events.

What this means if you work in marketing or eCommerce

The three-tier logic is, in practice, an invitation to stop treating “using AI” as a binary decision. A marketing team with sound judgment can now assign Luna to ticket classification or high-volume copy variant generation, Terra to day-to-day content production and campaign analysis, and reserve Sol for projects that truly require deep reasoning: a complex market analysis, a delicate data migration, or a security audit.

The question worth taking away from this article, in my view, is not “which model is the smartest?” but rather: “do we already have the judgment needed to decide, task by task, how much intelligence we need to pay for?

It is no coincidence that OpenAI described Terra, in its own words, as its “balanced model for day-to-day work.” Sol represents the far end of maximum capability, and also maximum cost, while Luna represents the far end of efficiency and, likewise, the limit of what it can solve well.

The temptation, with any AI launch, is to rush toward the most powerful model available, as if extra intelligence were automatically the right decision. Aristotle would say that this is a vice of excess, just as avoidable as the opposite vice of always choosing the cheapest option. Virtue and, I would argue, good business strategy lie in developing the judgment to know, task by task, what the right middle ground is in each case.

Image: GPT Images 2.0

Other articles related to

Published by

Stay up to date!

Únete a nuestro canal de Telegram

All you need to know!

Sign up for our newsletter and receive our best articles on eCommerce and digital marketing in your email for free.