IA 360
Current Affairs

GPT-5.6 arrives as a family of three models

OpenAI introduces Sol, Terra, and Luna with access varying by product and plan, and one warning: more capability needs better controls.

7 min read AI-generated Leer en español
GPT-5.6 arrives as a family of three models

On July 9, 2026, OpenAI released GPT‑5.6 for general availability as a family comprising Sol, Terra and Luna. The product followed a limited preview, with rollout varying by surface and plan.

GPT‑5.6 is not one model or a button that appears in the same way for everyone. OpenAI presents Sol as its flagship, Terra as a balanced option for everyday work and Luna as the fastest and most affordable variant. The distinction matters because choosing a model also means choosing how much capability a task actually needs.

The launch places the family in coding, research, professional work, science, cybersecurity, computer use and design. A list of capabilities does not replace a use decision, however. An everyday task may need a fast response, a complex analysis may benefit from more reasoning effort, and a sensitive workflow may also need human review, limited permissions and traceability.

Three tiers, three decisions

The official launch page defines three durable tiers. Sol concentrates maximum capability for demanding work. Terra seeks a balance between intelligence and cost. Luna prioritizes speed and price. They are not interchangeable names: they represent different tradeoffs among quality, latency and budget.

This structure forces teams to ask the question a model picker often hides: what happens if the answer is wrong? Explaining a function, classifying a list or preparing a first draft can tolerate later correction. Changing infrastructure, acting on customer data or supporting a clinical decision requires more verification, even when a model scores well in a table.

Task size and difficulty also need separating. A long document does not always require the most capable model when the goal is to extract well-defined fields. A short question can demand deep reasoning if it combines incomplete evidence, several constraints and a consequential outcome. Choosing by the visible length of the request confuses volume with complexity.

Availability depends on product and plan

In standard ChatGPT conversations, GPT‑5.6 Sol is rolling out gradually to eligible plans. OpenAI documentation keeps GPT‑5.5 Instant as the default for fast responses and assigns Sol to Medium, High and Extra High reasoning levels depending on the plan. Not seeing the model does not necessarily signal a fault: access can depend on rollout, subscription or a managed-workspace setting.

Terra and Luna are not selectable in standard ChatGPT conversations. OpenAI offers them in Work, Codex and the API under different conditions: in Codex, for example, Terra is available to Free and Go, while Plus, Pro, Business and Enterprise can access all three variants. Developers can use Sol, Terra and Luna through the API.

The lesson is not to memorize a table that will change. It is to check the specific surface every time. ChatGPT, Work, Codex and the API do not necessarily share the same picker, permissions, limits or administrative controls. Designing a workflow on the assumption that “GPT‑5.6 is available” is not enough: specify which variant, in which product and under which access policy.

The same family takes another form outside OpenAI: Microsoft Foundry’s July 9 announcement offers all three variants with its own deployment modes and pricing. It independently confirms the lesson: a model name alone does not determine the product, region, quota or cost an organization will encounter.

Capability with limits

OpenAI accompanied the launch with a GPT‑5.6 system card. The document classifies Sol, Terra and Luna as High capability in Biological and Chemical risk and in Cybersecurity, and below High in AI Self-Improvement. It also describes safeguards tailored to each model’s profile.

That assessment is not a universal guarantee. The system card itself says its evaluations provide a lower bound on potential capability: other prompts, fine-tuning, longer rollouts or different scaffolding could elicit behavior the tests did not observe. This is an important doorway into any safety report: a test documents specific conditions and results; it does not enumerate every future use.

METR’s independent pre-deployment evaluation shows why those conditions matter: the lab did not consider its task-horizon estimate robust because the result changed sharply depending on how attempts to exploit evaluation-environment flaws were treated. An unstable number should not become a product promise.

The practical rule is proportionality. Summarizing a document does not always need the highest reasoning effort. For a workflow involving sensitive data, response quality is not the only priority: permissions, review and the ability to stop an action also matter. GPT‑5.6 expands the available choices; it does not remove the duty to make those choices carefully.

Choose by task, not prestige

Before connecting a model to a tool, a team should define five things: the result it needs, acceptable latency, budget, the consequence of an error and the person or system that will review the output. Only then does comparing variants make sense.

ARC Prize’s verified results, published on the same July 9, compare the three variants and several reasoning levels on a particular benchmark suite. They are useful for that domain; by themselves they do not measure latency, cost, privacy or the impact of an error in a real workflow.

In programming, a fast response may be enough to explain a function or prepare a test. A change touching several services needs more context, verification and review. Research is similar: locating sources is different from evaluating a claim or deciding whether evidence supports a conclusion. The family lets capacity be adjusted, but it does not transfer responsibility for framing the right question to the model.

Compare variants without buying the ranking

The ARC Prize table reveals something a headline often hides: each variant was tested at several reasoning levels, and results change as that effort changes. This supports comparison on a shared suite, but it does not make the suite’s winner the best choice for every job. The benchmark measures its tasks and environment, not an organization’s documents, tools or risks.

The Microsoft Foundry launch table adds another axis: each variant has a different price. Even so, price per million tokens is not the cost of a completed task. Retries, latency, human review and corrections caused by errors also count.

A useful comparison can start with a small set of real cases and criteria written in advance. Run the same tasks with the same context, cap retries, and measure reviewed successes, elapsed time and final cost. The choice then stops being “which model tops the table?” and becomes “which configuration completes the work with sufficient quality and control?”

Rollout and governance

Before designing a workflow, an organization should check which variant is enabled, what data it can receive, what actions it can execute and how results are recorded. For work affecting code, customer data or operational decisions, the principle remains simple: test within a limited scope, review outputs and retain a clear way to stop the process.

A responsible deployment is not measured only by how many prompts it processes. It is useful to observe time saved, human-reviewed quality, correction rates, cost per task and cases where the model should not have intervened. These measures prevent the choice among Sol, Terra and Luna from becoming a branding preference.

The best variant is not automatically the one at the top of the family. It is the one that provides sufficient quality for a specific task, with latency, cost and controls appropriate to its impact.

This article was produced with artificial intelligence under human editorial oversight.

Share this article

This website uses cookies to improve the browsing experience. Cookie policy.

↑↓ navigate ↵ open esc close