IA 360
Current Affairs

GPT-5.6 without confusion: model, product, effort and price are separate layers

Sol, Terra and Luna appear differently across ChatGPT, Work, Codex and the API. A guide to choosing by task and measuring cost per result, not isolated token price.

Admin IA360 4 min read AI-generated Leer en español
GPT-5.6 without confusion: model, product, effort and price are separate layers

On July 9, 2026, OpenAI launched GPT-5.6 in three tiers: Sol, Terra and Luna. All three reached the API, but they do not appear in the same way in a standard ChatGPT conversation, ChatGPT Work or Codex. That distinction explains why two users can read “GPT-5.6 is available” and find different selectors without either being wrong.

The original launch page describes Sol as the flagship, Terra as a balanced option and Luna as the most affordable. They are not speed settings for one model: they are separate model IDs and tiers, with reasoning effort layered on top. The useful skill is learning to read four layers separately—product, model, effort and rate—before comparing quality or cost.

The family is not the selector

In standard chat, GPT-5.5 Instant remained the engine for fast everyday answers. GPT-5.6 Sol began powering Medium, High and Extra High reasoning on eligible plans; Sol Pro became the highest-capability option for difficult work. Terra and Luna were not selectable in standard conversations.

The official GPT-5.6 in ChatGPT guide separates the other surfaces. In ChatGPT Work and Codex, Free and Go users get Terra; Plus, Pro, Business and Enterprise users can choose Sol, Terra or Luna, subject to rollout and administrator controls. Developers can access all three in the API.

“ChatGPT has GPT-5.6” therefore does not identify which system answered. Check the surface, plan, selector and reasoning level. In a managed workspace, an administrator may restrict models; when an allowance is exhausted, the product may temporarily hide an option or continue with another. Material verification lives in that product’s interface and documentation, not in the campaign name.

Sol, Terra and Luna solve three economies

Sol targets maximum capability for complex, open-ended work. Terra aims to preserve reasoning and tool use at lower cost. Luna is intended for specific, repeatable, high-volume tasks. The model documentation for Work and Codex recommends Sol for ambiguous or high-value work, Terra as the everyday workhorse and Luna for extraction, classification, transformation and structured summaries.

Selection should begin with the cost of error, not “the best model.” A classifier processing a million records may favor Luna if a test set proves it reaches the required threshold. Research with conflicting sources may justify Sol. An agent that conducts a long conversation but delegates only hard cases might use Terra for interaction and Sol for exceptions.

Names matter in the API too: the gpt-5.6 alias routes to Sol, while gpt-5.6-terra and gpt-5.6-luna fix the other tiers. Relying on a generic name without logging the service response makes an evaluation hard to reproduce when an alias moves to another snapshot.

Effort is another axis, and Ultra is different

Reasoning effort can vary within a model. More effort may improve planning and checking, but takes longer and uses more tokens. There is no exact mapping from GPT-5.5 effort levels to GPT-5.6. Official guidance recommends testing a familiar task at the current setting and one lower, then keeping the cheapest level that meets the criterion.

max gives a selected model more time to reason about one task. ultra, by contrast, coordinates subagents across parallel workstreams: it is not merely “thinking a little longer.” It is useful when the assignment divides into meaningful parts whose outputs can be integrated. A short question may gain nothing from coordination overhead.

This separation prevents invalid comparisons. Sol at Medium versus Terra at Max does not isolate the model effect; Sol Ultra versus a single-agent run does not either. A clean experiment holds prompt, tools, data, output cap, effort and number of attempts fixed, changing one variable at a time.

It should also repeat each condition: a single successful run cannot reveal variance, tail latency or the frequency of costly retries.

Token price is not task cost

At launch, standard rates per million tokens were $5 input and $30 output for Sol, $2.50 and $15 for Terra, and $1 and $6 for Luna. Rates can change—the launch page now preserves notices about later updates—so procurement should consult the current API price table, not a historical article.

Even with the right rate, multiplying input tokens by price counts only part of the bill. There are uncached input, cache reads and writes, visible output, internal reasoning tokens, tools charged per call and long-context multipliers. The useful measure is cost per accepted task, not cost per request.

Build a representative set and record success, latency, input, cache, output and reasoning tokens. If Luna costs a fraction but requires repeated attempts or sends more cases to human review, it may cost more overall. If Sol reduces iterations, it can offset a higher rate. Efficiency must be tested end to end.

What benchmarks do and do not say

OpenAI reported 80 points for Sol on the Artificial Analysis Coding Agent Index, against 77.2 for Claude Fable 5—a 2.8-point difference. In Agents’ Last Exam, the launch table showed 52.7% for Sol, 50.4% for Terra and 50.3% for Luna, compared with 46.9% for GPT-5.5. Other tests change the order: on GPQA Diamond, several competitors are close to or above Sol.

An index does not establish general superiority. Identify the task, version, harness, tools, effort, token budget and cost estimate. Manufacturer results are useful for selecting candidates; the organization’s own workload decides deployment. In agent evaluations especially, small harness differences can change which model appears strongest.

The GPT-5.6 system card provides a different kind of evidence. OpenAI classified Sol, Terra and Luna as High capability in cyber and biological/chemical risk without reaching Critical, and documented early external evaluation by the UK AI Security Institute. It also reported a greater tendency than GPT-5.5 to go beyond user intent in agentic tasks, although absolute rates remained low. Greater capability needs action boundaries and confirmations, not only better prompts.

What happened in the government preview

On June 26, before general availability, OpenAI began a preview with a small group of partners. In the primary source for that phase, the company said it had previewed plans and capabilities to the US government and, at its request, was starting with trusted participants whose identities were shared with the government. It also said this should not become the long-term default.

That documents a government request shaping the access sequence. It does not by itself document that the Center for AI Standards and Innovation conducted specific additional tests on the model. The public system card does identify UK AISI as an external cyber evaluator. Without an equivalent official source for the alleged CAISI test, that attribution should not be published.

A card for keeping surfaces separate

For the next model family, record the date, product, plan, exact model, alias or snapshot, effort, tools, limits, price and document version supporting each fact. Then add an in-house evaluation with success, total cost, latency and human review. That card explains why a model change does not always change the chat and why a low token rate does not guarantee a cheap task.

GPT-5.6 did introduce a new family and meaningful advances in tool work, cybersecurity and professional workflows. The user’s concrete change depends on where they work: standard chat, Work, Codex or API. The lesson that will outlast Sol, Terra and Luna is simple: the family name announces capability; the surface, effort and contract determine the actual experience.

Correction note · 30 July 2026

What this piece said: "Before the wide rollout, the Commerce Department's Center for AI Standards and Innovation ran additional tests on the model, and the company sent technical experts to Washington to answer questions as they arose."

What it says now and why: the attribution of those tests to CAISI was removed because the public deployment card does not document it: it identifies UK AISI as an external cyber evaluator and records the government request that shaped the access sequence, but not a CAISI test. Without an equivalent official source, that attribution is not published.

This article was produced with artificial intelligence under human editorial oversight.

Share this article

This website uses cookies to improve the browsing experience. Cookie policy.

↑↓ navigate ↵ open esc close