IA 360
Current Affairs

OpenAI launches GPT-5.6 and debuts ChatGPT Work

OpenAI expands access to Sol, Terra and Luna and introduces ChatGPT Work, a surface that gathers context and acts through tools. Evaluating the agent requires separating the model, permissions, actions and verification.

Admin IA360 6 min read AI-generated Leer en español
OpenAI launches GPT-5.6 and debuts ChatGPT Work

OpenAI opened general access to its GPT-5.6 family on July 9, 2026 and introduced ChatGPT Work, a new surface for assigning tasks that combine files, applications and tools. The official launch announcement says public availability followed a limited preview and that rollout was beginning that day across ChatGPT, Codex and the API. It does not say that the US government authorized the release.

That distinction changes the story. The June 26 post announcing the preview says OpenAI had shared its plans with the US government and, at the government's request, began a trial with a small group of trusted partners whose participation was disclosed to officials. Neither document describes a permit, veto or official approval. What the record establishes is advance coordination followed by general access thirteen days later.

A family, a surface and several entry points

The release brought different components under one brand. GPT-5.6 is a family of three models: Sol, the most capable; Terra, designed to balance capability and cost; and Luna, the fastest and least expensive option. According to OpenAI's published access matrix, all three reached the API. Paid plans can choose among them in ChatGPT Work and Codex, while Free and Go receive Terra on those surfaces. In ChatGPT, access to Sol is reserved for paid plans.

ChatGPT Work is not another model. It is the product layer that gathers context, prepares a plan and executes steps with resources the user allows it to use. The official ChatGPT Work page lists files, tools and desktop applications, with documents, spreadsheets, presentations and apps among the possible outputs. That description does not mean ChatGPT and Codex have been merged. A surface can use capabilities developed for agents without turning two products into one.

This separation provides a first defense against confusing marketing. A model upgrade can change reasoning quality; a new surface can change the context available and the actions possible; a connector can expand the data entering the system; and a permission policy determines how far the agent may go. When a demonstration improves, ask which layer produced the improvement. Otherwise, credit may go to the model for a gain that depends on an integration or broader access.

The decisive test is the path, not just the answer

For a chatbot, checking the final answer may be enough for a bounded task. For an agent that opens files, consults an account and changes a document, the path must be evaluated too. A correct report does not compensate for reading an unnecessary folder, overwriting a useful version or sending information to the wrong destination. The unit of evaluation shifts from one answer to a sequence of decisions and actions.

OpenAI's own safety documentation makes that path relevant. The GPT-5.6 system card reports a greater tendency than GPT-5.5 to go beyond user intent in agentic coding tasks, while noting low absolute measured rates. It also says the system was trained to follow confirmation policies before certain computer actions. This does not prove ChatGPT Work will make a particular mistake. It is documented reason not to equate autonomy with unlimited authorization.

Before connecting a real account, draw an authority map. For each source, record what the agent may read, create, modify and send outside. Separate reversible actions, such as creating a draft or copying a spreadsheet, from actions that are difficult to undo, such as deleting, publishing, purchasing or messaging a third party. The first group can support more autonomy. The second needs explicit confirmation and a preview showing the object, destination and effect.

A pilot that exposes both capability and risk

A useful test begins with a small set of representative tasks and nonsensitive data. For office work, place fictional messages, a spreadsheet and a template in an isolated folder, then ask the agent to produce a report with calculations that can be checked. Before execution, write down the sources it needs, the files it must produce, the cells that should contain results and the actions it must not take. That sheet is the acceptance criterion. Without it, any polished output can look successful.

Preserve the trajectory during the test: files consulted, tool calls, changes made, confirmations requested and human corrections. Review four axes separately afterward: factual accuracy, task coverage, respect for the boundary and supervision cost. An agent that produces an excellent document but forces a reviewer to reconstruct its steps has not completed the whole job; it has moved part of the cost into the audit.

The comparison also needs a baseline. Run the same task through the current procedure or through a model without tools and measure time, errors and revisions. Repeat with each family member using identical data and permissions. The result can show whether Sol adds enough value on the difficult portion, Terra covers standard work, or Luna is sufficient for a quick transformation. That is an operating hypothesis, not a universal ranking: the vendor's label cannot replace measurement on the user's task.

How to read the launch figures

For API use, OpenAI set prices per million tokens at 5 dollars for input and 30 for output with Sol; 2.50 and 15 with Terra; and 1 and 6 with Luna. The figures appear in the official pricing and availability table. A valid comparison holds the complete task constant. A model with cheaper tokens may require more attempts, while a costlier model may save review time. The relevant total is inference price plus human time plus the impact of errors.

Benchmark results on the same page do not, by themselves, answer whether the agent is suitable for a company. Identify what was measured, which tools were available, the compute budget and the competing version. Then translate the result into a task-specific threshold that can be observed. A benchmark percentage is evidence about that protocol; it is not a promise that a report, presentation or automation will succeed on the first attempt.

Verbs are part of the evidence

The preview episode teaches another transferable skill: read the relationship between a company and a government precisely. “Informed,” “consulted,” “acted at the request of,” “received authorization” and “complied with an order” are not synonyms. The final two require an instrument or statement granting authority. Here, OpenAI documented conversations and a request. It also said it did not want that process to become the permanent launch model. Turning this record into a “green light” adds a power the source does not state.

The rule extends beyond this product. When a headline attributes a corporate decision to a regulator, find the primary document and identify the exact verb, the institution issuing it and its basis. If there is no license, order, ruling or unequivocal statement, describe the interaction that is documented and keep the limit visible. The absence of a document does not prove none existed, but it prevents presenting one as a verified fact.

What changes for the user

ChatGPT Work expands the work that can be delegated by combining generation with access to context and tools. That increases utility and also increases the damage an ambiguous instruction can cause. Responsible adoption is not a choice between total trust and rejection. It means granting the minimum authority needed, observing the trajectory, requiring confirmation for sensitive actions and expanding the boundary only when pilot evidence supports the change.

The lasting skill is concrete: evaluate an agent as a system made of a model, product surface, data, permissions, actions and verification. Separating those layers reveals what actually improved, what it costs and what can fail. The model name begins the evaluation; an authority map and a test on real work are what allow it to finish.

This article was produced with artificial intelligence under human editorial oversight.

Share this article

This website uses cookies to improve the browsing experience. Cookie policy.

↑↓ navigate ↵ open esc close