On 9 July 2026, following a Sol preview on 26 June, OpenAI announced GPT-5.6 Sol, Terra and Luna. OpenAI moved GPT-5.6 from limited Sol preview to general availability as a three-model family spanning flagship to low-cost work.
This release brief was checked against first-party material on 10 August 2026. The date above is the public announcement date, not the date a repository was created or a third-party provider added the model. Where access or weights arrived later, that distinction is recorded below.
Release record
| Field | Verified detail |
|---|---|
| Announcement | 9 July 2026, following a Sol preview on 26 June |
| Availability or weight release | 9 July 2026 across ChatGPT, Codex and the OpenAI API |
| Release type | three-tier proprietary frontier-model family |
| Access | hosted OpenAI products and API; no open weights |
| Architecture | an undisclosed model family with Sol, Terra and Luna tiers plus max and multi-agent ultra effort |
| Maximum stated context | the tier-specific limits documented in OpenAI platform materials |
What changed
GPT-5.6’s release argument is performance per completed task. Sol set strong scores across coding, browsing, computer use and cybersecurity while using fewer tokens than several comparisons. Terra and Luna move much of that agent capability into lower price tiers. Programmatic Tool Calling lets the model filter and transform tool outputs in memory, and ultra coordinates four agents by default, so developers must record whether a score came from one model call or a multi-agent budget.
The practical comparison is therefore not simply whether GPT-5.6 Sol, Terra and Luna has the largest headline score. Teams need to ask whether its architecture, access terms, latency, tool behaviour and evaluation setup match the workload they actually intend to run. A model can lead one harness while losing on cost, refusal behaviour, multilingual quality or repeatability in another.
Benchmarks worth retaining
| Evaluation | Reported result | How to read it |
|---|---|---|
| Terminal-Bench 2.1 | 88.8 Sol max; 91.9 Sol ultra | Single-model and multi-agent settings are distinct |
| BrowseComp | 92.2% for Sol | Agentic browsing result at launch |
| ExploitBench | 73.5% versus 47.9% for GPT-5.5 | Controlled cyber capability comparison |
These are release-time results, not independently reproduced guarantees. OpenAI’s tables distinguish reasoning effort and ultra multi-agent runs; prices also changed on 30 July for Terra and Luna, after the launch figures were published. Scores should remain attached to the disclosed effort setting, agent harness, tool access, timeout, context-management policy and judge model. Moving a number into a procurement sheet without those conditions creates false comparability.
Architecture and access
GPT-5.6 Sol, Terra and Luna is described as an undisclosed model family with Sol, Terra and Luna tiers plus max and multi-agent ultra effort with the tier-specific limits documented in OpenAI platform materials of stated context. Its access position at verification time is hosted OpenAI products and API; no open weights. That wording matters: open weights, source-available weights, an API, a product preview and a research demonstration give adopters very different rights and different levels of reproducibility.
Before deployment, record the exact model identifier or checkpoint, inference stack, quantisation, reasoning setting, region, price schedule and supplier terms. If the release uses a custom licence, read the licence itself rather than relying on the word “open” in launch copy. If it is API-only, preserve the dated documentation and change-notice route because the served snapshot can change without a downloadable artefact.
What an evaluation should test next
For GPT-5.6 Sol, Terra and Luna, a credible internal gate should include:
- a frozen set of representative tasks with pass, fail and abstain criteria;
- a matched baseline using the same tools, timeout, prompt budget and reviewer rubric;
- repeated runs to expose variance rather than reporting a single best attempt;
- latency, token use and total task cost alongside task success;
- adversarial, multilingual and long-context cases relevant to the real deployment; and
- rollback evidence showing the previous model can be restored safely.
The wider model change-control guide explains how to keep model, prompt, tool and corpus changes reconstructable. The AI dependency inventory guide covers the release and supplier records needed after deployment.
AIEngine verdict
GPT-5.6 is the window’s central proprietary model-family launch. Evaluate each tier separately and preserve the post-launch price update rather than assuming Sol results or July 9 economics apply to all three.
This is a launch assessment, not a certification. Benchmark leadership is useful evidence of where to test; it is not authorization to place the model in a high-impact workflow without domain evaluation, security review and an accountable owner.
Primary sources
- OpenAI GPT-5.6 general-availability announcement
- OpenAI GPT-5.6 Sol preview
- OpenAI GPT-5.6 system card
Image provenance
Hero image: OpenAI official release artwork. The locally served WebP is a crop of the first-party release or model-card asset recorded in the repository provenance manifest.



