Omi Iyamu · Personal DossierVol. XVII · 2026 Edition
Omi Iyamu.
← All essays
2026 · 07 · 094 min read

OpenAI ships GPT-5.6 to the public after 30-day US cyber review

# GPT-5.6 shipped, and the interesting story is not the benchmark

OpenAI released GPT-5.6 to the public today. Three models — Sol, Terra, Luna — priced at a spread that finally makes the routing question honest.

Sol is the frontier tier. It ships with two reasoning-effort settings and a new Ultra mode that spawns subagents to work a task in parallel. On Terminal-Bench 2.1 — a six-week-old benchmark of iterative command-line workflows — Sol Ultra scored 91.9%, plain Sol scored 88.8%. For comparison, GPT-5.5 was 88.0%, Claude Mythos 5 was 84.3%, Claude Opus 4.8 was 78.9%, Gemini 3.1 Pro Preview was 70.7%. Sol lists at $5 per million input tokens, $30 per million output.

Terra is the everyday tier. Same claimed performance as GPT-5.5 at roughly half the cost. $2.50 in, $15 out.

Luna is the cheap tier. $1 in, $6 out.

Three things about this release actually matter.

First: the Ultra-versus-Sol gap. The cleanest evidence yet that spending more inference to spawn coordinated subagents beats letting one big model think harder. That is a 3.1-point spread on a benchmark the industry cares about, and it is the pattern I have been telling teams to structure around for six months. If you have a Sol subscription and you are still calling it as a single agent, you are leaving accuracy on the table. Fair warning: Ultra is expensive per call. You should not be running it against your top-200 queries unless you know why.

Second: the price ladder. This is the release where OpenAI stopped pretending everyone should pay frontier prices. Terra at $2.50/$15 is priced to be the default. Luna at $1/$6 is priced to be the answer to 'why are you calling Sol for that.' I said in Brief 45 that most teams are paying frontier prices for their third tier. Terra and Luna make it much harder to keep doing that with a straight face. If your stack does not have a router that can pick between three tiers on the same provider, you have a rewrite to schedule.

Third — and this is the actually interesting one: GPT-5.6 is the first model of consequence to ship through the Trump administration's June 2 cybersecurity executive order. The EO asked labs to voluntarily submit their most powerful models for a 30-day cyber review before public release. OpenAI complied. The model sat in limited preview with government-vetted partners for roughly two weeks. Today, the review closed and the model shipped.

I want to be careful here. This is a voluntary framework, and the review looks light-touch compared to what an earlier draft of the EO reportedly proposed. But it did happen. A frontier lab held a launch for a US government cyber review, and the review took time, and the release calendar shifted around it. That is a fact about how AI companies operate now that was not true in April.

The precedent matters more than the mechanism. Every subsequent frontier release from a US lab is now measured against 'did they submit or not.' Sonnet 5 shipped on June 30 without any pre-release review; that is fine — Sonnet is not the frontier tier. But the next Opus, the next Gemini Ultra, the next Grok top model — they will all have to answer this question. Some will submit. Some will not. The ones that do not will get asked why, publicly, on the record. That is a shift.

For engineering leaders, the short version:

- If you are on Sol, decide whether Ultra earns its price on your workload. Run the eval. Do not vibe it. - If you are on GPT-5.5, Terra is your migration target, not Sol. Do the migration before your next billing cycle. - If you are using a frontier model for classification, extraction, or short-answer completion, Luna is the price benchmark you now have to beat. Beat it or move. - If you sell into US regulated markets, add 'was this model submitted for the cyber review' to your vendor questionnaire. It is a real signal.

One thing I will be watching over the next two weeks: whether Ultra mode's numbers hold up on evals people write themselves. Terminal-Bench 2.1 is a benchmark from six weeks ago. Subagent architectures are exactly the kind of thing that overfits to a fresh benchmark. If Ultra is 91.9% on Terminal-Bench and 78% on your team's boring-cases eval, you learned something important — but only if you ran the eval.

I will write more on the EO angle in Brief 48. If you are navigating a submission-review process for a model of your own, or trying to figure out how to describe your model's cyber posture to a customer, my inbox is open. This is going to be a busy quarter for that conversation.

If this was useful, the weekly Brief covers shorter ideas like this every Wednesday.
Read the Briefs →
© Omi Iyamu · MMXXVIContact → · linkedin.com/in/omiiyamu