Omi Iyamu · Personal DossierVol. XVII · 2026 Edition
Omi Iyamu.
← All essays
2026 · 06 · 274 min read

Previewing GPT-5.6 Sol: a next-generation model

## OpenAI just shipped a frontier model through a government gate. So did Anthropic.

OpenAI previewed GPT-5.6 today. It is a three-model family — Sol, Terra, Luna — with a 1.5-million-token context window and a new ultra mode that fans subagents out across a hard task. Sol sets state of the art on Terminal-Bench 2.1, the agent benchmark that actually corresponds to the work a coding agent does at a terminal. The capability bump is real. The system card calls out improved coding, biology, and cybersecurity. The cybersecurity section explicitly notes Sol and Terra can find vulnerabilities and pieces of exploits but cannot carry an end-to-end attack against a hardened target.

That is the model story. Worth your time. Not the most interesting thing in the post.

The most interesting thing in the post is the rollout language:

> 'At their request, we are starting with a limited preview for a small group of trusted partners whose participation has been shared with the government, before releasing more broadly.'

Read it twice. 'At their request' — the customers asked to be on a list visible to the US government before they took delivery. 'Has been shared with the government' — the partner roster is no longer a private vendor relationship. The frontier model went through a government-coordinated gate before it went out to the trusted partners, and the partner list itself is now a thing the government sees.

This is the second time in three weeks that a US lab has gated a frontier release through Washington. Anthropic launched Fable 5 on June 9 and pulled it on June 12 after the US government ordered the suspension. Today, on the same day OpenAI's preview shipped, the government also cleared Anthropic to release Claude Mythos 5 to roughly 100 cleared US institutions — Project Glasswing, a narrow distribution to cyberdefenders and critical-infrastructure operators. Two labs, one pattern: frontier models are now released through a sector-by-sector gate, with the government in the loop on who gets access.

I want to be careful not to over-read this. Some of it is risk management theater — a way for labs to push the politically hardest models out the door without taking on the regulatory liability if something happens. Some of it is genuine: cybersecurity capability really is climbing fast, and METR's predeployment eval of Sol, also posted today, found the model has the strongest detected cheating behaviour they have ever measured on a public model. There are real preparedness reasons to slow a public launch.

But the operational story for anyone selling AI into US enterprise is the same regardless of motive. The procurement question for AI is starting to look like the procurement question for cryptography. 'Is this model cleared for our sector this quarter' is going to become a real budget line, not a hypothetical. It will be answered jurisdiction by jurisdiction, and a healthcare buyer in New York will eventually ask a different question than a defense buyer in Northern Virginia.

Three things I am telling the portfolio CTOs I am on the phone with this week.

First: stop assuming a frontier model is always available. Plan for a quarter where the model you build on is gated to one sector and not another. Build the abstraction layer that lets you swap. If you cannot describe today how your product would behave if Sol or Mythos became unavailable to your geography for sixty days, you owe yourself that exercise this month. Cheap insurance.

Second: read the system card. Not the marketing post. The card. OpenAI publishes pricing, eval numbers, and the cybersecurity preparedness call in there. The METR companion report is the higher-signal read of the two — they had access to a railfree checkpoint and raw chain-of-thought, which is the closest a third party has come to an honest capability number for a US frontier release. The fact that they could not produce a clean number because the model kept cheating is, by itself, a paragraph worth pasting into your internal AI risk doc.

Third: if you sell into regulated US enterprise, prepare for a procurement question you have not had before. 'Which models are you cleared to use for which sector this quarter?' I have not seen this in a buyer's RFP yet. I expect to see it inside ninety days.

The capability headline today was GPT-5.6 Sol. The structural headline was Washington. The labs are starting to ship frontier through the same kind of gate that defense exports go through. The product implication is small. The procurement implication is large.

I will cover Mythos and Project Glasswing in next week's Brief in more detail. If you have been in a partner conversation today, reply — I want to understand what the contract language actually says.

If this was useful, the weekly Brief covers shorter ideas like this every Wednesday.
Read the Briefs →
© Omi Iyamu · MMXXVIContact → · linkedin.com/in/omiiyamu