Resources

The Month AI Models Had to Go to the DMV

By Ali Shan NasirJuly 20, 20267 min read


Something strange happened to AI this summer. The models got a lot better, and also you were not allowed to use half of them. In the United States, the best new models had to clear the government first, up to 30 days of national-security review before anyone could touch them. In China, a startup gave away the biggest model ever built, for free, and then had to turn people away because too many showed up. Launch day used to be a demo. In June and July of 2026 it looked more like a trip to the DMV.

Two things are true at once here, and you have to hold both. The capability jump is real. And the funniest story of the summer was not what the models could do, it was who was allowed to do it. You now wait in line for a frontier model twice: once for the government, and once for the GPUs.

The models that needed a permission slip

Start with the part nobody had a mental model for. Following an executive order in early June, the strongest new US frontier models got limited to government-approved customers during a national-security review that can run up to 30 days, per CNBC. OpenAI's flagship and Anthropic's Mythos, the two most hyped models of the moment, were the ones you literally could not just log in and use.

Then, on July 9, OpenAI launched the GPT-5.6 family, three models named Sol, Terra, and Luna, alongside a "ChatGPT Work" agent, with the flagship Sol pointed at high-end reasoning and coding, per TechCrunch. A frontier launch used to be a keynote and a big glowing "try it now" button. This one came with a review period. The button said pending.

The best line about all of it came from Fireship, and it is annoyingly accurate: frontier models now have to go to the DMV before you can drive them. If you want the vibe of the month in four minutes, this is it.

Fireship, 'OpenAI is so back... GPT 5.6 Sol first look.' Irreverent, fast, and the source of the DMV joke.

Under the comedy, the thing is genuinely more capable. Here is OpenAI's own short demo of GPT-5.6 building a working card game from a single prompt. Watch it and then remember that the headline was not the demo, it was the waiting room.

OpenAI's official 'Meet GPT-5.6' demo, building a card game from a prompt.

Meanwhile, China gave the biggest model ever away for free

While the US was inventing a permission slip, a Beijing lab called Moonshot AI did the opposite. It shipped Kimi K3, the largest open-weight model ever released: 2.8 trillion total parameters as a mixture of experts, 896 experts with 16 active at a time, a one-million-token context window, and native multimodal understanding, per Tom's Hardware. Open weights. Download it. No permission slip.

And it was not a vanity release. In a blind Frontend Code Arena eval it landed at number one, ahead of Anthropic's Claude Fable 5, though Moonshot itself is honest that it still trails Fable 5 and GPT-5.6 Sol on overall performance, per Simon Willison. The catch was never money. It was capacity. Demand was so heavy that Moonshot paused new signups to protect compute for the users it already had, which Axios reported the same week. A free model, sold out.

SemiAnalysis on the Kimi K3 release and what a free, giant, open model means.

The most-hyped US models were the ones you could not use without a government sign-off. The biggest model on earth was free to anyone, until the servers melted. That was the whole month.

What people actually made with them

While the internet argued about permission slips, everyone else just used the things, and some of what came out was genuinely wild. The clearest example came straight from Kimi K3. People handed it a rough sketch and it wrote a full, playable 3D open-world game that runs in the browser, built on Three.js and WebGPU, with a procedurally generated world and even a rider on a horse. The part that made people sit up was that when the game threw runtime errors, the model read them and fixed its own code without being asked, per SiliconANGLE.

AI IXX: Kimi K3 generating a playable browser 3D game and then fixing its own runtime bugs.

And it was not one lucky clip. For a week, X filled up with people handing Kimi K3 a single prompt and getting back things that should not fit in one prompt. Someone built a working game console and played Ace Attorney on it. Someone else had it design a chip, make a 3D video about it, cut a teaser trailer, and run a playable emulator inside a generated 3D model, all in one go.

Crystal: a working game console built with Kimi K3 in a single prompt.
MTS: Kimi K3 designing a chip, making a 3D video, cutting a trailer, and running an emulator.

On the OpenAI side the viral moment was less a single demo and more a wave of people posting one-shot builds, whole working web apps from a single prompt, that would have fallen apart on the model from a month earlier. And the ChatGPT Work agent turned "do this task" into finished deliverables: real spreadsheets, slides, and working apps handed back after hours of work, instead of a chat transcript. A lot of people spent launch week trying to delegate their actual job to it.

Tech Friend AJ putting the ChatGPT Work agent through a real workload.

None of this needed a 47-tweet thread. Someone pointed a new model at real work and it did more than the last one could. That is the whole event, every time.

The discourse ran its usual play

Every time a model drops, the internet runs the same script, and it fits a very old shape. The people just discovering it and the people deepest in it agree on the same boring thing: it is better, use it. The people in the exact middle write the 47-tweet thread about why it does not really count.

The midwit take on every new AI modelnew model better.i use it.new model better.i use it.you can't even USE it yet,it's in a 30-day federal review,Kimi's 2.8T only runs 16 of 896,anyway... thread 1/47

This summer the midwit had extra ammunition, and all of it was technically true. You could not use the US model yet. The Chinese one only lights up 16 of its 896 experts at a time. The benchmarks might be contaminated. The geopolitics of the compute moat, thread one of forty-seven. All real, all beside the point, which was that the thing got noticeably better and you could feel it the moment you used it on something you actually cared about.

Strip away the drama

Underneath the executive orders and the sold-out free model, the trend is boring and clear: the models keep getting more capable, faster than the takes can keep up. That is the signal. The permission slips and the melted servers are the noise of a technology arriving faster than the world built to handle it.

The part that does not change is what you do with it. A better model is a better engine, and an engine is not a car. It still needs everything built around it, the part that survives real users, real data, and the worst day of the quarter, which is the unglamorous work I wrote about in shipping AI agents that actually work. The model getting better makes that work more valuable, not less.

So here is the only advice that survives a month this weird. The models will keep going to the DMV, and the servers will keep melting, and the thread will keep getting longer. Ignore all of it. Go try the new thing on your own real work. As someone put it on Reddit while Kimi's servers were on fire, the ultimate benchmark is just usage. If the model is good, you will use it. Everything else is a waiting room.

Further reading

The sources behind the month, if you want the primary accounts.

Weighing an agent or an AI system of your own? Tell us what it would do and what it costs you when it gets something wrong. We will give you a straight read on whether it is worth building.

Talk through a project