AI is about to get fast, and it’s never going to slow down
On June 9, Anthropic shipped Claude Fable 5, and every scoreboard agreed at once. #1 on Artificial Analysis.
On June 9, Anthropic shipped Claude Fable 5, and every scoreboard agreed at once. #1 on Artificial Analysis. #1 on LMArena’s text, web-dev, and agent arenas. First on SWE-Bench Pro. More than double the previous best on Cognition’s FrontierCode. Andrej Karpathy called it “SOTA on everything by a margin.”. He would say that, he joined Anthropic in May, but the independent numbers back the brag. The last model to sweep everything was GPT-4, and that lead lasted a year.
Now read the other column. The most powerful model the public has ever been handed ranks #64 of 152 on output speed: 60 tokens per second, and a 108-second wait for the first one, against a field median under three seconds. On price it sits at #139 of 152. Simon Willison called it “something of a beast. It’s slow, expensive” and, $110.42 of tokens later, capable.
For three years, model intelligence was the bottleneck, so we bought it with time. That trade just closed. The scarce resource now is the hour, and every lab on earth is retooling to win it. From here forward, the speed column picks the winners.
The only question was “is it smarter?”
Since ChatGPT arrived in late 2022, you have judged every release the same way. You opened the benchmark table, found the bold column, and asked whether the new model could do something the old one couldn’t. Speed was the discount tier. “Mini,” “flash,” and “turbo” were polite words for cheaper and worse, the variants you routed boring traffic to while the real model handled everything that mattered. The waiting was rational: a model that took 90 seconds and got the architecture right beat one that answered instantly and confidently wrong. Three and a half years of that trained everyone to treat tokens per second as an implementation detail.
Then the models stopped failing at the work.
The bottleneck moved
The bottleneck did not move on June 9. It had been moving for months. Something flipped in December, I argued back in February, when the models picked up the long-term coherence and tenacity to hold hours-long agentic sessions in Claude Code, Codex, and other harnesses without falling apart. Fable 5 raised the ceiling, and the early testers hit the same two notes. Ethan Mollick, the Wharton professor who got early access, called it “a genuine jump in capability: I could feed it a 15 page design document for a project and it would work for 9+ hours and deliver terrific results.”. Dan Shipper, whose team at Every tested it for a week, found it “routinely uses 500k to 1M tokens on tasks” and is “very slow, token-hungry. Using this thing for regular knowledge work is like squashing an ant with a rocket launcher.” The question that defined the last three years, “can the model do this at all?”, is retiring. The sentence replacing it is half brag, half complaint: “my agent opened a genuinely impressive refactor PR, and it took 20 hours.”
Engineers inside the labs feel the inversion too. Sherwin Wu, an engineer at OpenAI, wrote in February that GPT-5.3-Codex
Intelligence leads, meanwhile, stopped lasting. GPT-4 held the top of the leaderboard for about a year. Stanford’s AI Index has six labs within 79 Elo points, and the best Chinese model within 2.7% of the best American one despite a 23-to-1 US investment advantage. A capability lead is a wasting asset measured in months. Speed is the axis nobody has saturated.
The speed race has already started
Google’s fast model beats Google’s smart model
Google led its newest generation with the speed tier. Gemini 3.5 Flash, shipped at I/O in May, outperforms Google’s own frontier model on most agentic benchmarks while running, in DeepMind CTO Koray Kavukcuoglu’s words, “four times faster than comparable frontier models.” It costs triple what its predecessor did. This week, Xiaomi announced MiMo-V2.5-Pro-UltraSpeed: a trillion-parameter model that Xiaomi says clears 1,000 tokens per second on a single 8-GPU node, priced at 3x the standard rate for roughly 10x the speed. The launch tweet:
DeepSeek, Alibaba, and Z.ai shipped fast variants the same quarter. The fast tier used to be where intelligence went to die. This spring it became the marquee release.
Paying more for the same model, sooner
In February, Anthropic launched fast mode for Claude Opus at six times the standard price, since cut to 2x. The docs are explicit that the only thing you are buying is time: “Fast mode is not a different model… You get identical quality and capabilities with faster responses.” OpenAI sells the same inversion: GPT-5.5 spans a 5x price range from Flex to Priority with the model held constant. Cursor prices its coding agent the same way: Composer’s fast variant has identical intelligence at six times the price per token. Every pricing page you have ever seen sold the premium tier on more: more features, more seats, more storage. Premium AI in 2026 sells the same capability, sooner.
Thariq, an engineer on the Claude Code team, framed the fast product on launch day: “Use it when you want to be locked in, like when iterating on a design or fixing an incident vs multi-Clauding in the background. It’s more expensive because it uses more compute, but it’s the exact same intelligence.”
Twenty billion dollars says the next constraint is silicon
In December, Nvidia paid a reported $20 billion, the largest deal in its history, to license Groq’s low-latency inference technology and hire its founder. By March it had canceled a GPU it announced six months earlier and put a Groq-derived inference rack in its place. OpenAI signed a reported $10 billion deal for 750 megawatts of Cerebras wafer-scale systems dedicated to low-latency inference; Cerebras went public in May and jumped 70% on day one. OpenAI’s Sachin Katti, announcing the deal: “Cerebras adds a dedicated low-latency inference solution to our platform. That means faster responses, more natural interactions, and a stronger foundation to scale real-time AI to many more people.” A Toronto startup called Taalas etched a compressed 8B Llama model directly into silicon and demoed it at a claimed 17,000 tokens per second. Google markets its newest TPU as its first built “for the age of inference”; Microsoft brands Maia 200 “the AI accelerator built for inference.”
The talent market moved with the money. The 2025 poaching war was over researchers, with $250 million packages on the table. The marquee poach of June 2026 was a chip engineer: Anthropic hired one of the first engineers OpenAI ever put on its custom silicon program, while Apple pays emergency retention bonuses to keep its hardware people away from OpenAI. The engineer Anthropic poached describes his new role as “perplexity per picojoule.”
That’s a job title now.
Why speed compounds
Cerebras ran the same agentic request against Kimi K2.6 twice: 163.7 seconds through the model’s official endpoint, 5.6 seconds on its wafer-scale silicon. Same weights, same answer, 29 times sooner. That is how much speed is already on the table without anyone training a smarter model.
Then there’s the loop that decides the race: AI is starting to build AI. Five days before Fable 5 launched, Anthropic published When AI builds itself, reporting that Claude now authors more than 80% of the code Anthropic merges, and concluding that in a world of automated research, the pace of progress “becomes determined entirely by the availability of compute (or the speed of discovering various efficiencies in algorithmic training or inference).” Dario Amodei put it more plainly in a June 2026 policy essay: “The iterative ability of AI to build even better AI may supercharge that growth even further.” In May, Anthropic hired Andrej Karpathy onto pre-training specifically to use Claude to accelerate the research that builds the next Claude. If that loop is real, the lab whose researcher-agents finish in one hour instead of twenty takes twenty times as many turns at it.
Run the arithmetic on that. Two labs start the year with the same model: same benchmarks, same floor, same ceiling. One closes a research loop in an hour, the other in twenty. A week in, the score is 168 experiments to 8. And the cycles stack, because every loop ends with a slightly better model running the next one: by Friday the fast lab is iterating with a model the slow lab won’t meet until spring. Hold that pace for a quarter and the two labs no longer share a ceiling. Same weights in January. Same ideas, same talent. The clock made one of them the frontier lab and the other a fast follower.
That is the leapfrog. A speed lead at equal intelligence converts into an intelligence lead, and it keeps converting for as long as the loop runs. It also reprices the silicon deals: twenty billion for Groq and ten billion for Cerebras is what turns at the loop cost in 2026. Speed stops being a user-experience feature and becomes the rate constant on the most important feedback loop in the industry.
By the end of 2026
Tokens per second becomes a launch-day headline number. Sundar Pichai already quoted one on stage at I/O. Speed guarantees become contract terms: Azure already sells a tier promising 99% of GPT-5.5 requests above 100 tokens per second.
Speed may also turn out to be the first moat in this industry that holds. Intelligence crosses the Pacific in months; Epoch AI puts the open-weight lag at four months. Speed hasn’t crossed: the same DeepSeek V4 weights anyone can download run about 48 tokens per second on DeepSeek’s own API and three times faster on Nvidia’s newest racks. Benchmark scores can be distilled across a border. Tokens per second have to be manufactured on one side of it. And the lab that holds the speed moat re-mints its intelligence lead every cycle the loop turns.
Anthropic is so capacity-constrained right now that it will pull Fable 5 from subscription plans on June 22, thirteen days after launch. That is what it looks like when demand for intelligence outruns anyone’s ability to serve it. The 20-hour PR is an artifact of this exact moment: models smart enough to finish the work, infrastructure too slow to finish it while you’re still at your desk.
Watch the other column
Anthropic just proved it can build the smartest model in the world: first out of 152. It ships at 64th in speed. Somewhere in the gap between those two rankings is the next phase of this industry, and every lab, chipmaker, and sovereign wealth fund can see it. The next launch that matters may not move the intelligence leaderboard at all. It will decide who moves it next.
When it lands, skip the benchmark table. Look up the tokens per second.