BTC $63,085 ▼0.50% ETH $1,879 ▲0.28% SOL $75.22 ▲ 0.50% XRP $1.02 ▲ 0.90%
Artificial Intelligence

Cerebras Powers GPT-5.6 Sol "Ultrafast" at 750 Tokens Per Second

Cerebras Powers GPT-5.6 Sol "Ultrafast" at 750 Tokens Per Second
📑 Table of Contents

What if a frontier model answered like a commodity API — so fast you stop noticing the wait? That's the pitch behind Ultrafast, a new preview tier in the OpenAI API that runs GPT-5.6 Sol on Cerebras hardware at up to 750 output tokens per second. OpenAI says the tier processes Sol up to 14x faster than its Standard tier, and Cerebras claims it does so with no quality compromise. Announced on August 13, Ultrafast is initially a limited preview for select customers, with access expected to expand as capacity grows.

The headline numbers are only part of the story. The bigger claim — and the one that has developers paying attention — is that speed no longer has to be traded against intelligence.

Eleven hours versus three days

The most visceral evidence comes from Humanity's Last Exam, the notoriously punishing benchmark of 2,500 questions designed to stump frontier models. Cerebras reports that an Ultrafast run worked through all 2,500 questions in 11 hours and 11 minutes, versus 78 hours and 27 minutes for Claude Fable 5 — roughly seven times faster, at comparable accuracy. In practical terms, a job that took more than three days now finishes in less than half of one.

For teams running long evaluation suites, that's not a nice-to-have; it's a different way of working. Model evals that were overnight batch jobs become something you can run between meetings, and iteration cycles that took a week can compress into a day.

A wafer the size of a dinner plate

The trick behind the speed is physical. Where conventional GPUs shuttle model weights between memory chips and compute cores — a constant bottleneck that grows with model size — Cerebras builds a Wafer-Scale Engine: a single silicon wafer, roughly the size of a dinner plate, with 44 GB of on-chip SRAM. That keeps the entire model's weights on the chip itself, eliminating the memory-bandwidth bottleneck that dominates GPU inference. No shuttling, no waiting on memory buses — the silicon just computes, continuously.

It's the same logic that has made Cerebras a fixture of the AI-infrastructure world: instead of networking thousands of small chips together, build one enormous chip that doesn't need to talk to anyone.

Where 750 tokens a second matters

Latency isn't an abstract metric in production. Cerebras and OpenAI point to a cluster of use cases where output speed is the difference between useful and unusable:

Early customers are already measuring the difference. John Crepezzi of Jane Street called the Ultrafast tier "impressive," while Podium's voice-AI product lead, Courtland Lykins, put it more directly:

"The speed completely changes the call experience."

For voice agents especially, the stakes are obvious: a model that thinks at 750 tokens per second can hold a conversation that feels human, instead of one that makes callers wait through awkward silences.

The numbers, carefully attributed

Speed claims in this space deserve precise attribution, because the benchmarks don't all use the same yardstick:

All of these are vendor-generated figures. Neither OpenAI nor Cerebras has published independently audited results yet, so the exact multiples will likely shift as third parties get access — but the direction of travel is consistent across every measurement.

The commercial stakes

The announcement cements Cerebras' position as an OpenAI inference partner, a relationship that has been quietly deepening since the two companies' January 2026 partnership around 750 MW of low-latency compute. Cerebras hardware already ran OpenAI's GPT-5.3-Codex-Spark at more than 1,000 tokens per second in February; Ultrafast extends that relationship from coding models to a general frontier model.

The competitive signal is aimed as much at the GPU-centric incumbents as at developers. If a frontier model can run at commodity-class latency on wafer-scale silicon, the argument that "you need GPUs for serious inference" starts to erode — and the economics of AI infrastructure get a lot more interesting. For now, the question is simple: once developers taste 750 tokens per second, what feels acceptable to go back to?

Sources

J

Jai

Jai covers trending tech, AI developments, and the cultural impact of emerging technologies at Veritya Daily. When he's not tracking viral stories, he's probably doom-scrolling through AI research papers.