GPT-5.6

GPT-5.6 Sol Ultrafast: Cerebras Pushes Inference to 750 Tokens per Second

GPT-5.6 Sol Ultrafast, powered by Cerebras wafer-scale hardware, hits 750 output tokens per second. The hardware, the HLE and GDP-Val results, and what real-time speed unlocks for agents.

GPT-5.6 Sol Ultrafast: Cerebras Pushes Inference to 750 Tokens per Second — article cover

Most model announcements are about being smarter. The GPT-5.6 Sol Ultrafast service tier that Cerebras and OpenAI announced on August 13, 2026 is about something else: being faster. The new tier, launching first in the OpenAI API and powered by Cerebras hardware, pushes GPT-5.6 Sol to up to 750 output tokens per second — with, per the announcement, “no quality compromise.”

For builders of real-time applications, speed is a new capability in itself. When response time compresses from minutes toward real time, product shapes that were impractical suddenly work.

What 750 Tokens per Second Means

Against output speeds tracked by Artificial Analysis, GPT-5.6 Sol Ultrafast runs 11x faster than Anthropic’s Claude Fable 5 and 5x faster than Opus 4.8 on Fast mode. At 750 tokens per second, a 3,000-token response finishes in about four seconds — inside the latency range that feels conversational.

Cerebras is careful about the framing: the point is not a peak-speed number game, but frontier intelligence delivered at real-time speed for the first time. Capability and inference speed used to be separate roads; they are starting to merge.

Wafer-Scale Engineering: Keep the Weights On-Chip

The speed comes from the hardware architecture. Cerebras’ Wafer-Scale Engine puts 44 GB of SRAM on each wafer-sized chip. Model weights stay on-chip, and tokens flow through the layers pipelined across wafers. The bottleneck in conventional GPU inference is shuttling weights in and out of HBM memory; if the weights never leave the chip, that bottleneck simply disappears.

Cerebras also says the architecture scales smoothly as models grow — a promise about future generations that is worth tracking, not taking on faith.

The Numbers: 7x on HLE, 5.6x on GDP-Val

Two measured results stand out. First, on Humanity’s Last Exam, Ultrafast completed all 2,500 questions in 11 hours 11 minutes, versus 78 hours 27 minutes for Claude Fable 5 — roughly a 7x speedup at comparable accuracy (tested through Codex and Claude Code at xhigh reasoning, July 10–15, 2026). Second, on GDP-Val it delivered a 5.6x end-to-end speedup with no quality degradation.

“Comparable accuracy” is the load-bearing phrase: this is not speed bought with distillation or a smaller model. It is the same model running on hardware where bandwidth is no longer the constraint.

Real-Time Agents Are the Big Winner

The use cases Cerebras names are pointed: root-causing production outages, responding to cyberattacks in real time, agent workflows that need live interaction, and legal, financial, and engineering work where response time carries direct value. The common thread is “cannot wait” — every extra minute of latency is a measurable loss.

For now, Ultrafast remains a limited preview open to a select group of customers, expanding as capacity grows, and pricing has not been announced. For product planners, the useful move today is identifying which of your workflows are latency-sensitive, so migration is straightforward when access opens up.

Sources

AI-assisted summary compiled from the sources above, reviewed by a human before publishing.

FOUND_THIS_USEFUL?

Support more practical AI articles, tutorials, and build notes.

BUY_ME_A_COFFEE
SHAREXEMAIL