NVIDIA

Akamai Buys Thousands of Blackwell GPUs for Edge Inference

Akamai acquired thousands of NVIDIA Blackwell GPUs for a distributed inference platform spanning 4,400+ edge locations, claiming up to 2.5x latency gains and up to 86% savings versus hyperscalers.

Akamai Buys Thousands of Blackwell GPUs for Edge Inference — article cover
On this page6 SECTIONS
  1. The Inference-Era Thesis
  2. Anatomy of a “Decentralized Nervous System”
  3. The Numbers: Up to 2.5x Latency Gains, Up to 86% Cheaper
  4. From CDN and Linode to the Inference Cloud
  5. What It Means for Developers and Platform Teams
  6. Sources

On March 3, 2026, Akamai (NASDAQ: AKAM) announced it has acquired thousands of NVIDIA Blackwell GPUs and will spread them across its global distributed cloud, aiming to build one of the world’s most widely distributed AI platforms. The story is not about training. It is about inference.

The CDN-turned-cloud company is betting on a specific thesis: the center of gravity in AI is shifting from centralized training in a handful of mega data centers toward inference that runs close to users. Akamai cites MIT Technology Review finding that 56% of organizations see latency as the primary barrier to deploying AI at scale. Data Center Dynamics followed the next day with one more detail: the exact GPU count was not disclosed.

The Inference-Era Thesis

Akamai’s argument is blunt. The first wave of AI concentrated training in a few hyperscale regions. The industry has now hit the point where inference matters as much as training. For models to leave the lab and enter autonomous delivery, smart grids, surgical robotics, and real-time fraud prevention, decisions have to happen at real-world speed — and the latency and egress cost of round-tripping a distant hyperscaler region becomes the new bottleneck.

Adam Karon, Akamai’s COO and general manager of its Cloud Technology Group, frames it this way: hyperscalers will keep pushing the frontier of training, but Akamai is focused on the “inference era.” Centralized AI factories remain necessary; what scaling AI into applications needs is a “decentralized nervous system” that puts inference compute at “the street corner and the hospital bed: where the work is happening, where the data lives, and where the ROI is realized.”

Anatomy of a “Decentralized Nervous System”

The concrete stack: NVIDIA RTX PRO servers running RTX PRO 6000 Blackwell Server Edition GPUs, paired with BlueField-3 DPUs, on top of Akamai’s own distributed cloud and a global edge network spanning more than 4,400 locations. The platform is built to do three things: run predictable, high-performance inference on dedicated GPU clusters; fine-tune LLMs locally so data stays in-region for privacy and compliance; and run post-training optimization on private datasets.

In other words, this is not a matter of racking GPUs in a warehouse. It is an orchestration system that treats model placement as a variable, intelligently routing inference workloads to the location nearest the user.

The Numbers: Up to 2.5x Latency Gains, Up to 86% Cheaper

Akamai’s headline claims: versus traditional hyperscaler infrastructure, running inference on NVIDIA AI infrastructure reduces latency by up to 2.5x and cuts costs by as much as 86%. Treat both as marketing until you benchmark your own workload — the baselines and traffic patterns are not published. But the direction is credible: inference rewards proximity and minimal egress, and moving bits closer to users is the business a CDN has run for two decades. The only new part is moving the GPUs along with them.

The demand signal matters too. Akamai says its initial RTX PRO 6000 Blackwell deployments have seen “strong demand,” and it will keep adding GPU capacity. Edge inference is no longer a proof of concept; it is a product with a queue.

From CDN and Linode to the Inference Cloud

The groundwork was laid years ago. Akamai bought cloud provider Linode in 2022 for roughly $900 million to build out IaaS capabilities, then announced Akamai Inference Cloud in October 2025 to move inference workloads closer to users. This large-scale Blackwell buy extends the same line: CDN-grade global connectivity and caching, layered with GPU clusters and DPUs, forming an alternative architecture that competes head-on with hyperscalers for inference demand.

What It Means for Developers and Platform Teams

Three observations. First, inference is getting its own vendor spectrum: training is a battle between hyperscalers and sovereign clusters, but inference can be carved up by distributed-network players like Akamai — especially latency-sensitive, data-residency-sensitive workloads such as agentic AI, physical AI, and compliance-driven local fine-tuning. Second, for architects this adds an option in the model-placement era: LLM routing will increasingly mean choosing not just a model but an inference location. Third, cost structures shift — the 86% saving will not land wholesale on your bill, but egress fees and latency penalties are two of the most invisible lines in current AI bills, and both are worth testing yourself before committing.

Sources

AI-assisted summary compiled from the sources above, reviewed by a human before publishing.

FOUND_THIS_USEFUL?

Support more practical AI articles, tutorials, and build notes.

BUY_ME_A_COFFEE
SHAREXEMAIL