LLM

DeepSeek Makes Its 75% V4 Pro Price Cut Permanent

DeepSeek confirmed May 22, 2026 its 75% V4 Pro discount is permanent: $0.87 per million output tokens, a quarter of list price; Flash is $0.28. What it says about API economics.

DeepSeek Makes Its 75% V4 Pro Price Cut Permanent — article cover

On May 22, 2026, DeepSeek announced on X that the 75% discount on its flagship V4 Pro model “is now permanent.” The promotion, originally set to expire at 15:59 UTC on May 31, will simply not end: official pricing is being adjusted to one quarter of the list price. Bloomberg followed the next day under the headline “DeepSeek to Make Permanent 75% Discount on Flagship AI Model,” framing the move as pricing strategy rather than a marketing sprint.

This is not an ordinary price cut. According to DeepSeek’s pricing documentation, the discounted output price works out to $0.87 per million tokens; making it permanent retires the list price for good. When a lab known for open weights turns a promotional price into the official one, the action itself is the signal.

The Promotion That Refused to End

DeepSeek’s pricing page is unusually blunt: the 75% V4 Pro promotion ends May 31, 2026 at 15:59 UTC, and “after the promotion ends, Pro pricing will be officially adjusted to one quarter of the original price.” In other words, DeepSeek is not extending a sale — it is converting the discounted price into the standing price. When the promotion expires, nothing changes on the invoice.

Developer reaction was immediate. A Hacker News thread linking straight to the pricing docs drew 621 points and 549 comments, revolving around two questions: how long can this price hold, and do competitors have to follow?

The Numbers on the Table

Three figures from the May pricing page are worth remembering:

  • V4 Pro: $0.435 per million input tokens (cache miss), $0.87 output, cache hits as low as $0.003625, with a concurrency cap of 500
  • V4 Flash: $0.14 input, $0.28 output, cache hits at $0.0028, concurrency cap 2,500; both models offer 1M context
  • As of April 26, cache-hit pricing had already dropped to one tenth of launch pricing

Working backward from the discount, V4 Pro’s list output price was $3.48 per million tokens; the permanent cut takes it to a quarter of that. Stack the one-tenth cache-hit pricing on top and the marginal cost of cache-heavy workloads approaches rounding error — for agent workloads that re-read the same system prompts and documents every turn, this is a deliberately built sweet spot.

Why DeepSeek Can Afford Permanence

Making a discount permanent usually means one of three things: costs genuinely fell, demand needs stimulation, or both. DeepSeek’s case touches all three. Open-weight models build their ecosystem on API revenue, and price is the most direct lever for winning developers; a cache-friendly inference architecture makes the unit economics plausible at these levels.

The more telling difference is in supply structure. Pro carries a concurrency cap of just 500, versus 2,500 for Flash. The cheap, high-concurrency Flash absorbs scaled traffic, while Pro keeps the flagship badge at a low price but with controlled supply. This is capacity management, not charity. Teams that put critical services on Pro will hit the concurrency cap long before they feel the price.

What It Means for Developers and Product Teams

First, budgeting gets simpler. The uncertainty of “will the promo end?” is gone, and long-horizon cost models can be built directly on the standing price. Second, the room for model routing widens: the output-price gap between Flash and Pro ($0.28 versus $0.87) comfortably supports a tiered architecture that sends easy work to Flash and escalates hard problems to Pro. Third, this puts direct pressure on commercial APIs in the same band — when an open-weights flagship prices output below $1 per million tokens, everyone else’s pricing cushion shrinks.

We noted in our opening-of-year outlook how fast model turnover and the open-source split were moving in 2026; this permanent cut is the newest data point on the same curve. For product teams the question is no longer whether to use cheap models — it is whether your routing architecture is ready to let you.

The arithmetic is worth spelling out. A workload generating a billion output tokens a month costs $870 at the standing Pro price, against $3,480 at list — a gap that dwarfs most engineering optimizations a team could ship in a quarter. And because the price is now official rather than promotional, that saving survives procurement review and finance sign-off instead of expiring with a promo timer.

Sources

AI-assisted summary compiled from the sources above, reviewed by a human before publishing.

FOUND_THIS_USEFUL?

Support more practical AI articles, tutorials, and build notes.

BUY_ME_A_COFFEE
SHAREXEMAIL