Talk to an Expert

August 28, 2026 | Reading Time 5 mins

DeepSeek’s Price Increase Puts AI Inference on Utility Time

TL;DR DeepSeek, the vendor whose January 2025 launch pricing reset what the market believed LLM inference should cost, has raised prices. On August 16, 2026, its flat per-token rate card became a peak/off-peak grid: every cell above the old flat rate, off-peak set at half of peak. Time-of-day metering, the electric utility’s mechanism, is now live in AI inference. It entered this market as a discount in February 2025 and returned as a raise. Any cost model built on a flat token price now has a clock in it.


For a year and a half, DeepSeek was the floor. Its January 2025 launch pricing became the number every AI cost projection quietly leaned on, the proof that inference cost little and would cost less. That floor just moved up.

The increase sits on the vendor’s own rate card, captured in our snapshots, dated. The raise is the headline. The clock is the story: the price of a token is now a function of what time the request completes.

The DeepSeek Price Increase, Dated

The full sequence is tracked in our AI Pricing Observatory, captured from DeepSeek’s own pricing pages on the dates shown. The table below reads from that record:

Source: SPP Pricing Observatory · 8 tracked DeepSeek moves · last verified Sep 01, 2026.
DateThe move
DeepSeek-R1 API is priced at $0.14 per million input tokens on cache hit. SourceRate card
DeepSeek-R1 API is priced at $2.19 per million output tokens. SourceRate card · 2nd consecutive
DeepSeek introduced time-of-day API pricing on 2025-02-26, splitting its rate card into a standard price (UTC 00:30-16:30) and an off-peak discount price (UTC 16:30-00:30): deepseek-chat (V3) discounted 50% and deepseek-reasoner (R1) discounted 75% in the nightly window, landing both models at identical off-peak rates ($0.035 cache-hit input / $0.135 cache-miss input / $0.55 output per 1M tokens), with the tier decided by each request's completion timestamp. The page still showed a single flat rate card with no time-of-day tiers in Wayback captures of 2025-02-24 and 2025-02-25. SourceRate cut
DeepSeek cut its API prices by more than 50%, effective immediately, with the release of DeepSeek-V3.2-Exp. SourceRate cut · 2nd in 7 months
DeepSeek posted a footnote on its API pricing page announcing the API service 'will soon adopt a peak/off-peak pricing policy' with peak-hour prices at 2x the regular prices across all billing items, peak hours defined as 9:00-12:00 and 14:00-18:00 Beijing time (UTC+8) daily, and the effective date 'subject to the official announcement'. The footnote is absent from SPP's 2026-07-26 capture and present in the 2026-08-02 capture (pricing_page_watch snapshot 189); listed prices were unchanged. SourceTerms change · first
DeepSeek warned on its API pricing page that it plans 'to raise the overall pricing for DeepSeek API services in the near future, with a significant increase expected', advising users to plan usage accordingly; the warning replaced the prior peak/off-peak pre-announcement footnote while listed prices stayed unchanged. Captured in SPP's 2026-08-09 snapshot (pricing_page_watch snapshot 231, window 2026-08-02 to 2026-08-09); the South China Morning Post reported the same wording on 2026-08-06, placing the footnote live by that date. SourceRate increase · first
DeepSeek replaced flat token pricing with peak/off-peak tiers and raised prices significantly: cache-hit input tokens went from $0.0028/$0.003625 to $0.007–$0.014/$0.022–$0.044 per 1M, cache-miss input from $0.14/$0.435 to $0.22–$0.44/$0.66–$1.32 per 1M, and output from $0.28/$0.87 to $0.66–$1.32/$1.98–$3.96 per 1M, while also renaming DeepSeek-V4-Pro to DeepSeek-V4-Pro-0813 and removing the footnote warning about a planned future price increase. SourceRate increase · 2nd consecutive
1 more DeepSeek move on the Observatory →

The August 16 grid is the first outright increase on a model vendor’s per-token rate card in our Observatory record. The cuts had names and press cycles. The raise had a footnote, a warning, and then a Sunday.

Read those two footnotes as one move rather than as two notices. The first named the mechanism and its multiple and left the date open. The second dropped the mechanism, kept the direction, and told customers to plan. None of it was hidden: all of it was published on the vendor’s own pricing page, in the register of a policy note rather than a price change. Each notice moved the surprise a little further from the day the number moved, and a customer reading only the rate card met the increase for the first time as an increase.

Time-of-Day Pricing Is Utility Economics Applied to LLM Inference

Peak/off-peak pricing is not a software invention. Electric utilities have priced this way for decades, and peer-reviewed work on utility-computing pricing identified the conditions under which it pays: demand that cycles predictably, fixed capacity, and capacity that is wasted when idle. GPU inference satisfies all three. Daytime traffic saturates the cluster, overnight hours run slack, and an idle accelerator earns nothing.

Be precise about what this is not. It is not surge pricing, and it is not dynamic pricing in the marketplace sense. Nothing here floats with live demand. The grid is printed, deterministic, and identical for every buyer: the same request at the same hour costs the same amount, and anyone can verify the cell. The clock sets the price, not an auction.

The commercial consequences run in opposite directions. Surge pricing makes costs unpredictable. Time-of-day pricing makes them schedulable, which is a polite way of saying it makes them your problem. The vendor has handed its capacity curve to its customers with prices attached. Two identical jobs, hours apart, now carry different unit costs.

The Raise Used a Mechanism Installed as a Discount

When the clock first appeared in February 2025, it only ever lowered a price. The daytime rate stayed where the old flat rate had been, and the overnight window cut it: half off for the chat model, three quarters off for the reasoner. Nobody objects to cheaper nights. The billing logic shipped, buyers learned that the hour prices the token, and the apparatus ran for a year and a half as a discount.

In August 2026 the same mechanism came back pointed the other way. The two-to-one ratio survived into the new grid. The base did not.

We are not claiming the sequence was planned from the start. We do not need to. The mechanism arrived in its generous configuration, was accepted without friction, and later reversed without any structural change. Peer-reviewed research on price fairness has long observed that buyers judge the loss of a discount far more gently than an equivalent surcharge, and judge demand-driven increases most harshly of all. The ordering is the tell: a meter that arrives as a discount has already cleared the fairness hurdle it will need if it ever reverses.

The read travels to every vendor minting a usage currency. Credits, points, and compute units are surrogate units, and a conversion table has the same property as DeepSeek’s clock: a multiplier introduced as a bonus is a rate mechanism on probation. The lineage of minted currencies shows how old that lever is. Watch the mechanism, not the direction it currently points.

What a Clock in the Rate Card Breaks

The floor was a posture

The low token price was never a floor. It was a promotional posture, held as long as the vendor’s economics supported it. The September 2025 cut lasted under a year. Every projection that extrapolated the discount era forward is now wrong in a direction nobody modeled.

Your cost model has a clock in it

If your unit economics assume one price per million tokens, your model now has a missing variable: when the workload runs. Batchable work, overnight embedding runs and agentic pipelines that tolerate delay, can chase the off-peak window; interactive work cannot. A copilot answering your customers during their business day runs when they work, and if those hours overlap the vendor’s peak, that is a doubling you do not control. The exposure is not the average price. It is the correlation between your demand curve and the vendor’s.

Pass-through now runs in reverse

When model prices fell, vendors with cost-linked meters donated the deflation to their customers automatically. The same wiring conducts increases. Pass model costs through and your customers just inherited a clock they never agreed to; absorb, and your margin did. This is the subsidy cliff with a timestamp attached. Where you sit on the pass-through spectrum decides which side of you takes the hit.

The metric the clock cannot touch

A value metric anchored to what customers receive is indifferent to all of this. The customer buys resolved conversations, active users, or completed runs, and the hour DeepSeek’s cluster fills is the vendor’s cost problem, managed like any other input. That indifference is the argument for pricing AI on value rather than on cost, and it is worth pressure-testing with a pricing expert before the next rate card surprises you.

The test question travels: if the same job costs twice as much at 10:00 as at 02:00, which line of your cost model notices? Which clause in your customer contracts does?


DeepSeek drove token prices down, taught the market to meter by the clock while the clock only helped, then used the clock to raise prices. The arc is dated and public. If your pricing, packaging, or cost model leans on a flat token price anywhere, talk to an expert. Describe where model costs enter your unit economics and where they exit into your prices, and a pricing expert replies with where the exposure sits.


Is Your AI Vendor Dependency Built on a Promotional Posture?

If your licensing, packaging, and pricing decisions were calibrated to a token price that was always temporary, the architecture underneath them needs examination. Find out where the exposure sits.

FAQs



Linkedin X (Twitter) Facebook

Ready for profitable growth?

Hit the ground running and learn how to fix your pricing.

Book A Demo Contact Us