Author
TL;DR DeepSeek, the vendor whose January 2025 launch pricing reset what the market believed LLM inference should cost, has raised prices. On August 16, 2026, its flat per-token rate card became a peak/off-peak grid: every cell above the old flat rate, off-peak set at half of peak. Time-of-day metering, the electric utility’s mechanism, is now live in AI inference. It entered this market as a discount in February 2025 and returned as a raise. Any cost model built on a flat token price now has a clock in it.
For a year and a half, DeepSeek was the floor. Its January 2025 launch pricing became the number every AI cost projection quietly leaned on, the proof that inference cost little and would cost less. That floor just moved up.
The increase sits on the vendor’s own rate card, captured in our snapshots, dated. The raise is the headline. The clock is the story: the price of a token is now a function of what time the request completes.
The DeepSeek Price Increase, Dated
The full sequence is tracked in our AI Pricing Observatory, captured from DeepSeek’s own pricing pages on the dates shown. The table below reads from that record:
The August 16 grid is the first outright increase on a model vendor’s per-token rate card in our Observatory record. The cuts had names and press cycles. The raise had a footnote, a warning, and then a Sunday.
Read those two footnotes as one move rather than as two notices. The first named the mechanism and its multiple and left the date open. The second dropped the mechanism, kept the direction, and told customers to plan. None of it was hidden: all of it was published on the vendor’s own pricing page, in the register of a policy note rather than a price change. Each notice moved the surprise a little further from the day the number moved, and a customer reading only the rate card met the increase for the first time as an increase.
Time-of-Day Pricing Is Utility Economics Applied to LLM Inference
Peak/off-peak pricing is not a software invention. Electric utilities have priced this way for decades, and peer-reviewed work on utility-computing pricing identified the conditions under which it pays: demand that cycles predictably, fixed capacity, and capacity that is wasted when idle. GPU inference satisfies all three. Daytime traffic saturates the cluster, overnight hours run slack, and an idle accelerator earns nothing.
Be precise about what this is not. It is not surge pricing, and it is not dynamic pricing in the marketplace sense. Nothing here floats with live demand. The grid is printed, deterministic, and identical for every buyer: the same request at the same hour costs the same amount, and anyone can verify the cell. The clock sets the price, not an auction.
The commercial consequences run in opposite directions. Surge pricing makes costs unpredictable. Time-of-day pricing makes them schedulable, which is a polite way of saying it makes them your problem. The vendor has handed its capacity curve to its customers with prices attached. Two identical jobs, hours apart, now carry different unit costs.
The Raise Used a Mechanism Installed as a Discount
When the clock first appeared in February 2025, it only ever lowered a price. The daytime rate stayed where the old flat rate had been, and the overnight window cut it: half off for the chat model, three quarters off for the reasoner. Nobody objects to cheaper nights. The billing logic shipped, buyers learned that the hour prices the token, and the apparatus ran for a year and a half as a discount.
In August 2026 the same mechanism came back pointed the other way. The two-to-one ratio survived into the new grid. The base did not.
We are not claiming the sequence was planned from the start. We do not need to. The mechanism arrived in its generous configuration, was accepted without friction, and later reversed without any structural change. Peer-reviewed research on price fairness has long observed that buyers judge the loss of a discount far more gently than an equivalent surcharge, and judge demand-driven increases most harshly of all. The ordering is the tell: a meter that arrives as a discount has already cleared the fairness hurdle it will need if it ever reverses.
The read travels to every vendor minting a usage currency. Credits, points, and compute units are surrogate units, and a conversion table has the same property as DeepSeek’s clock: a multiplier introduced as a bonus is a rate mechanism on probation. The lineage of minted currencies shows how old that lever is. Watch the mechanism, not the direction it currently points.
What a Clock in the Rate Card Breaks
The floor was a posture
The low token price was never a floor. It was a promotional posture, held as long as the vendor’s economics supported it. The September 2025 cut lasted under a year. Every projection that extrapolated the discount era forward is now wrong in a direction nobody modeled.
Your cost model has a clock in it
If your unit economics assume one price per million tokens, your model now has a missing variable: when the workload runs. Batchable work, overnight embedding runs and agentic pipelines that tolerate delay, can chase the off-peak window; interactive work cannot. A copilot answering your customers during their business day runs when they work, and if those hours overlap the vendor’s peak, that is a doubling you do not control. The exposure is not the average price. It is the correlation between your demand curve and the vendor’s.
Pass-through now runs in reverse
When model prices fell, vendors with cost-linked meters donated the deflation to their customers automatically. The same wiring conducts increases. Pass model costs through and your customers just inherited a clock they never agreed to; absorb, and your margin did. This is the subsidy cliff with a timestamp attached. Where you sit on the pass-through spectrum decides which side of you takes the hit.
The metric the clock cannot touch
A value metric anchored to what customers receive is indifferent to all of this. The customer buys resolved conversations, active users, or completed runs, and the hour DeepSeek’s cluster fills is the vendor’s cost problem, managed like any other input. That indifference is the argument for pricing AI on value rather than on cost, and it is worth pressure-testing with a pricing expert before the next rate card surprises you.
The test question travels: if the same job costs twice as much at 10:00 as at 02:00, which line of your cost model notices? Which clause in your customer contracts does?
DeepSeek drove token prices down, taught the market to meter by the clock while the clock only helped, then used the clock to raise prices. The arc is dated and public. If your pricing, packaging, or cost model leans on a flat token price anywhere, talk to an expert. Describe where model costs enter your unit economics and where they exit into your prices, and a pricing expert replies with where the exposure sits.
Is Your AI Vendor Dependency Built on a Promotional Posture?
If your licensing, packaging, and pricing decisions were calibrated to a token price that was always temporary, the architecture underneath them needs examination. Find out where the exposure sits.