TL;DR The claim that AI retires volume discounting has the logic inverted. Near-zero marginal cost is what made the seat-era volume discount indefensible: it cost the vendor nothing to give, so it bought nothing back. Linear cost restores what a quantity discount was invented to do: pay for something that lowers cost to serve. Across eighteen AI vendors, not one prices a volume step on tokens; every one discounts deeply on another axis: scheduling freedom, cache reuse, committed capacity. Discount the buyer’s size and the vendor funds it; discount utilization and the discount pays for itself.
- Near-Zero Marginal Cost Is What Made the Seat-Era Discount Indefensible
- Linear Cost Restores What a Quantity Discount Was Invented to Do
- What the Published Rate Cards Show
- Where the Room to Discount Comes From
- Discount the Behavior, Not the Buyer
- Calibrate Against Margin, Not Revenue
- The Commit Layer: Where AI Volume Discounts Live
- Deviation: Where the Schedule and the Deal Desk Diverge
- FAQs
A position is circulating that AI ends volume discounting, on the grounds that inference costs do not fall with scale. The premise is correct. The conclusion runs backwards. Costs that scale linearly are the condition quantity discounts were built for.
What software lost when marginal cost went to zero was the reason a discount could be earned at all. AI software pricing hands that reason back. The market settled this in public. Every major AI vendor discounts deeply, and none indexes the discount to how much the buyer buys.
Near-Zero Marginal Cost Is What Made the Seat-Era Discount Indefensible
In classic software the next unit cost nothing to produce. A volume discount cost the vendor nothing to give and bought nothing back. It was revenue arithmetic, and the only discipline available was willpower.
That is the world the case against volume discounting was written for. In that world the case held. Even the shape was borrowed. Tier-step volume discounting answered four real problems in manufacturing, and software inherited the staircase without inheriting any of the rationales, as Margin-Calibrated Discounting sets out in full.
The criticism was arithmetic, not opinion. It also expired with the condition it was built on.
Linear Cost Restores What a Quantity Discount Was Invented to Do
On an AI meter the next unit costs real money. Inference is a cost of goods. It moves with models, workload mix, and prompt growth, none of which move on your contract calendar.
A discount now has to be earned by something that genuinely lowers what it costs to serve, which is what a quantity discount did before software made cost disappear. The cost line finally pushes back.
There is no quantity-driven scale economy at the buyer level. Buying more tokens does not make the next one cheaper to produce, which is precisely why no vendor publishes a volume step on a rate card. But serving cost does fall with utilization and with the predictability of the fleet behind it. A commitment is a purchase of predictability.
So “AI costs do not fall with scale” is not an argument to stop discounting. It is an argument to discount on utilization rather than on account size.
A discount schedule set against list price alone can cross under cost at exactly the volumes it was designed to reward. The answer is to index the discount to something that moves cost in your favor, and to calibrate whatever remains against gross margin at every commitment level on the surface.
What the Published Rate Cards Show
We checked eighteen vendors across the inference layer and the credit-denominated application layer, against primary sources only: pricing pages, published rate cards, vendor documentation. Anyone can repeat the check.
The one axis nobody prices on
No inference provider publishes a per-token rate that steps down as monthly volume rises. Two published ladders resemble volume discount schedules and are not: OpenAI’s usage tiers and Anthropic’s named tiers both move rate limits. Neither moves price. A buyer who climbs them earns throughput headroom, not a better rate.
The axes everyone prices on
Batch or asynchronous processing runs at 50 percent off list at essentially every major provider. The buyer surrenders scheduling control. The vendor fills the troughs in its own fleet. Cache reads run at 90 percent off, because a cache read skips the compute outright, and at OpenAI and Anthropic the two discounts stack to roughly 95 percent off list input.
Commitment is priced the same way, in public, by the largest infrastructure vendors in the market. AWS Bedrock’s provisioned throughput documentation states plainly that a longer commitment means a lower hourly price, reaching 50 percent at six months. Azure capacity reservations reach 64 percent monthly and 70 percent annual, and Google’s provisioned throughput ladder puts a one-year commitment about 26 percent below a one-month commitment. Together AI, a GPU cloud rather than a hyperscaler, runs a term ladder to 33 percent on the same logic.
Commitment as the discount driver is not a preference. It is published, in-force, checkable behavior at Google, Microsoft, AWS and Together AI. The room to move is enormous. Across one vendor’s published input rates the surface runs from about $0.25 to $10.00 per million tokens, and no volume step appears anywhere inside that span.
Classic volume discounting, alive in the credit layer
Higher up the stack, where the unit is a credit rather than a token, volume discounting is published under its own name. Snowflake applies AI Credit discounts indexed to annual contract value automatically, with no action required by the customer. Salesforce shipped Agentforce stating that standard volume discounts apply. Microsoft publishes savings of up to 20 percent on Copilot credit commitments. GitHub’s individual plan ladder buys credits at 33 to 50 percent below its own list credit price, and markets the saving.
Read closely, these are commitment discounts too: annual contract value is a commitment, and a credit pack bought up front is a forecast the customer has agreed to absorb. The application layer pools serving cost across a portfolio, so the commitment can be denominated in dollars instead of throughput.
Does Your Rate Card Survive Linear Cost Compression at Scale?
Eighteen published rate cards reveal a pattern: volume discounts reappear wherever input costs scale linearly. A few targeted questions will show whether your licensing, packaging, and pricing structure is positioned for that reality.
Where the Room to Discount Comes From
An application vendor is not only a seller on this market. It is also a buyer. The two roles are the same mechanic pointed in opposite directions.
Every axis the model providers publish is available to the application vendor as a buyer: batch what can wait, cache what repeats, commit to the capacity it can forecast. Working that structure fully is what earns the price break on the input side, and that margin is the space the vendor has to build a discount schedule on the output side. A vendor that pays list for inference has no room to give, and a vendor that has done the work on its own cost line has room that its competitors do not.
That space is worth more than the margin it holds, because of what a schedule does to the buyer. A customer deciding whether to route a whole use case through your product is deciding how much financial risk to absorb on a bill they cannot yet predict. A published schedule with a defined path from a small commitment to a large one is what makes going all in survivable. Without it, the rational move is to keep the use case small and split it, which is the outcome that costs the vendor most.
And the schedule is not priced against one customer. The vendor is monetizing the mix: light users, heavy users, the ones still deciding, and the ones whose usage is about to change shape. The schedule has to hold margin across that whole portfolio at once. Moving one band is never a local decision.
This is genuinely hard. It is not spreadsheet work. It means modelling the alternatives against each other, trading margin at one point on the surface against volume at another, and doing it while the underlying usage patterns move.
That last part is what makes this moment different. Usage patterns across a customer base are changing faster now than they have in a long time, so a schedule calibrated on last year’s shape is calibrated on a distribution that no longer exists. Record revenue on a stale schedule is not evidence the architecture holds; it is the window between the usage mix turning and the invoice showing it. The instrumentation exists to model this, and SPP built LevelSetter to do it, but the tool is the support underneath the judgment, not a replacement for it.
The test question is this. If the mix of usage across your customer base shifted meaningfully next quarter, would your schedule still hold margin at every band, and how would you know?
Discount the Behavior, Not the Buyer
Every published AI discount is indexed to a property of the workload, not to the size of the account. Batch buys scheduling freedom. Cache skips compute. Commitment moves utilization risk off the vendor and onto the customer.
A classic volume discount rewards the size of the buyer, so it accumulates against the vendor as the account grows. A discount indexed to utilization rewards a behavior that lowers the vendor’s own cost, so it is self-funding.
Peer-reviewed work in B2B markets supplies the permission: bigger deals do not always earn lower unit prices, and capacity and cost constraints legitimately produce higher unit rates at some volumes. Bigger-is-cheaper is a reflex, not a law.
Calibrate Against Margin, Not Revenue
A consumption pricing model should produce a target net price at every commitment level, engineered against gross margin at every point on the pricing surface. There is no commitment level at which the serving cost stops applying. That target is the scheduled net price; where the rep closed is the landed net price, and the gap between them is where margin leaks, deal by deal.
Averages hide that leak from both sides. An average discount rate blends the disciplined deals with the ones draining the base.
Wherever we have watched it accrue, software usage is never normally distributed: a small share of customers consumes disproportionately, which drags the average away from every customer in the base. The average stops being a typical value. It becomes an artifact of the heavy tail, and a schedule built on it misprices both ends. The heavy end is where the inference bill lives.
Comp closes the loop. Plans that pay on revenue reward contract size; plans that pay on proximity to scheduled net price reward margin contribution. Revenue-paid comp pays the rep best for landing the largest commitments at the deepest rates, which is where the inference bill is largest. Rebuilding the plan around margin is its own discipline, covered in sales compensation on gross profit.
Designed schedules hold. Where our pattern library can read a published volume schedule against the deals written under it, essentially every banded deal lands exactly at the designed rate. The only deviations leave designed discount on the table. Discount chaos is not a sales-behavior inevitability. It is what fills the space where a designed schedule is missing.
Where your own surface should sit, and against which margin targets, is judgment work on your cost structure and your market. That is a conversation to have with a pricing expert, with your own deal patterns in hand.
The Commit Layer: Where AI Volume Discounts Live
Committed consumption is the dominant AI deal shape. The customer commits to volume or spend, earns the discounted rate, and the contract defines what happens above it. The volume discount and the commitment structure are one design, and a schedule that ignores the commit layer prices deals that never occur.
Seat volume was a headcount, auditable at any moment. Consumption volume is a forecast, and the buyer writes it. Commits inflate to reach a better rate. The true-down conversation arrives at the end of the term when usage lands below the number. A schedule that anticipates that conversation, with a defined path back onto the surface, keeps the account on the pricebook; one that improvises turns every miss into a bespoke deal.
Past the commitment, the question is whether overage reads as continuation or as punishment. Peer-reviewed work on tiered plans with usage allowances finds that buyers facing an uncertain bill select allowances defensively, weighing the exposure rather than their expected usage. A penalty rate past the line is more of that exposure, so the rational response next cycle is a smaller commitment. That shrinks the very thing the discount was built to reward. Underneath both directions sits boundedness. An unbounded metric under a deeply discounted rate is exposure at any schedule shape.
A credit is a surrogate unit, so where the discounted unit is a credit, the volume discount and the conversion table are two levers on the same net price. On 2026-06-01 GitHub retired premium requests into AI credits and raised model multipliers the same day: the headline rate held still while the meaning of the unit moved underneath it.
A discount on a credit is worth exactly as much as the credit’s engineering. That means a published conversion table, definition stability, and any change in what a credit buys treated as a price change.
Deviation: Where the Schedule and the Deal Desk Diverge
The gap between the published schedule and where deals land is pricebook deviation. On a consumption meter it compounds. A discount granted off-schedule persists across every unit the account ever consumes. A seat-era exception leaked margin on a fixed count; a consumption-era exception leaks it on a meter that grows.
Three questions test whether your schedule protects anything. If your largest account doubled its consumption next quarter, would the rate it then earns keep that account margin-positive? Can you name the scheduled net price for a commitment level between your published breaks? When a rep lands below the schedule, does anything in your comp plan notice?
If any of the three has no answer, the schedule is a suggestion the deal desk edits. The volume discount is one of several levers that shape your pricing surface, and it is not a concession program. Vendors who treat it as one discover their AI margin a quarter at a time.
If your volume schedule predates your consumption meter, the margin math underneath it has already moved. Talk to an expert: describe your schedule, your meter, and where deals are landing, and a pricing architecture expert replies with where the calibration should start.