Talk to an Expert

October 3, 2026 | Reading Time 8 mins

Inference Cost: What It Does to Software Pricing

TL;DR Inference cost is what a software vendor pays every time a model produces an output, and it lands in cost of goods sold: the first real marginal cost software has carried in decades. The number moves when the vendor changes nothing, because five variables set the cost of a call and agentic workflows multiply the calls. The value metric decides whether revenue runs ahead of that cost or behind it. A cost-aligned unit caps the ceiling; a metric denominated in customer value lets revenue outrun the cost. Five tests at the end locate the exposure.

Inference cost recurs on every use, scales with the work produced, and sits underneath the price on every transaction. Most discussions of AI cost stop at the token rate.

Three separate problems travel under one name. Inference cost is the vendor’s cost of running a model. Inference pricing is what a model supplier charges for it. Token-based pricing is a licensing choice a software company makes downstream of both.


What Inference Cost Is, and Where It Sits in a Software P&L

What is inference cost in AI software?

A generative AI model produces output one token at a time, and every token costs compute. The model supplier charges for those tokens, typically pricing output above input because generating text costs more than reading it, and pricing cached input far below fresh input. Research on the economics of foundation model inference confirms both gaps.

Inference cost and training cost are different obligations

Training cost is large and finite: a company trains a model, pays the bill, and moves on. Inference cost is smaller per event and permanent. Every customer interaction, every API call, every automated run adds to the tab.

Where inference cost lands in the P&L

Inference cost lands in cost of goods sold. Software running on pure distribution had a marginal cost close to zero; software running on LLM inference has a cost that appears at the level of a single customer action. Every pricing decision in an AI product sits on top of that fact.

Most often a software company cannot produce account-level COGS at all, because inference spend arrives in aggregate from the supplier and stays aggregate in the books. Bringing it to the account takes instrumentation across every agentic layer in the product’s codebase. Until that exists the cost sits in an engineering line item where the pricing model cannot reach it.


Why the Cost Moves When You Change Nothing

What sets the cost of a single call

Five variables drive the cost of a single inference call. They are which model handles the request, how much context it carries, how much output it generates, how many retries it requires, and whether the input was cached. A vendor can hold its product constant and still watch per-request cost shift. A model supplier reprices a model tier, a new model version changes the token count for the same output, or a user’s workflow starts producing longer inputs.

Why agentic workflows multiply the call count

Research on agentic workflows finds they can consume orders of magnitude more tokens than a single-turn interaction because tasks decompose recursively. The agent breaks one request into subtasks, calls the model for each, checks the output, and may call again if the result is incomplete.

The harness multiplies it again, and the harness is the vendor’s own design. The code that chains model calls, routes each step to a model, checks the output and retries sets how many calls one job takes. It also sets which models serve them and how many verification passes sit between them. Call count per job is as much a product decision as a customer behavior. Of the costs above, it is the one the vendor can engineer down most directly without touching what the customer asked for.

Why the same request can cost two different amounts

Two customers issuing identical requests can generate very different inference costs, depending on the context each carries into the session and how their workflows decompose. A seat-based metric ignores exactly this property: it counts users and never sees the work.

Who pays for a run that returns nothing usable is a contract question. When the Agent Is Wrong covers it.


Where Inference Cost Enters Your Pricing

The value metric decides whether revenue scales ahead of inference cost, alongside it, or behind it.

What a cost-aligned unit guarantees

A value metric denominated in the supplier’s unit, tokens or compute units, covers the bill on every transaction and caps the ceiling. When the metric is pinned to cost, revenue cannot grow faster than the cost to serve, and the vendor’s margin is an infrastructure margin, the spread left on resold compute, and nothing more.

A credit as a surrogate unit is the exception. A surrogate unit is a vendor-defined billing unit, most often a credit, that stands in for the value metrics underneath it. The vendor writes the exchange rate between the credit and the underlying cost, so revenue per unit of work can run ahead of the cost to serve. That gap is the margin lever, and it is also the trust problem, because the hand that sets the rate can move it.

Credit-based pricing escapes the ceiling only to the degree the rate is set independently of the token; a credit that simply shadows the token stays pinned to it.

What boundedness means in a value metric

A bounded value metric, one whose unit is bound to the value the customer receives rather than to the supplier’s cost, lets revenue outrun the cost. Under that metric, the customer who derives more value from the product pays more, however many tokens the vendor spent serving them. The inference cost becomes a floor beneath the price rather than a ceiling above it.

Economic analysis of AI workflow optimization finds that the value of a token depends on where it sits in a workflow rather than on its count, a difference per-token billing cannot capture.

When does passing the cost through make sense?

Passing a cost through makes sense when buyers can compare the price of it easily. Pass-through exposes the price to a comparison the customer can make directly: the vendor’s rate against the supplier’s published rate card. Upcharging a cost the buyer can look up reads as nickel and diming and has nothing to do with value.

The pattern repeats across our corpus: a vendor wins a deal in part because a cost the buyer could look up appears as its own line item. In one case it was postage for mailers, which every competitor had bundled in with margin on top. A company does not have to mark up a pass-through cost. It has to deliver its profitability goals across the whole approach, with each product it sells still carrying its own price.

Who carries the supplier’s metered cost

Which party carries a model supplier’s metered cost is its own architectural question, and Pass-Through or Recast covers it. How licensing, packaging, and pricing sequence around that cost is covered in AI Software Pricing.


Does Your Revenue Scale Ahead of Inference Cost, or Behind It?

Your value metric determines whether each new inference call widens margin or erodes it. Answer a few questions to see whether your licensing, packaging, and pricing put revenue ahead of, alongside, or behind the compute your customers consume.

Falling Per-Token Prices Do Not Automatically Reach You

Who receives a price cut at the model layer?

A price change at the model layer lands on whoever the value metric points at. Where a vendor bills in the supplier’s unit, a supplier cut passes to the customer on the next invoice and the vendor’s margin is unchanged. Where the vendor bills in a customer-value unit, the change accrues to the vendor. AI Price Cuts and Deflation Capture covers that mechanism.

The shape of a move tells you more than its direction

The commercially interesting property of model-layer pricing is the shape of the moves, and the published record shows them in both directions. Flat token rates have become peak and off-peak grids, promotional rates have carried expiry dates, announced increases have been cancelled, and cuts have reached one rung of a model range while the rest held.

Source: SPP Pricing Observatory · 9 tracked DeepSeek moves · last verified Oct 09, 2026.
DateThe move
The V3.2-Exp release cut API prices by more than half, effective immediately, and the vendor's stated reason was architectural: sparse attention lowering the cost of inference. SourcePricing · Rate cut · 2nd in 7 months
The reversal announces itself in a footnote. SourcePricing · Terms change · first
A week later the footnote dropped the mechanism and kept the direction: 'We plan to raise the overall pricing for DeepSeek API services in the near future, with a significant increase expected. SourcePricing · Rate increase · first
The raise landed August 16, 2026 at 16:00 UTC, exactly as the August 13 release note scheduled it. SourcePricing · Rate increase · 2nd consecutive
Two footnotes changed in one update, moving in opposite directions. SourcePricing · Unit swap · first
DeepSeek collapsed its separate Flash text and vision-experimental SKUs into a single deepseek-flash model and cut its off-peak Flash per-token rates for cache-hit input, cache-miss input and output. SourcePackaging + Pricing · Rate cut · 3rd in 18 months
3 more DeepSeek moves on the Observatory →

The DeepSeek sequence above is the clearest specimen. Nothing here depends on where token prices go next. A vendor that prices on the assumption of continued cost reductions has built a subsidy into its structure, and AI Pricing and the Subsidy Cliff covers what happens when that assumption fails.


What Inference Cost Does to the Price Floor

Why the cost floor used to be irrelevant

For traditional software, marginal cost was near zero and the hard cost floor sat far beneath any price a company would accept. The floor that constrained pricing was a margin target set much higher: a policy floor living in a discount approval matrix.

That policy floor erodes deal by deal when nothing enforces it. Each exception arrives with a local justification. The exceptions accumulate into a realized price that has drifted from the pricebook, and every step of that drift is a decision someone approved.

Policy floors and structural floors

Generative AI raises the cost floor and moves it closer to the price. A policy floor is a target with nothing structural behind it, and it erodes. A structural floor is engineered into the pricing surface so that no commitment level produces a scheduled net price that erodes margin below target. It holds because the surface itself prevents the erosion, where a policy floor relies on deal-by-deal discipline.

Modeling work on capped-usage subscriptions establishes that exposure lives in the tail of the usage distribution, and across our engagements software usage has never been normally distributed. Data collected under a cap understates true consumption, so a policy floor is being set against an understated cost.

Theoretical work on how model suppliers should price finds that menus combining a fixed access fee with usage-based charges are revenue-optimal, and that higher-volume buyers rationally receive lower per-unit rates on a published schedule. That is the supplier layer’s rationale for tiered rate cards, and the same logic runs downstream when a vendor constructs its own commitment tiers.

Across our corpus the policy floor fails later than vendors expect. Customers sign the proposal and then, in the days before go-live, ask for a discount. The enterprise discount negotiation often runs well past signature, and a floor that held through the deal gives way at the last step.


Tests to Run Against Your Own Inference Exposure

The spend-doubling test

If your inference spend doubled next quarter, which part of your architecture responds? Does revenue move with it, or does margin compress?

The supplier-unit test

Which of your units is denominated in your supplier’s unit? A token, a compute unit, or a credit pegged to a token is a cost-aligned unit.

The cost-visibility test

Can you see cost per account, per edition, and per Customer Group, or only per month?

The unusable-run test

What happens commercially when a run consumes real compute and returns something the customer cannot use? Your contract already answers this question, and the answer may not be the one you intend.

The rate-cut test

If your supplier cut its rates tomorrow, who receives that? The answer is in the value metric.

If any of those questions surfaces a gap, a conversation with an SPP pricing expert is the fastest way to read what the exposure does to your price structure.


FAQs



Linkedin X (Twitter) Facebook

Ready for profitable growth?

Hit the ground running and learn how to fix your pricing.

Book A Demo Contact Us