Author
TL;DR Inference cost is what a software vendor pays every time a model produces an output, and it lands in cost of goods sold: the first real marginal cost software has carried in decades. The number moves when the vendor changes nothing, because five variables set the cost of a call and agentic workflows multiply the calls. The value metric decides whether revenue runs ahead of that cost or behind it. A cost-aligned unit caps the ceiling; a metric denominated in customer value lets revenue outrun the cost. Five tests at the end locate the exposure.
Inference cost recurs on every use, scales with the work produced, and sits underneath the price on every transaction. Most discussions of AI cost stop at the token rate.
Three separate problems travel under one name. Inference cost is the vendor’s cost of running a model. Inference pricing is what a model supplier charges for it. Token-based pricing is a licensing choice a software company makes downstream of both.
What Inference Cost Is, and Where It Sits in a Software P&L
What is inference cost in AI software?
A generative AI model produces output one token at a time, and every token costs compute. The model supplier charges for those tokens, typically pricing output above input because generating text costs more than reading it, and pricing cached input far below fresh input. Research on the economics of foundation model inference confirms both gaps.
Inference cost and training cost are different obligations
Training cost is large and finite: a company trains a model, pays the bill, and moves on. Inference cost is smaller per event and permanent. Every customer interaction, every API call, every automated run adds to the tab.
Where inference cost lands in the P&L
Inference cost lands in cost of goods sold. Software running on pure distribution had a marginal cost close to zero; software running on LLM inference has a cost that appears at the level of a single customer action. Every pricing decision in an AI product sits on top of that fact.
Most often a software company cannot produce account-level COGS at all, because inference spend arrives in aggregate from the supplier and stays aggregate in the books. Bringing it to the account takes instrumentation across every agentic layer in the product’s codebase. Until that exists the cost sits in an engineering line item where the pricing model cannot reach it.
Why the Cost Moves When You Change Nothing
What sets the cost of a single call
Five variables drive the cost of a single inference call. They are which model handles the request, how much context it carries, how much output it generates, how many retries it requires, and whether the input was cached. A vendor can hold its product constant and still watch per-request cost shift. A model supplier reprices a model tier, a new model version changes the token count for the same output, or a user’s workflow starts producing longer inputs.
Why agentic workflows multiply the call count
Research on agentic workflows finds they can consume orders of magnitude more tokens than a single-turn interaction because tasks decompose recursively. The agent breaks one request into subtasks, calls the model for each, checks the output, and may call again if the result is incomplete.
The harness multiplies it again, and the harness is the vendor’s own design. The code that chains model calls, routes each step to a model, checks the output and retries sets how many calls one job takes. It also sets which models serve them and how many verification passes sit between them. Call count per job is as much a product decision as a customer behavior. Of the costs above, it is the one the vendor can engineer down most directly without touching what the customer asked for.
Why the same request can cost two different amounts
Two customers issuing identical requests can generate very different inference costs, depending on the context each carries into the session and how their workflows decompose. A seat-based metric ignores exactly this property: it counts users and never sees the work.
Who pays for a run that returns nothing usable is a contract question. When the Agent Is Wrong covers it.
Where Inference Cost Enters Your Pricing
The value metric decides whether revenue scales ahead of inference cost, alongside it, or behind it.
What a cost-aligned unit guarantees
A value metric denominated in the supplier’s unit, tokens or compute units, covers the bill on every transaction and caps the ceiling. When the metric is pinned to cost, revenue cannot grow faster than the cost to serve, and the vendor’s margin is an infrastructure margin, the spread left on resold compute, and nothing more.
A credit as a surrogate unit is the exception. A surrogate unit is a vendor-defined billing unit, most often a credit, that stands in for the value metrics underneath it. The vendor writes the exchange rate between the credit and the underlying cost, so revenue per unit of work can run ahead of the cost to serve. That gap is the margin lever, and it is also the trust problem, because the hand that sets the rate can move it.
Credit-based pricing escapes the ceiling only to the degree the rate is set independently of the token; a credit that simply shadows the token stays pinned to it.
What boundedness means in a value metric
A bounded value metric, one whose unit is bound to the value the customer receives rather than to the supplier’s cost, lets revenue outrun the cost. Under that metric, the customer who derives more value from the product pays more, however many tokens the vendor spent serving them. The inference cost becomes a floor beneath the price rather than a ceiling above it.
Economic analysis of AI workflow optimization finds that the value of a token depends on where it sits in a workflow rather than on its count, a difference per-token billing cannot capture.
When does passing the cost through make sense?
Passing a cost through makes sense when buyers can compare the price of it easily. Pass-through exposes the price to a comparison the customer can make directly: the vendor’s rate against the supplier’s published rate card. Upcharging a cost the buyer can look up reads as nickel and diming and has nothing to do with value.
The pattern repeats across our corpus: a vendor wins a deal in part because a cost the buyer could look up appears as its own line item. In one case it was postage for mailers, which every competitor had bundled in with margin on top. A company does not have to mark up a pass-through cost. It has to deliver its profitability goals across the whole approach, with each product it sells still carrying its own price.
Who carries the supplier’s metered cost
Which party carries a model supplier’s metered cost is its own architectural question, and Pass-Through or Recast covers it. How licensing, packaging, and pricing sequence around that cost is covered in AI Software Pricing.
Does Your Revenue Scale Ahead of Inference Cost, or Behind It?
Your value metric determines whether each new inference call widens margin or erodes it. Answer a few questions to see whether your licensing, packaging, and pricing put revenue ahead of, alongside, or behind the compute your customers consume.
Falling Per-Token Prices Do Not Automatically Reach You
Who receives a price cut at the model layer?
A price change at the model layer lands on whoever the value metric points at. Where a vendor bills in the supplier’s unit, a supplier cut passes to the customer on the next invoice and the vendor’s margin is unchanged. Where the vendor bills in a customer-value unit, the change accrues to the vendor. AI Price Cuts and Deflation Capture covers that mechanism.
The shape of a move tells you more than its direction
The commercially interesting property of model-layer pricing is the shape of the moves, and the published record shows them in both directions. Flat token rates have become peak and off-peak grids, promotional rates have carried expiry dates, announced increases have been cancelled, and cuts have reached one rung of a model range while the rest held.
The DeepSeek sequence above is the clearest specimen. Nothing here depends on where token prices go next. A vendor that prices on the assumption of continued cost reductions has built a subsidy into its structure, and AI Pricing and the Subsidy Cliff covers what happens when that assumption fails.
What Inference Cost Does to the Price Floor
Why the cost floor used to be irrelevant
For traditional software, marginal cost was near zero and the hard cost floor sat far beneath any price a company would accept. The floor that constrained pricing was a margin target set much higher: a policy floor living in a discount approval matrix.
That policy floor erodes deal by deal when nothing enforces it. Each exception arrives with a local justification. The exceptions accumulate into a realized price that has drifted from the pricebook, and every step of that drift is a decision someone approved.
Policy floors and structural floors
Generative AI raises the cost floor and moves it closer to the price. A policy floor is a target with nothing structural behind it, and it erodes. A structural floor is engineered into the pricing surface so that no commitment level produces a scheduled net price that erodes margin below target. It holds because the surface itself prevents the erosion, where a policy floor relies on deal-by-deal discipline.
Modeling work on capped-usage subscriptions establishes that exposure lives in the tail of the usage distribution, and across our engagements software usage has never been normally distributed. Data collected under a cap understates true consumption, so a policy floor is being set against an understated cost.
Theoretical work on how model suppliers should price finds that menus combining a fixed access fee with usage-based charges are revenue-optimal, and that higher-volume buyers rationally receive lower per-unit rates on a published schedule. That is the supplier layer’s rationale for tiered rate cards, and the same logic runs downstream when a vendor constructs its own commitment tiers.
Across our corpus the policy floor fails later than vendors expect. Customers sign the proposal and then, in the days before go-live, ask for a discount. The enterprise discount negotiation often runs well past signature, and a floor that held through the deal gives way at the last step.
Tests to Run Against Your Own Inference Exposure
The spend-doubling test
If your inference spend doubled next quarter, which part of your architecture responds? Does revenue move with it, or does margin compress?
The supplier-unit test
Which of your units is denominated in your supplier’s unit? A token, a compute unit, or a credit pegged to a token is a cost-aligned unit.
The cost-visibility test
Can you see cost per account, per edition, and per Customer Group, or only per month?
The unusable-run test
What happens commercially when a run consumes real compute and returns something the customer cannot use? Your contract already answers this question, and the answer may not be the one you intend.
The rate-cut test
If your supplier cut its rates tomorrow, who receives that? The answer is in the value metric.
If any of those questions surfaces a gap, a conversation with an SPP pricing expert is the fastest way to read what the exposure does to your price structure.