Author
- Per-call pricing worked because every call was the same work
- The call stopped being CRUD when business logic moved in
- Connectors became applications, and applications do not price like interfaces
- The API pricing model decision: pass the cost through, or recast the metric
- The same decision, one layer up
- FAQs
TL;DR Per-call API pricing was correct, not naive. While a call meant create, read, update, or delete, every call was the same work, and the call was a real value metric. When business logic moved into the layer, calls stopped being interchangeable, and every API vendor faced a binary no rate card could dodge: pass the underlying cost through in the unit you buy, or recast the metric on what your layer now does. Per-call pricing survived only in the categories where the work stayed uniform. The same decision is playing out again one layer up, in products denominated in a model vendor’s tokens.
For about a decade, API pricing was the simplest rate card in software. A gateway counted requests, volume tiers thinned the unit price, and the bill arrived in pennies per call. That rate card did not fail because vendors priced badly. It failed because the thing being counted changed underneath it, and the pricing conversation never caught up.
Every API vendor whose layer thickened faced the same choice: keep billing in the unit underneath, or recast the price on what the layer had come to do. Our AI software pricing pillar makes the category-level argument that most of what the industry calls pricing models are payment wrappers around an unexamined unit. The API era is where software first learned that lesson at scale, and asking for an API pricing model without asking what the call has become repeats the era’s original mistake.
Per-call pricing worked because every call was the same work
The early API economy billed the way it did because the unit deserved it. A call meant create, read, update, or delete: CRUD, the four verbs of moving a record. One call fetched a customer record, the next wrote one back, and nothing about the second call did more work than the first. Cost to serve was flat across requests. Value delivered was flat across requests. When the work is uniform, the count of that work is a legitimate value metric, and per-call pricing was exactly that: a real metric matched to real work.
Peer-reviewed work on the pricing of information goods backs the instinct the era ran on. Flat-fee and usage-based structures each win under identifiable conditions, and uniform, countable work is precisely the condition that favors metering usage. The standard advice that pay-as-you-go suits API calls because every action can be counted and priced is true on exactly that condition: counting prices well only when the counted things are alike.
What made an API call a countable unit of value?
Three properties, all quiet enough to go unnoticed until they disappeared. The calls were interchangeable: any request cost the vendor about the same to serve. They were legible: a buyer could look at last month’s volume and estimate next month’s bill within a rounding error. And consumption tracked delivery: the customer received exactly what the counter counted, a record moved. When those three hold, the argument between vendor and buyer is only ever about the rate, never about the unit. That is what a working value metric feels like from the inside: invisible.
Where per-call API pricing still holds today
Per-call API pricing did not die; it retreated to the territory it was built for. Payment authorization, address validation, identity verification, message delivery, tax and rate lookups: categories where a call is still one uniform, completed piece of work. The survivorship pattern is the whole argument in miniature. The unit outlived its era only where the work stayed uniform, and where you can still buy an API by the penny, you are looking at a call that never stopped being CRUD.
The call stopped being CRUD when business logic moved in
Somewhere in the platform era, the layer between systems started doing more than ferrying records. Validation moved in, then transformation, then enrichment, then orchestration: the same request now called three systems, reconciled their answers, and handed back a decision instead of a row. The request on the wire still resembled a call. The work behind it no longer did.
Why one price per call breaks under heterogeneous work
Picture, as an illustration, two requests hitting the same platform at the same posted rate. One updates a phone number on a record. The other takes an inbound lead, verifies it, enriches it from two sources, scores it, and routes it to the right owner. Same price, radically different cost, radically different value. A vendor in that position is averaging. The light calls subsidize the heavy ones, the rate is set against a mix, and the mix belongs to the customer, not the vendor. Buyers with light workloads discover they are overpaying for plumbing and negotiate or leave. Buyers with heavy workloads discover the bargain and lean in. Margin now moves with a distribution the vendor neither sets nor sees in advance, which is another way of saying the unit no longer describes the work.
The symptoms: call classes, rate limits, policy retrofits
The market showed its tells before anyone named the disease. Rate cards fractured into call classes, with a read priced differently from a search and a search differently from an enrichment. Endpoint-specific pricing appeared, a quiet admission that a call is no longer one thing. Rate limits started doing pricing’s job, capping consumption the price had failed to bound. And flat plans sprouted fair usage policies, the boundary a plan grows after the fact when its price never carried one. Each symptom is the same event: a unit built for uniform work, patched to survive heterogeneous work.
Connectors became applications, and applications do not price like interfaces
The second act was quieter and more consequential. The layers that validated, enriched, and orchestrated did not stop there. A connector that reconciles invoices across two systems of record, flags the exceptions, and posts the clean entries has not integrated an accounting workflow. It has become one. Customers stopped buying access to systems and started buying finished work, and the vendors who noticed stopped describing themselves as integrations at all.
At that point the pricing question changes species. An interface prices access; an application prices capability. A vendor still billing per call for what has become an application describes a shrinking fraction of its own product on every invoice, and no discount schedule fixes a unit that measures the wrong thing.
The replacement test: connector or hire?
The test question we put to teams in this position is short: if your layer disappeared tomorrow, would the customer replace it with a cheaper connector, or would they hire someone? If the answer is a hire, the layer is an application, whatever the architecture diagram calls it.
Where Does Your Pricing Architecture Actually Stand?
A few questions return your pricing architecture score and show which of your licensing, packaging, and pricing decisions needs attention first. Real diagnosis, not a mailing-list toll.
The API pricing model decision: pass the cost through, or recast the metric
Ask the industry for an API pricing model and the answer arrives as a menu: flat fee, per-unit, editions with usage allowances, usage with overage, credit-based, freemium, hybrid. The menu is real, and it is the wrong first question. Every entry on it is a wrapper around the value metric decision, the choice of which unit the price attaches to, and the menu never surfaces it. We make that argument in full in pricing model vs. value metric, so here it is enough to say the wrapper is not the decision. The same goes for the checklist advice that model choice depends on a handful of market factors: every factor on those lists operates downstream of the unit, and the unit is the choice the API era forced into the open.
Vendors whose call stopped being CRUD met that choice in one of two forms.
What pass-through pricing actually prices
Pass-through keeps billing in the unit underneath: the call, the record, the compute the layer consumes, marked up to carry margin. It describes cost and says nothing about value. It is also legible and easy to defend on day one, which is why thin layers rightly choose it. The strain arrives with thickness. The invoice describes what the vendor consumes while the product delivers something the unit never mentions, and a buyer who can see the underlying meter will eventually price the layer as plumbing plus markup, because that is what the unit told them it was. The industry’s current consensus answer, a subscription base with metered overage, does not resolve this. That shape is the oldest structure in pricing theory, a fixed fee with a usage charge behind it, and it works exactly as well as the unit inside the meter. A base-plus-meter shape wrapped around the wrong unit is the wrong unit with a floor under it.
What recasting the value metric requires
Recasting means choosing a unit denominated in what the layer itself does: the record enriched, the exception resolved, the workflow completed. The frames for judging such a metric are buyer-side, and the API era proved them. The buyer has to understand the unit without a diagram. The buyer has to be able to estimate their own volume before signing; peer-reviewed field evidence shows that buyers facing usage uncertainty gravitate toward predictable plans and will pay for that predictability, so a metric a buyer cannot forecast pushes deals into pilots and stalls them there. And the unit has to diverge from your cost structure, because a metric that restates cost is pass-through with extra steps. Notice what is absent from that list: your infrastructure. The observation that per-seat licensing makes no sense for an API because nobody sits in a seat is true and settles nothing; the missing seat opens the metric question rather than answering it, a thread we pull in per-seat licensing variations. Which unit is right for a given layer is diagnosis, not doctrine. We walk the selection reasoning in AI pricing model selection, and the metric decision sits at the top of the larger architecture of licensing, packaging, and pricing we map in software monetization.
Is a credit layer a recast? Usually it is a deferral
Credits present as a third door, and mostly they are not. A credit is a surrogate unit: it denominates accounting rather than value, folding many kinds of work into one purchasable number while a conversion table decides how fast each action burns it. That can be sound engineering layered over a chosen value metric, and it can equally be a way to avoid choosing one, with the conversion table quietly carrying the pass-through underneath. The unit is not the problem; the deferral is. We take the full argument up in credit-based pricing for AI. For this history it is enough to say that a vendor who cannot name the value metric underneath its credits has not recast anything.
The same decision, one layer up
The reason to tell the API era’s story now is that it is running again under new vocabulary. The thickening layer of the moment wraps retrieval, planning, tool use, and verification around a large language model, and the default pricing shape for that layer is denomination in the model vendor’s tokens: the supplier’s unit, passed through with margin, sometimes hidden behind credits, which moves the accounting without moving the decision. The vocabulary is new enough that teams experience the question as unprecedented, and the going advice makes it worse by recommending that a product built on a model vendor’s API should price the way the model vendor does. That is not a strategy. It is the pass-through door taken by default, pricing the LLM inference underneath instead of the work the layer finishes.
Whether pass-through is wrong for any given product is not the claim; thin wrappers exist now as they did then, and for a thin wrapper the supplier’s unit fits the transaction. The claim is that the decision is the same decision, and the API era already showed how it ends. The unit survives where the work stays uniform. Everywhere else, the vendors who recast onto their own metric become application companies, and the rest get priced as markup on a meter everyone can see. The portable test costs one sentence: whose unit denominates your price, yours or your supplier’s? If the answer is your supplier’s, and your product long ago started finishing work that unit never describes, talk to a pricing expert and run the whose-unit test on your own rate card before a renewal cycle runs it for you.
The API era’s tuition has already been paid, and there is no reason to pay it twice. If your price is denominated in a unit you buy rather than a unit you deliver, the decision this history describes is already on your desk. Talk to a pricing expert: describe the layer you are pricing, and a pricing expert replies with a read on which side of the binary it sits.