Talk to an Expert

August 4, 2026 |

When the Agent Is Wrong: Risk Allocation in Outcome and Consumption Pricing

Author

TL;DR An AI agent burns real compute and returns something the customer cannot use. Somebody pays for that run. Every consumption or outcome contract for an agent already decides who, and most decide it by accident. The cost of a wrong answer lands in one of three places: on the customer, on the vendor, or split through an intermediate unit. Where the meter sits allocates it. What you charge per unit only sets the size of it. And a meter that counts effort cannot tell a right answer from a wrong one, so on an effort meter the bill rises in exactly the case where value fell.


An agent is dispatched to do a piece of work. It plans, calls tools, retries a failed step, burns tokens, and returns an answer that is confidently wrong. Pressed on it, the agent produces the most honest sentence in the entire transaction: “Honest assessment: I completely hallucinated that.” The customer throws the answer away. The invoice, which hallucinated nothing, arrives anyway: LLM inference the vendor bought and compute somebody rented.

Somebody pays for that. There is no version of the contract where nobody does. Who pays when an AI agent fails is a question the contract settled long before anyone argued about it.

Every Agent Contract Already Decides Who Pays for a Failed Run

Pricing teams rarely meet this as a design question. They meet it as three complaints from the field. Retries: one customer-visible task generates a long tail of invisible work underneath it, and whether those attempts reach the invoice was decided by a developer picking what was convenient to instrument. Partial completion: the agent covered most of the distance and stopped, and nobody wrote down what fraction of the work that represents, so the meter counted it as everything or nothing. Abandonment: the customer watched the run go sideways and killed it. The compute was consumed. The value was not.

Every consumption or outcome contract for an agent already decides who pays for a failed run, and most of them decide it by accident. Whether a customer disputes a successful outcome, arguing your share of a result down at renewal, is a different problem, treated separately in designing outcome-based pricing. Attribution is a fight about who gets credit for a win. This is about who absorbs a loss.

Who Pays When an AI Agent Fails: Three Places the Cost Can Land

There are exactly three, and every agent product sits in one of them whether or not the team chose it.

The customer pays for the attempt

Meter on effort and you charge for the attempt. Tokens, runs, agent-minutes, compute consumed: all count work performed rather than work completed. Revenue tracks cost almost perfectly, and the failure lands entirely on the customer, who paid full price for an answer they discarded. It is the market default because effort is what the instrumentation already emitted, not because anyone chose it.

Coarse units go a step further: they can put margin on the miss. Salesforce opened Agentforce at $2 per conversation, and a conversation that collapses after two turns bills the same $2 as one that runs to resolution, so the failed attempt consumed less inference than the unit priced. Round the meter up far enough and failure is not merely passed through to the customer. It is marked up. Replit ran this allocation in public. In June 2025 its agent moved from a flat rate per checkpoint to effort-based pricing, variable checkpoint pricing that scales with how hard the agent worked, with simple tasks landing below the old flat rate and complex ones above it. Pricing the effort prices the failures with it: a run that loops, errors, or unwinds its own mistakes is expensive precisely because it worked hard, so the meter hands the customer the full variance of the agent’s reliability. It is the purest version of paying for the attempt yet shipped. The reception proved the allocation: users were billed for runs that failed or looped on the agent’s own errors, and within a month Replit conceded the transition did not meet its standards, auto-refunded a day of mis-computed charges, and issued blanket credits. The remediation ran through the price lever while the meter stood, which is the distinction this article exists to make: a refund apologizes for the bill; only the meter decides who pays for the next failure.

The vendor pays for the miss

Meter on a verified result and you charge only when the result arrives: a resolved case, a completed workflow, a qualified record. The customer’s failure cost falls to zero, which buyers notice and say so loudly. What they do not see is where the exposure went. The vendor now carries the variance of its own model’s reliability plus the cost of every attempt that never resolved. Meters that count only successful invocations are already shipping, and some agent products state plainly that they do not bill when the agent cannot answer. Those vendors did not eliminate the failure cost. They bought it. The claim usually made for it is that it removes consumption guesswork for the buyer. It does. Removing guesswork for one party is the same act as adding variance to the other.

The intermediate unit

The third option prices something between attempt and result: a completed step, a verified intermediate artifact, or a surrogate unit whose conversion table treats failed and successful consumption differently. The vendor lane argues credits are justified when an agent’s output cannot yet be expressed as a countable value metric. Half right. A surrogate unit is a legitimate variance buffer when it is deliberately engineered. Left unengineered, it becomes a dial the vendor re-rates whenever reliability moves, which is a re-rating lever rather than a risk allocation.

Each of the three puts the cost of failure somewhere specific. None removes it. A vendor who believes an outcome model eliminated the customer’s exposure has moved it onto its own income statement, which may be the right trade and is never neutral. Charging for the attempt is fair when the customer directed work that genuinely consumed value, and hard to defend when they directed nothing except a request the software failed to satisfy. Where a product sits between those cases is the question to settle.

Where the Meter Sits Decides Who Absorbs the Error

Risk allocation for agent failure is settled by the licensing model, not the pricing model. Choosing what to count determines who carries the failure. Choosing what to charge per unit only determines how much it costs them. That ordering governs every decision in AI software pricing: licensing model, then packaging model, then pricing model.

SPP’s term for the vendor-side version is execution risk: what a vendor takes on when the value metric sits so far downstream that the customer’s own behaviour, not the software’s performance, decides whether the vendor gets paid. Agent failure adds a second party, because the model’s own reliability now sits in that same position.

The frame is the metric spectrum, running from pure activity to pure outcome. Every position on it is a risk-allocation choice, not a maturity ladder to climb. The three decisions applied to agents are worked through in agentic AI pricing strategy; what this piece adds is that the failed run is one of the things the licensing decision silently allocates. Readiness tests for outcome pricing ask whether the value metric is aligned, attributable and predictable. None asks what happens when the agent fails.

Copying a competitor’s agent pricing copies their risk allocation. Their model was fitted to their reliability, their cost structure, and their customers’ tolerance for variance, and the fitting step is the part that does not come along. If you cannot say which allocation their model makes, talk to an expert before it becomes yours too.

Why a discount does not fix a risk-allocation mistake

Price is the fastest lever a team owns, so it is the one they reach for. It is the wrong one. A discount compensates one customer for one bad month; it does not change which party the architecture assigned the failure to, so the same conversation returns next quarter under a different account name. Price makes the wrong allocation cheaper to live with, which buys quiet and changes nothing.

Where Does Your Pricing Architecture Actually Stand?

A few questions return your pricing architecture score and show which of your licensing, packaging, and pricing decisions needs attention first. Real diagnosis, not a mailing-list toll.

A Meter That Counts Effort Cannot See a Wrong Answer

A confidently incorrect answer consumes the same compute as a correct one. Often more, because wrong paths run longer before they terminate. On an effort meter, correct and incorrect are the same line item, so the bill goes up in precisely the circumstance where the customer’s value went down.

The result meter has the inverse problem. It can separate success from failure only if a verified definition of success exists and both parties trust it, which pushes the question back to who verifies and on what evidence. That verification is a real cost, and it belongs inside the outcome model.

A system that cannot trace a charge back to the request that produced it certainly cannot classify that request as failed. A pricing model requiring a distinction the product cannot make is a model the invoice cannot defend, and a value metric the product cannot evidence is one the company has not actually adopted.

Reliability Is a Variable, Not a Constant

Agent reliability is not a fixed property of the product. It moves as models improve, as prompts and tools get tuned, and occasionally backwards when a model change lands badly. A contract written when the agent resolved a modest share of requests is a different economic deal once it resolves most of them, and neither party controls that drift.

What happens to the bill as the agent improves

The direction depends entirely on the meter, and it runs silently either way. On an effort meter, rising reliability transfers value to the customer: fewer wasted attempts means more usable result per dollar, a price cut the vendor never decided to give. On a result meter, direction depends on what bounds the count. Where the unit is bounded by the customer’s own demand, per resolved incident in a queue only their business can grow, a more reliable agent converts more of that fixed demand and every increment replaces cost the customer was already carrying. The bill rises toward a ceiling the customer controls, which is expansion, not repricing; the residual conversation is spend predictability, which commits and caps exist to settle. Where the unit is something the agent itself can multiply, actions, steps, artifacts, units where the agent decides how many billable results one request becomes, improvement can raise the count for the same intent. A customer writes in with a wrong invoice, a broken login, and a refund request; an agent that runs its own triage files three tickets, resolves each, and bills three resolutions for one contact. The split is defensible support practice, which is exactly the point: the agent, not the customer’s demand, decided how many billable units the incident became. That is the silent price increase, and it is a boundedness defect in the metric definition, not a rate change.

This is the durability question for agent pricing, and the argument for building the architecture so the allocation can be revisited on a known cadence rather than discovered at renewal. It belongs inside the rollout discipline covered in managing risk in software pricing. The vendor who prepares for this owns the renewal conversation: know the answer, before anyone asks, to what happens to the customer’s bill as the agent improves.

What the research on outcome contracts actually supports

Peer-reviewed modelling of performance-based payment structures found that moving payment closer to the buyer’s outcome relocates the incentive problem rather than removing it, because the party who controls whether the outcome occurs is not the party being paid. An outcome meter makes the vendor accountable for a result the customer’s own inputs partly determine, which is the source of the disagreement both sides eventually have. Related modelling of contracts under unobservable effort found the best structure depends on who can influence the outcome and whose effort can be observed, so none is correct for everyone.

It cuts the other way too. Empirical research on performance-based service contracts found that tying payment to a delivered result improved the reliability of what was delivered. That evidence comes from a very different industry, so it transfers as a direction rather than a magnitude, but it argues for outcome structures rather than against them. Structure changes behaviour on both sides of the table and supplies no universal answer about which to pick. Neither does this article.

The Questions to Settle Before the Deal Is Written

Five questions. If your team cannot answer them from memory, the contract answered them and the company did not.

What does your meter charge for when the agent fails? Real cost consumed, nothing usable returned. If nobody can answer without reading the billing code, the code decided it for you.

Can your system tell a failed run from a successful one at invoice resolution? Not in a dashboard. On a line the customer will read.

What happens to the customer’s bill if the agent gets meaningfully better, and what happens to yours? One of those numbers moves without anyone deciding it should.

Who directed the failed work? A run the customer scoped badly and a run the model handled badly are different events, and a meter that cannot tell them apart assigned both to one party.

What is your answer to a customer who says they will not pay for a wrong answer? If your reply is that the contract does not address it, then it does, and the answer is that they will.

Notice what these questions are not. There is no ratio, no allowance, no correctness level at which a rule flips. Those are sizing decisions, and sizing belongs to the engagement rather than the frame. A fixed fee wrapped around unbounded agent consumption is a real exposure too, but that is a boundedness problem, and it belongs with AI cost control architecture.

None of the three allocations is wrong. Which one fits depends on your reliability, your cost structure, and what your customers tolerate. Choosing one on purpose is the discipline, and choosing it before the meter is built is what keeps the choice available.

If you are designing an agent pricing model now, or repairing one already generating these conversations, settle the failure case before launch rather than after the first invoice dispute. Talk to an expert, describe where your agent fails and what you bill for it, and we will tell you which allocation your architecture already made.

FAQs

Ready for profitable growth?

Hit the ground running and learn how to fix your pricing.