Talk to an Expert

August 20, 2026 | Reading Time 9 mins

Bring Your Own Model Is a Licensing Decision, Not a Discount

TL;DR When a customer supplies the model behind your product, a value metric with LLM inference blended into it keeps counting and stops meaning anything. Bring your own model does not remove a markup; it nets a cost component out of the unit you license. The decision underneath bring your own model pricing is a licensing decision: what the unit grants, and who supplies the compute behind it. Price it before you answer the request, not at the first invoice.


An enterprise buyer is three weeks from signature when procurement adds a condition: the product will run against the customer’s own Azure OpenAI deployment, on the customer’s model contract, under model governance the customer has already approved. The vendor’s deal desk processes the request the only way its commercial model knows how, as a price concession, and takes something off the platform fee to close. The first invoice arrives and the credits count exactly as they did in the pilot. What changed does not appear in the count: the LLM inference those credits were built to recover is now billed to the customer, by the customer’s model provider. The unit kept counting. It stopped covering.

Bring your own model arrives as a discount request

Bring your own model, BYOM from here on, is an arrangement where the customer supplies the AI model the vendor’s software calls and pays the model provider directly for inference, while the vendor keeps licensing the application around it. Read structurally, the request changes what the unit grants and who supplies the compute behind it. It does not change the price. Vendors process it as a price concession anyway, because a concession is the only category their commercial model has for a request that arrives inside a deal, attached to a number.

The enterprise market has already run this shape once. In the customer-hosted era, vendors deployed their software into the customer’s own AWS environment instead of running it in their own. The customer absorbed the infrastructure cost, brought governance it had already built, and the software was priced on its own terms. Nobody folded the customer’s AWS bill into the license price, and nobody read customer-hosted deployment as a discount. The buyer-side drivers behind BYOM are identical, cost ownership plus already-approved governance, which is also why it arrives as an enterprise-driven motion rather than a one-off ask. A vendor who would never read “we host it in our VPC” as a discount request should not read “we supply the model” as one.

The reason so many vendors do is upstream of the deal, in how AI software is priced in the first place: when the licensing decision was never made separately from the pricing decision, every structural request gets translated into the only vocabulary available, the price.

The test: when a customer supplies the model, does your agreement already have an answer for what your unit covers, or do you find out at the invoice?

The metric was carrying the model cost, and nobody wrote that down

Most AI products are licensed on a value metric with inference blended into it: a per-request unit, a per-action unit, or a credit whose consumption rate varies by model class. That last shape is the surrogate-unit layer, an accounting unit that folds several underlying metrics into one, with a conversion table the vendor controls. A unit whose rate card changes when the model changes is a cost pass-through presenting as a value metric.

Salesforce runs the clearest public version of this structure. The Einstein Request, the unit Agentforce consumption is denominated in, carries multipliers tiered by model capability, and the bring-your-own-LLM multiplier is lower than the standard one. Salesforce’s own rate-card FAQ, effective October 2025, lowered the bring-your-own-LLM multiplier from 7 to 4, in its words to offer more competitive pricing.

Read that rate card for what it says structurally. The multiplier is indexed to the model class, which means the metric was carrying the inference cost the whole time. Bring your own model did not remove a markup Salesforce was taking on tokens it never touched. It netted a cost component out of the unit itself. That is the mechanics of most bring your own model pricing across the market: not a concession, a recalculation.

The test: if your metric’s rate card changes when the model changes, whose cost is the meter tracking?

Bring your own model breaks the cost linkage, and the metric does not notice

A token-derived or inference-indexed value metric stops tracking your cost the moment the customer supplies the model. Nothing in the counting machinery registers the event. The unit keeps counting; it just stops meaning what it meant. And no downstream computation repairs it, because the pricing model is the wrapper that computes against the unit, not the thing that gives the unit meaning.

The exposure this opens is boundedness, not price shape. An unbounded metric tethered to a cost you paid had a brake built into it: your own model bill rose with the count, and that tension forced conversations before the numbers ran away. An unbounded metric tethered to a cost someone else pays is a different instrument entirely, and your agreement is the only place the difference gets recorded. If it is not recorded there, it is not recorded anywhere.

An outcome metric assumes you control the thing producing the outcome

If your unit is a resolved case, a completed task, or a delivered outcome, and the model producing that outcome now belongs to the customer, the denominator of your value metric has moved outside your control. Model swaps, provider changes, and capability drift now move the count your price attaches to, through decisions you do not make. That is a metric-design fact, not a legal allocation, and it exists whether or not anyone writes a clause about it.

The test: if your largest customer brought their own model tomorrow, which line on your invoice would still be defensible on its own terms?

Where Does Your Pricing Architecture Actually Stand?

A few questions return your pricing architecture score and show which of your licensing, packaging, and pricing decisions needs attention first. Real diagnosis, not a mailing-list toll.

Splitting the bill is not the same as deciding what you sell

The market’s standard answer to bring your own model pricing is two line items: a platform fee for the software, inference passed through at cost or carried by the customer directly. That answer describes an invoice format, nothing more; whether it prices the product correctly is the question this article exists to ask. GitHub already prices per-token model usage at parity with the model providers’ own rates and supports bring-your-own-key arrangements across its Copilot plans; its credit-and-passthrough architecture shows the same separation shipping at scale. And the separation itself is not novel. In customer-hosted deployment nobody commingled the customer’s infrastructure bill with the license fee. BYOM is that separation returning at the model layer.

A scoping note before the argument: BYOM is only on the table when the model is genuinely swappable. In some products the inference is the value story, not a cost line under it: the model is tuned to the workflow, the outputs are the product, and moving the work onto a customer’s account or a different model changes what comes out and how reliably it comes out. A customer cannot bring their own model to a product whose behavior is the model. This article is about the other kind of product, where the application layer carries the value and the model behind it can change without changing what the customer bought. Knowing which kind you are selling is itself a licensing question: it is a statement about where the value your unit grants actually lives.

But the split is an output, not the decision. It answers what the customer pays and leaves untouched what the unit grants. Economic theory of two-part pricing has long held that the fixed component is where value gets captured while the variable component tracks cost. BYOM does not create that structure for a vendor. It reveals that the vendor already had one, and had been hiding the fixed half inside the variable price.

The test: strip the inference component out of your price. Is what remains a number you can defend on value, without reference to cost? If the answer takes longer than a sentence, that is a conversation to have with a pricing expert before the next BYOM request lands, not after it closes.

The cost story is the one argument you cannot get back

Cost justification is a fairness argument, and BYOM permanently removes yours. Peer-reviewed research on price fairness has found that buyers judge a price change carrying a cost story very differently from an identical change with no cost story behind it. What we observe in enterprise renewals follows the same asymmetry: the cost narrative is the argument those conversations lean on first.

Once you no longer buy the tokens, that argument is gone, at this renewal and every renewal after it. The customer can see your marginal cost on their account, because they are paying it, to someone else. Every argument you have left is a value argument, and value arguments only work when the unit denominates value.

The test: what is the sentence you say at the first post-BYOM renewal, when your costs did not move and you still need the number to?

Run bring your own model through the grant, one dimension at a time

A software licensing model is two decisions: the value metric names the unit you sell, and the grant defines what that unit entitles the buyer to do, across four dimensions. BYOM lands on all four.

Assignment. Who, or what, may invoke a customer-supplied model through your product, and whether that right is scoped the way your existing user and agent definitions assume.

Deployment. Where the model runs and whose infrastructure it sits on. This is where the customer-hosted precedent lands formally: running vendor software in the customer’s cloud was a deployment decision, and it was priced as one. BYOM is the same decision arriving at the model layer.

Duration. Whether the arrangement survives a renewal, a model swap, or a provider change, or whether each of those reopens the commercial question.

Boundedness. Whether the counted quantity still has a ceiling once the cost that implicitly bounded it belongs to someone else.

Which dimension each BYOM question lands in can be named from the outside. Where each answer lands is architecture work, specific to what you sell and to whom. When the grant is silent, the meter does the arguing, and under BYOM the meter is arguing from a number that no longer means anything.

The test: which of the four does your paper already answer for a customer-supplied model?

The questions to run before you say yes

Three questions, none of them fixable by revising a number.

First, open the agreement behind your largest AI deal and find the sentence that says what a unit covers when the model behind it is not yours. If the sentence does not exist, the operative answer currently lives in your product telemetry, which means it was never decided.

Second, subtract the inference component from your unit price and defend the remainder to a CFO who knows your model bill for their account is zero. What survives that conversation is your software’s value story, or the absence of one.

Third, take the four grant dimensions and ask which your paper already answers for a customer-supplied model, and which the next enterprise buyer will answer for you.

A bring your own model request is not a threat. It is the market handing you the licensing question you should have answered at design time, with a deadline attached. If one is sitting in your pipeline now, sequence matters more than speed: the grant decision has to be made before the price conversation can mean anything. Talk to a pricing expert, describe the unit you sell and the request on the table, and a pricing architect will read what your metric has been carrying before you commit to what it carries next.

FAQs



Linkedin X (Twitter) Facebook

Ready for profitable growth?

Hit the ground running and learn how to fix your pricing.

Book A Demo Contact Us