Author
TL;DR: A software license grant resolves into four dimensions: assignment, deployment, duration, and boundedness. Software license deployment rights are the ones most agreements answer by template rather than by decision, which is why a request to run the product somewhere else arrives at the deal desk as a discount conversation. Where the software runs and where the model runs are now two separate grant questions. Most agreements predate generative AI in the product. They answer only the first. The second is settled at the invoice, not in the grant.
Three weeks before signature, the customer’s security review comes back with one condition. The product has to run inside their own environment, on their infrastructure, under governance they have already approved. Not a different price. A different place.
The account executive routes it to the deal desk, the deal desk routes it to engineering, engineering says it is possible with a few weeks of work, and somebody takes a slice off the platform fee to cover the inconvenience. Nobody asks what the license now permits that it did not permit before, because the agreement carries one sentence about deployment and that sentence came with the template. Every part of what just happened was a licensing decision. None of it was made as one.
Deployment Is a Grant Dimension, and Most Agreements Answer It by Template
A software licensing model is two decisions rather than one. The value metric names the unit you sell. The grant defines what the holder of that unit may do. Strip the boilerplate out of the second half and it resolves into four dimensions: assignment, deployment, duration, and boundedness.
Three of the four come up for debate: assignment the first time someone asks whether a seat follows a person or floats across a population, duration when perpetual licensing gave way to term, boundedness the moment a large account asks for the word unlimited. Deployment arrives pre-answered. The template either hosts the software or permits the customer to install it on equipment it controls, and neither posture was chosen against what the company sells. It was inherited with the file.
That is why the request above lands where it lands. A deployment term conceded inside a deal is a licensing decision made by the customer’s procurement team on the vendor’s behalf. The trouble is not that procurement made the call. The trouble is that your side had no answer ready, so the absence was priced at whatever closed the quarter.
Licensing is the first of the three architecture decisions for a structural reason. Packaging groups what the licensing model made a unit of, and the pricing model computes against that unit. A deployment right nobody decided propagates into both.
The Four Questions Deployment Rights Ask
Deployment reads like one question. It is four, and they resolve differently.
Private deployment license
A private deployment license grants the right to run the software, and increasingly the model behind it, on infrastructure the customer controls. The decision underneath is physical versus virtual: the customer’s own hardware, their own cloud account, or a single-tenant instance you operate on their behalf. Agreements fold all three into one phrase, but they are three grants with three different verification problems: your compliance provisions now have to check something you do not run.
Model vendors already treat this as an explicit grant. Mistral publishes self-hosted deployment as a supported option for Mistral Medium 3. Application vendors more often permit it implicitly, which is permitting it on request.
In the patterns our library holds, the first private-deployment request is handled as an exception, the second as a precedent, and by the fourth the company has a deployment policy it never wrote down.
Geo-fencing restrictions
Geo-fencing is a usage and deployment restriction: a region, an address range, a time zone, a named territory. Two separate questions hide inside it. Usage location asks where the software may be accessed from. Deployment location asks where it may run. A customer with users in eleven countries and one hosting region has answered one and left the other open.
The restriction becomes a pricing question the moment the answer changes what the customer receives. OpenAI’s ChatGPT Enterprise plan supports data residency in ten regions. Residency there is a stated property of the plan. Treat hosting region as an implementation detail in your own agreement and you will meet that difference in a competitive deal, as a price problem.
Where export control or data-residency law binds, that belongs to counsel. What stays yours is whether the restriction is priced or given away.
Model access license
When your product invokes a model, whose model and on whose contract is a grant question before it is a cost question. A model access license decides which models the grant permits, and which party holds the contract with the model provider.
The market already draws that boundary in product. GitHub excludes open-weight models, and models not covered by GitHub’s data retention agreement, from default enablement regardless of policy setting. The permitted set is drawn on the model, and the criterion is which contract covers the data.
When the customer supplies the model instead, the arrangement is bring your own model, and what moves is what your unit covers rather than what it costs. That article carries the argument; it belongs here because a BYOM request and a private-deployment request come through the same door.
API access license
An API access license grants the right to access and call an API. The deployment question it raises is narrow: what does the grant permit programmatically, and what bounds it.
An API grant limited by nothing but fair use is the boundedness dimension arriving through the deployment door. The software has not moved anywhere. What changed is that something other than a person can consume the unit at machine speed, and the only brake in the paper is a phrase that means whatever the party reading it needs it to mean. Vendors find this out the first time a customer points AI agents at an endpoint scoped for a nightly sync.
Where Does Your Pricing Architecture Actually Stand?
A few questions return your pricing architecture score and show which of your licensing, packaging, and pricing decisions needs attention first. Real diagnosis, not a mailing-list toll.
Where the Software Runs and Where the Model Runs Are Two Grants
Deployment used to be a hosting preference. It is now two questions stacked on each other: where the software runs, and where LLM inference happens. Most agreements predate generative AI in the product. They answer only the first.
Sovereignty, data residency, and private inference are three vocabularies for one decision: whose infrastructure the computation happens on, and whose contract covers it. Industry analysts already expect sovereign AI investment to produce regional leaders rather than one global winner, the demand side of the same decision arriving at market scale. Answer it in the grant and the cost structure follows. Leave it open and the first enterprise buyer with an approved model governance program answers it for you, inside a deal, as a condition of signature.
The infrastructure half is not a legacy remnant your roadmap can wait out. The FinOps Foundation reported in February 2026 that 57% of practitioners manage private cloud, up from 39% the prior year. Customer-controlled infrastructure is expanding at the same moment LLM inference is becoming a material variable cost inside the product.
Peer-reviewed work with business buyers of cloud services found that governance, meaning security, control over where data sits, and the commitments made around it, forms a distinct component of the value those buyers perceive, one absent from consumer models of the same purchase. Buyers asking to run your product in their own environment are not primarily asking to spend less. They are asking for something they value on its own terms, and vendors keep answering with a discount.
The counter-argument is that regional hosting and single-tenant deployment are expectations now, so pricing them is pricing air. Take the premise as given. What a buyer expects to be available and what a buyer expects to be free are two different expectations, and only one of them is yours to concede. The collapse happens upstream, in how AI software is priced: the licensing decision was never made separately from the pricing decision.
What the Grant Should Answer Before the Deal Desk Sees It
When the grant is silent on deployment, the meter does the arguing, and it argues from a position you never chose. Whatever the product happens to check becomes the operative definition of what was sold.
The enforcement layer will not rescue this. The tooling category is right that a central license and entitlement management system is agnostic to deployment model, and that is the problem. It enforces a deliberate deployment grant and an accidental one with equal fidelity. Systems execute decisions. They cannot make them.
A common move is to vary the pricing model itself by geographic market or hosting environment: one model for the hosted product, a second for private deployment, a third for a region with its own rules. That multiplies the architecture instead of deciding it. The pattern that holds up across the companies we work with is a single licensing model whose grant names the deployment options it supports, with prices differing inside it.
Add a partner and the question travels with the unit. Channel, OEM, and white-label grants each add a hop, and a deployment right that stayed implicit in a direct agreement becomes something a second company exercises on infrastructure you have never seen.
None of this is a drafting exercise. The words belong to counsel. What belongs to you is the set of answers counsel drafts from, and whether those answers were chosen or inherited. If you are deciding which deployment options your product should support and what each one is worth, talk to a pricing expert and describe the agreement you sell today.
Tests to Run Against Your Own Agreement
Open the agreement you sell, find the section that says where the product may run, and ask three questions.
- If a customer asked to run this in their own environment tomorrow, does your agreement say what changes about the price, or only what changes about the deployment? An agreement that describes the move without pricing it has already given it away.
- Does your grant distinguish where the software runs from where the model runs? If it does not, a bring-your-own-model request and a private-deployment request arrive as the same conversation and receive the same improvised answer.
- Is your API grant bounded by anything other than fair use? If the only ceiling is a phrase, the unit you sell has no ceiling, and your customers’ usage will establish what the phrase meant.
Most companies answer the first and stall on the other two. Deployment is the dimension decided last and priced never. If your pipeline carries three deals asking to run the product somewhere your agreement never addresses, the sentence in your paper is not the problem: the decision behind it was never made. Talk to a pricing expert about what your grant says and what your deals keep asking it for, or see how we build pricing architecture.