TL;DR A benchmark is the average of other people’s outcomes with the context stripped off, and the context is the part that decided the outcome. The companies inside one count different value metrics, package capabilities into different things a customer can buy, serve different Customer Groups, and carry different cost to serve, so the median describes no architecture, including the publisher’s. Most benchmarks report published positions rather than realized net prices, and they sample only the survivors. Pricing ground truth is the alternative already in your systems: your deal record, won and lost, at line-item resolution. The losses are what no average can hold.
A pricing decision is due. Someone opens a slide with a median on it and the room relaxes. That number came from outside the building, nobody present has to defend it, and for the next hour it is the only evidence that cannot be attacked on political grounds. Benchmarks come to set prices because they arrive first and they arrive clean.
The reflex is understandable. The instrument is weak. What outranks it is already in your own systems, in the deals you won and the deals you lost. That record is what we call pricing ground truth, and the difference between it and a benchmark is one of kind: the difference between an average and a pattern.
Why Teams Reach for the Benchmark
The first problem it solves is political. Internal pricing evidence is contested: sales reads the loss data one way, finance reads margin another, and every internal number arrives attached to whoever produced it. A benchmark has no author in the room, which makes it the easiest thing in the meeting to agree on, and that is why benchmarks travel further than better evidence.
The second is speed. Building a real read of your own record takes work: pulling deals at line-item resolution, joining won to lost, reconciling what the pricebook said against what was signed. A median is a search away, and when the decision is due Thursday the fast instrument wins.
Both are reasons a benchmark gets used. Neither is a reason it should set a price.
What a Benchmark Is an Average Of
An average across architectures describes no architecture
The companies inside a pricing benchmark are not variations on one design. They count different value metrics: per named seat, per active user, per transaction processed, per environment deployed. They package capabilities into different things a customer can buy, and that structure varies more than the word packaging suggests: an edition ladder, a set of modules, a platform with apps around it, add-ons on a base, one all-in-one offer. They serve different Customer Groups, and they carry different cost to serve, which for anything with generative AI inside it now varies by an order of magnitude.
Average across that and you have averaged architectures, not prices. Two companies charging the same headline number per unit are not charging the same price if the unit is not the same unit, which is what sits underneath every claim of being below the industry median on price per seat. It is also why the first move in value-based pricing strategy is establishing what you count.
Published positions are not realized net prices
The second defect is the one practitioners feel. Benchmark data is collected from published sources, so it captures list price and not net price, and the distance between those two numbers is company-specific.
The method is usually disclosed, and the disclosure is more damning than any critique. One family of regional state-of-SaaS-pricing reports says it plainly: they visited each company’s website, collected the pricing details by hand, and reviewed the result with their in-house experts. That is a careful description of scraped list prices. Every condition that produced them is gone: no realized price, no discount taken, no deal outcome, no record of who the buyer was. What survives is a wall of published positions, sold on the number of pricing pages examined, answering questions like how many plans belong on a pricing page. Volume does not repair it, because how many editions belong in your packaging is a question about how many distinct Customer Groups you serve.
A benchmark reports the survivors
A benchmark samples the companies still in market at the moment of collection, so the architectures that failed are not in the distribution. You are looking at outcomes with the failures filtered out, which is a different object from a distribution of choices.
The consequence for a decision is precise. A benchmark can tell you that you are unusual. It cannot tell you whether being unusual is the problem or the point. Peer-reviewed work on price discrimination and self-selection is clear that differentiated structures are supposed to sit apart from an undifferentiated average, because separating buyers by what they choose is the whole mechanism. Being outside the benchmark is not evidence of a pricing error. It is the outside view of a differentiated architecture.
A Benchmark Strips Conditions. A Pattern Carries Them.
There is an obvious objection to all of that, and a careful reader reaches it in one step. SPP reads a company’s deal record against a pattern library built across decades of B2B software pricing work. Is that library not itself an aggregate of other companies, and if aggregates are weak, why is that one strong?
Fair question. The answer is not that ours is bigger. Scale is the benchmark’s own defense, and borrowing it concedes the frame this argument rejects.
A benchmark and a pattern are different objects. A benchmark reports a central tendency and strips the conditions that produced it: what the companies were counting, who they sold to, what they competed against, what the buyer refused before agreeing. A pattern is a conditional regularity that carries its conditions with it. It does not say what the number is. It says how a mechanism behaves and under what circumstances, which is exactly the part a median deletes. When we observe that a particular discount behavior tends to follow a particular pricebook structure, the useful content is not an average discount. It is the relationship and the conditions it held under, and that is something you can look for in your own record.
That decides the unit of comparison, which is the whole argument. A benchmark invites you to price against another company’s number. A pattern helps you read your own. Your record stays the thing being measured. The library supplies interpretation, never a target, and never another company’s specifics. Nothing surfaced in your read references anyone else’s deals, which is a structural property of working from derived patterns rather than a promise about handling.
So the claim is not that SPP holds a better average. It is that averages are the wrong instrument here and conditional regularities are the right one. That cuts both ways, and we say it publicly so it can be held against us: if we ever publish a data artifact of our own, it will carry patterns with the conditions under which they held, not medians for anyone to price against.
Talk to an Expert About the Pricing Problem in Front of You
Describe what you’re facing and a pricing expert will reply with a concrete read on your licensing, packaging, and pricing architecture.
Pricing Ground Truth Is Your Deal Record, Won and Lost
Pricing ground truth is your own deal record at line-item resolution, won and lost, carrying the configuration that was offered, the discount that was actually taken, and the outcome that followed. Not summary data. Not a pricing page. Every line on every deal, including the deals that died.
Three things people call pricing evidence
Three other things routinely get asked to do this job.
What buyers say. Stated willingness to pay, collected through surveys and price-sensitivity instruments. The peer-reviewed record on willingness-to-pay measurement is consistent that stated values run above actual payment behavior, and the treatment lives in why willingness-to-pay surveys fail B2B software companies.
What won customers were invoiced. Billing exhaust, a record written entirely by buyers who accepted a price, as argued in why billing data is not decision evidence.
What other companies reportedly charge. The benchmark, covered above.
These are not four grades of the same evidence. They are four different things, and three of them are routinely asked to do a job only the fourth can do.
The losses are the difference
A billing system has no row for a deal that died. A deal record does, and that single difference is why one can calibrate a pricing decision and the other cannot. The configuration a buyer walked away from, the discount your deal desk declined, the quote that went quiet, the packaging you lost on last quarter: each marks a boundary where your pricing stopped working, and boundaries are what a pricing decision is about.
Against a benchmark, the deal record holds four things no aggregate can: your customers, your configurations, your competitive set, and the shape of the deals your architecture failed to close. It also holds the distance between what your pricebook said and what your reps signed, which is its own diagnostic and the subject of pricebook deviation.
Ground truth is not a report, not an outside opinion about your company, and not a synonym for analytics. It is the record itself, read properly, and reading it is a starting point rather than an answer. It shows where the architecture issues are and roughly what fixing each is worth, which is to say it tells you which questions deserve the money. It does not hand you the redesign, which is why the Pricing Ground Truth engagement is scoped the way it is.
This does not reverse the position on billing data
We recently argued that billing data cannot calibrate a pricing decision. Now we are arguing that your own transaction record is the strongest evidence you hold. Those are the same position, not opposite ones.
Billing exhaust is a subset of the deal record, and it is the winners-only subset. It shows how customers who already said yes consume and what net prices your closed deals realized, and it is structurally blind to refusal. Naming the deal record as ground truth does not promote billing exhaust to referee. It names the larger record that billing exhaust is a fragment of, and the fragment is missing precisely the events a price gets set against. For a read on which parts of that record your company already holds and which it never collected, talk to a pricing expert. That mapping is a working conversation, not a data export.
The One Thing Your Own Record Cannot Do
Here is the limit, stated without hedging, because an article that names only its own instrument’s strengths is an advertisement.
Your record covers the markets you have sold into, the configurations you have offered, and the prices you have asked for. It cannot tell you what would have happened at a price you never quoted, or in a market you never entered. Peer-reviewed work on demand estimation is clear on the point: behavior observed under one price schedule identifies behavior under that schedule, not at prices that were never on the table. That is also the boundary of any read on price elasticity from your own deals.
The gap is real, and it is exactly the gap a benchmark is reaching for. The reflex is answering a legitimate need badly, not inventing one.
What fills it properly is patterns rather than averages. Patterns carry the conditions under which they held. Averages carry none. Reading your record against patterns accumulated across decades of B2B software pricing work is a different act from comparing your numbers to a median, because the pattern brings circumstances and the median brings a number. Doing that as a standing practice is the subject of continuous monetization.
One further limit. A company with very few closed deals does not yet have a record to read, and pretending otherwise is its own error.
When a Benchmark Earns Its Place
None of this makes benchmarks worthless. It makes them context rather than basis, and three uses hold.
Order-of-magnitude sanity checks. If a proposed price sits an order of magnitude away from everything else in the category, you want to know that before the meeting, and a benchmark tells you at almost no cost.
Trained buyer expectations. Benchmarks shape what buyers expect whether or not they are accurate, and peer-reviewed work on price elicitation shows that a disclosed reference number moves the values people state. That makes a circulated benchmark real evidence about the negotiation environment you sell into, even when it is poor evidence about price.
Board and investor conversations, where a comparison is being asked for and refusing costs more than supplying a caveated one.
All three treat the benchmark as context. None treats it as the basis for a price, and the slide from the first to the second is what this article exists to stop.
The questions to ask before a benchmark sets a price
Three of them are the defects above, turned around. What are the companies in it counting, and is their value metric yours? Is the number published or realized? What happened to the companies that are not in the sample?
Two more matter as much. Ask what your own record already answers, because most teams reaching for a benchmark have not first asked what their won and lost deals say, and the answer is usually available. Then ask what you are trying to learn that your record genuinely cannot tell you. A clear answer there is a real gap, and a real gap deserves a better instrument than an average.
The objection to benchmarks is not that other companies’ numbers are uninteresting. It is that a decision about your architecture has to be calibrated on evidence generated by your architecture.
Most companies find that the evidence they were missing was never missing. It was sitting in the sales system, unjoined, with the losses in a table nobody had opened. To find out what your own deal record says before anyone quotes a median at you, talk to a pricing expert and describe the decision in front of you. A pricing architect will read the situation with you and say what your record can and cannot settle.