TL;DR: Three stacked editions is the most common packaging in B2B software. The case for it stands on prevalence plus lab psychology, and the psychology is selection evidence: how buyers choose from an existing menu. Packaging is a construction decision, what belongs on the menu at all, and the research that speaks to construction derives the edition count from Customer Group structure and the cost of operating each version. Sometimes that yields three. It arrives there by derivation, not by template.
Good-better-best pricing is a packaging structure that groups a product’s capabilities into three stacked editions at ascending price points. Good lands the deal, Better is the one buyers standardize on, Best is the aspiration. It is also the reflex answer in software packaging. Ask the advice circuit why, and two answers come back: everyone in your category prices this way, and the psychology literature shows a middle option pulls buyers toward it.
We went back to the primary literature: the choice experiments the standard playbook cites, the pooled analyses that later tested them at scale, and the product-line economics the playbook skips. The psychology is real but conditional, and the conditions run against B2B. The evidence that addresses how many editions to build points at your Customer Groups, the foundation of the packaging decision itself.
- Where Good-Better-Best Pricing Draws Its Authority
- Selection Evidence Is Not Construction Evidence
- Even the Selection Evidence Is Conditional, in Exactly the B2B Direction
- Choice Overload Fails at the Count Level and Survives at the Comparability Level
- What Construction Evidence Looks Like: Versioning Economics
- Four Symptoms of Template-First Packaging
- When Good-Better-Best Genuinely Wins
- Choice Architecture Finishes. Packaging Founds.
- FAQs
Where Good-Better-Best Pricing Draws Its Authority
The prevalence leg is accurate as observation, and the template spread the way templates spread. Consultancies arrive force-fitting a perceived best practice, and software companies copy each other’s pricing pages. Common is not the same as correct, and the prevalence record carries an awkward companion.
The same advisory sources that install the template also describe buyers confused by the structures it produced. They document vendors repackaging within a year or two, and run consolidation projects collapsing sprawling edition grids back to a few. When the template underperforms, the diagnosis is that the company strayed from the principles; a recommendation that can never fail on its own terms is a belief, not evidence.
The psychology leg borrows genuine findings: the compromise effect, the decoy effect, and choice overload all exist in the peer-reviewed record. The trouble is what kind of evidence they are, and what question they were designed to answer.
Selection Evidence Is Not Construction Evidence
Every compromise and decoy experiment shares one design property: the products are fixed and the choice set varies. Researchers add or remove a third alternative and measure how share shifts among the survivors. In the original peer-reviewed experiments, repositioning an option as the middle of a three-way set won it a meaningfully larger share of choices, and adding a clearly inferior decoy pushed buyers toward the option that beat it. Real effects, repeatedly measured.
Those are selection findings. They describe how buyers choose from a menu that someone already built.
Packaging asks a prior question: what should exist on the menu? The experiments are silent on it by construction, because in the lab the third option is free. A decoy in a study is a row in a table, one attribute nudged down. A real edition is a roadmap commitment, a support obligation, a sales motion, and a migration path at every renewal.
The versioning literature calls this the menu cost: each version must earn back the cost of building and operating it, an economics no choice experiment models. The lab never bills for the decoy. Your engineering and support organizations do.
The standard citation uses evidence about choosing from menus to settle a question about writing them. Even on its own ground, the selection evidence is conditional.
Even the Selection Evidence Is Conditional, in Exactly the B2B Direction
Three boundary conditions recur across the replication literature, and each one runs against enterprise software.
The decoy needs buyers without formed preferences
The decoy effect operates on buyers still constructing preferences mid-choice. When researchers later tested it against real commercial choice data, the effect largely disappeared: buyers who had formed decision rules through repeat purchasing were mostly immune. B2B buyers are repeat evaluators with requirement lists, which is the immune profile.
The dominance has to be instantly legible
The effect reproduces reliably only when attributes are numeric and the inferiority of the decoy is instantly visible. A broad program of reproduction attempts across many buying scenarios succeeded in only a small minority of them, and roughly at chance whenever options carried qualitative descriptions. Real B2B feature grids are qualitative almost by definition: “advanced analytics” and “priority support” do not compute into dominance.
A decoy in the wrong region backfires
A decoy positioned where buyers have no interest can redirect attention toward competitors instead of lifting the target, a repulsion pattern documented in the same reproduction work. The most famous demonstration in the genre, a magazine-subscription decoy, failed to reproduce when independent teams re-ran it at far larger scale than the small original.
The middle-pull is also contingent on the menu’s own shape. Recent peer-reviewed work finds that when a choice set carries several options at each level, the way real pricing pages do once variants and add-ons accumulate, the preference for middles measurably weakens. Even the compromise effect is not robust to the menu it lives on.
One condition cuts the other way: buyers who must justify a choice to others lean harder on defensible positions, and the middle option is one. Uncertainty pushes the same direction; buyers under perceived threat reach for safe middles, a pattern demonstrated causally in consumer research during the pandemic. B2B buying is justification all the way up, so enterprise contexts do not erase context effects; they filter which ones survive. We made the adjacent argument about consumer anchoring in pricing to value: check the boundary conditions before importing the finding.
And one thing the borrowed evidence never says is three. Where the compromise effect holds, it holds in larger sets too: middles gain share in five-option menus much as they do in three. Taken at full strength, the psychology argues for a defensible middle, not for a count.
Which Pricing Model Fits Your Software — and Which One Breaks?
The model matters less than the metric underneath it. We can assess whether your licensing metric actually scales with how customers derive value. Describe what you’re facing and a pricing expert will reply.
Choice Overload Fails at the Count Level and Survives at the Comparability Level
The other borrowed prop says three is the ceiling because more options overwhelm buyers. Here the pooled record is blunt. When researchers gathered the published experiments on choice overload and analyzed them together, the average effect was indistinguishable from zero: no reliable relationship between option count and buyer paralysis at any set size tested.
The same analysis found a publication skew: journal articles reported overload more often than unpublished studies of the same question. The famous supermarket tasting-booth result sits at the extreme of that distribution, not its center. And buyers with clear prior preferences or domain expertise handle larger sets comfortably, often to their benefit.
A later, larger pooling of the same research explains the variance: overload is real when specific conditions are present. The strongest include options that resist comparison because they differ on incommensurable attributes, decisions that must be justified, time pressure, and buyers in active purchase mode rather than browsing. A clearly recommended or dominant option reduces overload.
Read that list against enterprise software. Required justification is the B2B norm, so overload is a live risk there, and claiming psychology is irrelevant to B2B packaging overcorrects. But the operative lever is comparability, not count. Editions that progress along one axis a buyer can follow tolerate breadth; editions that differ on incommensurable feature bundles confuse at any count, including three.
The design principle that survives the pooled record is alignability, coverage, and simplicity: a Customer Group test, not a rule of three. And force-fitting groups that do not naturally stack into three stacked editions manufactures the very non-comparability the overload research warns about. The template produces the confusion its advocates keep diagnosing.
What Construction Evidence Looks Like: Versioning Economics
There is a literature that answers the construction question directly, and the standard playbook rarely cites it.
The foundational quality-screening result, now close to fifty years old, models a seller whose buyers value quality differently and cannot be identified in advance. The seller’s answer is a menu of quality-price pairs buyers sort themselves into. The menu’s shape falls out of the demand’s shape: where buyer valuations cluster, the model collapses adjacent versions into one, a result the literature calls bunching. The edition count is an output of how value distributes across the customer base; nothing in the theorem privileges three.
A later strand of the same literature extends this to deliberately limited lower versions, what economists call damaged goods. A stripped-down edition can profitably serve a group that would otherwise go unserved, and under specific conditions every party ends up better off, premium buyers included. The conditions bind: the case is strongest when Customer Groups differ in kind, in use case, and much harder when buyers differ only in intensity of the same use.
Competitive product-line models add the market-structure term. An incumbent facing a low-end entrant sometimes adds a lower edition and sometimes prunes its line upward, depending on whether the low end is a genuinely distinct group. And versioning-optimality work on information goods shows that when every buyer type values the versions proportionally, one version beats several.
Different models, one shared property: each derives the version count from Customer Group structure and the economics of operating each version. This is the evidence class packaging should rest on. Our packaging decision framework operationalizes it, starting from Customer Groups rather than personas.
Four Symptoms of Template-First Packaging
In our engagements we observe four recurring symptoms of packaging where the template came before the Customer Group analysis.
Symptom 1: the grid-filler Best edition
A third edition built to complete the grid, with differentiation nobody can articulate. Sales cannot say who it is for, and it ends life as either a shadow of Better or the anchor a discount is negotiated down from. The tell is a sales team that sells Better by default and reaches for Best mainly to make Better sound reasonable.
Symptom 2: bolt-ons before the Good or after the Best
Something new needs monetizing, and the architecture offers exactly two slots: before the Good or after the Best. The new capability lands as an awkward bolt-on at one end because the decision set is the template’s artifact, not the market’s.
Symptom 3: volume-trough editions
Editions “differentiated” by usage thresholds rather than capabilities: the value metric melded into the packaging. A usage threshold is a metric decision; the metric’s job is expansion, more units of what the customer already has, while packaging’s job is upsell, a richer edition. Parked inside an edition boundary, the threshold asks customers to expand and upsell in one motion, and most do neither.
The packaging repair is to pull the threshold back out of the edition boundary. Whether that threshold is the right value metric at all is a separate question, owned by the licensing decision, and the packaging repair does not answer it.
Symptom 4: feature caps that produce hyper-gearing
The pattern a widely used marketing platform popularized: “up to X” caps on individual features within each edition. The price stays perfectly legible, so the structure looks harmless. The damage is fit. With enough independently capped dimensions, almost no buyer’s usage profile maps cleanly onto any edition, and an endless corridor of buyers almost fit.
The sales dialogue fragments into feature-by-feature haggling. Reps discount under a partial-use rationale, the customer only needs half the dashboards, and the deal desk approves case by case because each instance sounds reasonable. That is hyper-gearing, and its output is structured, self-justifying pricebook deviation.
We hear it live on renewal calls: the buyer arrives with a feature-by-feature list of what they never used. The discounting is not reps gaming the system; the architecture hands every rep a defensible partial-use story.
When Good-Better-Best Genuinely Wins
None of this argues that three stacked editions is wrong everywhere. Stacked editions are the right container when Customer Groups genuinely stack: a real progression buyers recognize, each step adding capabilities one identifiable group needs. In that case the selection psychology works with the structure instead of against it. The progression does not have to be a sophistication ladder. Editions can be independent of each other, can ladder by maturity, or can follow some other sequence natural to the product and market.
The test is the one the construction evidence implies. Every Customer Group is covered, each edition maps to a buying pattern visible in your transaction data, and the structure is simple enough for a rep to defend without a spreadsheet. The count that passes might be three; it might be two, four, or a different archetype altogether: modular add-ons, platform plus apps, a single all-in-one offer.
Whether your groups genuinely stack is discoverable from won/lost patterns and usage data. It is exactly the question to talk to a pricing expert about before a repackaging, not after one.
Choice Architecture Finishes. Packaging Founds.
The selection research has a legitimate place: after the set is derived. Ordering, anchoring, defaults, and a clearly recommended edition are finishing decisions the lab evidence genuinely informs; a recommended option demonstrably reduces overload, and presentation order moves share at the margin.
The construction decision runs upstream, on different evidence: how value distributes across Customer Groups, which capabilities move between editions, and what each version costs to operate. Derive the set from that, then let the psychology arrange it.
Psychology arranges the menu. It doesn’t write it.
If your editions came from a template and the four symptoms read like your deal desk’s last quarter, the repair is architectural, not cosmetic. Talk to a pricing expert: describe the situation and an expert replies directly. Or work the edition boundaries live against your own data: book a working session.