Mechanisms of Conversion
The funnel arithmetic, the testing statistics, the checkout evidence, and the dead science underneath LUCE_08 and LUCE_18
48 min read
Lineage: new for Lumen, August 2026. This is not a twin of a spine module — it is the mechanism layer beneath LUCE_08 (Store CRO) and LUCE_18 (CRO Advanced), in the same relation M1 holds to Advertising and M2 to Finance. Those modules tell you what to do. This one derives why, computes the arithmetic from scratch, and demotes every heuristic that does not survive. Research base: Ronny Kohavi (Microsoft/Airbnb, Trustworthy Online Controlled Experiments), Ron Berman (Wharton), Garrett Johnson & Randall Lewis (advertising measurement), Baymard Institute, Nielsen Norman Group, Peter Fader & Bruce Hardie (customer-base analysis), the FTC's dark-patterns and reviews rulemaking, and the replication literature.
HOW TO READ THIS MODULE
Two labels sit next to load-bearing claims, and they are never merged. One says how good the evidence is. One says who is saying it and what they get if you believe them. This is the same two-tier discipline the rest of the course uses.
Evidence tier
| Tier | Meaning |
|---|---|
| E1 | Meta-analysis, or a multi-lab / multi-site pre-registered replication |
| E2 | A single peer-reviewed randomised experiment, or a large field experiment with a real counterfactual |
| E3 | Observational or quasi-experimental (matching, DiD, IV) — causal identification is an argument, not a design |
| E4 | Large-sample benchmark with no counterfactual at all |
| E5 | Single case study, or a number with no traceable method |
Source tier
| Tier | Meaning |
|---|---|
| S1 | Academic journal, regulator, or a researcher with nothing to sell you |
| S2 | Independent research firm whose method is published and whose revenue is research, not the tactic |
| S3 | A platform or vendor reporting on its own product |
| S4 | Agency, affiliate, or app-store listing selling the thing the number is about |
An E4/S3 claim is not worthless. It is just not evidence that the thing causes the outcome, and it must never enter a P&L assumption.
This module's own limits. Direct page fetching was blocked during compilation; several primaries — Baymard's own pages, the Cornell menu-pricing PDF, the Berman p-hacking PDF — returned 403 to every attempt. Where a figure comes from a search index's summary rather than the source read end to end, it is marked UNVERIFIED inline and again in §9. The arithmetic in §2 and §3 was computed rather than cited, and anyone can re-derive it with a calculator.
1. FUNNEL ARCHITECTURE AS A REAL SUBJECT
1.1 The honest taxonomy
A funnel is a sequence of pages, each asking for one increment of commitment. The taxonomy that actually exists is not seven canonical types; it is a two-dimensional space, and every named "funnel" is a point in it.
Axis one: how much belief has to be manufactured before the ask. A commodity bought on price needs none. A product whose mechanism is unfamiliar, whose category is unfamiliar, or whose price sits above the buyer's reference point needs a lot.
Axis two: how much the visitor must self-identify before you can make a specific offer. A single-SKU brand needs none. A 40-SKU catalogue where the right SKU depends on skin type, hair type, sleep position or dog weight needs a lot.
| Architecture | Belief | Self-identification | What it actually is |
|---|---|---|---|
| Direct-to-PDP | None | None | Ad → product page → cart. The default |
| Advertorial | High | None | An article-shaped page doing the belief work before the PDP |
| Listicle | Medium | Weak | An advertorial whose structure is a ranked list; the product wins |
| Quiz funnel | Low–medium | High | A self-segmentation instrument that outputs a personalised offer |
| VSL | High | None | An advertorial where the medium is video and the seller controls pacing |
| Webinar | Very high | None | A VSL with a start time and a synchronous ask. Rare in physical goods |
| Two-step | Low | Low | Splitting the first commitment into a trivial one, then the real one |
1.2 What the evidence says — which is very little
There is no body of independent controlled evidence establishing when one architecture outperforms another for physical goods. Not "the evidence is mixed." There is essentially none comparing advertorial-mediated against direct-to-PDP with random assignment at adequate scale.
Advertorials. The entire claim base is agency and tool-vendor content. "Advertorials instantly lifted conversion by almost 40%" appears as a case study with no control group, no confidence interval, no audit (E5/S4). The academic native-advertising literature studies something else — disclosure and persuasion knowledge, not revenue — and its finding cuts both ways: disclosure decreases persuasion by activating persuasion knowledge and simultaneously increases it by raising perceived transparency (Journal of Interactive Advertising 2019) (E2/S1). So the honest statement about a disclosure label is: legally required, net effect theoretically ambiguous.
Quiz funnels. The circulating numbers are selection effects wearing a lab coat. "300% average lift" comes from a quiz vendor (E4/S3). "Quiz takers are 4.9× more likely to buy" is a single brand case study (E5/S3). The mechanism problem is glaring: taking a five-question quiz is itself an act of high intent, so comparing quiz-takers to all visitors compares people who volunteered three minutes of attention to people who did not. The only figure with a plausible causal shape is the within-store, same-cohort comparison — orders arriving through a quiz running roughly 11–15% higher in value, holding in about 7 of 10 stores (RevenueHunt) (E4/S3). Much smaller, much more believable, still confounded.
VSLs. The worst claim environment in the module. A widely circulated figure — "a 2025 Unbounce benchmark of 44,000 landing pages found VSL pages averaged 12.7% vs 4.8% text-only" — corresponds to nothing in Unbounce's published work. Their Conversion Benchmark Report analyses ~41,000 pages, 464 million visits, and reports a 6.6% median segmented by industry, traffic source and goal — not by whether the page had video. The statistic appears to be fabricated and has entered circulation. Treat it as a live specimen of how funnel folklore is manufactured. UNVERIFIED, and probably false.
Two-step. This one has a real mechanism: it is a foot-in-the-door procedure, and FITD is genuinely established (Burger 1999, PSPR) (E1/S1). Two things get left out of the sales pitch. The effect is modest — roughly r ≈ 0.15, with nearly half of experimental groups showing no effect or a reversal (E1/S1; the specific r is UNVERIFIED). And Burger's central contribution is that FITD is multiply determined — self-perception, commitment, consistency, reactance, conformity and attribution all fire, several in opposite directions. Reactance means an initial request that feels manipulative makes the second one less likely to be granted. A two-step that reads as a trick is a two-step that costs you money.
1.3 What survives
- Architecture is downstream of a diagnosis, not a preference. Belief deficit → advertorial or VSL. Selection deficit → quiz. Neither → direct to PDP, and stop reading blogs about funnels.
- The strongest real case for pre-sell pages is indirect. Cold paid traffic arrives with near-zero context, and §2 shows every step carries identical elasticity. A page moving engaged → add-to-cart from 16% to 21% is worth more than anything you will ever do to a checkout button. Advertorials plausibly do that. "Plausibly" is the correct strength of claim.
- Every large claimed number here traces to a vendor of the thing. Ask what the control group was, who paid, and whether the person quoting it sells the tactic. The answers are usually "none", "the vendor", and "yes".
2. THE MATHEMATICS OF A FUNNEL
This section is arithmetic, so it is the only one that cannot be wrong.
2.1 The chain
CVR = r₁ × r₂ × … × rₖ, and the number that pays you is
π = CVR × AOV × m − CPC
where m is contribution margin after landed COGS, fulfilment and payment fees — the same definition the rest of the course uses.
A worked store, used throughout:
| Step | Rate | Cumulative |
|---|---|---|
| Ad click → LP rendered | 90.00% | 90.00% |
| LP rendered → meaningful engagement | 55.00% | 49.50% |
| Engaged → add to cart | 16.00% | 7.92% |
| ATC → checkout initiated | 62.00% | 4.91% |
| Checkout initiated → purchase | 55.00% | 2.70% |
At AOV $68, margin 55%, CPC $0.80: RPV $1.84, contribution/visitor $1.01, profit/visitor $0.21, blended ROAS 2.30.
2.2 Where the sensitivity is — the result everyone gets wrong
Take logs: log CVR = Σ log rᵢ. Differentiate. The elasticity of revenue with respect to every step is exactly 1.0. A 10% relative improvement anywhere produces exactly a 10% revenue improvement:
| Step improved +10% relative | CVR | Revenue | Profit |
|---|---|---|---|
| Ad click → LP rendered | 2.70 → 2.97% | +10.0% | +48.1% |
| LP → engagement | 2.70 → 2.97% | +10.0% | +48.1% |
| Engaged → ATC | 2.70 → 2.97% | +10.0% | +48.1% |
| ATC → checkout | 2.70 → 2.97% | +10.0% | +48.1% |
| Checkout → purchase | 2.70 → 2.97% | +10.0% | +48.1% |
Identical. Every row. This kills the most common piece of funnel advice — "find your biggest drop-off and fix it" — because the biggest drop-off has no special claim on your attention. A 90% step and a 16% step are worth precisely the same per unit of relative improvement.
2.3 What actually differs: headroom
What differs is how much relative improvement is physically available. A step at 90% has a ceiling 11% above it. A step at 16% has a ceiling 525% above it. Hold the absolute gain constant at +5 percentage points and the picture inverts:
| Step, +5pp absolute | Relative change | New CVR | Profit change |
|---|---|---|---|
| 90% → 95% | +5.6% | 2.85% | +26.7% |
| 55% → 60% | +9.1% | 2.95% | +43.7% |
| 16% → 21% | +31.2% | 3.54% | +150.3% |
| 62% → 67% | +8.1% | 2.92% | +38.8% |
| 55% → 60% | +9.1% | 2.95% | +43.7% |
This is the real result. Sensitivity is not about where the leak is or which step matters most. It is about which step has room to move multiplicatively — and the lowest-rate step in a DTC funnel is nearly always interest → add-to-cart, which is a function of the offer, the price, and the belief the page manufactures. Not the checkout.
People optimising checkout buttons work on a step at 55–62% where a heroic gain is +8% relative. People rewriting the offer work on a step at 16% where an ordinary gain is +30%. Same effort, four times the return, and this is arithmetic rather than opinion.
2.4 The profit multiplier
A 10% CVR gain became a 48% profit gain above, because CPC is a fixed subtrahend. The leverage ratio is (CVR·AOV·m) / (CVR·AOV·m − CPC), which goes to infinity at breakeven:
| CPC | Profit/visitor | After a 10% CVR lift | Change |
|---|---|---|---|
| $0.80 | +$0.2101 | +$0.3111 | +48% |
| $1.00 | +$0.0101 | +$0.1111 | +1,003% |
| $1.10 | −$0.0899 | +$0.0111 | loss → profit |
| $1.20 | −$0.1899 | −$0.0889 | loss → smaller loss |
Two consequences. The closer to breakeven you are, the more a small CVR gain is worth — exactly backwards from how most operators allocate attention, since near breakeven they panic about media buying instead. And the same 10% gain that is a rounding error at ROAS 3.0 is the whole business at ROAS 1.7. Value conversion work at the CPC you actually pay, never at an abstract "10% more revenue".
2.5 Allocating effort by expected value
Stop ranking ideas by intuition:
EV = P(win) × E[relative lift | win] × annual PROFIT − cost
Profit, not revenue, per §2.4. For the worked store at 1.2M annual visitors — $252,083 annual profit:
| Project | P(win) | E[lift] | Cost | EV | EV/cost |
|---|---|---|---|---|---|
| New advertorial pre-sell for cold traffic | 30% | 25% | $6,000 | $12,906 | 3.2× |
| Rewrite offer + guarantee on LP | 35% | 18% | $4,000 | $11,881 | 4.0× |
| Express wallets at cart + checkout | 60% | 6% | $600 | $8,475 | 15.1× |
| Show shipping cost on PDP and cart | 55% | 5% | $400 | $6,532 | 17.3× |
| Checkout 4 steps → 3 | 40% | 4% | $2,500 | $1,533 | 1.6× |
| Button colour / microcopy test | 15% | 1% | $300 | $78 | 1.3× |
EV answers "what if I had unlimited hands." EV/cost answers "what first." Wallets and shipping transparency are 15–17× on capital and take an afternoon; the offer rewrite is the biggest absolute prize and takes a month. Do both, in that order. Do not do the button test at all — its $78 EV is smaller than the cost of the meeting in which you discussed it.
The three inputs are subjective, and that is fine; the discipline is being forced to write down P(win) before you start. Anyone claiming 80% confidence in a button-colour test has revealed something useful about themselves. Calibrate against the base rate in §3.1: roughly one in three well-designed changes at a mature company improve the metric they targeted.
2.6 Conservation of conversion — the trap in the arithmetic
The model assumes step rates are independent. They are not. Many interventions move volume between steps rather than creating it:
- A more aggressive ATC button raises ATC and lowers ATC→checkout by the same people.
- A discount pop-up raises checkout initiation and lowers margin — invisible in a CVR-only model.
- Removing shipping cost from the PDP raises ATC and craters checkout completion (§5).
The only metric that can be optimised safely is the one at the end: contribution per visitor. Step metrics are diagnostics, not objectives. Kohavi's term for the single end-metric is the OEC, and choosing it badly is how experimentation programmes produce years of "wins" and no revenue (Kohavi, Tang & Xu, Trustworthy Online Controlled Experiments, CUP 2020) (E1/S1).
3. TESTING SCIENCE, DONE PROPERLY
3.1 The base rate you are testing against
At Microsoft, of well-designed experiments intended to improve a key metric, about one third succeeded (Kohavi, Online Experimentation at Microsoft) (E1/S1). Reported failure rates elsewhere run 70–90% (E4/S3, secondary — UNVERIFIED).
Independently, Berman, Pekelis, Scott & Van den Bulte's analysis of 2,101 Optimizely experiments estimates roughly 75% of tested effects are truly null (E3/S1).
Hold that number. It is the prior you carry into every test, and it means a "significant" result in a low-powered test is far more likely to be noise than most operators believe.
3.2 Sample size and MDE at real traffic
Two-proportion test, α = 0.05 two-sided, 80% power, two arms, baseline CVR 1.4% — the honest Shopify average this course uses everywhere:
| Relative MDE | Absolute | n per arm | Total N |
|---|---|---|---|
| 2% | 0.028pp | 2,791,166 | 5,582,332 |
| 5% | 0.070pp | 453,119 | 906,238 |
| 10% | 0.140pp | 115,999 | 231,997 |
| 15% | 0.210pp | 52,762 | 105,523 |
| 20% | 0.280pp | 30,356 | 60,712 |
| 30% | 0.420pp | 14,093 | 28,185 |
| 50% | 0.700pp | 5,504 | 11,009 |
Now invert it. Given your traffic, what is the smallest effect you can detect?
| Weekly sessions | 4 weeks | 8 weeks | 12 weeks |
|---|---|---|---|
| 2,000 | 59.8% | 40.7% | 32.7% |
| 5,000 | 36.1% | 24.9% | 20.1% |
| 10,000 | 24.9% | 17.3% | 14.0% |
| 25,000 | 15.4% | 10.8% | 8.8% |
| 50,000 | 10.8% | 7.6% | 6.2% |
| 100,000 | 7.6% | 5.3% | 4.3% |
Read the top-left cell. A store doing 2,000 sessions a week, running a test for a month, can only detect a change lifting conversion 60% relative — 1.4% to 2.24%. Nothing you will ever do to a page does that.
3.3 Peeking, and why fixed-horizon tests break
Fixed-horizon inference is valid only if N was fixed in advance and you look once. Simulated — a true A/A test, no effect, baseline 1.4%, 40,000 sessions, declaring a winner at the first look where p < 0.05:
| Looks | Fixed-horizon FPR | Peeking FPR |
|---|---|---|
| 1 | 4.4% | 4.4% |
| 5 | 5.4% | 15.2% |
| 10 | 4.8% | 20.1% |
| 40 | 4.7% | 29.8% |
Peek daily for six weeks and three in ten of your "winners" are pure noise, before anything else goes wrong.
This is not hypothetical. Berman et al. found about 73% of experimenters stop just as a positive effect reaches 90% confidence, raising the false discovery rate among those experiments from 33% to 42% (E3/S1). Kohavi has publicly argued the framing is overstated — the disagreement is about magnitude, not about whether the mechanism is real. Contested in size, robust in direction.
3.4 What sequential inference buys
The fix is a procedure designed for continuous monitoring, not a promise not to look.
Always-valid inference (Johari, Pekelis & Walsh, arXiv:1512.04922; KDD 2017) defines p-values and intervals valid at every point in time, so you may stop whenever you like. The machinery is a mixture sequential probability ratio test (E1/S1 for the method; S3 for any vendor's implementation).
Group-sequential designs (O'Brien–Fleming alpha spending) are the clinical-trials alternative: a few pre-specified looks with a boundary conservative early and spending remaining alpha at the end, at very little sample-size inflation (E1/S1).
The trade is explicit: sequential methods do not give a free lunch, they give the option to stop early in exchange for slightly less power at any fixed N. Claims that they roll out winners "30–50% faster" are true only for large true effects — the case where you did not need statistics. UNVERIFIED as a general figure.
3.5 Bayesian vs frequentist, in practice
With a weak prior and a large sample the two give nearly the same answer with different labels. What matters:
- "Probability to beat baseline" is not peek-immune. It reads as though peeking is fine because it is a posterior — but the posterior depends on the prior, and vendor "non-informative" priors usually are not (Georgiev's critique) (E3/S2). If a tool says "87% chance B beats A" and will not tell you the prior, you cannot interpret the 87%.
- Bayesian methods genuinely help with the decision, not the inference. A posterior over the magnitude — combined with the cost of the change and the §3.1 base rate — beats a binary significance flag. That is a real advantage and not the one usually advertised.
- Neither rescues you from insufficient traffic. The uncertainty does not go away when you rename it.
3.6 Multiple comparisons, which is worse than peeking
Every variant, metric and post-hoc segment is another hypothesis. At α = 0.05, sixty tests → three false discoveries guaranteed by chance. Slice a finished test by device × source × new/returning and you have manufactured twenty hypotheses from one experiment.
Benjamini–Hochberg bounds the expected proportion of false discoveries; Bonferroni bounds the probability of any, and is brutal enough to make DTC testing impossible. The operationally sufficient rule for a small store: pre-register one primary metric and one primary segment before the test starts, and treat everything else as hypothesis generation, never as a result.
3.7 The winner's curse — the finding that should change your behaviour
In an underpowered test, conditional on significance, the observed effect is systematically and severely exaggerated. Simulated, baseline 1.4%:
| True lift | Total N | Power | Mean observed lift given significance | Exaggeration | P(wrong sign) |
|---|---|---|---|---|---|
| 2% | 40,000 | 5.7% | 12.0% | 6.0× | 25.5% |
| 5% | 40,000 | 8.9% | 20.2% | 4.0× | 5.6% |
| 5% | 200,000 | 25.8% | 10.0% | 2.0× | 0.2% |
| 10% | 40,000 | 21.6% | 23.0% | 2.3× | 0.4% |
| 10% | 200,000 | 73.9% | 11.8% | 1.2× | 0.0% |
| 20% | 200,000 | 99.9% | 20.1% | 1.0× | 0.0% |
Read row one. If the truth is a 2% lift and you ran 40,000 sessions, then on the 5.7% of occasions you got significance, the average result you saw was 12% — six times reality — and a quarter of the time it pointed the wrong way entirely.
This is why underpowered programmes report a stream of double-digit wins that never reach the P&L. They are not lying; they are reporting Type-M error. It also explains part of the "novelty effect" folk diagnosis — novelty is real and carefully documented (E2/S1), but some of what gets attributed to it is an inflated initial estimate regressing to the truth.
3.8 Variance reduction, and why it will not save you
CUPED (Deng, Xu, Kohavi & Walker, WSDM 2013) (E1/S1) regresses out the predictable component of each user's metric, leaving less residual variance. Reported reductions cluster at 30–50% (E4/S3 — UNVERIFIED).
The catch that makes it nearly useless here: CUPED requires identified users with pre-period history. For a store whose traffic is mostly first-time cold paid visitors, there is no covariate to regress on and the reduction is zero. It helps a returning-customer-heavy business. It does not help a cold-traffic dropshipper.
3.9 The honest conclusion for a store with modest traffic
Below roughly 25,000 sessions per week, A/B testing cannot be your primary method of improvement. At 10,000/week you need eight weeks to detect a 17% relative change. Most real changes are worth 2–8%. You will run tests powered only to detect effects that do not exist, and the ones that come back "significant" will be exaggerated three- to six-fold.
What to do instead, in order:
- Make large swings, not small ones. The MDE table is a design specification, not just a constraint. If you can only detect 25% changes, only make changes big enough to plausibly produce 25% — a new offer, a new price architecture, a new pre-sell page. Not a headline variant.
- Use judgement anchored to mechanism, and to the checkout evidence in §5, which is the one place large-sample independent research already exists. You do not need to re-test guest checkout. It has been tested.
- Test what is cheap to be wrong about; ship what the evidence already settles.
- Run before/after only for changes so large a seasonal confound cannot explain them — and say out loud that this is judgement, not evidence.
- Accept that you will make some changes that lose money and never know. That is the actual cost of operating below testing scale. Pretending otherwise is how a "CRO programme" becomes an expensive random number generator.
4. LANDING PAGE AND OFFER MECHANICS
4.1 Form-field reduction — the most over-claimed rule in CRO
The canonical claim ("four fields to three lifted conversions 50%") traces to a single HubSpot 2012 analysis (E3/S3) and has been generalised into a law it never supported.
- Unbounce's own database found conversion highest at one field, declining with each addition — then rising again, with ten-field forms beating three-field forms (E4/S3). The mechanism is intelligible: field count proxies two opposing things, effort (bad) and perceived seriousness (good).
- Michael Aagaard's Unbounce case study found reducing fields produced a 14% drop (E5/S3).
- CXL reports most modern field-count tests produce lift under 5%, several negative (E4/S2).
The rule: field count is not the mechanism. Perceived cost of the field is. A phone number is expensive; a postcode is cheap. Removing three cheap fields buys nothing; removing one expensive one buys a lot. Checkout is different and better-evidenced — §5.3.
4.2 Social proof and reviews
Real, quantified, independently established. The Floyd et al. (2014) meta-analysis in Journal of Retailing covers 26 studies and 443 sales elasticities: valence ≈ 0.78, volume ≈ 0.41, stronger on third-party sites and high-involvement products (E1/S1; exact elasticities UNVERIFIED).
Two things follow. Valence outweighs volume roughly two to one — fewer higher-rated reviews beat more mediocre ones. And third-party placement carries more weight than on-site, which is the incentive-compatibility mechanism: reviews you control are discounted by the reader.
The legal boundary is now hard. The FTC's Rule on the Use of Consumer Reviews and Testimonials, effective 21 October 2024, prohibits buying or selling fake reviews, incentivising reviews conditioned on sentiment, undisclosed insider reviews, company-controlled "independent" review sites, review suppression, and fake social indicators. Penalties reach $53,088 per violation (E1/S1). Suppression enforcement is not theoretical — Fashion Nova paid $4.2 million for holding back sub-four-star reviews (FTC, Jan 2022) (E1/S1).
Note what this makes illegal that many stores do routinely: an automated request offering a discount for a five-star review is a violation. Offering a discount for a review is not.
4.3 Urgency and scarcity — real effect, serious legal exposure
The effect is real. The Barton, Zlatevska & Oppewal meta-analysis, Journal of Retailing (2022) pools 416 effect sizes from 131 studies: scarcity cues raise purchase intention, with type mattering — demand-based works best for utilitarian products, supply-based for experiences, time-based for high-involvement (E1/S1; pooled magnitude UNVERIFIED). Consistent evidence also shows limited-quantity beats limited-time, despite limited-time being used about three times as often.
And it is where operators get sued. Three regimes apply:
- FTC Act §5. The staff report Bringing Dark Patterns to Light (Sept 2022) names, by example, the exact tactics in every funnel course: "Only 1 left in stock" when it isn't; "20 other shoppers have this in their cart" when they don't; and a countdown clock that resets or disappears when it expires (E1/S1).
- The FTC's Unfair or Deceptive Fees Rule (effective 12 May 2025) bans drip pricing — only for live-event ticketing and short-term lodging. General e-commerce is not covered by that rule. It is covered by §5 and by state junk-fee laws in force or pending in nearly half of US states (E1/S1). Anyone telling you the junk-fees rule covers your Shopify store is wrong; anyone telling you that means you are safe is also wrong.
- Prevalence is documented, so "everyone does it" is a bad defence. Mathur et al., Dark Patterns at Scale (CSCW 2019) crawled ~53,000 product pages across 11,000 sites and found 1,818 deceptive instances, including 393 countdown timers across 361 sites (E2/S1). That is a machine-readable enforcement roadmap, public since 2019.
The operating rule: scarcity works, so use scarcity that is true. A drop that genuinely sells out. A price that genuinely rises. A counter wired to actual inventory. Everything about the mechanism survives being honest — what does not survive is fabrication, and fabrication carries the $53,088-per-violation exposure.
4.4 Risk reversal and guarantees
Less straightforward than the sales literature suggests.
- The signalling account (only high-quality sellers can afford a guarantee) is undermined by ubiquity: MBGs are standard even among retailers with high return likelihood (E3/S1).
- Three field experiments at a large European retailer found a money-back guarantee increased returns, while product reviews decreased them (Electronic Markets, 2017) (E2/S1). This belongs in every P&L conversation and is in none: a guarantee is not free conversion, it is conversion bought with return rate.
- Generosity and believability interact — generosity raises patronage intentions via believability (E2/S1). An unbelievable guarantee converts worse than a credible one.
Model it as a price, not a feature. Worth having only if ΔCVR × m > Δreturn rate × (AOV + reverse logistics). For a $68 product with $12 round-trip shipping and no resale value on returns, that inequality fails surprisingly often.
4.5 Price presentation and anchoring
Anchoring is robust. The precision-of-anchor extension is not. Janiszewski & Uy (2008) found less adjustment away from precise anchors (E2/S1), and this became "always price at $1,997." Subsequent work has been much less kind: a pre-registered field experiment found expertise moderates it substantially, and two high-powered pre-registered experiments (N = 729) tested round vs just-below vs precise prices directly (Frontiers in Behavioral Economics, 2026) (E1/S1; direction UNVERIFIED). The defensible position: anchoring, yes; precision-of-anchor as a reliable retail lever, no.
Reference-price anchoring (a struck-through "was") is the robust version — and it is regulated. FTC and state discount-pricing rules require the reference to be a price actually and recently offered in good faith (E1/S1). A permanent "compare at" that was never charged is a deceptive pricing claim.
4.6 Decoy / asymmetric dominance — replication status stated honestly
Claimed: add a deliberately inferior third option to steer choice to your target tier.
What the evidence says: real in the paradigm it was discovered in, largely evaporating in the conditions you would deploy it.
- With abstract two-attribute numeric stimuli it replicates and is not disputed (E1/S1).
- Frederick, Lee & Baskin (2014) and Yang & Lynn (2014) found it largely disappears with realistic stimuli — qualitative descriptions, images, more than two attributes — and sometimes reverses, the decoy pulling share away from the option it was meant to help (E1/S1). Across 91 attempts covering 23 categories, a minority produced reliable effects (the "11 of 91" count is UNVERIFIED).
- The original authors pushed back — "Let's Be Honest About the Attraction Effect". Genuinely contested, not settled.
What works instead on a pricing page is making the tier you want easy to justify — per-unit price, a free-shipping threshold crossed, a subscription discount the buyer can compute. Real reasons survive the buyer thinking about them. A decoy does not.
4.7 Payment friction: wallets and BNPL, with real costs
Digital wallets. The mechanism is unambiguous — a wallet button eliminates every address and card field at once, the single largest form-field reduction available. The measurement is much weaker than the mechanism.
- Stripe reports ~2× conversion when Apple Pay is surfaced early (E4/S3 — Stripe sells this).
- "Shop Pay lifts conversion up to 50%" is Shopify-commissioned (E4/S3), comparing self-selecting Shop Pay users to everyone else. The selection problem is fatal.
- The strongest evidence is academic: Unal & Park, "Fewer Clicks, More Purchases," Management Science (2023) tracked 977 customers of a retailer launching one-click, finding +28.5% spending, +18.5% visits, +43.3% category breadth (E3/S1) — adopters self-selected, matched DiD, effect on adopters.
Real cost: wallets are usually processing-neutral. The cost is data — express wallets often deliver a wallet email and address with no marketing consent, degrading the retention list §6 says the economics live in. Quantify that before celebrating.
BNPL. Here the academic evidence is unusually good.
- Berg, Burg, Keil & Puri, NBER w33152 / JFE (2025): BNPL increases sales ~20%, concentrated among low-creditworthiness customers, functioning as price discrimination by willingness to pay (E3/S1).
- Di Maggio, Williams & Katz, NBER w30508: access raises total spending ~$130 at first use, sustained over 24 weeks — a "liquidity flypaper effect" (E3/S1).
Real cost: merchant fees commonly quoted at 3–6% plus a fixed fee (UNVERIFIED). Against §2.1's 55% margin, a 5% BNPL fee is 9% of contribution. It pays only if BNPL produces incremental orders rather than cannibalising card orders — and Berg et al. say incrementality concentrates in low-creditworthiness buyers, who are also your worst return and chargeback cohort. Turning BNPL on and watching AOV rise is not evidence; those buyers had higher baskets anyway.
4.8 Page speed
Milliseconds Make Millions (Deloitte/55, commissioned by Google, 2020) monitored 37 sites and 30 million sessions and reports 0.1s mobile improvement associates with +8.4% retail conversion (E3/S3).
Read the design before quoting it. Observational, on naturally occurring hourly load-time variation, commissioned by the company that sells speed as a ranking factor. Hourly load time correlates with time of day, device mix, network quality and traffic source — all of which independently predict conversion. The direction is almost certainly right. The magnitude should not enter a business case.
5. CHECKOUT SPECIFICALLY
The one area where large, independent, methodologically transparent research exists. Use it instead of testing it.
5.1 The abandonment number, and what it actually is
The headline ~70% (Baymard's current figure is 70.22%) is not a measurement Baymard made. It is an average of dozens of separately published studies (Baymard's list) (E4/S2), inheriting every definitional inconsistency in its inputs. Device split ~80.0% mobile vs ~66.4% desktop (E4/S2).
Do not benchmark against it. Its only legitimate use is order-of-magnitude framing.
5.2 The measured reasons
| Reason | Share |
|---|---|
| "Just browsing / not ready to buy" | ~43% |
| Extra costs too high (shipping, tax, fees) | ~48% of the remainder |
| Site wanted me to create an account | ~19–26% |
| Too long / complicated checkout | ~18% |
| Couldn't see total cost up front | ~17–21% |
(E4/S2.) Percentages differ between survey waves and baymard.com returned 403 to every fetch attempt — quote as approximate ranges, never exact. UNVERIFIED. What is stable across waves: unexpected cost is reason #1 by a wide margin, and forced account creation is top three.
The most important line is the first. Roughly four in ten "abandoners" were never buyers. Your addressable abandonment is not 70% — it is 70% minus the browsers, which changes the size of the prize by more than any tactic in this module.
5.3 Form fields and flow length
- The average US checkout shows ~23.5 form elements (~14.9 fields); an ideal guest flow is 12–14 elements / 7–8 fields, card included (Baymard) (E4/S2).
- ~18% report abandoning solely because checkout was too long (E4/S2).
- Baymard estimates ~35.3% achievable conversion increase from checkout usability work (E4/S2 — a modelled estimate from usability findings, not a measured randomised lift, produced by an organisation that sells checkout research. Treat as an upper bound.)
5.4 Guest checkout
The highest-confidence single recommendation in this module.
- Baymard: forced account creation is a top-three stated reason (E4/S2); 62% of sites still fail to make guest checkout the most prominent option (E4/S2 — UNVERIFIED).
- Nielsen Norman Group's Shopping Carts, Checkout & Registration — 137 recommendations from five rounds of testing across 350+ sites in five countries — recommends guest checkout above the fold and above sign-in, and notes the under-appreciated mechanism: account holders who forgot their password often find guest checkout easier than password recovery on mobile (E2/S2).
Two independent research organisations, different methods, same conclusion, over a decade. Ship it. Do not test it.
5.5 Number of steps
Thinner than the confident blog posts suggest.
- Baymard's position: step count is not the driver — field count and specific friction points are (E4/S2). For multi-step flows they recommend collapsing completed steps into editable summaries (E2/S2).
- "A 15-field form across three steps beats a 10-field one-pager by 11–14%" and "one-page converts 7–15% better under $150 AOV" circulate widely with no traceable study behind either. UNVERIFIED. Do not use these numbers.
Honest summary: no good independent evidence that one-page beats multi-step or the reverse. Good evidence that unexpected cost revealed late and forced registration hurt regardless of layout. Optimise those; treat step count as a stylistic choice.
5.6 Cost transparency — the highest-leverage checkout change
Because unexpected cost is reason #1, the correct intervention is not to hide the cost longer or reveal it marginally earlier. It is to make the total knowable before the buyer invests effort.
The economics are well-evidenced: Lewis, Singh & Fay, Marketing Science (2006) used a retailer experimenting across many shipping-fee schedules and found consumers highly sensitive to shipping charges, with free shipping and threshold-based free shipping both strongly effective (E3/S1). Threshold free shipping is the version that pays for itself, because it converts the fee into an AOV lever.
And note the §2.6 trap: removing shipping cost from the PDP raises add-to-cart and destroys checkout completion. Net contribution per visitor: negative. Show it early.
6. POST-PURCHASE AND THE SECOND ORDER
6.1 Why the second order decides the business
At CAC-limited scale, first-order contribution rarely clears CAC by much. Everything after order one is nearly pure margin. Worked — AOV $68, contribution 55% ($37.40/order), CAC $28:
| Cumulative contribution per acquired customer | Net of CAC | |
|---|---|---|
| After order 1 | $37.40 | +$9.40 |
| After order 2 | $46.75 | +$18.75 |
| After order 3 | $51.05 | +$23.05 |
| After order 5 | $54.96 | +$26.96 |
| After order 8 | $56.67 | +$28.67 |
The second order doubles lifetime profit. Orders four through eight add less than orders one through two. That asymmetry is why "increase LTV" as a general programme is usually a mistake and "get the second order" is usually right.
6.2 The retention curve — and the folk statistic it generated
A statistic in every DTC deck: after one purchase, ~27% chance of a second; after two, ~45–54% of a third; after three, ~54% of a fourth. No primary source for it can be found — variously attributed to Adobe and RJMetrics with no citable study (E5/S4. UNVERIFIED.) It is also routinely given a causal reading that is almost certainly false: "each purchase builds loyalty."
Here is what actually generates the pattern. Simulate a population where individual repeat propensity never changes — pure heterogeneity, zero loyalty formation. 70% of first-time buyers repeat with probability 10%; 30% with probability 60%:
| Order n | Share of cohort still buying | Observed P(order n+1 | order n) |
|---|---|---|
| 1 | 100.0% | 25.0% |
| 2 | 25.0% | 46.0% |
| 3 | 11.5% | 57.0% |
| 4 | 6.6% | 59.5% |
| 6 | 2.3% | 60.0% |
25% → 46% → 57%. The famous escalation, reproduced exactly, with no loyalty created and no individual's behaviour changing at any point. The rising curve is sorting: each order filters out the low-propensity segment, so survivors are increasingly drawn from the high-propensity one and the observed rate converges on the best segment's rate.
This is the central result of the Fader–Hardie customer-base analysis literature — the shifted-beta-geometric and BG/NBD models exist precisely to separate heterogeneity from duration dependence, and their repeated finding is that observed retention rises over time even when individual churn probabilities are constant or increasing (E1/S1).
What this changes:
- Stop treating "get them to order twice and they're loyal" as a mechanism. Repeat buyers were mostly always going to be repeat buyers. Your second-order campaign is partly identifying them, not creating them.
- The real second-order lever is acquisition mix, not retention tactics. If your cohort is 70/30 low/high propensity, the highest-value change is acquiring a cohort that is 50/50 — a targeting, offer and product-selection decision, not an email flow. This is the single most important cross-reference in the module: it sends you back to LUCE_03 and LUCE_13, not to LUCE_20.
- Never compute cohort LTV as AOV × average order count. Heterogeneity means the average customer does not exist, and that formula overstates the marginal acquired customer, sometimes badly. (M2 derives LTV properly as a survival-curve integral.)
6.3 Post-purchase upsells
The mechanism is exceptionally clean: a true post-purchase upsell fires after payment authorisation, so it cannot reduce base conversion. Downside risk is bounded at zero excluding brand cost and support load. That is rare enough to emphasise — most levers trade one step against another (§2.6); this one does not.
The numbers are all vendor-reported and mutually inconsistent: ReConvert ~4.7% average acceptance, +5.6% AOV (E4/S3); Zipify ~16.2% (E4/S3); agencies 8–15% (E5/S4). The spread is not noise — it is different definitions and different self-selected user bases. Use them to decide whether to try it (yes, bounded downside). Never to forecast with.
6.4 Subscription mechanics
Subscription converts §6.2's sorting problem into a contractual one — you stop inferring who the high-propensity segment is and ask them to declare it.
Benchmarks, all aggregator-reported (E4/S4): good DTC monthly churn 5–7%, excellent under 5%, top quartile under 3%. Replenishment categories ~4–7%; curated boxes 10–15%; meal kits highest at ~12.7%. Involuntary churn — failed payments — is typically 30–40% of total, reaching 50% for low-AOV brands.
That last figure is the one to act on, because it is the only churn category with a purely mechanical fix — card-account-updater services, retry schedules, pre-dunning notices. No persuasion, no product change. UNVERIFIED as to the exact share; the direction is consistent across sources.
The legal position changed twice recently. The FTC's "click-to-cancel" Negative Option Rule was vacated in its entirety by the Eighth Circuit on 8 July 2025, days before its main provisions took effect (Cooley) (E1/S1). But ROSCA, FTC Act §5 and state auto-renewal laws all still apply and are all still enforced, and the FTC restarted rulemaking on 30 January 2026 (E1/S1). The rule died; the obligation did not — clear disclosure before charging, affirmative consent, cancellation at least as easy as signup.
7. ATTRIBUTION, AND THE THING THAT RUINS ALL OF THIS
7.1 Funnel-step attribution is mostly fiction
Every number in §2 assumes you can observe the steps. Increasingly you cannot, and the failures are not random.
- ITP/ATT, cross-device journeys and in-app browsers break the identity chain differentially by platform, device and traffic source. Your funnel is a measurement artefact whose distortion correlates with the very thing you are comparing.
- Last-click is not merely imprecise; it is wrong in a direction that costs money. Berman, "Beyond the Last Touch," Marketing Science 37(5) shows it over-incentivises ad exposures and produces lower advertiser profits (E2/S1).
- But MTA does not fix it. MTA assumes an unbiased conversion-prediction model that generalises to counterfactual journeys. In real advertising, exposures are algorithmically assigned by predicted propensity, so that assumption fails by construction (E2/S1). In the Advantage+ era this is worse: the platform selects who sees the ad using the same signals that predict conversion. An MTA model fitted to that data measures the platform's targeting, not your creative.
The decisive evidence is Gordon, Zettelmeyer, Bhargava & Chapsky, Marketing Science 38(2) (2019): 15 randomised Facebook experiments, 500 million user-experiment observations, 1.6 billion impressions, comparing RCT results against observational estimates. Observational methods frequently failed to recover the experimental effect even after conditioning on extensive covariates (E2/S1). That is the state of the art, run by the platform, on the platform's own data, with more covariates than you will ever have. It did not work.
7.2 What incrementality testing actually costs
People do not run these tests because they are genuinely expensive, and the foundational paper says so.
Lewis & Rao, "The Unfavorable Economics of Measuring the Returns to Advertising," QJE 130(4) (2015) ran 25 large field experiments, most reaching millions of customers, $2.8m of digital ad spend (E2/S1):
- The median confidence interval on ROI was over 100 percentage points wide.
- Individual sales have a coefficient of variation around 10 relative to per-capita ad cost.
- An informative experiment can easily require more than ten million person-weeks.
- Because the true effect is small relative to noise, selection bias is crippling for observational methods — the same conclusion Gordon et al. reached by a different route.
Ghost ads (Johnson, Lewis & Nubbemeyer, JMR 2017) (E2/S1) identify, in the control group, users who would have been served the ad by logging auction outcomes without serving anything. Strictly better than PSA control, because PSA tests break under algorithmic delivery — the optimiser targets the PSA to people likely to engage with the PSA. The catch: you cannot implement ghost ads yourself. Only the platform can.
Geo experiments (Vaver & Koehler, Google; TBR extension; open-source code) (E1/S1) randomise non-overlapping regions. The TBR variant exists specifically for few geographic units — the small-advertiser case. The most accessible rigorous method, because it needs no platform cooperation.
Design parameters, all vendor/agency folklore (E4/S4, UNVERIFIED) but probably roughly right: hold out 20–30% of revenue-weighted markets; require weekly-revenue correlation >0.8 between matched markets; run 4–6 weeks.
Platform lift tools are user-level holdouts the platform runs for you. Third-party sources report Meta minimums around $30,000 spend, ≥10% of audience per cell, ~100+ conversions/week, 4–6 weeks (E4/S4 — not from Meta documentation, UNVERIFIED). Note the conflict: the platform grades its own homework and chooses the counterfactual.
Calibration on what you are hunting: across 184 conversion-outcome studies, median ad lift was 8.1% with a 90% inter-quantile range of [−8.9%, +83.4%] (Johnson, Lewis & Nubbemeyer) (E2/S1). A median of 8% inside a range including negative values is precisely why Lewis & Rao's sample sizes are so large.
7.3 When a small operator should not bother
Do not attempt incrementality testing if:
- You spend under roughly $30k/month on the channel. Meta's own stated minimum is about your entire monthly spend.
- You have fewer than a few hundred weekly conversions. A coefficient of variation of ~10 means your interval will be wider than the range of decisions you are choosing between.
- The decision is binary and obvious. If the realistic output is "ROAS is somewhere between 0.4 and 3.1," you paid for a holdout to learn nothing.
Instead, in order of usefulness:
- Blended MER against contribution margin, as a time series. Crude, unbiased in expectation, immune to attribution modelling entirely. It cannot tell you which channel worked; it can tell you whether the business worked. At your scale this is not a fallback — it is the best available measure, which is why LUCE_06 puts it first.
- One deliberate whole-channel on/off test per year. Turn a channel fully off for 2–4 weeks and watch blended revenue. Contaminated by seasonality and carryover, but it answers the only question that matters with a real counterfactual, and costs one month of one channel rather than a measurement programme.
- Post-purchase survey. Badly biased, and differently biased from pixel data — which is what makes it useful. When pixel and survey disagree by a factor of three, you have learned something even though neither is right.
- Stop reconciling platform-reported conversions to Shopify. They measure different things with different windows and different logic. Reconciliation is not achievable and the hours have negative expected value.
The connection back to §2 that most people miss: every funnel-step number in your analytics is subject to the same failure. The visitors you can follow through five steps are a biased sample of all visitors — they accepted cookies, used one device, avoided in-app browsers. Another argument for optimising contribution per visitor at the end of the chain rather than step rates in the middle. The end-of-chain number is measured against money, and money does not have a tracking-prevention setting.
8. THE DEAD SCIENCE LIST
Every entry is something an operator will hear from a course, an agency, or a thread. Claim, evidence, what to do instead.
8.1 Ego depletion ("decision fatigue in the funnel")
Claimed: each choice drains a finite reservoir; long funnels exhaust buyers.
Evidence: ego depletion has failed multi-lab pre-registered replication twice. Hagger et al. (2016), Perspectives on Psychological Science: 23 laboratories, N = 2,141, standardised protocol — effects small, and 95% confidence intervals for most labs included zero (E1/S1).
Instead: the practical advice — reduce choices, shorten checkout — is right for entirely different and better-evidenced reasons (§5). Keep the advice. Delete the theory. This matters because a false mechanism generates false predictions: "decision fatigue" wrongly implies a rest between funnel steps helps. It does not.
8.2 Subliminal priming
Evidence: traces to Vicary's 1957 cinema experiment. Vicary admitted he falsified the data, and critics doubt it was ever run (Snopes) (E1/S1). A confessed hoax, for sixty years.
Adjacent casualty — social priming. In Many Labs, the only two of thirteen findings with essentially zero mean effect were the two social priming effects (E1/S1). Doyen et al. failed to replicate Bargh's elderly-priming study (E2/S1). On money priming, Rohrer, Pashler & Harris report zero evidence across four large experiments (E1/S1).
Instead: stop hunting subtle unconscious cues. Supraliminal, explicit information — price, shipping cost, delivery date, return policy, what the product does — has large replicated effects. Buyers respond to things they can see and read.
8.3 "Colour psychology" and the red button
Evidence: the psychology literature says the opposite of the marketing claim. Elliot & Maier's colour-in-context work finds red primes avoidance motivation (E2/S1) — and the core theoretical claim is that colour effects are context-dependent and learned, not universal, directly contradicting "use colour X for buttons."
The famous "red won" results are real individual tests where the winning variant had higher contrast against its surroundings. Change the background and the winner changes. There is no colour effect; there is a salience effect wearing a colour's clothes.
Instead: maximise contrast between the primary CTA and everything around it, one primary CTA per view, then stop. §2.5 priced that project at $78 a year.
8.4 The "$" removal myth
Evidence: the source is real — Yang, Kimes & Sessarego, Cornell Hospitality Report (2009) — and narrower than the meme. Guests spent more with bare numerals than with a dollar sign, and no differently from prices spelled out in words (E2/S1; sample size and magnitude UNVERIFIED — the PDF returned 403).
Why it does not transfer: one restaurant, at lunch, on a menu, in a social dining context, in 2009, one currency. The mechanism — reducing pain-of-paying salience — is real but requires the price be ambiguous enough to de-emphasise. In e-commerce the buyer is entering a card number, sees a currency-denominated total, and often needs the symbol to know which currency they are being charged in. Removing it is a usability defect and, cross-border, arguably a disclosure problem.
Instead: the transferable finding is reduce pain-of-paying salience, and the e-commerce version is payment timing and framing — exactly the mechanism the BNPL literature identifies (§4.7), which has actual field evidence.
8.5 Precise-number anchoring overreach
Evidence: the original finding is real but was demonstrated in estimation and negotiation tasks, not retail purchases (E2/S1); pre-registered work finds substantial moderation by expertise (E2/S1); two high-powered pre-registered experiments (N = 729) now test it directly (E1/S1; results UNVERIFIED).
Instead: precise pricing carries a real and opposite risk the advice ignores — precision signals negotiability and calculation, which is why "$1,997" reads as an internet-marketing price to a large fraction of consumers, damaging exactly the premium positioning where it is most often deployed. Choose price endings for positioning consistency, not a contested lab effect.
8.6 Trust badges "increase conversions by 42%"
Evidence: Baymard did not measure a 42% lift from trust badges. Their actual research is qualitative usability testing plus a survey of which seals users say they trust — e.g. 35.4% of 3,516 respondents chose Norton as most trusted (E4/S2). A survey of stated trust preference is not a conversion experiment. UNVERIFIED and almost certainly false as stated.
Baymard's genuinely interesting finding is that adding any visual security icon raises perceived security, even when the badge means nothing — the mechanism is reassurance signalling, not verification. Counter-finding: six or more seals triggers suspicion.
Instead: one recognisable seal, at the payment fields, not above the fold. And note the ethical edge — if it works even for meaningless badges, displaying a meaningless badge is a misrepresentation, squarely in §4.3's FTC territory.
8.7 "94% of first impressions are design-related"
Evidence: the figure comes from Sillence, Briggs, Fishwick & Harris, CHI 2004 — not Stanford, as usually attributed. The study observed fifteen women aged roughly 41–60 facing a menopause-related health decision over four weeks. The 94% is the share of stated reasons for distrusting and rejecting a health website that referenced design (E3/S1).
So the real claim is: among 15 women, when rejecting health-advice sites, 94% of stated rejection reasons concerned design. It has become a universal law about first impressions across all commerce. The separate "50 milliseconds" figure is a different study measuring a different thing; the two are routinely fused and attributed to a third institution.
Instead: the underlying point — visual credibility matters, especially for unknown brands making health-adjacent claims — is defensible. Quoting "94%" is not.
8.8 Statistics that circulate with no traceable source
| Circulating claim | Status |
|---|---|
| "27% → 45% → 54%" repeat-purchase escalation | UNVERIFIED, no primary source, mechanistically explained by heterogeneity alone (§6.2) |
| "VSL pages convert 12.7% vs 4.8%, per Unbounce's 44,000-page study" | Does not exist. The report covers ~41,000 pages and does not segment by video |
| "34.7% VSL upsell acceptance, study of 1,847 businesses" | UNVERIFIED; no such study findable; the false precision is the tell |
| "Trust badges increase conversion 42% (Baymard)" | Misattributed (§8.6) |
| "Viewers retain 95% of video vs 10% of text" | UNVERIFIED; a mutation of the discredited learning pyramid |
| "11 fields to 4 increased conversions 120%" | Single vendor case study (E5/S3), contradicted by Unbounce's own database (§4.1) |
| "A 1-second delay costs 7% of conversions" | Attributed to Amazon/Google; no published study with that design is findable |
| "One-page checkout converts 7–15% better under $150 AOV" | UNVERIFIED (§5.5) |
| "Red CTAs convert 21%/34% better" | Salience artefact (§8.3) |
The reusable rule: before quoting any conversion statistic, find the study, the sample size, the control group, and who paid. If any of the four is missing, the number is decoration. If it ends in a suspiciously specific decimal and the sample is described but not linked, it is probably fabricated.
9. FLAGGED UNVERIFIED
Nothing below may be quoted publicly, in a lesson, or in a P&L assumption. Direct HTTP fetching was blocked during compilation, so sources were reached through search-index summaries; claims confirmed only that way are listed here even where the source is highly credible.
- Baymard's exact abandonment-reason percentages (48% / 19–26% / 18% / 17–21%) — differ across waves, baymard.com 403'd. Quote as ranges.
- "62% of sites fail to make guest checkout most prominent" — secondary only.
- Baymard's "35.26% achievable checkout increase" — a modelled estimate, not a measured lift.
- Experiment failure rates by company (Bing ~85%, Google ~90%, Netflix ~90%, Airbnb ~90%).
- Floyd et al. exact elasticities (valence 0.78, volume 0.41) — abstract only.
- Barton et al. pooled scarcity magnitude — moderators confirmed, effect size not read.
- Yang & Lynn "11 of 91 attempts" — count not confirmed.
- Foot-in-the-door r ≈ 0.15 — secondary source.
- Frontiers (2026) price-ending results direction — design and N confirmed, results not read.
- Cornell menu-pricing sample size and magnitude — PDF 403'd.
- CUPED's 30–50% variance reduction — practitioner-reported.
- "Sequential testing rolls out winners 30–50% faster" — vendor claim.
- Meta Conversion Lift minimums ($30k, ≥10%/cell, ~100 conv/week) — third-party blogs.
- Geo-test design parameters (20–30% holdout, r > 0.8, 4–6 weeks).
- BNPL merchant fees of 3–6% plus fixed fee.
- Subscription churn benchmarks by category, and "involuntary churn is 30–40% of total".
- All post-purchase upsell acceptance rates (4.7% / 16.2% / 8–15%).
- All quiz-funnel lift figures (300%, 4.9×, 11–15% AOV delta).
- "Advertorials lifted conversion by almost 40%".
- The entire §8.8 table, by construction.
- Multi-step vs one-page checkout percentages.
THE TEN THINGS TO ACTUALLY DO
Ordered by expected value per unit of effort, using §2.5, and only claims surviving §8 and §9.
- Show total landed cost — shipping, tax, fees — before the buyer invests effort. Highest-confidence intervention here (§5.2, §5.6). Do not test it.
- Make guest checkout the most prominent option. Two independent research bodies, a decade of consistent findings (§5.4). Do not test it.
- Add express wallets at cart and checkout — and instrument what it costs you in email capture before celebrating (§4.7).
- Work on the lowest-rate step, almost always interest → add-to-cart, which is a function of the offer, not the button (§2.3).
- Price every optimisation project at your actual CPC, because near breakeven a 10% CVR gain is a 1,000% profit gain (§2.4).
- Under ~25,000 sessions/week, stop A/B testing and start making large changes on mechanism (§3.9). Write down P(win) anyway.
- Build the second order into the acquisition decision, not the retention decision — the rising repeat curve is sorting, not loyalty (§6.2).
- Add post-purchase upsells, because the downside is structurally bounded at zero (§6.3). Do not forecast with vendor numbers.
- Use blended MER against contribution margin as your primary measure, and run at most one coarse channel-off test a year (§7.3).
- Delete every scarcity claim that is not literally true. The effect is real, so a true one works nearly as well; a false one is $53,088 per violation and already on the FTC's published list of examples (§4.3).
CROSS-REFERENCES
- LUCE_08 (Store CRO) — the 18-element product page. §2.3 says which of those elements is worth your week.
- LUCE_18 (CRO Advanced) — testing science and checkout. §3 supersedes any sample-size guidance there that assumed you could detect small effects.
- LUCE_06 (MER) and LUCE_16 (MEO) — §7 is the derivation of why MER-first is correct rather than a simplification.
- LUCE_03 / LUCE_13 (Product Selection) — §6.2 sends the second-order problem here, not to retention.
- LUCE_20 (Email & SMS) — §6.4's involuntary-churn finding is the cheapest retention win and lives in the dunning flow, not the campaign calendar.
- M1 (Mechanisms of Demand) — auction maths and memory science; §4.6 and §8 share its replication discipline.
- M2 (Mechanisms: Economics) — LTV as a survival-curve integral, which §6.2 depends on.
- LUCE_21 (Legal, Tax & Payments Armor) — §4.2, §4.3 and §6.4 are the conversion-side legal surface; 21 is the rest of it.
Up next
Legal, Tax & Payments Armor
The Survival Layer: Entities, Sales Tax, Processors, Chargebacks, FTC Law, Liability, IP, Fraud, and Income Tax — Explained by Mechanism, Not Checklist
69 min