Statistics 1

Condensed sheet

Everything, on one sheet

Every method, every named trap, and every reference card in Statistics 1 — pulled straight from the lessons, so it can never drift out of sync with them.

13 lessons · 620 min, condensed

Read this once, then stop reading it. Re-reading a summary raises how familiar the material feels without changing how much of it you can produce, which is why it feels like studying and mostly isn’t. Use lookup mode when you need a specific fact. Use self-test mode — where the answers stay covered until you’ve tried to say them — for everything else.

Spec 1.1

1 lesson

Mathematical modelling in probability and statistics

"Assume the die is fair" is not a fact about the die — it's a decision about which maths you're allowed to do. Spec item 1.1 is genuinely the thinnest, quietest topic on this paper: across all fourteen WST01 series reviewed for this course, not one question tests "the basic ideas of mathematical modelling" on its own — every direct hit for the word "model" turns out to be about a regression line, a different spec point entirely. And yet the assessment objective this content maps to, AO3, is worth 15 to 20 of this paper's 75 marks — nearly a fifth of it, sitting quietly inside almost every other question on the script. This lesson exists to make that content visible on its own terms, since no single past-paper question ever will.

The card

A model = a simplified, assumption-based description used because it's tractable — never a proven fact.
Assumption → distributional family → parameters → calculation. Wrong assumption = wrong numbers, even with perfect arithmetic.
"Fair die/coin/spinner" = each outcome assumed equally likely → discrete uniform. Values needn't be consecutive integers.
Spec 1.1 has no standalone past-paper question in the reviewed record — its AO3 marks (15-20/75) sit inside other questions.
A "critique this model" or refinement answer must name the SPECIFIC evidence given, not a generic "nothing's perfect" statement.
One probability changing forces others to change too (sum to 1, spec 5.2) — without more data, you can't say which, or by how much.

Why it works — The licensing chain: assumption → distributional family → parameters → calculation

Every calculation in this unit that starts from a modelling assumption follows the same four-step chain, whether the question ever spells the steps out or not. Step 1, the ASSUMPTION: a specific, statable claim about the situation — "the die is fair," "the events are independent," "the underlying variable is Normally distributed." This is a choice, made either explicitly by the question or as a standard convention this unit works under unless told otherwise; it is never something the arithmetic itself can prove. Step 2, the FAMILY the assumption licenses: fairness (every outcome equally likely) licenses the discrete uniform distribution; Normality licenses the Normal distribution; independence licenses P(AB)=P(A)P(B)P(A \cap B) = P(A)P(B) in place of the more general P(AB)=P(A)P(BA)P(A \cap B) = P(A)P(B|A) (this exact substitution — reaching for independence without it being licensed — is part of what this course's own research records as the richest, most consistently-tested sub-topic in the whole unit, appearing in a different guise in all five examiner reports reviewed, and has its own full lesson: conditional-probability-independence-vs-mutually-exclusive.ts). Step 3, the PARAMETERS that pin the family down to one specific member: for a fair nn-sided die, that's just nn itself and the values printed on its faces; for a Normal model, it's μ\mu and σ2\sigma^2. Step 4, the CALCULATION: everything this unit actually asks you to compute — a probability, an expectation, a standardised z-value — is arithmetic performed on the family and parameters step 1 licensed, not on the real situation directly. The chain only runs forward from a genuine assumption; it never runs the other way. A model can be undermined by evidence found AFTER a calculation is done — a box plot showing skew where Normality was assumed (spec 2.4, outliers-boxplots-comparing-distributions.ts's own territory), a regression line asked to predict miles outside the range of xx it was fitted to (spec 4.2, regression-gradient-interpretation-extrapolation.ts's own territory) — but that evidence never rewrites the arithmetic itself; it only tells you the assumption that licensed the arithmetic in the first place needs revisiting. Confusing "the calculation was done correctly" with "the model was appropriate" is exactly the gap this whole spec point exists to close, and it's a distinct question every single time: was the METHOD right, and separately, was the ASSUMPTION the method rested on actually reasonable for this situation?

Traps — 3

spec-1.1-expected-as-a-standalone-question
This course's own research pass is direct about what it found (and didn't find) searching all fourteen reviewed WST01 series for the words "model"/"modelling": no standalone question tests spec 1.1 in isolation anywhere in that record — every direct hit turns out to be about a regression model instead (spec 4.1/4.2). This isn't an examiner-report-documented misconception the way most trap-taxonomy items on this course are; it's a genuine, checkable finding about how this content is actually examined, and it produces a real exam-navigation trap in its own right: revising this spec point by looking for a past "define a mathematical model" question to practise on is a search that will come back empty every time, because the AO3 marks this content maps to (15-20 of 75) are distributed as sub-parts riding on top of other questions, not concentrated in a question of their own.
assumption-treated-as-a-provable-fact
VERIDIAN-original naming for a genuine failure mode this lesson's own material is built to guard against, not a quoted examiner misconception (none exists in the reviewed record for this specific spec point). "The die is fair" and "growth is Normally distributed" are grammatically identical to plain statements of fact, and it's easy to read them that way — as something the question has already established, rather than something it's asking you to accept for the purpose of the calculation that follows. The practical cost: a student who reads an assumption as a fact has nothing to say when a later part of the same question hands them evidence against it (see the marked-solution's part (c) above), because they never registered there was anything provisional about the claim in the first place.
refinement-answer-not-connected-to-the-specific-evidence-given
VERIDIAN-original naming, illustrated directly by this lesson's own marked-solution common-wrong-path above. A true-sounding, generically applicable statement — "real spinners are never perfectly fair," "real dice have manufacturing tolerances" — earns nothing on a "critique this model" or "explain what this evidence suggests" part, however scientifically reasonable it sounds, if it isn't tied to the SPECIFIC observation the question actually reported. The credited answer names what was specifically observed (sector 5 over-represented, not the spinner in general) and states the specific consequence for the model (which probability moves, and why the total must still sum to 1) — not a general truth that would apply equally well to any question of this shape.

Say it out loud

Out loud, from memory, no notes: explain the licensing chain: assumption → distributional family → parameters → calculation to someone who has never seen this topic — where does your explanation get vague or hand-wavy? That's the exact spot to re-study, and it only works if you check it: read back over the mechanism above the moment you finish talking and mark precisely where you drifted from it.

Spec 2.1

2 lessons

Reading Data Representations, and Comparing Distributions in Words

A histogram's bars can lie to you at a glance. The vertical axis reads frequency DENSITY, not frequency — and the two are only the same number when every class happens to be the same width, which a real exam question is under no obligation to give you. The single most repeated error the mark schemes for this section record is exactly this: reading a bar's height straight off the axis and writing it down as the answer, skipping the one multiplication — by class width — that turns a density into an actual count of people. A stem-and-leaf diagram hides a quieter version of the same problem: read it from the wrong end, and Q1 and Q3 swap places without anything on the page telling you they have. And once you've read the real figures off either diagram, a "compare the two distributions" question is asking you to use them — name a statistic, give both figures, answer what was actually asked — not just to have found them.

The card

Histogram: frequency = frequency density × class width — AREA, not height, is frequency.
Different-width bars can only be compared by area. A taller bar can still represent FEWER data points.
Self-check: your frequencies should sum to the total the question states.
Back-to-back stem-and-leaf: on BOTH sides, leaves ascend outward from the stem — the left side just prints that in reverse.
Quartile position: Q1 = (n+1)/4th value, median = (n+1)/2th, Q3 = 3(n+1)/4th — count from the correct end.
Comparing distributions: name a statistic, give both figures, state the direction, answer what's actually asked.

Traps — 3

histogram-height-read-as-frequency
Confirmed, in the one examiner report reviewed in this research pass that treats a histogram-reading question directly: "the majority of those that were unsuccessful failed to realise that the area of bar represented the frequency and so simply gave frequencies of 25 and 5" (Jan 2023, Q1(a)). Frequency density is a RATE — data points per unit of x-axis — not a count; multiplying by the class width is what turns it into one, and skipping that step is the entire error, not a rounding slip layered on top of an otherwise correct method.
stem-and-leaf-read-in-the-wrong-direction
Confirmed: "some students read the stem and leaf diagram the wrong way round and so incorrectly identified Q1 = 31 and Q3 = 51" (Jun 2024, Q1(b)) — swapping which end of the ordered leaves is the smaller quartile and which is the larger. On a back-to-back diagram specifically, this usually comes from reading the LEFT-hand side in plain left-to-right print order instead of recognising it mirrors the right-hand side's own ascending-outward rule (see the teach block above for the mechanism). This same trap is named in the Outliers, Box Plots and Comparing Distributions lesson's own trap-taxonomy, since a misread stem-and-leaf diagram is exactly as capable of feeding a wrong Q1/Q3 into an outlier-fence calculation as it is of feeding one into this lesson's own reading-practice content — it is cited in both places because it is genuinely load-bearing for both skills, not duplicated by accident.
comparison-lacks-a-named-statistic-figures-or-the-right-focus
Three separate confirmed patterns, all under the same "compare two distributions in words" guidance point, and all already given full standalone treatment (their own concept-ladder, marked-solution part and trap-taxonomy entries) in the Outliers, Box Plots and Comparing Distributions lesson — restated here only because they are exactly as relevant when the two distributions being compared were read off a histogram or a stem-and-leaf diagram, which is this lesson's own subject. (1) The single most common failure: treating the mean as inherently "more accurate" than the median — "A very common misconception was that the mean is more accurate than the median... showing no appreciation that the mean is just one measure of average and the median is another" (Jun 2022, Q1(e)). (2) Naming no actual figures: "surprisingly too many students failed to give supporting figures... a reference to a named statistic and supporting figures [is required]" (Jun 2024, Q1(e)). (3) Answering a different question from the one asked: "few comments referring to the distribution... others commented on the similarities [when asked for differences]" (Oct 2021, Q3(e)). A comparison that names a statistic, states both figures, and addresses exactly what the question asked for is the only shape of answer that scores against any of the three.

Say it out loud

Out loud, from memory, no notes: explain reading data representations, and comparing distributions in words to someone who has never seen this topic — where does your explanation get vague or hand-wavy? That's the exact spot to re-study, and it only works if you check it: read back over the mechanism above the moment you finish talking and mark precisely where you drifted from it.

Outliers, Box Plots, and Comparing Distributions

A box plot's whiskers only ever touch real data. The fence that decides who counts as an outlier is arithmetic you compute and then throw away — it never appears as a mark on the finished plot — and one of the most consistently documented box-plot errors on this paper is drawing a whisker straight to that invisible number instead of stopping at the last real value before it. The formula for the fence isn't fixed, either: the specification insists every question state its own rule, precisely so a memorised "1.5 × IQR" can't quietly become the wrong number for the one sitting in front of you.

The card

Five-number summary: min, Q1, median, Q3, max. Box = Q1 to Q3 (the IQR); line inside = median.
IQR = Q3 − Q1 — not in the booklet, memorise it.
The outlier rule is ALWAYS given in the question. Use exactly what's stated, never a remembered formula.
Fences point outward: lower = Q1 − (k × IQR); upper = Q3 + (k × IQR). Never the reverse.
A whisker stops at the most extreme NON-outlier value — never at the fence, never at the outlier.
Mean uses every value (outliers pull it); median uses only position (outliers barely move it).
Comparing distributions: name a statistic, give both figures, state the direction. No numbers, no marks.

Why it works — Where 1.5 × IQR fences come from, and why the direction is not optional

The IQR measures the width of the middle 50% of the data — how spread out the 'ordinary' half of it is. An outlier rule takes that width as its own unit of measurement and asks: how many IQRs outside the box does a value have to sit before it stops looking like a natural continuation of the data and starts looking like something genuinely unusual? A rule like 1.5×IQR1.5 \times IQR answers that by marking a fence at 1.5 IQRs beyond EACH quartile, extending outward, away from the box. That direction isn't a convention to memorise; it's forced by what a fence is for. The lower fence has to sit below Q1Q_1, so it's Q1Q_1 minus the distance — subtracting. The upper fence has to sit above Q3Q_3, so it's Q3Q_3 plus the distance — adding. A fence built the other way round, added to Q1Q_1 or subtracted from Q3Q_3, moves INTO the box rather than away from it, and for a large enough multiplier can end up on the wrong side of the median entirely — at which point it would start flagging perfectly ordinary values near the middle of the data as outliers, exactly the opposite of what the rule exists to do. The confirmed error record shows precisely this: "there were a surprising number of errors... with some multiplying the quartiles by 1.5 and others using Q1 + 1.5×IQR for the lower limit" (Jan 2021, Q2) — both are the same underlying mistake, arithmetic that produces a real number but not the number the fence is defined to be. There's a second reason not to treat 1.5 as a fixed fact at all: the specification's own guidance on this spec item states plainly that "any rule to identify outliers will be specified in the question" — 1.5 is the multiplier used throughout this lesson because it's the one seen most often in the reviewed record, not because it's fixed. A real question can hand you a different multiplier, and when it does, that stated rule overrides whatever you remember from practice. The confirmed record shows exactly what happens when a student substitutes the remembered version anyway: "a small number of candidates opted to apply the more commonly used outlier formula of Q3 + 1.5×(Q3−Q1) rather than the one quoted in the question" (Jun 2022, Q1(b)). What has to transfer between questions is the outward-facing DIRECTION of the rule, not a specific number attached to it.

Traps — 8

outlier-lower-fence-direction-reversed
Confirmed directly, and the research bank's own phrasing makes clear this was not a rare slip: "there were a surprising number of errors... with some multiplying the quartiles by 1.5 and others using Q1 + 1.5×IQR for the lower limit" (Jan 2021, Q2). The lower fence has to SUBTRACT from Q1, moving further below it — the mechanism block above derives why, rather than asking you to remember it as a rule with no reason behind it.
whisker-drawn-to-fence-not-to-data
The research bank's own account of the same Jan 2021 Q2 finding (its documented summary of the examiner report, not a direct quotation) is that box-plot whiskers were commonly drawn out to the outlier boundary itself (98, in that series) rather than to the actual highest non-outlier value in the data (97). A fence is a number used to test values against; it's never itself a point that gets drawn. See the diagram block above for the same error, reproduced with this lesson's own numbers (74.5 versus the real value, 61).
remembered-formula-overrides-the-rule-given-in-the-question
Confirmed: "a small number of candidates opted to apply the more commonly used outlier formula of Q3 + 1.5×(Q3−Q1) rather than the one quoted in the question" (Jun 2022, Q1(b)). This matters because the specification itself states that "any rule to identify outliers will be specified in the question" (spec S1.3, item 2.4, guidance) — there is deliberately no single fixed rule to memorise, and treating 1.5×IQR as a universal constant is itself the error, even on the (common) occasions where the question happens to specify 1.5 anyway.
show-that-outliers-not-listed
Confirmed, on a "show that there are 3 outliers" question: "having gained the correct limits some did not list the 3 outliers in this part in order to show there are 3 outliers" (Oct 2021, Q3(c)). Finding the correct fences is necessary but not sufficient — a "show that N outliers exist" question needs the N values actually named, not just the machinery that would find them.
mean-assumed-more-accurate-than-the-median
Confirmed, and stated by the examiner report as a genuinely common pattern rather than a rare one: "very few candidates scored this mark... A very common misconception was that the mean is more accurate than the median. A large number of responses simply explained how to calculate a mean or said that it was because the mean is the average, showing no appreciation that the mean is just one measure of average and the median is another" (Jun 2022, Q1(e)). The credited answers were specific and short: "the mean uses all the data" / "the mean includes the outliers." Neither statistic is "more accurate" — they are different measures of the same idea, and the concept-ladder above works through exactly why an outlier moves one and not the other.
comparison-missing-supporting-figures
Confirmed: "surprisingly too many students failed to give supporting figures... a question like this will require some context..., a reference to a named statistic and supporting figures" (Jun 2024, Q1(e)). "Branch A had a higher average" names nothing and gives no numbers; "Branch A had a higher median (34 minutes against 28)" does both, and only the second shape of answer scores.
comparison-answers-the-wrong-question
Confirmed: "few comments referring to the distribution of ages were seen... others commented on the similarities [when asked for differences]" (Oct 2021, Q3(e)). Read what the question is actually asking for — differences, similarities, a specific statistic — before writing the comparison, since a technically-true observation about the wrong aspect of the data scores nothing.
stem-and-leaf-read-in-the-wrong-direction
Confirmed, on a real quartile-from-stem-and-leaf question: "some students read the stem and leaf diagram the wrong way round and so incorrectly identified Q1 = 31 and Q3 = 51" (Jun 2024, Q1(b)) — they swapped which end of the ordered leaves is the lower quartile and which is the upper. Whatever representation a box plot is built from — a stem-and-leaf diagram, a table, a raw list — check which end you're counting from before quoting a quartile out of it.

Say it out loud

Out loud, from memory, no notes: explain where 1.5 × iqr fences come from, and why the direction is not optional to someone who has never seen this topic — where does your explanation get vague or hand-wavy? That's the exact spot to re-study, and it only works if you check it: read back over the mechanism above the moment you finish talking and mark precisely where you drifted from it.

Spec 2.2

1 lesson

Measures of location and dispersion — mean, coding, and standard deviation

Coding turns an awkward mean — 255.2 — into an easy one, 0.1. That saving isn't free: add a constant and the mean slides by exactly that constant, but multiply by one and the mean scales by it once while the *variance* scales by its square — and has to be undone by squaring again on the way back. One real WST01 examiner's report records three separate ways candidates got that one decoding step wrong, in a single sitting. The standard deviation formula fails in an even more literal way: two different exam series produced malformed versions of it — missing a division, or missing a square root in the place it actually belongs — and both examiner reports trace the damage back to the same habit, rounding a value mid-calculation instead of carrying it through exactly. None of this is really about arithmetic. It's about which few numbers this exam expects you to reconstruct from memory, because the formula booklet — deliberately — will not hand them to you.

The card

Mean: x̄ = Σx/n (list) or Σfx/Σf (table/grouped, using midpoints) — divide by the total frequency, never the number of rows.
Standard deviation: √(Σfx²/n − x̄²) — ÷n before the subtraction, √ wraps the whole result. Not in the booklet: know it from memory.
Coding Y = b(X − a) or (X − a)/b: decode the mean by reversing the steps; decode the variance by the SQUARE of the constant, and the OPPOSITE operation.
Recovering Σx² from a given SD: Σx² = n(SD² + mean²) — square the SD first, always.
Interpolation position: n/4, n/2, 3n/4 (or the n+1 versions) — both accepted, but can diverge right at a class boundary.
Keep exact values through a calculation; round only the final answer, to 3 s.f. unless told otherwise.
A manifestly implausible final answer is a reason to re-check the method, not just the arithmetic.

Why it works — Why interpolation works — and why the n-vs-(n+1) convention can (very slightly) disagree

Linear interpolation for a quartile rests on one assumption, stated plainly rather than left implicit: within a class, the data is spread EVENLY across it. Nobody actually knows where, inside the 30m<4030 \le m < 40 class, each of its 15 packages' true masses fall — only that 15 of them do — so interpolation assumes they're spaced out uniformly along that 10 kg interval, and estimates Q3Q_3's position accordingly. That's a genuine modelling assumption, the same kind of assumption spec 1.1 names directly ("the basic ideas of mathematical modelling as applied in probability and statistics") — it's very often close enough to be useful, and it's never claimed to be exact, which is precisely why Q1Q_1 and Q3Q_3 found this way are described as ESTIMATES of the quartiles, not their true values. The formula itself just encodes that assumption in numbers: Q=L+positionFfclass×hQ = L + \dfrac{\text{position} - F}{f_{\text{class}}} \times h, where LL is the class's lower boundary, FF is the cumulative frequency reached before that class starts, fclassf_{\text{class}} is the class's own frequency, and hh is its width — position, minus what's already been passed, as a FRACTION of the class, scaled up to the class's actual width. Spec 2.3's own guidance flags this directly: "simple interpolation may be required" — and it isn't printed in the formula booklet any more than mean or standard deviation are, so the shape of it has to be remembered, not looked up. Two conventions exist for finding the POSITION itself — n/4n/4 or (n+1)/4(n+1)/4 for Q1Q_1, and correspondingly 3n/43n/4 or 3(n+1)/43(n+1)/4 for Q3Q_3 — and a real WST01 mark scheme is confirmed to tolerate either. They usually land in the same class and differ only in the last decimal place of the final answer. This lesson's own worked masses table above is a case where they don't quite agree even that much: n/4=60/4=15n/4 = 60/4 = 15 lands EXACTLY on the cumulative frequency 15 reached at the boundary between the 10m<2010 \le m < 20 and 20m<3020 \le m < 30 classes, giving Q1=20Q_1 = 20 flat with no fraction of a class to interpolate across; (n+1)/4=61/4=15.25(n+1)/4 = 61/4 = 15.25 lands just PAST that same boundary, inside the next class, giving Q1=20+0.2520×10=20.125Q_1 = 20 + \frac{0.25}{20} \times 10 = 20.125 instead — a genuinely different class used for the interpolation, not just a rounding difference. (This mechanism — that the two conventions can disagree specifically at a class-boundary cumulative frequency — is VERIDIAN's own reasoning about why the finding below holds, not a claim the source material itself makes.) The verified finding itself, from a real series: "those that worked with n rather than n + 1 were usually more successful" — plausibly because (n+1)/4(n+1)/4's tiny extra push past a boundary like this one is exactly the kind of ambiguity that trips a rushed interpolation up, landing a student in the wrong class without them noticing. Neither convention is marked wrong on its own — but this is worth knowing before choosing one out of habit, especially since the two-quartile version of this effect compounds: this lesson's masses table gives IQR=16.7IQR = 16.7 under the nn convention and 17.017.0 under (n+1)(n+1) (3 s.f. each) — different enough to matter, from data that never changed at all. One further, related trap belongs here by name even though its own full worked treatment lives in this course's "Show that" answer discipline lesson: on a "show that the IQR equals [given value]" question, a real examiner report records candidates "making up two ages... which differed by 16 as their 'working'" (Oct 2021, Q3(b)) — inventing two numbers with the right GAP between them instead of genuinely interpolating either quartile. That's a different failure from choosing the wrong convention; it's not attempting the method at all.

Traps — 10

mean-denominator-is-class-count-not-total-frequency
Confirmed, and stated by the examiner report as a genuinely disappointing pattern, not a rare one: "it was disappointing to see that some students were unsure on how to calculate a mean from a frequency table. Some added the frequencies and divided by 5 whilst others used the sum of class width multiplied by the frequencies" (Jan 2023, Q1(c)). The denominator of a mean is always Σf — the total number of data values — never the number of rows the table happens to be split across.
mean-numerator-uses-class-width-instead-of-midpoint
The second half of the same confirmed finding: some candidates used Σ(class width × frequency) as the numerator instead of Σ(midpoint × frequency). Worth noticing explicitly: for a table where every class has the SAME width (as in the marked-solution example above), this particular wrong shortcut always collapses to exactly the class width itself, no matter what the actual frequencies are — Σ(h×f)/Σf = h when h is constant — which is itself a giveaway that something's gone wrong, since a genuine mean has to depend on where the data actually sits, not just on how the table happens to be split into intervals.
coding-decode-forgot-to-square-constant
Confirmed, real, and directly quoted: candidates "scaling using 0.5 instead of 0.5²" when decoding a coded variance (Jun 2022, Q3(e)). Coding's effect on variance is always squared, because variance is itself built from squared deviations — see the concept-ladder above for the full mechanism, not just the rule.
coding-decode-multiplied-instead-of-divided
A more subtle version of the same real question, quoted directly: candidates "knowing they should use 0.5² but multiplying instead of dividing" (Jun 2022, Q3(e)). Getting the CONSTANT right (squaring it) but the OPERATION wrong moves the variance in the wrong direction entirely — decoding should always make the variance BIGGER than the coded value, since the original data is more spread out on its own (uncoded) scale than the easy numbers coding produced.
coding-not-decoded-at-all
Confirmed, real, and distinct from either error above: the same examiner report records candidates "not decoding" at all — reporting the coded variance, or a value still carrying "the assumed mean of 255" folded in incorrectly, as if it were the answer for the original data. A coded answer is a WORKING TOOL, never the final one; every coded calculation has to be translated back before it answers the question that was actually asked.
malformed-standard-deviation-formula
Confirmed as a real pattern from a genuine examiner report — a formula that isn't shaped like a standard deviation at all, rather than a numerically close slip. WST01-verified-facts.md describes (its own summary, not a Pearson direct quotation) two specific malformed shapes candidates produced in the same series, Jan 2023 Q1(d): one missing the ÷n step before the square root, one missing the square root entirely while also dividing the whole subtraction by n instead of just the first term. Either produces a wildly implausible number — see the marked-solution's commonWrongPath above for exactly how implausible, with real numbers attached. (A related but distinct error — squaring a given standard deviation before using it, rather than misshaping the SD formula itself — is confirmed independently in Jan 2021 Q6(c); see the separate forgot-to-square-sd-recovering-sum-of-squares entry below.)
mean-rounded-mid-calculation
Confirmed, and directly quoted: "students should be encouraged to work with exact answers in their working of calculations" (Jan 2023, Q1(d)) — rounding the mean before it feeds into a variance calculation was a documented, real source of inaccuracy on this exact question. Carry the exact value through every intermediate step; round only the final answer, to the paper's own stated 3 significant figures unless told otherwise.
forgot-to-square-sd-recovering-sum-of-squares
Confirmed, real, and directly quoted: "forgot to square the standard deviation of 2 when finding Σy² for the second group" (Jan 2021, Q6(c)). Recovering Σx² from a given mean and standard deviation requires rearranging Var = Σx²/n − mean² to Σx² = n(SD² + mean²) — the SD has to be squared into a variance first, the same squaring rule that governs coding above, showing up again in a differently-shaped question.
interpolation-convention-can-shift-the-class-at-a-boundary
Confirmed, real, and directly quoted: "those that worked with n rather than n + 1 were usually more successful" (Jan 2023, Q1(e)). Both conventions for the quartile position — n/4 or (n+1)/4 — are genuinely accepted by the mark scheme; the risk with (n+1)/4 is specifically at a class boundary, where a small extra push past a boundary can land the position inside a DIFFERENT class than n/4 would, changing which numbers the interpolation formula actually uses. See the mechanism block above for exactly this happening with this lesson's own masses data (VERIDIAN's own explanation for why the finding holds, clearly separated there from the verified finding itself).
interpolation-answer-reverse-engineered-to-match-a-given-target
A distinct, real trap on "show that" interpolation questions specifically, confirmed and directly quoted: "the main errors, since the answer was given, were making up two ages... which differed by 16 as their 'working'" (Oct 2021, Q3(b)). This is not a wrong CONVENTION or a wrong FORMULA — it's inventing numbers with the right gap between them instead of genuinely interpolating either quartile. Its full treatment, including why a mark scheme structurally cannot credit this kind of working, lives in this course's "Show that" answer discipline lesson; it's named here because it's specifically an interpolation trap, and this lesson would be incomplete without at least naming it.

Say it out loud

Out loud, from memory, no notes: explain why interpolation works — and why the n-vs-(n+1) convention can (very slightly) disagree to someone who has never seen this topic — where does your explanation get vague or hand-wavy? That's the exact spot to re-study, and it only works if you check it: read back over the mechanism above the moment you finish talking and mark precisely where you drifted from it.

Spec 2.3

2 lessons

Measures of location and dispersion — mean, coding, and standard deviation

Coding turns an awkward mean — 255.2 — into an easy one, 0.1. That saving isn't free: add a constant and the mean slides by exactly that constant, but multiply by one and the mean scales by it once while the *variance* scales by its square — and has to be undone by squaring again on the way back. One real WST01 examiner's report records three separate ways candidates got that one decoding step wrong, in a single sitting. The standard deviation formula fails in an even more literal way: two different exam series produced malformed versions of it — missing a division, or missing a square root in the place it actually belongs — and both examiner reports trace the damage back to the same habit, rounding a value mid-calculation instead of carrying it through exactly. None of this is really about arithmetic. It's about which few numbers this exam expects you to reconstruct from memory, because the formula booklet — deliberately — will not hand them to you.

The card

Mean: x̄ = Σx/n (list) or Σfx/Σf (table/grouped, using midpoints) — divide by the total frequency, never the number of rows.
Standard deviation: √(Σfx²/n − x̄²) — ÷n before the subtraction, √ wraps the whole result. Not in the booklet: know it from memory.
Coding Y = b(X − a) or (X − a)/b: decode the mean by reversing the steps; decode the variance by the SQUARE of the constant, and the OPPOSITE operation.
Recovering Σx² from a given SD: Σx² = n(SD² + mean²) — square the SD first, always.
Interpolation position: n/4, n/2, 3n/4 (or the n+1 versions) — both accepted, but can diverge right at a class boundary.
Keep exact values through a calculation; round only the final answer, to 3 s.f. unless told otherwise.
A manifestly implausible final answer is a reason to re-check the method, not just the arithmetic.

Why it works — Why interpolation works — and why the n-vs-(n+1) convention can (very slightly) disagree

Linear interpolation for a quartile rests on one assumption, stated plainly rather than left implicit: within a class, the data is spread EVENLY across it. Nobody actually knows where, inside the 30m<4030 \le m < 40 class, each of its 15 packages' true masses fall — only that 15 of them do — so interpolation assumes they're spaced out uniformly along that 10 kg interval, and estimates Q3Q_3's position accordingly. That's a genuine modelling assumption, the same kind of assumption spec 1.1 names directly ("the basic ideas of mathematical modelling as applied in probability and statistics") — it's very often close enough to be useful, and it's never claimed to be exact, which is precisely why Q1Q_1 and Q3Q_3 found this way are described as ESTIMATES of the quartiles, not their true values. The formula itself just encodes that assumption in numbers: Q=L+positionFfclass×hQ = L + \dfrac{\text{position} - F}{f_{\text{class}}} \times h, where LL is the class's lower boundary, FF is the cumulative frequency reached before that class starts, fclassf_{\text{class}} is the class's own frequency, and hh is its width — position, minus what's already been passed, as a FRACTION of the class, scaled up to the class's actual width. Spec 2.3's own guidance flags this directly: "simple interpolation may be required" — and it isn't printed in the formula booklet any more than mean or standard deviation are, so the shape of it has to be remembered, not looked up. Two conventions exist for finding the POSITION itself — n/4n/4 or (n+1)/4(n+1)/4 for Q1Q_1, and correspondingly 3n/43n/4 or 3(n+1)/43(n+1)/4 for Q3Q_3 — and a real WST01 mark scheme is confirmed to tolerate either. They usually land in the same class and differ only in the last decimal place of the final answer. This lesson's own worked masses table above is a case where they don't quite agree even that much: n/4=60/4=15n/4 = 60/4 = 15 lands EXACTLY on the cumulative frequency 15 reached at the boundary between the 10m<2010 \le m < 20 and 20m<3020 \le m < 30 classes, giving Q1=20Q_1 = 20 flat with no fraction of a class to interpolate across; (n+1)/4=61/4=15.25(n+1)/4 = 61/4 = 15.25 lands just PAST that same boundary, inside the next class, giving Q1=20+0.2520×10=20.125Q_1 = 20 + \frac{0.25}{20} \times 10 = 20.125 instead — a genuinely different class used for the interpolation, not just a rounding difference. (This mechanism — that the two conventions can disagree specifically at a class-boundary cumulative frequency — is VERIDIAN's own reasoning about why the finding below holds, not a claim the source material itself makes.) The verified finding itself, from a real series: "those that worked with n rather than n + 1 were usually more successful" — plausibly because (n+1)/4(n+1)/4's tiny extra push past a boundary like this one is exactly the kind of ambiguity that trips a rushed interpolation up, landing a student in the wrong class without them noticing. Neither convention is marked wrong on its own — but this is worth knowing before choosing one out of habit, especially since the two-quartile version of this effect compounds: this lesson's masses table gives IQR=16.7IQR = 16.7 under the nn convention and 17.017.0 under (n+1)(n+1) (3 s.f. each) — different enough to matter, from data that never changed at all. One further, related trap belongs here by name even though its own full worked treatment lives in this course's "Show that" answer discipline lesson: on a "show that the IQR equals [given value]" question, a real examiner report records candidates "making up two ages... which differed by 16 as their 'working'" (Oct 2021, Q3(b)) — inventing two numbers with the right GAP between them instead of genuinely interpolating either quartile. That's a different failure from choosing the wrong convention; it's not attempting the method at all.

Traps — 10

mean-denominator-is-class-count-not-total-frequency
Confirmed, and stated by the examiner report as a genuinely disappointing pattern, not a rare one: "it was disappointing to see that some students were unsure on how to calculate a mean from a frequency table. Some added the frequencies and divided by 5 whilst others used the sum of class width multiplied by the frequencies" (Jan 2023, Q1(c)). The denominator of a mean is always Σf — the total number of data values — never the number of rows the table happens to be split across.
mean-numerator-uses-class-width-instead-of-midpoint
The second half of the same confirmed finding: some candidates used Σ(class width × frequency) as the numerator instead of Σ(midpoint × frequency). Worth noticing explicitly: for a table where every class has the SAME width (as in the marked-solution example above), this particular wrong shortcut always collapses to exactly the class width itself, no matter what the actual frequencies are — Σ(h×f)/Σf = h when h is constant — which is itself a giveaway that something's gone wrong, since a genuine mean has to depend on where the data actually sits, not just on how the table happens to be split into intervals.
coding-decode-forgot-to-square-constant
Confirmed, real, and directly quoted: candidates "scaling using 0.5 instead of 0.5²" when decoding a coded variance (Jun 2022, Q3(e)). Coding's effect on variance is always squared, because variance is itself built from squared deviations — see the concept-ladder above for the full mechanism, not just the rule.
coding-decode-multiplied-instead-of-divided
A more subtle version of the same real question, quoted directly: candidates "knowing they should use 0.5² but multiplying instead of dividing" (Jun 2022, Q3(e)). Getting the CONSTANT right (squaring it) but the OPERATION wrong moves the variance in the wrong direction entirely — decoding should always make the variance BIGGER than the coded value, since the original data is more spread out on its own (uncoded) scale than the easy numbers coding produced.
coding-not-decoded-at-all
Confirmed, real, and distinct from either error above: the same examiner report records candidates "not decoding" at all — reporting the coded variance, or a value still carrying "the assumed mean of 255" folded in incorrectly, as if it were the answer for the original data. A coded answer is a WORKING TOOL, never the final one; every coded calculation has to be translated back before it answers the question that was actually asked.
malformed-standard-deviation-formula
Confirmed as a real pattern from a genuine examiner report — a formula that isn't shaped like a standard deviation at all, rather than a numerically close slip. WST01-verified-facts.md describes (its own summary, not a Pearson direct quotation) two specific malformed shapes candidates produced in the same series, Jan 2023 Q1(d): one missing the ÷n step before the square root, one missing the square root entirely while also dividing the whole subtraction by n instead of just the first term. Either produces a wildly implausible number — see the marked-solution's commonWrongPath above for exactly how implausible, with real numbers attached. (A related but distinct error — squaring a given standard deviation before using it, rather than misshaping the SD formula itself — is confirmed independently in Jan 2021 Q6(c); see the separate forgot-to-square-sd-recovering-sum-of-squares entry below.)
mean-rounded-mid-calculation
Confirmed, and directly quoted: "students should be encouraged to work with exact answers in their working of calculations" (Jan 2023, Q1(d)) — rounding the mean before it feeds into a variance calculation was a documented, real source of inaccuracy on this exact question. Carry the exact value through every intermediate step; round only the final answer, to the paper's own stated 3 significant figures unless told otherwise.
forgot-to-square-sd-recovering-sum-of-squares
Confirmed, real, and directly quoted: "forgot to square the standard deviation of 2 when finding Σy² for the second group" (Jan 2021, Q6(c)). Recovering Σx² from a given mean and standard deviation requires rearranging Var = Σx²/n − mean² to Σx² = n(SD² + mean²) — the SD has to be squared into a variance first, the same squaring rule that governs coding above, showing up again in a differently-shaped question.
interpolation-convention-can-shift-the-class-at-a-boundary
Confirmed, real, and directly quoted: "those that worked with n rather than n + 1 were usually more successful" (Jan 2023, Q1(e)). Both conventions for the quartile position — n/4 or (n+1)/4 — are genuinely accepted by the mark scheme; the risk with (n+1)/4 is specifically at a class boundary, where a small extra push past a boundary can land the position inside a DIFFERENT class than n/4 would, changing which numbers the interpolation formula actually uses. See the mechanism block above for exactly this happening with this lesson's own masses data (VERIDIAN's own explanation for why the finding holds, clearly separated there from the verified finding itself).
interpolation-answer-reverse-engineered-to-match-a-given-target
A distinct, real trap on "show that" interpolation questions specifically, confirmed and directly quoted: "the main errors, since the answer was given, were making up two ages... which differed by 16 as their 'working'" (Oct 2021, Q3(b)). This is not a wrong CONVENTION or a wrong FORMULA — it's inventing numbers with the right gap between them instead of genuinely interpolating either quartile. Its full treatment, including why a mark scheme structurally cannot credit this kind of working, lives in this course's "Show that" answer discipline lesson; it's named here because it's specifically an interpolation trap, and this lesson would be incomplete without at least naming it.

Say it out loud

Out loud, from memory, no notes: explain why interpolation works — and why the n-vs-(n+1) convention can (very slightly) disagree to someone who has never seen this topic — where does your explanation get vague or hand-wavy? That's the exact spot to re-study, and it only works if you check it: read back over the mechanism above the moment you finish talking and mark precisely where you drifted from it.

"Show that" answer discipline

A "show that" question already gives you the answer. That is not a hint about how easy the question is — it is a warning about how the mark scheme has to be built. When the final number is printed on the page, an examiner cannot give you credit for reaching it, because copying a printed number proves nothing about whether you can produce it yourself. Every mark in a "show that" question is attached to the *steps between* the start of the question and the number you were already handed — and Pearson has said, in three separate examiner reports spanning three years, that this is the single thing students most often get wrong on this paper, in every topic it touches.

The card

"Show that X" means the answer is given — the marks are for the steps to X, not for X itself.
Never write the target value first. Show the method, and let the target value appear as your last line.
Work to one more d.p./s.f. than the target throughout — matching its precision is how the last mark is lost.
"Show that there are N of something" needs the N things named, not the rule that would find them.
M0 A1 is impossible: no visible method means the accuracy marks have nothing to attach to, whatever the final line says.

Why it works — Why the mark scheme genuinely cannot pay for the printed answer

This follows directly from how M, A and B marks work, not from an arbitrary examiner preference (WST01-verified-facts.md §4; this exact "General Instructions for Marking" wording opens at least 6 of the 14 reviewed mark schemes, byte-for-byte — confirmed directly against the real Jan 2024 mark scheme PDF; Jun 2024's own preamble is shorter but keeps the same M-before-A dependency rule). An M mark is "given for a correct method or an attempt at a correct method" — it has to be visible to be given, because there is nothing else for the examiner to assess it against. An A mark is "dependent... and can only be awarded if the previous M mark has been earned. E.g. M0 A1 is impossible." On an ordinary question this dependency is nearly invisible, because a correct final answer is itself strong evidence a correct method was used. On a "show that" question that evidence is destroyed on purpose: the final answer is printed in the question, so a correct final answer is evidence of nothing — it is equally consistent with genuine working and with simply copying what was already given. The mark scheme's only defence is to refuse to award the A mark unless the M mark — the actual method, shown on the page — is there first. This is also why "the answer is printed on the paper" has its own dedicated symbol in the mark schemes' general abbreviations list: it is common enough, specifically because "show that" questions are common enough, to need its own shorthand. The Statistics-specific marking note reinforces the same point from the examiner's side: "Any correct method should gain credit. If you cannot see how to apply the mark scheme but believe the method to be correct then please send to review" — the examiner is instructed to credit what they can see, not what they can infer might have happened off the page.

Traps — 4

answer-only-no-intermediate-step
The single most literal version of the trap, on a real Normal-distribution "show that" question (Oct 2021, Q6(a)): "many students wrote down the standardisation followed by 0.15 missing out the intermediate step and the accurate answer. Some students simply wrote down 0.1500… with no working at all." Two separate failures are named in that one quote — some students showed the standardisation but skipped the accurate table value in between; others showed nothing at all and just wrote the given probability. Both lose marks, because both leave the examiner unable to distinguish "did the work, wrote it up badly" from "copied the given answer."
reverse-engineered-working
A more subtle version, on a real interquartile-range "show that" question (Oct 2021, Q3(b)): "the main errors, since the answer was given, were making up two ages... which differed by 16 as their 'working'." This is worse than showing nothing, not better — it produces two numbers that are correct only in that they happen to be 16 apart, with no interpolation, no cumulative frequency, no class boundary in sight. A "working" that was clearly reverse-engineered from the given answer, rather than derived from the data, does not read as method at all.
shown-result-not-fully-stated
Confirmed on a "show that there are 3 outliers" question (Oct 2021, Q3(c)): "having gained the correct limits some did not list the 3 outliers in this part in order to show there are 3 outliers." The calculation (the outlier boundaries) was right; what was missing was the last, cheapest step — actually pointing at the three data values that satisfy it. When the question asks you to show a count, a stated boundary is not the same as a demonstrated count, and the gap between them is one line of writing.
insufficient-accuracy-in-a-show-that
A precision-specific version, on a real regression-line "show that" question (Jun 2024, Q4(c) — the actual target line was g=42.3+0.722dg = -42.3 + 0.722d, 3 s.f.): "Many students were able to show the given regression line but too often students lost the final A mark as they failed to give values to the required degree of accuracy. b = 12105.12/16769.78 = 0.722 was not accurate enough to gain the final mark due to it being a 'show that' question. Students should be encouraged in these types of questions to give answers to at least one more decimal place than the given value." The candidate's 0.722 was correct — 12105.12 ÷ 16769.78 really does round to 0.722 — but stopping the working at the same 3 significant figures as the printed line cannot prove the calculation reached it rather than having been rounded to match; the extra digit (0.7218…) is the only thing that can.

Say it out loud

Out loud, from memory, no notes: explain why the mark scheme genuinely cannot pay for the printed answer to someone who has never seen this topic — where does your explanation get vague or hand-wavy? That's the exact spot to re-study, and it only works if you check it: read back over the mechanism above the moment you finish talking and mark precisely where you drifted from it.

Spec 2.4

2 lessons

Outliers, Box Plots, and Comparing Distributions

A box plot's whiskers only ever touch real data. The fence that decides who counts as an outlier is arithmetic you compute and then throw away — it never appears as a mark on the finished plot — and one of the most consistently documented box-plot errors on this paper is drawing a whisker straight to that invisible number instead of stopping at the last real value before it. The formula for the fence isn't fixed, either: the specification insists every question state its own rule, precisely so a memorised "1.5 × IQR" can't quietly become the wrong number for the one sitting in front of you.

The card

Five-number summary: min, Q1, median, Q3, max. Box = Q1 to Q3 (the IQR); line inside = median.
IQR = Q3 − Q1 — not in the booklet, memorise it.
The outlier rule is ALWAYS given in the question. Use exactly what's stated, never a remembered formula.
Fences point outward: lower = Q1 − (k × IQR); upper = Q3 + (k × IQR). Never the reverse.
A whisker stops at the most extreme NON-outlier value — never at the fence, never at the outlier.
Mean uses every value (outliers pull it); median uses only position (outliers barely move it).
Comparing distributions: name a statistic, give both figures, state the direction. No numbers, no marks.

Why it works — Where 1.5 × IQR fences come from, and why the direction is not optional

The IQR measures the width of the middle 50% of the data — how spread out the 'ordinary' half of it is. An outlier rule takes that width as its own unit of measurement and asks: how many IQRs outside the box does a value have to sit before it stops looking like a natural continuation of the data and starts looking like something genuinely unusual? A rule like 1.5×IQR1.5 \times IQR answers that by marking a fence at 1.5 IQRs beyond EACH quartile, extending outward, away from the box. That direction isn't a convention to memorise; it's forced by what a fence is for. The lower fence has to sit below Q1Q_1, so it's Q1Q_1 minus the distance — subtracting. The upper fence has to sit above Q3Q_3, so it's Q3Q_3 plus the distance — adding. A fence built the other way round, added to Q1Q_1 or subtracted from Q3Q_3, moves INTO the box rather than away from it, and for a large enough multiplier can end up on the wrong side of the median entirely — at which point it would start flagging perfectly ordinary values near the middle of the data as outliers, exactly the opposite of what the rule exists to do. The confirmed error record shows precisely this: "there were a surprising number of errors... with some multiplying the quartiles by 1.5 and others using Q1 + 1.5×IQR for the lower limit" (Jan 2021, Q2) — both are the same underlying mistake, arithmetic that produces a real number but not the number the fence is defined to be. There's a second reason not to treat 1.5 as a fixed fact at all: the specification's own guidance on this spec item states plainly that "any rule to identify outliers will be specified in the question" — 1.5 is the multiplier used throughout this lesson because it's the one seen most often in the reviewed record, not because it's fixed. A real question can hand you a different multiplier, and when it does, that stated rule overrides whatever you remember from practice. The confirmed record shows exactly what happens when a student substitutes the remembered version anyway: "a small number of candidates opted to apply the more commonly used outlier formula of Q3 + 1.5×(Q3−Q1) rather than the one quoted in the question" (Jun 2022, Q1(b)). What has to transfer between questions is the outward-facing DIRECTION of the rule, not a specific number attached to it.

Traps — 8

outlier-lower-fence-direction-reversed
Confirmed directly, and the research bank's own phrasing makes clear this was not a rare slip: "there were a surprising number of errors... with some multiplying the quartiles by 1.5 and others using Q1 + 1.5×IQR for the lower limit" (Jan 2021, Q2). The lower fence has to SUBTRACT from Q1, moving further below it — the mechanism block above derives why, rather than asking you to remember it as a rule with no reason behind it.
whisker-drawn-to-fence-not-to-data
The research bank's own account of the same Jan 2021 Q2 finding (its documented summary of the examiner report, not a direct quotation) is that box-plot whiskers were commonly drawn out to the outlier boundary itself (98, in that series) rather than to the actual highest non-outlier value in the data (97). A fence is a number used to test values against; it's never itself a point that gets drawn. See the diagram block above for the same error, reproduced with this lesson's own numbers (74.5 versus the real value, 61).
remembered-formula-overrides-the-rule-given-in-the-question
Confirmed: "a small number of candidates opted to apply the more commonly used outlier formula of Q3 + 1.5×(Q3−Q1) rather than the one quoted in the question" (Jun 2022, Q1(b)). This matters because the specification itself states that "any rule to identify outliers will be specified in the question" (spec S1.3, item 2.4, guidance) — there is deliberately no single fixed rule to memorise, and treating 1.5×IQR as a universal constant is itself the error, even on the (common) occasions where the question happens to specify 1.5 anyway.
show-that-outliers-not-listed
Confirmed, on a "show that there are 3 outliers" question: "having gained the correct limits some did not list the 3 outliers in this part in order to show there are 3 outliers" (Oct 2021, Q3(c)). Finding the correct fences is necessary but not sufficient — a "show that N outliers exist" question needs the N values actually named, not just the machinery that would find them.
mean-assumed-more-accurate-than-the-median
Confirmed, and stated by the examiner report as a genuinely common pattern rather than a rare one: "very few candidates scored this mark... A very common misconception was that the mean is more accurate than the median. A large number of responses simply explained how to calculate a mean or said that it was because the mean is the average, showing no appreciation that the mean is just one measure of average and the median is another" (Jun 2022, Q1(e)). The credited answers were specific and short: "the mean uses all the data" / "the mean includes the outliers." Neither statistic is "more accurate" — they are different measures of the same idea, and the concept-ladder above works through exactly why an outlier moves one and not the other.
comparison-missing-supporting-figures
Confirmed: "surprisingly too many students failed to give supporting figures... a question like this will require some context..., a reference to a named statistic and supporting figures" (Jun 2024, Q1(e)). "Branch A had a higher average" names nothing and gives no numbers; "Branch A had a higher median (34 minutes against 28)" does both, and only the second shape of answer scores.
comparison-answers-the-wrong-question
Confirmed: "few comments referring to the distribution of ages were seen... others commented on the similarities [when asked for differences]" (Oct 2021, Q3(e)). Read what the question is actually asking for — differences, similarities, a specific statistic — before writing the comparison, since a technically-true observation about the wrong aspect of the data scores nothing.
stem-and-leaf-read-in-the-wrong-direction
Confirmed, on a real quartile-from-stem-and-leaf question: "some students read the stem and leaf diagram the wrong way round and so incorrectly identified Q1 = 31 and Q3 = 51" (Jun 2024, Q1(b)) — they swapped which end of the ordered leaves is the lower quartile and which is the upper. Whatever representation a box plot is built from — a stem-and-leaf diagram, a table, a raw list — check which end you're counting from before quoting a quartile out of it.

Say it out loud

Out loud, from memory, no notes: explain where 1.5 × iqr fences come from, and why the direction is not optional to someone who has never seen this topic — where does your explanation get vague or hand-wavy? That's the exact spot to re-study, and it only works if you check it: read back over the mechanism above the moment you finish talking and mark precisely where you drifted from it.

"Show that" answer discipline

A "show that" question already gives you the answer. That is not a hint about how easy the question is — it is a warning about how the mark scheme has to be built. When the final number is printed on the page, an examiner cannot give you credit for reaching it, because copying a printed number proves nothing about whether you can produce it yourself. Every mark in a "show that" question is attached to the *steps between* the start of the question and the number you were already handed — and Pearson has said, in three separate examiner reports spanning three years, that this is the single thing students most often get wrong on this paper, in every topic it touches.

The card

"Show that X" means the answer is given — the marks are for the steps to X, not for X itself.
Never write the target value first. Show the method, and let the target value appear as your last line.
Work to one more d.p./s.f. than the target throughout — matching its precision is how the last mark is lost.
"Show that there are N of something" needs the N things named, not the rule that would find them.
M0 A1 is impossible: no visible method means the accuracy marks have nothing to attach to, whatever the final line says.

Why it works — Why the mark scheme genuinely cannot pay for the printed answer

This follows directly from how M, A and B marks work, not from an arbitrary examiner preference (WST01-verified-facts.md §4; this exact "General Instructions for Marking" wording opens at least 6 of the 14 reviewed mark schemes, byte-for-byte — confirmed directly against the real Jan 2024 mark scheme PDF; Jun 2024's own preamble is shorter but keeps the same M-before-A dependency rule). An M mark is "given for a correct method or an attempt at a correct method" — it has to be visible to be given, because there is nothing else for the examiner to assess it against. An A mark is "dependent... and can only be awarded if the previous M mark has been earned. E.g. M0 A1 is impossible." On an ordinary question this dependency is nearly invisible, because a correct final answer is itself strong evidence a correct method was used. On a "show that" question that evidence is destroyed on purpose: the final answer is printed in the question, so a correct final answer is evidence of nothing — it is equally consistent with genuine working and with simply copying what was already given. The mark scheme's only defence is to refuse to award the A mark unless the M mark — the actual method, shown on the page — is there first. This is also why "the answer is printed on the paper" has its own dedicated symbol in the mark schemes' general abbreviations list: it is common enough, specifically because "show that" questions are common enough, to need its own shorthand. The Statistics-specific marking note reinforces the same point from the examiner's side: "Any correct method should gain credit. If you cannot see how to apply the mark scheme but believe the method to be correct then please send to review" — the examiner is instructed to credit what they can see, not what they can infer might have happened off the page.

Traps — 4

answer-only-no-intermediate-step
The single most literal version of the trap, on a real Normal-distribution "show that" question (Oct 2021, Q6(a)): "many students wrote down the standardisation followed by 0.15 missing out the intermediate step and the accurate answer. Some students simply wrote down 0.1500… with no working at all." Two separate failures are named in that one quote — some students showed the standardisation but skipped the accurate table value in between; others showed nothing at all and just wrote the given probability. Both lose marks, because both leave the examiner unable to distinguish "did the work, wrote it up badly" from "copied the given answer."
reverse-engineered-working
A more subtle version, on a real interquartile-range "show that" question (Oct 2021, Q3(b)): "the main errors, since the answer was given, were making up two ages... which differed by 16 as their 'working'." This is worse than showing nothing, not better — it produces two numbers that are correct only in that they happen to be 16 apart, with no interpolation, no cumulative frequency, no class boundary in sight. A "working" that was clearly reverse-engineered from the given answer, rather than derived from the data, does not read as method at all.
shown-result-not-fully-stated
Confirmed on a "show that there are 3 outliers" question (Oct 2021, Q3(c)): "having gained the correct limits some did not list the 3 outliers in this part in order to show there are 3 outliers." The calculation (the outlier boundaries) was right; what was missing was the last, cheapest step — actually pointing at the three data values that satisfy it. When the question asks you to show a count, a stated boundary is not the same as a demonstrated count, and the gap between them is one line of writing.
insufficient-accuracy-in-a-show-that
A precision-specific version, on a real regression-line "show that" question (Jun 2024, Q4(c) — the actual target line was g=42.3+0.722dg = -42.3 + 0.722d, 3 s.f.): "Many students were able to show the given regression line but too often students lost the final A mark as they failed to give values to the required degree of accuracy. b = 12105.12/16769.78 = 0.722 was not accurate enough to gain the final mark due to it being a 'show that' question. Students should be encouraged in these types of questions to give answers to at least one more decimal place than the given value." The candidate's 0.722 was correct — 12105.12 ÷ 16769.78 really does round to 0.722 — but stopping the working at the same 3 significant figures as the printed line cannot prove the calculation reached it rather than having been rounded to match; the extra digit (0.7218…) is the only thing that can.

Say it out loud

Out loud, from memory, no notes: explain why the mark scheme genuinely cannot pay for the printed answer to someone who has never seen this topic — where does your explanation get vague or hand-wavy? That's the exact spot to re-study, and it only works if you check it: read back over the mechanism above the moment you finish talking and mark precisely where you drifted from it.

Spec 3.1

1 lesson

Elementary probability, and conditional-probability notation

P(AB)P(A \mid B) is not P(A)P(A) divided by P(B)P(B) — the bar means something has already happened, and 'something has already happened' means the sample space itself just got smaller. Every genuine sample-space question, at bottom, is a counting question: list the possibilities carefully, count the ones you want, divide by the ones that are possible. Conditional notation asks you to do exactly that counting inside a SHRUNK sample space — restricted to whatever's on the right of the bar — and the single most repeatable way to get it wrong, confirmed on real WST01 papers, is skipping that restriction and dividing two numbers that were never counted from the same space to begin with.

The card

Ω = the full sample space; an event is a subset of it. Classical probability: P(A) = n(A)/n(Ω), for equally likely outcomes.
Combining independent stages: sizes MULTIPLY (4×4=16), not add. Multi-stage outcomes are usually ORDERED — (2,6) ≠ (6,2).
Words → notation: 'or' → ∪, 'and'/'both' → ∩, 'not' → ′, 'given that' → | (the event after 'given' goes on the RIGHT of the bar).
On the formula sheet: P(A∪B) = P(A)+P(B)−P(A∩B). Not on it: P(A′) = 1−P(A) — memorise it, one line.
P(A|B) = P(A∩B)/P(B) = n(A∩B)/n(B) for equally likely outcomes. NEVER P(A)/P(B) — that drops the intersection entirely.
P(A)/P(B) only ever equals P(A|B) when A is entirely contained in B. Otherwise it strictly overshoots — check: is the answer > 1?

Why it works — Why P(A|B) is never P(A) divided by P(B) — and the one case where they happen to agree

Write both calculations out in raw counts and the difference stops being a rule to memorise and becomes something you can see directly. P(AB)=n(AB)n(B)P(A \mid B) = \dfrac{n(A \cap B)}{n(B)} — the outcomes shared by AA and BB, out of BB's own count. The raw-ratio shortcut, by contrast, is P(A)P(B)=n(A)/n(Ω)n(B)/n(Ω)=n(A)n(B)\dfrac{P(A)}{P(B)} = \dfrac{n(A)/n(\Omega)}{n(B)/n(\Omega)} = \dfrac{n(A)}{n(B)}AA's ENTIRE count, out of BB's own count, with no reference anywhere to whether AA and BB actually overlap. These agree only when AB=AA \cap B = A — that is, when every single outcome in AA also happens to lie inside BB (formally, ABA \subseteq B). That's a real but narrow special case; the moment any part of AA sits OUTSIDE BB, n(AB)<n(A)n(A \cap B) < n(A), and the raw ratio overshoots. In fact this overshoot direction is guaranteed, not just typical: since ABA \cap B is always a subset of AA itself, n(AB)n(A)n(A \cap B) \le n(A) always holds, which means P(AB)P(A)P(B)P(A \mid B) \le \dfrac{P(A)}{P(B)} is true for EVERY pair of events, with equality exactly in that one subset case. The raw-ratio shortcut can never underestimate the true conditional probability — it can only match it, in that one narrow case, or inflate it, which is exactly why it so often produces a value that's suspiciously large, or that exceeds 1 outright and announces itself as impossible.

Traps — 2

conditional-probability-as-raw-ratio-of-marginals
Verified directly, quoted from a real WST01 conditional-probability question: "we occasionally still saw P(0.35∩0.4)/0.4 or 0.35/0.4" (Jan 2021, Q1(c)) — dividing two given probabilities directly, as though P(A|B) meant P(A) ÷ P(B) rather than P(A ∩ B) ÷ P(B). The mechanism block above proves this is never a harmless shortcut: it can only match the correct value (when A is entirely contained in B) or overshoot it — never undershoot — which is exactly why it so often lands above 1, or on a plausible-looking but wrong number just below it.
conditioning-probability-refolded-into-the-numerator
A second, verified instance of conditional-probability notation going wrong, checked directly against the real Pearson examiner-report PDF rather than taken on trust from a summary of it (Jan 2023, Q2(d), a tree-diagram question): candidates were finding P(A|B) from a tree diagram where the correct denominator, P(B) = 61/234, had already been found in an earlier part, and the correct numerator was the single branch product 5/9 × 4/8 × 8/13 = 20/117. The two malformed answers the examiner report actually quotes both use that 61/234 CORRECTLY as the denominator — the error is entirely in the numerator, where an extra copy of the same 61/234 gets folded in: one wrote (5/9 × 4/8 × 8/13 + 61/234)/(61/234), another wrote (5/9 × 4/8 × 8/13 × 61/234)/(61/234). Different arithmetic from the raw-ratio trap above, but the same root confusion: not treating P(A∩B) as one clean, self-contained quantity, separate from whatever's about to divide it — instead letting the denominator's own value leak back into the numerator.

Say it out loud

Out loud, from memory, no notes: explain why p(a|b) is never p(a) divided by p(b) — and the one case where they happen to agree to someone who has never seen this topic — where does your explanation get vague or hand-wavy? That's the exact spot to re-study, and it only works if you check it: read back over the mechanism above the moment you finish talking and mark precisely where you drifted from it.

Spec 3.2

3 lessons

Elementary probability, and conditional-probability notation

P(AB)P(A \mid B) is not P(A)P(A) divided by P(B)P(B) — the bar means something has already happened, and 'something has already happened' means the sample space itself just got smaller. Every genuine sample-space question, at bottom, is a counting question: list the possibilities carefully, count the ones you want, divide by the ones that are possible. Conditional notation asks you to do exactly that counting inside a SHRUNK sample space — restricted to whatever's on the right of the bar — and the single most repeatable way to get it wrong, confirmed on real WST01 papers, is skipping that restriction and dividing two numbers that were never counted from the same space to begin with.

The card

Ω = the full sample space; an event is a subset of it. Classical probability: P(A) = n(A)/n(Ω), for equally likely outcomes.
Combining independent stages: sizes MULTIPLY (4×4=16), not add. Multi-stage outcomes are usually ORDERED — (2,6) ≠ (6,2).
Words → notation: 'or' → ∪, 'and'/'both' → ∩, 'not' → ′, 'given that' → | (the event after 'given' goes on the RIGHT of the bar).
On the formula sheet: P(A∪B) = P(A)+P(B)−P(A∩B). Not on it: P(A′) = 1−P(A) — memorise it, one line.
P(A|B) = P(A∩B)/P(B) = n(A∩B)/n(B) for equally likely outcomes. NEVER P(A)/P(B) — that drops the intersection entirely.
P(A)/P(B) only ever equals P(A|B) when A is entirely contained in B. Otherwise it strictly overshoots — check: is the answer > 1?

Why it works — Why P(A|B) is never P(A) divided by P(B) — and the one case where they happen to agree

Write both calculations out in raw counts and the difference stops being a rule to memorise and becomes something you can see directly. P(AB)=n(AB)n(B)P(A \mid B) = \dfrac{n(A \cap B)}{n(B)} — the outcomes shared by AA and BB, out of BB's own count. The raw-ratio shortcut, by contrast, is P(A)P(B)=n(A)/n(Ω)n(B)/n(Ω)=n(A)n(B)\dfrac{P(A)}{P(B)} = \dfrac{n(A)/n(\Omega)}{n(B)/n(\Omega)} = \dfrac{n(A)}{n(B)}AA's ENTIRE count, out of BB's own count, with no reference anywhere to whether AA and BB actually overlap. These agree only when AB=AA \cap B = A — that is, when every single outcome in AA also happens to lie inside BB (formally, ABA \subseteq B). That's a real but narrow special case; the moment any part of AA sits OUTSIDE BB, n(AB)<n(A)n(A \cap B) < n(A), and the raw ratio overshoots. In fact this overshoot direction is guaranteed, not just typical: since ABA \cap B is always a subset of AA itself, n(AB)n(A)n(A \cap B) \le n(A) always holds, which means P(AB)P(A)P(B)P(A \mid B) \le \dfrac{P(A)}{P(B)} is true for EVERY pair of events, with equality exactly in that one subset case. The raw-ratio shortcut can never underestimate the true conditional probability — it can only match it, in that one narrow case, or inflate it, which is exactly why it so often produces a value that's suspiciously large, or that exceeds 1 outright and announces itself as impossible.

Traps — 2

conditional-probability-as-raw-ratio-of-marginals
Verified directly, quoted from a real WST01 conditional-probability question: "we occasionally still saw P(0.35∩0.4)/0.4 or 0.35/0.4" (Jan 2021, Q1(c)) — dividing two given probabilities directly, as though P(A|B) meant P(A) ÷ P(B) rather than P(A ∩ B) ÷ P(B). The mechanism block above proves this is never a harmless shortcut: it can only match the correct value (when A is entirely contained in B) or overshoot it — never undershoot — which is exactly why it so often lands above 1, or on a plausible-looking but wrong number just below it.
conditioning-probability-refolded-into-the-numerator
A second, verified instance of conditional-probability notation going wrong, checked directly against the real Pearson examiner-report PDF rather than taken on trust from a summary of it (Jan 2023, Q2(d), a tree-diagram question): candidates were finding P(A|B) from a tree diagram where the correct denominator, P(B) = 61/234, had already been found in an earlier part, and the correct numerator was the single branch product 5/9 × 4/8 × 8/13 = 20/117. The two malformed answers the examiner report actually quotes both use that 61/234 CORRECTLY as the denominator — the error is entirely in the numerator, where an extra copy of the same 61/234 gets folded in: one wrote (5/9 × 4/8 × 8/13 + 61/234)/(61/234), another wrote (5/9 × 4/8 × 8/13 × 61/234)/(61/234). Different arithmetic from the raw-ratio trap above, but the same root confusion: not treating P(A∩B) as one clean, self-contained quantity, separate from whatever's about to divide it — instead letting the denominator's own value leak back into the numerator.

Say it out loud

Out loud, from memory, no notes: explain why p(a|b) is never p(a) divided by p(b) — and the one case where they happen to agree to someone who has never seen this topic — where does your explanation get vague or hand-wavy? That's the exact spot to re-study, and it only works if you check it: read back over the mechanism above the moment you finish talking and mark precisely where you drifted from it.

Conditional probability, and independence vs. mutually exclusive

Two events overlapping tells you nothing about whether they are independent — and two events being independent tells you they cannot be mutually exclusive. All five WST01 examiner reports read for this course flag a version of the same confusion, in a different disguise each time: candidates assume independence instead of computing P(AB)=P(AB)/P(B)P(A \mid B) = P(A \cap B) / P(B), or they judge independence by eye — *"they are not independent as they overlap"* — instead of by the one calculation that actually settles it. Both errors come from treating two genuinely different questions (do these events share any outcomes? does knowing one change the probability of the other?) as though they were the same question.

The card

P(A|B) = P(A∩B)/P(B), requires P(B) > 0. Conditioning restricts the sample space to B.
Independent (memorise, not on the sheet): P(A|B)=P(A), P(B|A)=P(B), P(A∩B)=P(A)P(B) — any one implies the others.
Mutually exclusive: A∩B=∅, so P(A∩B)=0 and P(A∪B)=P(A)+P(B).
If P(A)>0 and P(B)>0: mutually exclusive and independent can never both hold. Overlap is necessary for independence but never sufficient — always run the product-rule check.
On the formula sheet: P(A∪B)=P(A)+P(B)−P(A∩B), and P(A∩B)=P(A)P(B|A). Not on it: P(A′)=1−P(A) and all three independence rules.

Why it works — Why mutually exclusive events (with nonzero probability) can never be independent — and why 'they overlap' isn't the whole story either

Run the two definitions against each other directly. If AA and BB are mutually exclusive, P(AB)=0P(A \cap B) = 0 — that's what mutually exclusive means. If AA and BB are independent, P(AB)=P(A)×P(B)P(A \cap B) = P(A) \times P(B) — that's what independent means. If both held at once, P(A)×P(B)P(A) \times P(B) would have to equal 00. But a product of two numbers is zero only if at least one of them is zero — and if P(A)>0P(A) > 0 and P(B)>0P(B) > 0, that's impossible. So for any two events that both have positive probability, mutually exclusive and independent are not just different — they are mutually incompatible: proving one is true is a proof the other is false. This is the precise version of the pattern an examiner report names directly, calling it chronic rather than occasional: 'there was the usual confusion between events being independent and events being mutually exclusive highlighted by statements such as "they are not independent as they overlap"' (Oct 2021, Q1(b)) — the word 'usual' there is Pearson's own, describing a recurring pattern rather than a one-off. But read that quoted student statement carefully, because the mechanism above actually cuts partway in the student's favour: NOT overlapping (being mutually exclusive) genuinely does rule out independence, whenever both probabilities are positive — that half of the reasoning is sound. Where it breaks is the unstated assumption running the other way: that overlapping is enough, on its own, to conclude independence. It isn't. Overlapping is necessary for independence but nowhere near sufficient, and the dice example shows exactly how a small change in the numbers exposes the gap: let BB' = 'the two dice sum to 8' instead of 7. AA = 'first die is 6' and BB' certainly overlap — 6-then-2 is a valid outcome in both. P(B)=536P(B') = \frac{5}{36} (five ways to make 8: 2+6, 3+5, 4+4, 5+3, 6+2), and P(AB)=136P(A \cap B') = \frac{1}{36} (only 6-then-2). Check independence: P(A)×P(B)=16×536=5216P(A) \times P(B') = \frac{1}{6} \times \frac{5}{36} = \frac{5}{216}, but P(AB)=136=6216P(A \cap B') = \frac{1}{36} = \frac{6}{216}. 62165216\frac{6}{216} \neq \frac{5}{216} — not independent, despite genuinely overlapping. So the full picture has three regions, not two: mutually exclusive events (with both probabilities positive) are never independent; independent events are never mutually exclusive; but 'overlaps and isn't mutually exclusive' is a large middle ground that still needs the actual product-rule check, because it contains both independent pairs (first-die-6 and sum-is-7) and dependent ones (first-die-6 and sum-is-8) side by side, indistinguishable without calculating.

Traps — 5

independence-and-mutual-exclusivity-conflated
The chronic, named error across the whole archive: "there was the usual confusion between events being independent and events being mutually exclusive highlighted by statements such as 'they are not independent as they overlap'" (Oct 2021, Q1(b)). Pearson's own word "usual" marks this as a recurring pattern, not a single script's slip. The mechanism block above shows exactly where the reasoning goes wrong: NOT overlapping does rule out independence (when both probabilities are positive) — but overlapping on its own proves nothing about independence, which needs the actual P(A∩B) = P(A)P(B) check.
conditional-probability-as-raw-ratio-of-marginals
Verified directly, quoted from a real conditional-probability question: "we occasionally still saw P(0.35∩0.4)/0.4 or 0.35/0.4" (Jan 2021, Q1(c)) — dividing two given probabilities directly, as though P(A|B) meant P(A) ÷ P(B) rather than P(A∩B) ÷ P(B). A fast self-check: if the "conditional probability" you've computed by dividing two marginals exceeds 1, you've made exactly this error — a real probability never can.
independence-assumed-instead-of-tested
"Candidates often assume independence when an appropriate conditional probability should be used instead" (Jun 2022, general examiner comment — stated as a paper-wide diagnosis, not tied to one question). A second, more specific instance on a Venn-diagram conditional-probability question: "common errors were to assume independence or to do 1 − 1/15 before dividing by 3/8" (Jun 2022, Q4(b)) — independence substituted in as a shortcut for the actual conditioning calculation the question required.
tree-diagram-numerator-missing-a-branch
"the most common error was using a product of 2 probabilities rather than 3 in the numerator" (Oct 2021, Q4(c)), on a three-branch tree-diagram conditional probability question. The worked-chain example above is built around this exact trap: forgetting one branch of a multi-stage path produces a wrong but perfectly plausible-looking fraction, with no obvious sign anything went wrong.
extra-probability-folded-into-the-numerator
A second, independently confirmed instance of the numerator/denominator confusion (Jan 2023, Q2(d)) — corrected 2026-08-28 against the real examiner-report PDF, since an earlier reading of this citation had the mechanism backwards: the conditioning denominator itself (61/234, carried from an earlier part) was correct in both wrong answers candidates gave. The actual slip was in the numerator — an extra, spurious copy of that same 61/234 got folded in, either added or multiplied: (5/9 × 4/8 × 8/13 + 61/234)/(61/234) and (5/9 × 4/8 × 8/13 × 61/234)/(61/234), instead of the correct numerator 5/9 × 4/8 × 8/13 = 20/117. The lesson here: once a probability from an earlier part is sitting on the page, it's tempting to reuse it a second time inside a new calculation where it doesn't belong — write P(A|B) = P(A∩B)/P(B) out in full first, work out P(A∩B) as its own clean fraction, and only then substitute, so there's a formula on the page to check the substitution against.

Say it out loud

Out loud, from memory, no notes: explain why mutually exclusive events (with nonzero probability) can never be independent — and why 'they overlap' isn't the whole story either to someone who has never seen this topic — where does your explanation get vague or hand-wavy? That's the exact spot to re-study, and it only works if you check it: read back over the mechanism above the moment you finish talking and mark precisely where you drifted from it.

The Normal distribution — standardisation, table precision, and conditional probability

Every one of the five examiner reports checked for this course treats one skill on this topic as the most discriminating part of the paper — and it is not really a Normal-distribution skill at all. It is whether you notice, underneath a completely ordinary standardisation, that the question has quietly become a conditional probability — and by the time you are three lines into the wrong calculation, the diagram that would have shown you is nowhere on the page.

The card

Z = (X − μ)/σ. Standardise first, always — not in the booklet, memorise it.
Φ(z) = P(Z < z), read straight from the table. P(Z > z) = 1 − Φ(z) — sketch before you subtract.
Percentage Points table: inverse lookup, gives z to 4 d.p. for a stated tail probability. Quote in full — 1.6449, never 1.64.
Conditional probability with Normal: same rule as always, P(A|B) = P(A∩B)/P(B). If A sits inside B, P(A∩B) = P(A).
'Show that' a probability: the standardisation must appear on the page, not just the final decimal.

Why it works — Why 'greater than' always costs you a subtraction — and why it fails in both directions

The Normal Distribution Function table stores exactly one kind of number, on every page: Φ(z) = P(Z < z), the area under the standard Normal curve to the LEFT of z. It does not store P(Z > z) anywhere — because it does not need to. The total area under the whole curve is exactly 1 (one of the two shape facts spec 6.1 explicitly expects you to know, not derive), so whatever fraction of that area sits to the left of z, everything else — the entire right-hand region — is 1 minus that fraction: P(Z > z) = 1 − Φ(z). That single subtraction is the single most consistently mis-fired step across the whole of this topic, confirmed in the facts bank failing in BOTH directions across three separate series. Under-subtracting: a January 2023 examiner report records that 'a significant number of students lost 2 marks as they failed to subtract from one the value obtained from the normal tables.' Over-subtracting: a June 2024 report records candidates who 'went on to subtract the correct answer from 1, which of course is P(X > 18) and not P(X < 18) which is what was required.' Both failures share one root cause: reaching for the table before deciding, from the question's own wording, which side of z is actually being asked about. Pearson's own fix, stated explicitly in the same January 2023 report: 'a simple diagram would have helped many to avoid this error.'

Traps — 5

subtract-from-1-direction-confusion
Confirmed in three separate series, failing in BOTH directions — the same underlying gap producing opposite mistakes depending on the question. Under-subtracting: "a significant number of students lost 2 marks as they failed to subtract from one the value obtained from the normal tables. A simple diagram would have helped many to avoid this error" (Jan 2023, Q5(a)). Over-subtracting: "a few lost the final mark as they went on to subtract the correct answer from 1, which of course is P(X > 18) and not P(X < 18) which is what was required" (Jun 2024, Q5(a)). A third series confirms the skill is fragile even when it goes right: "most students standardising correctly and the majority realising that they then needed to subtract the value found in the tables from 1" (Jan 2021, Q3(a)) — implying, in Pearson's own words, that a real minority did not. The fix in every case is the same: sketch which side of z the question describes before opening the table at all.
rounded-z-value-instead-of-4dp-table-value
The exact same mechanism, confirmed in three independent series, each with its own quoted wrong value: "many students were using a z value of 1.03 or 1.04 rather than the value 1.0364 from the 'Percentage Points of the Normal Distribution' table" (Jan 2021, Q3(b)); "not using an inaccurate value such as 1.64 instead of 1.6449" (Oct 2021, Q6(b)); "the most common error included the use of an inaccurate z value... students should be reminded that when values are required from the tables, they need to be 4 decimal places. A common error was to use z value = 0.25" (Jun 2024, Q5(b)). This is not carelessness in isolation — it is exam technique stated verbatim on every WST01 question-paper front page: "Values from the statistical tables should be quoted in full" (verified, facts bank §2).
sign-error-reversing-standardisation
Confirmed in two series, once in each direction of the sign: "others gained this mark but were using −1.0364, an error that could probably have been avoided if they had drawn a suitable diagram" (Jan 2021, Q3(b)) — where the value needed was positive; "using the wrong sign for 1.6449 appropriate to their standardisation giving 34.1, a value higher than the upper limit" (Oct 2021, Q6(b)) — where an unflipped sign produced an answer that was, on inspection, impossible. That second detail is the real lesson: the wrong answer was checkable as wrong using nothing but the question's own numbers, and the check was skipped anyway.
conditional-probability-not-recognised-with-normal
The most consistently documented trap in the whole facts bank — present, in a different concrete form, in every one of the 5 examiner reports reviewed. "Many did not realise that a conditional probability was required... a common error P(W<18)/0.85... but there were a number of correct attempts of the form (0.85−0.5)/0.85 which usually led to the correct answer" (Jan 2021, Q3(c)). "Many students did not realise that the ratios only applied to the middle 80% of the data" (Oct 2021, Q6(c)). "Like question 2, a common error was that students failed to realise that a conditional probability was required. A common error was to find P(L≤5) and go no further" (Jan 2023, Q5(e)). Stated as a paper-wide diagnosis, not a topic-specific aside: "Candidates often assume independence when an appropriate conditional probability should be used instead" (Jun 2022, general comment). One series' report names this sub-skill, in these words, as the hardest thing on the entire paper: a Jun 2022 question is flagged as "the final part of the paper... also the most discriminating part."
show-that-standardisation-not-shown
A cross-cutting instruction-following failure, not a maths error, confirmed directly on a Normal-distribution question: "many students wrote down the standardisation followed by 0.15 missing out the intermediate step and the accurate answer. Some students simply wrote down 0.1500… with no working at all" (Oct 2021, Q6(a)). Stated as a general rule in a later series: "if asked to use standardisation then the standardisation should be shown" (Jun 2024, general introduction). This costs marks even when the final decimal is completely correct — a 'show that' mark scheme has nothing to attach a mark to if the standardisation step itself never appears on the page.

Say it out loud

Out loud, from memory, no notes: explain why 'greater than' always costs you a subtraction — and why it fails in both directions to someone who has never seen this topic — where does your explanation get vague or hand-wavy? That's the exact spot to re-study, and it only works if you check it: read back over the mechanism above the moment you finish talking and mark precisely where you drifted from it.

Spec 3.3

1 lesson

Conditional probability, and independence vs. mutually exclusive

Two events overlapping tells you nothing about whether they are independent — and two events being independent tells you they cannot be mutually exclusive. All five WST01 examiner reports read for this course flag a version of the same confusion, in a different disguise each time: candidates assume independence instead of computing P(AB)=P(AB)/P(B)P(A \mid B) = P(A \cap B) / P(B), or they judge independence by eye — *"they are not independent as they overlap"* — instead of by the one calculation that actually settles it. Both errors come from treating two genuinely different questions (do these events share any outcomes? does knowing one change the probability of the other?) as though they were the same question.

The card

P(A|B) = P(A∩B)/P(B), requires P(B) > 0. Conditioning restricts the sample space to B.
Independent (memorise, not on the sheet): P(A|B)=P(A), P(B|A)=P(B), P(A∩B)=P(A)P(B) — any one implies the others.
Mutually exclusive: A∩B=∅, so P(A∩B)=0 and P(A∪B)=P(A)+P(B).
If P(A)>0 and P(B)>0: mutually exclusive and independent can never both hold. Overlap is necessary for independence but never sufficient — always run the product-rule check.
On the formula sheet: P(A∪B)=P(A)+P(B)−P(A∩B), and P(A∩B)=P(A)P(B|A). Not on it: P(A′)=1−P(A) and all three independence rules.

Why it works — Why mutually exclusive events (with nonzero probability) can never be independent — and why 'they overlap' isn't the whole story either

Run the two definitions against each other directly. If AA and BB are mutually exclusive, P(AB)=0P(A \cap B) = 0 — that's what mutually exclusive means. If AA and BB are independent, P(AB)=P(A)×P(B)P(A \cap B) = P(A) \times P(B) — that's what independent means. If both held at once, P(A)×P(B)P(A) \times P(B) would have to equal 00. But a product of two numbers is zero only if at least one of them is zero — and if P(A)>0P(A) > 0 and P(B)>0P(B) > 0, that's impossible. So for any two events that both have positive probability, mutually exclusive and independent are not just different — they are mutually incompatible: proving one is true is a proof the other is false. This is the precise version of the pattern an examiner report names directly, calling it chronic rather than occasional: 'there was the usual confusion between events being independent and events being mutually exclusive highlighted by statements such as "they are not independent as they overlap"' (Oct 2021, Q1(b)) — the word 'usual' there is Pearson's own, describing a recurring pattern rather than a one-off. But read that quoted student statement carefully, because the mechanism above actually cuts partway in the student's favour: NOT overlapping (being mutually exclusive) genuinely does rule out independence, whenever both probabilities are positive — that half of the reasoning is sound. Where it breaks is the unstated assumption running the other way: that overlapping is enough, on its own, to conclude independence. It isn't. Overlapping is necessary for independence but nowhere near sufficient, and the dice example shows exactly how a small change in the numbers exposes the gap: let BB' = 'the two dice sum to 8' instead of 7. AA = 'first die is 6' and BB' certainly overlap — 6-then-2 is a valid outcome in both. P(B)=536P(B') = \frac{5}{36} (five ways to make 8: 2+6, 3+5, 4+4, 5+3, 6+2), and P(AB)=136P(A \cap B') = \frac{1}{36} (only 6-then-2). Check independence: P(A)×P(B)=16×536=5216P(A) \times P(B') = \frac{1}{6} \times \frac{5}{36} = \frac{5}{216}, but P(AB)=136=6216P(A \cap B') = \frac{1}{36} = \frac{6}{216}. 62165216\frac{6}{216} \neq \frac{5}{216} — not independent, despite genuinely overlapping. So the full picture has three regions, not two: mutually exclusive events (with both probabilities positive) are never independent; independent events are never mutually exclusive; but 'overlaps and isn't mutually exclusive' is a large middle ground that still needs the actual product-rule check, because it contains both independent pairs (first-die-6 and sum-is-7) and dependent ones (first-die-6 and sum-is-8) side by side, indistinguishable without calculating.

Traps — 5

independence-and-mutual-exclusivity-conflated
The chronic, named error across the whole archive: "there was the usual confusion between events being independent and events being mutually exclusive highlighted by statements such as 'they are not independent as they overlap'" (Oct 2021, Q1(b)). Pearson's own word "usual" marks this as a recurring pattern, not a single script's slip. The mechanism block above shows exactly where the reasoning goes wrong: NOT overlapping does rule out independence (when both probabilities are positive) — but overlapping on its own proves nothing about independence, which needs the actual P(A∩B) = P(A)P(B) check.
conditional-probability-as-raw-ratio-of-marginals
Verified directly, quoted from a real conditional-probability question: "we occasionally still saw P(0.35∩0.4)/0.4 or 0.35/0.4" (Jan 2021, Q1(c)) — dividing two given probabilities directly, as though P(A|B) meant P(A) ÷ P(B) rather than P(A∩B) ÷ P(B). A fast self-check: if the "conditional probability" you've computed by dividing two marginals exceeds 1, you've made exactly this error — a real probability never can.
independence-assumed-instead-of-tested
"Candidates often assume independence when an appropriate conditional probability should be used instead" (Jun 2022, general examiner comment — stated as a paper-wide diagnosis, not tied to one question). A second, more specific instance on a Venn-diagram conditional-probability question: "common errors were to assume independence or to do 1 − 1/15 before dividing by 3/8" (Jun 2022, Q4(b)) — independence substituted in as a shortcut for the actual conditioning calculation the question required.
tree-diagram-numerator-missing-a-branch
"the most common error was using a product of 2 probabilities rather than 3 in the numerator" (Oct 2021, Q4(c)), on a three-branch tree-diagram conditional probability question. The worked-chain example above is built around this exact trap: forgetting one branch of a multi-stage path produces a wrong but perfectly plausible-looking fraction, with no obvious sign anything went wrong.
extra-probability-folded-into-the-numerator
A second, independently confirmed instance of the numerator/denominator confusion (Jan 2023, Q2(d)) — corrected 2026-08-28 against the real examiner-report PDF, since an earlier reading of this citation had the mechanism backwards: the conditioning denominator itself (61/234, carried from an earlier part) was correct in both wrong answers candidates gave. The actual slip was in the numerator — an extra, spurious copy of that same 61/234 got folded in, either added or multiplied: (5/9 × 4/8 × 8/13 + 61/234)/(61/234) and (5/9 × 4/8 × 8/13 × 61/234)/(61/234), instead of the correct numerator 5/9 × 4/8 × 8/13 = 20/117. The lesson here: once a probability from an earlier part is sitting on the page, it's tempting to reuse it a second time inside a new calculation where it doesn't belong — write P(A|B) = P(A∩B)/P(B) out in full first, work out P(A∩B) as its own clean fraction, and only then substitute, so there's a formula on the page to check the substitution against.

Say it out loud

Out loud, from memory, no notes: explain why mutually exclusive events (with nonzero probability) can never be independent — and why 'they overlap' isn't the whole story either to someone who has never seen this topic — where does your explanation get vague or hand-wavy? That's the exact spot to re-study, and it only works if you check it: read back over the mechanism above the moment you finish talking and mark precisely where you drifted from it.

Spec 3.4

1 lesson

Sampling with/without replacement, tree diagrams, and Venn diagrams

A tree diagram for sampling without replacement is not one bag drawn from twice — it is two different bags, drawn from once each. The first draw changes what is left, so every branch after it describes a smaller, different sample space than the one before it. The single most-repeated way to lose marks on this topic, confirmed directly from a real question, is writing the first bag's fraction onto the second bag's branch, as though nothing had been removed at all.

The card

Product law along a path: multiply the branches. P(A∩B) = P(A) × P(B|A).
Sum law across paths: add complete, mutually exclusive paths. Branches from one node always sum to 1.
Without replacement: total AND the count just removed both fall, every single draw. With replacement: nothing changes between draws.
Venn diagram: 4 regions (A only, B only, both, neither) sum to 1 (or to the total). Never leave a region blank — write 0 if it is 0.
P(A|B) = P(A∩B) / P(B). Overlap ⇒ not mutually exclusive. Independence is P(A)P(B) = P(A∩B) — a separate check, never assumed.

Why it works — Why branches multiply and paths add — a tree diagram is the product and sum laws, drawn

Rearranging the conditional-probability formula from spec 3.2, P(AB)=P(A)P(BA)P(A \cap B) = P(A)P(B \mid A), already says exactly what a tree diagram draws: the probability of following two branches in sequence — first AA, then BB — is the first branch's probability multiplied by the second branch's probability, where the second is conditioned on the first having happened. That is the product law, and a tree diagram does not add anything new to it; it is a picture of applying it once per draw. For a two-draw tree, the probability of reaching any single end-point is the product of every fraction on the path that leads there — three draws multiply three fractions, four multiply four, and so on, which is exactly why forgetting to update even one of them (the without-replacement trap above) corrupts the whole path, not just the branch where the mistake happened. The sum law answers a different question: not "what is the probability of this one path", but "what is the probability of this event, which several different paths could produce". Two different complete paths through a tree — say, black-then-white and white-then-black — cannot both happen on the same pair of draws, so they are mutually exclusive by construction, whatever the actual probabilities turn out to be. That is what licenses P(event)=P(path 1)+P(path 2)+P(\text{event}) = P(\text{path 1}) + P(\text{path 2}) + \ldots with no overlap term to subtract: add the probability of every complete path that counts as a "success", and nothing is double-counted, because no single pair of draws can walk two different paths at once. This is also the logic behind the branches-sum-to-1 rule above: the full set of paths through a tree is mutually exclusive and exhaustive, so summing every end-point's probability has to return exactly 1. Real WST01 mark schemes reward exactly this structure. The general marking guidance printed on every one of the 14 series' mark schemes reviewed for this course defines an 'M' mark as one "given for a correct method or an attempt at a correct method" — here, setting up the right multiplication or the right sum — and defines an 'A' mark as a dependent accuracy mark, awarded only once the method mark behind it has been earned. A note specific to Statistics papers adds a further rule: "For method marks, we generally allow or condone a slip or transcription error if these are seen in an expression. We do not, however, condone or allow these errors in accuracy marks." In practice, that means the structure of a tree-diagram calculation — which branches were multiplied, which paths were summed — is worth setting out visibly even under time pressure, because that structure is where the more forgiving mark lives.

Traps — 5

original-fraction-reused-after-removal
Confirmed directly on a real without-replacement counters question: "common errors seen were to repeat the 4/8 (given in the question) following the 1st was green onto the branches following the 1st was blue" (Jan 2023, Q2). The mechanism is not carelessness with arithmetic — the fraction 48\frac{4}{8} is not wrong in itself, it is simply the wrong bag's fraction, carried onto branches that describe a bag with one fewer counter in it. The fix is procedural, not conceptual: rebuild the composition of what remains before writing a single second-draw fraction, every time.
second-draw-branches-swapped
Confirmed directly on the same real question, immediately after the reused-fraction error above: "on the second branches 5/8 and 3/8 were sometimes given the wrong way round" (Jan 2023, Q2). This is a different mechanism from the reused-fraction error above, not a restatement of it: the bag was correctly reduced to 8 counters, so the right pair of fractions was on the page — but the two fractions were then swapped onto the wrong colour's branch, as if the reducing had been done correctly and then mis-filed. A tree diagram that has been rebuilt with the right numbers can still be wrong if those numbers are not checked against which branch they actually belong to.
denominator-kept-constant-across-repeated-draws
Confirmed on a real without-replacement counting question spanning four draws: "the most common response was scoring 1 mark for the special case using (62/88)⁴" (Jun 2022, Q3(d)) — raising a single fraction to the fourth power, which keeps the denominator fixed at 88 the whole way through, exactly as though every counter were replaced before the next draw. Without replacement, a bag that loses one counter per draw needs a different denominator — and often a different numerator — at every single stage; four genuine draws without replacement need four genuinely different fractions multiplied together, never the same one raised to a power.
venn-blank-region-assumed-zero
Confirmed directly on a real Venn-diagram question: "some candidates left out the 0 in the outside region of N and should be reminded that blank spaces are not assumed to be 0s in Venn diagrams" (Jun 2022, Q4(c)). Even when the correct value for a region genuinely is zero, it has to be written there — an empty space on the diagram reads as unfinished working, not as an answer.
independence-assumed-instead-of-tested
A real examiner report's own general, paper-wide comment states plainly: "Candidates often assume independence when an appropriate conditional probability should be used instead" — named as a pattern across the whole paper, not tied to one question. The same series' report names the concrete version of it on a Venn-diagram question: "common errors were to assume independence" (Jun 2022, Q4(b)). Independence is a conclusion a calculation reaches — comparing P(AB)P(A \cap B) with P(A)×P(B)P(A) \times P(B) — never a shortcut taken on the way to one; see the common wrong path attached to the marked Venn-diagram solution above for exactly what this costs.

Say it out loud

Out loud, from memory, no notes: explain why branches multiply and paths add — a tree diagram is the product and sum laws, drawn to someone who has never seen this topic — where does your explanation get vague or hand-wavy? That's the exact spot to re-study, and it only works if you check it: read back over the mechanism above the moment you finish talking and mark precisely where you drifted from it.

Spec 4.3

1 lesson

Correlation coefficient and regression — the calculation mechanics

This is the thinnest evidence base of any WST01 lesson in this batch, and it's worth saying plainly rather than dressing it up. Where the regression-interpretation lesson draws on five separate examiner reports, spec item 4.3 — the product moment correlation coefficient itself — is anchored by exactly two sources: one examiner report and one mark scheme. What those two sources hand over is precise and genuinely useful all the same: a real, evidenced trap on the r-calculation itself — an omitted square root in the denominator — a well-documented Pearson accuracy discipline for 'show that' questions this lesson extends to r's own square-root division, and a clean, quotable coding-invariance property Pearson credits outright. This lesson assumes the Sxx/Sxy machinery already taught elsewhere and adds exactly what that machinery didn't need — a third summary sum, and the coefficient it unlocks.

The card

PMCC: r = Sxy / √(Sxx × Syy). Syy = Σy² − (Σy)²/n — same shape as Sxx, with y in place of x.
r always lies between −1 and 1. Getting |r| > 1 out of a calculation means an arithmetic mistake, not an unusual dataset.
Sign of r always matches sign of b — same numerator, Sxy.
r unaffected by (linear) coding. b DOES scale with the coding constants — r doesn't. Never recompute r from original data after already finding it from coded data.
Keep a square-root denominator unrounded (or to 1+ extra figure) until the final line — rounding it early can change the final significant figure, not just its precision.

Why it works — Why r has no units when b does

Set b=Sxy/Sxxb = S_{xy}/S_{xx} and r=Sxy/SxxSyyr = S_{xy}/\sqrt{S_{xx}S_{yy}} side by side and the difference is exactly one factor: r=b×SxxSxxSyy=b×SxxSyyr = b \times \dfrac{S_{xx}}{\sqrt{S_{xx}S_{yy}}} = b \times \sqrt{\dfrac{S_{xx}}{S_{yy}}}. Every quantity on the right carries units — bb is measured in (units of yy) per (unit of xx); SxxS_{xx} is measured in (units of xx)2^2; SyyS_{yy} is measured in (units of yy)2^2 — and Sxx/Syy\sqrt{S_{xx}/S_{yy}} works out to (units of xx)/(units of yy), the exact reciprocal of bb's own units. Multiply the two together and every unit cancels, leaving a pure number. That isn't a coincidence baked into the formula by choice; it follows directly from rr being built out of a ratio of two MATCHING kinds of spread (SxxS_{xx} against SyyS_{yy}, both 'sum of squared deviations from a mean') rather than bb's ratio of an association (SxyS_{xy}) against just one variable's own spread (SxxS_{xx} alone). This is exactly why bb can be 156156, or 4.254.25, or 0.003-0.003 — any real number at all, in whatever units the question happens to use — while rr is always trapped between 1-1 and 11: nothing about bb's formula bounds it, and everything about rr's formula does.

Traps — 2

sqrt-omitted-in-pmcc-denominator
Confirmed on the real question this comes from — Jun 2024 Q4(b), the actual PMCC step of the same question whose part (c) is a 'show that' for the regression line (verified directly against the official Pearson mark scheme, WST01_01_2406_MS, and examiner report, WST01_01_2406_ER, during this lesson's own review, not just against the facts bank's secondhand write-up of it): *'Part (b) was answered well with many students able to calculate a correct value of the product moment correlation coefficient. Common error included the omission of the square root in the denominator.'* The failure isn't an accuracy slip — it's dividing by SxxSyyS_{xx}S_{yy} directly instead of SxxSyy\sqrt{S_{xx}S_{yy}}, which doesn't just lose precision, it produces a genuinely different (and, since the one step that keeps rr bounded between 1-1 and 11 has been skipped, often an out-of-range) number.
pmcc-recomputed-unnecessarily-after-coding
A real mark scheme credits a clean, quotable fact directly: *'r not affected by (linear) coding'* (Jan 2024 MS, Q2(d)). The trap is spending time — and sometimes marks — undoing a coding scheme that never needed undoing: recomputing SxxS_{xx}, SyyS_{yy} and SxyS_{xy} from the original, uncoded xx and yy values after already finding rr from the coded data, on the mistaken assumption that rr needs 'converting back' the way the regression coefficient bb genuinely does. bb scales with the coding constants; rr doesn't, by construction (see the mechanism and derivation earlier in this lesson) — and a script that recalculates rr from scratch for the uncoded data isn't doing extra-safe working, it's demonstrating it hasn't understood the property the mark scheme is actually crediting.

Say it out loud

Out loud, from memory, no notes: explain why r has no units when b does to someone who has never seen this topic — where does your explanation get vague or hand-wavy? That's the exact spot to re-study, and it only works if you check it: read back over the mechanism above the moment you finish talking and mark precisely where you drifted from it.

Spec 4.2

1 lesson

Regression — gradient interpretation and extrapolation/reliability

Every one of the five examiner reports read for this unit flags the same failure, worded a different way each time. Computing b=Sxy/Sxxb = S_{xy}/S_{xx} is not where the marks go missing — candidates who can barely finish a SxxS_{xx} calculation still reliably reach the right line. What goes missing is the sentence after it: saying, in words, what one extra unit of xx actually buys you in yy, and knowing the difference between a prediction the data has already tested and one it is only guessing at.

The card

Regression line of y on x: ŷ = a + bx. b = Sxy/Sxx, a = ȳ − bx̄. Sxx = Σx² − (Σx)²/n, Sxy = Σxy − ΣxΣy/n.
Never x̄² for (Σx)²/n inside Sxx. Never leave a or b as a fraction — decimals, 3 s.f., in the final line.
Gradient sentence: 'for each extra 1 unit of x, y changes by b units, on average.' Check what 1 unit of x actually is.
Interpolation = inside the recorded range = usually reliable. Extrapolation = outside it, either direction = name the value and the range.
The fitted line always passes through (x̄, ȳ) — a free check on any computed a, b.

Why it works — Why 'as x increases, y increases' is not an interpretation

A gradient interpretation question is not asking whether xx and yy move together — the scatter diagram and the sign of bb already answered that, and restating it earns nothing on its own; examiner reports say so independently, across several different series and contexts. What the question is checking is whether you can read bb as a genuine RATE: the amount yy changes, on average, for one whole extra unit of xx — a number with a size and a unit attached, not a description of direction. Two things then have to survive the translation from algebra into English, and both are independently confirmed as the specific things that go wrong. First, 'one extra unit of xx' has to match how xx was actually defined in the question, not whatever feels like a natural unit — a real report puts the cost of getting this wrong in exact numbers: the main error 'was not recognising that a single rise in the number of employees led to a rise in the amount spent on paper of $156 and not $1.56' (Oct 2021, Q2(d)), a hundred-fold slip from reading the gradient's raw value as if xx had been measured in ones. Another report shows the identical failure from the opposite direction, where xx was compressed into millions rather than expanded into ones: 'only the most able candidates successfully managed to write that the GDP increases by 31.2 billion dollars for every 1 million increase in the population' (Jun 2022, Q2(d)) — most answers that series gave the direction of the relationship and stopped, never converting the raw coefficient into a sentence about what one real unit of xx buys. Second, the sentence has to be anchored in the actual variables and their real meaning, not a generic template: one report records exactly that failure — *'too many students referred to positive correlation or "as x increases then y increases"... some students mixed up the units (grams and °C)'* (Jan 2023, Q6(a)) — and another records a sentence that survives everything except getting turned around: *'a few lacked the context required, whilst others gave the interpretation the wrong way round'* (Jun 2024, Q4(d)). A credited answer names both variables in their own terms, states the correct SIZE of one real unit of xx, states the direction, and does all three in a single sentence — not four separate half-marks' worth of fragments.

Traps — 8

sxx-computed-with-xbar-squared
Confirmed directly on a real SxxS_{xx} question: *'the most common error made was using x̄² in the calculation of Sxx'* (Oct 2021, Q2(b)). xˉ2\bar x^2 and (x)2n\frac{(\sum x)^2}{n} are not the same number — for this lesson's own dataset they're 36 against 180 — and the two formulas that use them, Sxx=x2xˉ2S_{xx} = \sum x^2 - \bar x^2 (wrong) against Sxx=x2(x)2nS_{xx} = \sum x^2 - \frac{(\sum x)^2}{n} (correct, and the one printed in the booklet), look similar enough on the page that the substitution can slip past unnoticed.
gradient-fraction-inverted
Confirmed on a real regression question: gradient errors came from *'the fraction the wrong way up'* (Oct 2021, Q2(c)) — computing Sxx/SxyS_{xx}/S_{xy} instead of Sxy/SxxS_{xy}/S_{xx}. The fix is naming the formula out loud before substituting: bb is 'the sum that mixes x and y' over 'the sum that's just x', in that order, every time.
intercept-from-raw-totals-not-means
Confirmed on the same question, with the exact numbers the report names: intercept errors came from substituting the raw column totals — the report gives 273 and 93 — straight into a=yˉbxˉa = \bar y - b\bar x in place of the actual means those totals needed to be divided by nn first to produce (Oct 2021, Q2(c)). x\sum x and xˉ\bar x differ by a factor of nn; using one where the formula asks for the other doesn't produce a slightly-off answer, it produces one wrong by exactly that factor.
stops-after-a-and-b-never-states-the-line
Confirmed on the same question again: candidates who correctly found both aa and bb but never wrote them back up as the actual equation of the line lost the final mark of the part for stopping one step early (Oct 2021, Q2(c)). Finding the two numbers is not the same task as answering 'find the equation of the regression line' — the sentence y=a+bxy = a + bx, with the values substituted in, is the actual deliverable.
fraction-left-in-final-line
Confirmed directly: *'it is also important to note that fractions are not accepted in a final regression equation; the values of a and b are both estimates and so a fractional answer is not appropriate'* (Jun 2022, Q2(c)). Once b=Sxy/Sxxb = S_{xy}/S_{xx} has actually been divided out, it stays a decimal (to at least 3 s.f.) for the rest of the question — reverting to the exact fraction it started as, e.g. 17040\frac{170}{40}, in the final line costs the mark even when the value is numerically correct.
gradient-interpreted-at-the-wrong-scale
The richest single trap in this topic, confirmed independently in two different subjects' worth of context. One report gives the cost in dollars: the main error 'was not recognising that a single rise in the number of employees led to a rise in the amount spent on paper of $156 and not $1.56' (Oct 2021, Q2(d)) — a hundred-fold misreading of the gradient's own scale. Another shows the same failure from the other direction: 'only the most able candidates successfully managed to write that the GDP increases by 31.2 billion dollars for every 1 million increase in the population' (Jun 2022, Q2(d)) — most answers gave the direction of change and stopped, never converting the raw coefficient into a sentence about what one real unit of the explanatory variable actually buys. Whatever xx is measured in — dollars, hundreds of dollars, millions of people — 'one unit of x' in the interpretation sentence has to mean that, not '1' read off the page with no units attached.
interpretation-gives-correlation-not-a-rate
Confirmed twice: *'too many students referred to positive correlation or "as x increases then y increases"... some students mixed up the units (grams and °C)'* (Jan 2023, Q6(a)); and separately, *'a few lacked the context required, whilst others gave the interpretation the wrong way round'* (Jun 2024, Q4(d)). Stating that xx and yy move together restates something the scatter diagram already showed, and earns nothing on its own — a full-credit sentence names both variables, states the correct-scale rate of change, and gets the direction right, all three, in one sentence.
reliability-judged-without-naming-the-range
Confirmed directly: *'it was rare to see responses which accurately assessed the reliability of the estimate found... many stated the estimate was unreliable, but they were unable to refer to the correct variable or value that is not in the range... it was also no surprise to find responses saying that the estimate was reliable, even with correct working earlier in the question [that showed it wasn't]'* (Jun 2022, Q2(e)). The credited form is specific: a separate report records that successful answers 'made reference to 90 being outside of the range and therefore unreliable' (Jan 2023, Q6(c)) — name the value, name the range, connect the two, or the comment doesn't score even when the underlying judgement (reliable / unreliable) happens to be right.

Say it out loud

Out loud, from memory, no notes: explain why 'as x increases, y increases' is not an interpretation to someone who has never seen this topic — where does your explanation get vague or hand-wavy? That's the exact spot to re-study, and it only works if you check it: read back over the mechanism above the moment you finish talking and mark precisely where you drifted from it.

Spec 4.1

2 lessons

Regression — gradient interpretation and extrapolation/reliability

Every one of the five examiner reports read for this unit flags the same failure, worded a different way each time. Computing b=Sxy/Sxxb = S_{xy}/S_{xx} is not where the marks go missing — candidates who can barely finish a SxxS_{xx} calculation still reliably reach the right line. What goes missing is the sentence after it: saying, in words, what one extra unit of xx actually buys you in yy, and knowing the difference between a prediction the data has already tested and one it is only guessing at.

The card

Regression line of y on x: ŷ = a + bx. b = Sxy/Sxx, a = ȳ − bx̄. Sxx = Σx² − (Σx)²/n, Sxy = Σxy − ΣxΣy/n.
Never x̄² for (Σx)²/n inside Sxx. Never leave a or b as a fraction — decimals, 3 s.f., in the final line.
Gradient sentence: 'for each extra 1 unit of x, y changes by b units, on average.' Check what 1 unit of x actually is.
Interpolation = inside the recorded range = usually reliable. Extrapolation = outside it, either direction = name the value and the range.
The fitted line always passes through (x̄, ȳ) — a free check on any computed a, b.

Why it works — Why 'as x increases, y increases' is not an interpretation

A gradient interpretation question is not asking whether xx and yy move together — the scatter diagram and the sign of bb already answered that, and restating it earns nothing on its own; examiner reports say so independently, across several different series and contexts. What the question is checking is whether you can read bb as a genuine RATE: the amount yy changes, on average, for one whole extra unit of xx — a number with a size and a unit attached, not a description of direction. Two things then have to survive the translation from algebra into English, and both are independently confirmed as the specific things that go wrong. First, 'one extra unit of xx' has to match how xx was actually defined in the question, not whatever feels like a natural unit — a real report puts the cost of getting this wrong in exact numbers: the main error 'was not recognising that a single rise in the number of employees led to a rise in the amount spent on paper of $156 and not $1.56' (Oct 2021, Q2(d)), a hundred-fold slip from reading the gradient's raw value as if xx had been measured in ones. Another report shows the identical failure from the opposite direction, where xx was compressed into millions rather than expanded into ones: 'only the most able candidates successfully managed to write that the GDP increases by 31.2 billion dollars for every 1 million increase in the population' (Jun 2022, Q2(d)) — most answers that series gave the direction of the relationship and stopped, never converting the raw coefficient into a sentence about what one real unit of xx buys. Second, the sentence has to be anchored in the actual variables and their real meaning, not a generic template: one report records exactly that failure — *'too many students referred to positive correlation or "as x increases then y increases"... some students mixed up the units (grams and °C)'* (Jan 2023, Q6(a)) — and another records a sentence that survives everything except getting turned around: *'a few lacked the context required, whilst others gave the interpretation the wrong way round'* (Jun 2024, Q4(d)). A credited answer names both variables in their own terms, states the correct SIZE of one real unit of xx, states the direction, and does all three in a single sentence — not four separate half-marks' worth of fragments.

Traps — 8

sxx-computed-with-xbar-squared
Confirmed directly on a real SxxS_{xx} question: *'the most common error made was using x̄² in the calculation of Sxx'* (Oct 2021, Q2(b)). xˉ2\bar x^2 and (x)2n\frac{(\sum x)^2}{n} are not the same number — for this lesson's own dataset they're 36 against 180 — and the two formulas that use them, Sxx=x2xˉ2S_{xx} = \sum x^2 - \bar x^2 (wrong) against Sxx=x2(x)2nS_{xx} = \sum x^2 - \frac{(\sum x)^2}{n} (correct, and the one printed in the booklet), look similar enough on the page that the substitution can slip past unnoticed.
gradient-fraction-inverted
Confirmed on a real regression question: gradient errors came from *'the fraction the wrong way up'* (Oct 2021, Q2(c)) — computing Sxx/SxyS_{xx}/S_{xy} instead of Sxy/SxxS_{xy}/S_{xx}. The fix is naming the formula out loud before substituting: bb is 'the sum that mixes x and y' over 'the sum that's just x', in that order, every time.
intercept-from-raw-totals-not-means
Confirmed on the same question, with the exact numbers the report names: intercept errors came from substituting the raw column totals — the report gives 273 and 93 — straight into a=yˉbxˉa = \bar y - b\bar x in place of the actual means those totals needed to be divided by nn first to produce (Oct 2021, Q2(c)). x\sum x and xˉ\bar x differ by a factor of nn; using one where the formula asks for the other doesn't produce a slightly-off answer, it produces one wrong by exactly that factor.
stops-after-a-and-b-never-states-the-line
Confirmed on the same question again: candidates who correctly found both aa and bb but never wrote them back up as the actual equation of the line lost the final mark of the part for stopping one step early (Oct 2021, Q2(c)). Finding the two numbers is not the same task as answering 'find the equation of the regression line' — the sentence y=a+bxy = a + bx, with the values substituted in, is the actual deliverable.
fraction-left-in-final-line
Confirmed directly: *'it is also important to note that fractions are not accepted in a final regression equation; the values of a and b are both estimates and so a fractional answer is not appropriate'* (Jun 2022, Q2(c)). Once b=Sxy/Sxxb = S_{xy}/S_{xx} has actually been divided out, it stays a decimal (to at least 3 s.f.) for the rest of the question — reverting to the exact fraction it started as, e.g. 17040\frac{170}{40}, in the final line costs the mark even when the value is numerically correct.
gradient-interpreted-at-the-wrong-scale
The richest single trap in this topic, confirmed independently in two different subjects' worth of context. One report gives the cost in dollars: the main error 'was not recognising that a single rise in the number of employees led to a rise in the amount spent on paper of $156 and not $1.56' (Oct 2021, Q2(d)) — a hundred-fold misreading of the gradient's own scale. Another shows the same failure from the other direction: 'only the most able candidates successfully managed to write that the GDP increases by 31.2 billion dollars for every 1 million increase in the population' (Jun 2022, Q2(d)) — most answers gave the direction of change and stopped, never converting the raw coefficient into a sentence about what one real unit of the explanatory variable actually buys. Whatever xx is measured in — dollars, hundreds of dollars, millions of people — 'one unit of x' in the interpretation sentence has to mean that, not '1' read off the page with no units attached.
interpretation-gives-correlation-not-a-rate
Confirmed twice: *'too many students referred to positive correlation or "as x increases then y increases"... some students mixed up the units (grams and °C)'* (Jan 2023, Q6(a)); and separately, *'a few lacked the context required, whilst others gave the interpretation the wrong way round'* (Jun 2024, Q4(d)). Stating that xx and yy move together restates something the scatter diagram already showed, and earns nothing on its own — a full-credit sentence names both variables, states the correct-scale rate of change, and gets the direction right, all three, in one sentence.
reliability-judged-without-naming-the-range
Confirmed directly: *'it was rare to see responses which accurately assessed the reliability of the estimate found... many stated the estimate was unreliable, but they were unable to refer to the correct variable or value that is not in the range... it was also no surprise to find responses saying that the estimate was reliable, even with correct working earlier in the question [that showed it wasn't]'* (Jun 2022, Q2(e)). The credited form is specific: a separate report records that successful answers 'made reference to 90 being outside of the range and therefore unreliable' (Jan 2023, Q6(c)) — name the value, name the range, connect the two, or the comment doesn't score even when the underlying judgement (reliable / unreliable) happens to be right.

Say it out loud

Out loud, from memory, no notes: explain why 'as x increases, y increases' is not an interpretation to someone who has never seen this topic — where does your explanation get vague or hand-wavy? That's the exact spot to re-study, and it only works if you check it: read back over the mechanism above the moment you finish talking and mark precisely where you drifted from it.

"Show that" answer discipline

A "show that" question already gives you the answer. That is not a hint about how easy the question is — it is a warning about how the mark scheme has to be built. When the final number is printed on the page, an examiner cannot give you credit for reaching it, because copying a printed number proves nothing about whether you can produce it yourself. Every mark in a "show that" question is attached to the *steps between* the start of the question and the number you were already handed — and Pearson has said, in three separate examiner reports spanning three years, that this is the single thing students most often get wrong on this paper, in every topic it touches.

The card

"Show that X" means the answer is given — the marks are for the steps to X, not for X itself.
Never write the target value first. Show the method, and let the target value appear as your last line.
Work to one more d.p./s.f. than the target throughout — matching its precision is how the last mark is lost.
"Show that there are N of something" needs the N things named, not the rule that would find them.
M0 A1 is impossible: no visible method means the accuracy marks have nothing to attach to, whatever the final line says.

Why it works — Why the mark scheme genuinely cannot pay for the printed answer

This follows directly from how M, A and B marks work, not from an arbitrary examiner preference (WST01-verified-facts.md §4; this exact "General Instructions for Marking" wording opens at least 6 of the 14 reviewed mark schemes, byte-for-byte — confirmed directly against the real Jan 2024 mark scheme PDF; Jun 2024's own preamble is shorter but keeps the same M-before-A dependency rule). An M mark is "given for a correct method or an attempt at a correct method" — it has to be visible to be given, because there is nothing else for the examiner to assess it against. An A mark is "dependent... and can only be awarded if the previous M mark has been earned. E.g. M0 A1 is impossible." On an ordinary question this dependency is nearly invisible, because a correct final answer is itself strong evidence a correct method was used. On a "show that" question that evidence is destroyed on purpose: the final answer is printed in the question, so a correct final answer is evidence of nothing — it is equally consistent with genuine working and with simply copying what was already given. The mark scheme's only defence is to refuse to award the A mark unless the M mark — the actual method, shown on the page — is there first. This is also why "the answer is printed on the paper" has its own dedicated symbol in the mark schemes' general abbreviations list: it is common enough, specifically because "show that" questions are common enough, to need its own shorthand. The Statistics-specific marking note reinforces the same point from the examiner's side: "Any correct method should gain credit. If you cannot see how to apply the mark scheme but believe the method to be correct then please send to review" — the examiner is instructed to credit what they can see, not what they can infer might have happened off the page.

Traps — 4

answer-only-no-intermediate-step
The single most literal version of the trap, on a real Normal-distribution "show that" question (Oct 2021, Q6(a)): "many students wrote down the standardisation followed by 0.15 missing out the intermediate step and the accurate answer. Some students simply wrote down 0.1500… with no working at all." Two separate failures are named in that one quote — some students showed the standardisation but skipped the accurate table value in between; others showed nothing at all and just wrote the given probability. Both lose marks, because both leave the examiner unable to distinguish "did the work, wrote it up badly" from "copied the given answer."
reverse-engineered-working
A more subtle version, on a real interquartile-range "show that" question (Oct 2021, Q3(b)): "the main errors, since the answer was given, were making up two ages... which differed by 16 as their 'working'." This is worse than showing nothing, not better — it produces two numbers that are correct only in that they happen to be 16 apart, with no interpolation, no cumulative frequency, no class boundary in sight. A "working" that was clearly reverse-engineered from the given answer, rather than derived from the data, does not read as method at all.
shown-result-not-fully-stated
Confirmed on a "show that there are 3 outliers" question (Oct 2021, Q3(c)): "having gained the correct limits some did not list the 3 outliers in this part in order to show there are 3 outliers." The calculation (the outlier boundaries) was right; what was missing was the last, cheapest step — actually pointing at the three data values that satisfy it. When the question asks you to show a count, a stated boundary is not the same as a demonstrated count, and the gap between them is one line of writing.
insufficient-accuracy-in-a-show-that
A precision-specific version, on a real regression-line "show that" question (Jun 2024, Q4(c) — the actual target line was g=42.3+0.722dg = -42.3 + 0.722d, 3 s.f.): "Many students were able to show the given regression line but too often students lost the final A mark as they failed to give values to the required degree of accuracy. b = 12105.12/16769.78 = 0.722 was not accurate enough to gain the final mark due to it being a 'show that' question. Students should be encouraged in these types of questions to give answers to at least one more decimal place than the given value." The candidate's 0.722 was correct — 12105.12 ÷ 16769.78 really does round to 0.722 — but stopping the working at the same 3 significant figures as the printed line cannot prove the calculation reached it rather than having been rounded to match; the extra digit (0.7218…) is the only thing that can.

Say it out loud

Out loud, from memory, no notes: explain why the mark scheme genuinely cannot pay for the printed answer to someone who has never seen this topic — where does your explanation get vague or hand-wavy? That's the exact spot to re-study, and it only works if you check it: read back over the mechanism above the moment you finish talking and mark precisely where you drifted from it.

Spec 5.2

1 lesson

Discrete random variables — the probability function and the discrete uniform distribution

A probability function isn't finished the moment every value looks reasonable on its own — it's finished when the whole list adds up to exactly 1. A real WST01 examiner report names the single most common way this goes wrong, and it isn't a miscalculation: it's forgetting that a value — often x=0x=0 specifically — belongs in the domain at all. This lesson builds the probability function and the cumulative distribution function from the one fact that makes that self-check work (p(x)=1\sum p(x) = 1, because XX is certain to take *some* value), then uses a genuinely rich real exam question — two four-sided dice, from the most recent series reviewed for this course — to show a discrete uniform distribution, a probability, a mean found "by symmetry," and a variance that has to be summed by hand all compose inside one real question.

The card

p(x) = P(X=x): every value 0 ≤ p(x) ≤ 1, and Σp(x) = 1 over the WHOLE domain — check this sum before moving on.
F(x₀) = P(X≤x₀) = Σp(x) for x≤x₀ — NOT in the formula booklet. Memorise it, and note the ≤, not <.
Discrete uniform: every value in the domain equally likely, p(x) = 1/n for n values — values needn't be consecutive integers.
Mean of a symmetric/evenly-spaced discrete uniform distribution: read off by symmetry, (min+max)/2. No shortcut exists for variance — sum every x²p(x) term.
Documented WST01 examiner-report error: leaving a value (often x=0) out of the domain entirely — check Σp(x)=1 to catch it (Oct 2021, Q4(d)).
Documented WST01 examiner-report error: solving for an unknown constant without ever writing down the Σp(x)=1 equation itself (Jun 2022, Q5(d)).

Why it works — Why $\sum p(x)$ has to equal exactly 1 — not approximately, not usually

Every discrete random variable's probability function has to satisfy two conditions, and neither is optional: p(x)0p(x) \ge 0 for every value (a probability can never be negative), and p(x)=1\sum p(x) = 1 over the whole domain. Where does the second condition actually come from? Start from what XX IS: a variable whose value is the outcome of a random process, which means XX is certain to take SOME value from its domain — that's simply what "random variable" means, not an extra assumption layered on top. The events "X=x1X=x_1", "X=x2X=x_2", and so on for every value in the domain, are mutually exclusive (XX cannot simultaneously equal two different values at once) and, together, exhaustive (there is no other outcome XX could have). For mutually exclusive events, probabilities add — the general addition rule P(AB)=P(A)+P(B)P(AB)P(A \cup B) = P(A)+P(B)-P(A \cap B) (spec 3.2, this course's own conditional-probability lesson covers it in full) collapses to plain addition the moment P(AB)=0P(A \cap B) = 0 — so P(X=x1 or X=x2 or )=p(x1)+p(x2)+P(X=x_1 \text{ or } X=x_2 \text{ or } \dots) = p(x_1)+p(x_2)+\dots. And the left-hand side of that equation is just P(X takes some value)P(X \text{ takes some value}): the probability of the certain event, which is 11 by definition. That's the whole derivation. p(x)=1\sum p(x) = 1 isn't a separate rule to memorise alongside p(x)=P(X=x)p(x)=P(X=x) — it falls straight out of what a random variable's domain actually IS: a list of outcomes that between them account for everything that could happen. Which is exactly why checking the sum is such a powerful self-check: if your working produces a total that isn't 11, either an arithmetic slip was made, or — the single most common version of this on the real exam — a genuinely possible value was left out of the domain before any arithmetic even started. A verified WST01 examiner report states this almost as a piece of exam advice in its own right: "had students checked whether the sum of their probabilities equalled 1 they may have realised that they had missed zero out" (Oct 2021, Q4(d)).

Traps — 2

possible-value-left-out-of-domain
The headline trap of this whole spec point, in Pearson's own words: "the most common error was to miss out the value x = 0... had students checked whether the sum of their probabilities equalled 1 they may have realised that they had missed zero out" (Oct 2021, Q4(d)). The defence is the mechanism block above, not repetition: Σp(x) = 1 because X is certain to take SOME value, and a total that falls short of 1 is a direct signal that a genuinely possible value never made it into the working at all — not just an arithmetic slip to hunt for.
sum-to-1-equation-not-written-down
The construction-side version of the same idea, confirmed directly: "many candidates scored the 2nd M mark but failed to write down the equation for the sum of probabilities = 1 for the first M mark" (Jun 2022, Q5(d)). This matters specifically on a "show that" item, where the target value is already printed on the page: reaching the correct final number without ever writing the equation that forces it isn't a shortcut, it's a derivation that never actually happened — see the marked-solution's WarrantCheck above for exactly this failure, built around this same real quote.

Say it out loud

Out loud, from memory, no notes: explain why $\sum p(x)$ has to equal exactly 1 — not approximately, not usually to someone who has never seen this topic — where does your explanation get vague or hand-wavy? That's the exact spot to re-study, and it only works if you check it: read back over the mechanism above the moment you finish talking and mark precisely where you drifted from it.

Spec 5.4

1 lesson

Discrete random variables — the probability function and the discrete uniform distribution

A probability function isn't finished the moment every value looks reasonable on its own — it's finished when the whole list adds up to exactly 1. A real WST01 examiner report names the single most common way this goes wrong, and it isn't a miscalculation: it's forgetting that a value — often x=0x=0 specifically — belongs in the domain at all. This lesson builds the probability function and the cumulative distribution function from the one fact that makes that self-check work (p(x)=1\sum p(x) = 1, because XX is certain to take *some* value), then uses a genuinely rich real exam question — two four-sided dice, from the most recent series reviewed for this course — to show a discrete uniform distribution, a probability, a mean found "by symmetry," and a variance that has to be summed by hand all compose inside one real question.

The card

p(x) = P(X=x): every value 0 ≤ p(x) ≤ 1, and Σp(x) = 1 over the WHOLE domain — check this sum before moving on.
F(x₀) = P(X≤x₀) = Σp(x) for x≤x₀ — NOT in the formula booklet. Memorise it, and note the ≤, not <.
Discrete uniform: every value in the domain equally likely, p(x) = 1/n for n values — values needn't be consecutive integers.
Mean of a symmetric/evenly-spaced discrete uniform distribution: read off by symmetry, (min+max)/2. No shortcut exists for variance — sum every x²p(x) term.
Documented WST01 examiner-report error: leaving a value (often x=0) out of the domain entirely — check Σp(x)=1 to catch it (Oct 2021, Q4(d)).
Documented WST01 examiner-report error: solving for an unknown constant without ever writing down the Σp(x)=1 equation itself (Jun 2022, Q5(d)).

Why it works — Why $\sum p(x)$ has to equal exactly 1 — not approximately, not usually

Every discrete random variable's probability function has to satisfy two conditions, and neither is optional: p(x)0p(x) \ge 0 for every value (a probability can never be negative), and p(x)=1\sum p(x) = 1 over the whole domain. Where does the second condition actually come from? Start from what XX IS: a variable whose value is the outcome of a random process, which means XX is certain to take SOME value from its domain — that's simply what "random variable" means, not an extra assumption layered on top. The events "X=x1X=x_1", "X=x2X=x_2", and so on for every value in the domain, are mutually exclusive (XX cannot simultaneously equal two different values at once) and, together, exhaustive (there is no other outcome XX could have). For mutually exclusive events, probabilities add — the general addition rule P(AB)=P(A)+P(B)P(AB)P(A \cup B) = P(A)+P(B)-P(A \cap B) (spec 3.2, this course's own conditional-probability lesson covers it in full) collapses to plain addition the moment P(AB)=0P(A \cap B) = 0 — so P(X=x1 or X=x2 or )=p(x1)+p(x2)+P(X=x_1 \text{ or } X=x_2 \text{ or } \dots) = p(x_1)+p(x_2)+\dots. And the left-hand side of that equation is just P(X takes some value)P(X \text{ takes some value}): the probability of the certain event, which is 11 by definition. That's the whole derivation. p(x)=1\sum p(x) = 1 isn't a separate rule to memorise alongside p(x)=P(X=x)p(x)=P(X=x) — it falls straight out of what a random variable's domain actually IS: a list of outcomes that between them account for everything that could happen. Which is exactly why checking the sum is such a powerful self-check: if your working produces a total that isn't 11, either an arithmetic slip was made, or — the single most common version of this on the real exam — a genuinely possible value was left out of the domain before any arithmetic even started. A verified WST01 examiner report states this almost as a piece of exam advice in its own right: "had students checked whether the sum of their probabilities equalled 1 they may have realised that they had missed zero out" (Oct 2021, Q4(d)).

Traps — 2

possible-value-left-out-of-domain
The headline trap of this whole spec point, in Pearson's own words: "the most common error was to miss out the value x = 0... had students checked whether the sum of their probabilities equalled 1 they may have realised that they had missed zero out" (Oct 2021, Q4(d)). The defence is the mechanism block above, not repetition: Σp(x) = 1 because X is certain to take SOME value, and a total that falls short of 1 is a direct signal that a genuinely possible value never made it into the working at all — not just an arithmetic slip to hunt for.
sum-to-1-equation-not-written-down
The construction-side version of the same idea, confirmed directly: "many candidates scored the 2nd M mark but failed to write down the equation for the sum of probabilities = 1 for the first M mark" (Jun 2022, Q5(d)). This matters specifically on a "show that" item, where the target value is already printed on the page: reaching the correct final number without ever writing the equation that forces it isn't a shortcut, it's a derivation that never actually happened — see the marked-solution's WarrantCheck above for exactly this failure, built around this same real quote.

Say it out loud

Out loud, from memory, no notes: explain why $\sum p(x)$ has to equal exactly 1 — not approximately, not usually to someone who has never seen this topic — where does your explanation get vague or hand-wavy? That's the exact spot to re-study, and it only works if you check it: read back over the mechanism above the moment you finish talking and mark precisely where you drifted from it.

Spec 5.3

1 lesson

E(aX+b) and Var(aX+b) for discrete random variables

Adding a constant to every outcome of a random variable slides its mean sideways and leaves its spread completely untouched; multiplying by a constant does the opposite — it scales the mean by that constant once, and the spread by that constant twice, because variance is built from a squared deviation in the first place. Two separate WST01 exam sittings, three years apart, produced the identical one-line slip three times over: multiplying the variance by aa instead of a2a^2. This lesson exists because that error is not a memory lapse to drill away — it is what happens when a formula is memorised without ever being asked why the square is there.

The card

E(aX+b) = aE(X) + b — NOT in the formula booklet. Memorise it.
Var(aX+b) = a²Var(X) — no b anywhere; a is always squared, whatever its sign.
Var(X) = E(X²) − [E(X)]² IS in the booklet — only the aX+b shortcuts aren't.
A negative variance is never a right answer — if you get one, you forgot to square a.
Documented WST01 examiner-report error, 3 times across 3 years: multiplying variance by a instead of a².

Why it works — Why $b$ vanishes completely, and why $a$ gets squared and never $b$ does

Start from the one formula that's actually on the sheet: E(g(X))=g(xi)P(X=xi)E(g(X)) = \sum g(x_i)P(X=x_i). Set g(x)=ax+bg(x) = ax+b, the general linear transformation. Then E(aX+b)=(axi+b)P(X=xi)=axiP(X=xi)+bP(X=xi)E(aX+b) = \sum (ax_i+b)P(X=x_i) = a\sum x_i P(X=x_i) + b\sum P(X=x_i) — splitting the sum is legal because aa and bb are constants, not random, so they can be pulled outside the summation. The first piece is a×E(X)a \times E(X) by definition. The second piece is bb times P(X=xi)\sum P(X=x_i), and every discrete probability distribution's probabilities sum to exactly 1 (spec 5.2), so that piece is just bb. Hence E(aX+b)=aE(X)+bE(aX+b) = aE(X) + b: the mean transformation is a direct, one-line consequence of linearity, nothing more exotic than that. Now variance. Variance is not itself a sum of XX — it is defined as E[(Xμ)2]E\left[(X - \mu)^2\right], the average SQUARED distance from the mean, which is a different kind of object with the mean already baked into it. To find Var(aX+b)\text{Var}(aX+b), first find the mean of aX+baX+b: it's aE(X)+baE(X)+b, from the paragraph just derived. So the deviation of aX+baX+b from ITS OWN mean is (aX+b)(aE(X)+b)=aXaE(X)=a(XE(X))(aX+b) - (aE(X)+b) = aX - aE(X) = a(X - E(X)) — and the bb has already cancelled, before any squaring has even happened. Squaring that deviation is what variance is: [a(XE(X))]2=a2(XE(X))2\left[a(X-E(X))\right]^2 = a^2(X-E(X))^2. Take the expectation of both sides and a2a^2, being a constant, pulls straight back out: Var(aX+b)=a2E[(XE(X))2]=a2Var(X)\text{Var}(aX+b) = a^2 E\left[(X-E(X))^2\right] = a^2\text{Var}(X). Read the two derivations side by side and the whole trap dissolves into one sentence: bb is gone from the variance formula because it was gone from the DEVIATION before the squaring step ever started — a pure shift moves every value and the mean by exactly the same amount, so the gap between them, which is all variance ever measures, is untouched. And aa comes out squared, not because someone chose to square it, but because variance is defined as an average of a SQUARED quantity, and squaring is what turns "the spread got multiplied by aa" into "the *squared* spread got multiplied by aa, twice."

Traps — 3

coefficient-not-squared-in-variance
The headline trap of this whole spec point, confirmed independently at least three times across three years — the same underlying mechanism, never squaring the multiplicative constant when finding Var(aX+b): "others forgot to square [a constant] when subtracting" (Jan 2021, Q4(b–d)); "the most common error seen was a × 4.14 = 66.24 rather than a² × 4.14 = 66.24" (Jun 2024, Q2(c)); "common errors included incorrect use of variance expressions and failure to realise that Var(aX) = a²Var(X)" (Jun 2024, Q3(d) — the same paper, a second independent instance). The defence is the derivation above, not repetition: the square is there because variance is defined as an average of a squared deviation, and a pure scaling of X scales that deviation by a, which the squaring then squares again.
e-x-squared-confused-with-var-x
Confirmed on the same question that produced the "forgot to square" trap above: "some still think E(X²) = Var(X)" (Jan 2021, Q4(b–d)). These are two different objects computed from two different formulae — E(X²) = Σx²P(X=x) is a raw second moment; Var(X) = E(X²) − [E(X)]² is that moment with the SQUARE OF THE MEAN subtracted back off. Skipping the subtraction step doesn't just lose a mark on a routine question — reused inside an aX+b question, it feeds the wrong number into everything that follows it.
formulae-not-known-from-memory
"others did not know the formulae for E(aX+b) and Var(aX+b)" (Jan 2021, Q4(b–d)) — and this is worth taking literally rather than as a synonym for the other two traps above. E(aX+b) and Var(aX+b) are explicitly NOT printed anywhere in the Mathematical Formulae and Statistical Tables booklet (verified in WST01-verified-facts.md §2a, against the spec's own statement that formulae "students are expected to know... will not appear in the booklet"). A student who forgets the base formulae for E(X) and Var(X) can look them up; a student who forgets the aX+b shortcuts has nowhere in the exam room to check.

Say it out loud

Out loud, from memory, no notes: explain why $b$ vanishes completely, and why $a$ gets squared and never $b$ does to someone who has never seen this topic — where does your explanation get vague or hand-wavy? That's the exact spot to re-study, and it only works if you check it: read back over the mechanism above the moment you finish talking and mark precisely where you drifted from it.

Spec 6.1

2 lessons

The Normal distribution — standardisation, table precision, and conditional probability

Every one of the five examiner reports checked for this course treats one skill on this topic as the most discriminating part of the paper — and it is not really a Normal-distribution skill at all. It is whether you notice, underneath a completely ordinary standardisation, that the question has quietly become a conditional probability — and by the time you are three lines into the wrong calculation, the diagram that would have shown you is nowhere on the page.

The card

Z = (X − μ)/σ. Standardise first, always — not in the booklet, memorise it.
Φ(z) = P(Z < z), read straight from the table. P(Z > z) = 1 − Φ(z) — sketch before you subtract.
Percentage Points table: inverse lookup, gives z to 4 d.p. for a stated tail probability. Quote in full — 1.6449, never 1.64.
Conditional probability with Normal: same rule as always, P(A|B) = P(A∩B)/P(B). If A sits inside B, P(A∩B) = P(A).
'Show that' a probability: the standardisation must appear on the page, not just the final decimal.

Why it works — Why 'greater than' always costs you a subtraction — and why it fails in both directions

The Normal Distribution Function table stores exactly one kind of number, on every page: Φ(z) = P(Z < z), the area under the standard Normal curve to the LEFT of z. It does not store P(Z > z) anywhere — because it does not need to. The total area under the whole curve is exactly 1 (one of the two shape facts spec 6.1 explicitly expects you to know, not derive), so whatever fraction of that area sits to the left of z, everything else — the entire right-hand region — is 1 minus that fraction: P(Z > z) = 1 − Φ(z). That single subtraction is the single most consistently mis-fired step across the whole of this topic, confirmed in the facts bank failing in BOTH directions across three separate series. Under-subtracting: a January 2023 examiner report records that 'a significant number of students lost 2 marks as they failed to subtract from one the value obtained from the normal tables.' Over-subtracting: a June 2024 report records candidates who 'went on to subtract the correct answer from 1, which of course is P(X > 18) and not P(X < 18) which is what was required.' Both failures share one root cause: reaching for the table before deciding, from the question's own wording, which side of z is actually being asked about. Pearson's own fix, stated explicitly in the same January 2023 report: 'a simple diagram would have helped many to avoid this error.'

Traps — 5

subtract-from-1-direction-confusion
Confirmed in three separate series, failing in BOTH directions — the same underlying gap producing opposite mistakes depending on the question. Under-subtracting: "a significant number of students lost 2 marks as they failed to subtract from one the value obtained from the normal tables. A simple diagram would have helped many to avoid this error" (Jan 2023, Q5(a)). Over-subtracting: "a few lost the final mark as they went on to subtract the correct answer from 1, which of course is P(X > 18) and not P(X < 18) which is what was required" (Jun 2024, Q5(a)). A third series confirms the skill is fragile even when it goes right: "most students standardising correctly and the majority realising that they then needed to subtract the value found in the tables from 1" (Jan 2021, Q3(a)) — implying, in Pearson's own words, that a real minority did not. The fix in every case is the same: sketch which side of z the question describes before opening the table at all.
rounded-z-value-instead-of-4dp-table-value
The exact same mechanism, confirmed in three independent series, each with its own quoted wrong value: "many students were using a z value of 1.03 or 1.04 rather than the value 1.0364 from the 'Percentage Points of the Normal Distribution' table" (Jan 2021, Q3(b)); "not using an inaccurate value such as 1.64 instead of 1.6449" (Oct 2021, Q6(b)); "the most common error included the use of an inaccurate z value... students should be reminded that when values are required from the tables, they need to be 4 decimal places. A common error was to use z value = 0.25" (Jun 2024, Q5(b)). This is not carelessness in isolation — it is exam technique stated verbatim on every WST01 question-paper front page: "Values from the statistical tables should be quoted in full" (verified, facts bank §2).
sign-error-reversing-standardisation
Confirmed in two series, once in each direction of the sign: "others gained this mark but were using −1.0364, an error that could probably have been avoided if they had drawn a suitable diagram" (Jan 2021, Q3(b)) — where the value needed was positive; "using the wrong sign for 1.6449 appropriate to their standardisation giving 34.1, a value higher than the upper limit" (Oct 2021, Q6(b)) — where an unflipped sign produced an answer that was, on inspection, impossible. That second detail is the real lesson: the wrong answer was checkable as wrong using nothing but the question's own numbers, and the check was skipped anyway.
conditional-probability-not-recognised-with-normal
The most consistently documented trap in the whole facts bank — present, in a different concrete form, in every one of the 5 examiner reports reviewed. "Many did not realise that a conditional probability was required... a common error P(W<18)/0.85... but there were a number of correct attempts of the form (0.85−0.5)/0.85 which usually led to the correct answer" (Jan 2021, Q3(c)). "Many students did not realise that the ratios only applied to the middle 80% of the data" (Oct 2021, Q6(c)). "Like question 2, a common error was that students failed to realise that a conditional probability was required. A common error was to find P(L≤5) and go no further" (Jan 2023, Q5(e)). Stated as a paper-wide diagnosis, not a topic-specific aside: "Candidates often assume independence when an appropriate conditional probability should be used instead" (Jun 2022, general comment). One series' report names this sub-skill, in these words, as the hardest thing on the entire paper: a Jun 2022 question is flagged as "the final part of the paper... also the most discriminating part."
show-that-standardisation-not-shown
A cross-cutting instruction-following failure, not a maths error, confirmed directly on a Normal-distribution question: "many students wrote down the standardisation followed by 0.15 missing out the intermediate step and the accurate answer. Some students simply wrote down 0.1500… with no working at all" (Oct 2021, Q6(a)). Stated as a general rule in a later series: "if asked to use standardisation then the standardisation should be shown" (Jun 2024, general introduction). This costs marks even when the final decimal is completely correct — a 'show that' mark scheme has nothing to attach a mark to if the standardisation step itself never appears on the page.

Say it out loud

Out loud, from memory, no notes: explain why 'greater than' always costs you a subtraction — and why it fails in both directions to someone who has never seen this topic — where does your explanation get vague or hand-wavy? That's the exact spot to re-study, and it only works if you check it: read back over the mechanism above the moment you finish talking and mark precisely where you drifted from it.

"Show that" answer discipline

A "show that" question already gives you the answer. That is not a hint about how easy the question is — it is a warning about how the mark scheme has to be built. When the final number is printed on the page, an examiner cannot give you credit for reaching it, because copying a printed number proves nothing about whether you can produce it yourself. Every mark in a "show that" question is attached to the *steps between* the start of the question and the number you were already handed — and Pearson has said, in three separate examiner reports spanning three years, that this is the single thing students most often get wrong on this paper, in every topic it touches.

The card

"Show that X" means the answer is given — the marks are for the steps to X, not for X itself.
Never write the target value first. Show the method, and let the target value appear as your last line.
Work to one more d.p./s.f. than the target throughout — matching its precision is how the last mark is lost.
"Show that there are N of something" needs the N things named, not the rule that would find them.
M0 A1 is impossible: no visible method means the accuracy marks have nothing to attach to, whatever the final line says.

Why it works — Why the mark scheme genuinely cannot pay for the printed answer

This follows directly from how M, A and B marks work, not from an arbitrary examiner preference (WST01-verified-facts.md §4; this exact "General Instructions for Marking" wording opens at least 6 of the 14 reviewed mark schemes, byte-for-byte — confirmed directly against the real Jan 2024 mark scheme PDF; Jun 2024's own preamble is shorter but keeps the same M-before-A dependency rule). An M mark is "given for a correct method or an attempt at a correct method" — it has to be visible to be given, because there is nothing else for the examiner to assess it against. An A mark is "dependent... and can only be awarded if the previous M mark has been earned. E.g. M0 A1 is impossible." On an ordinary question this dependency is nearly invisible, because a correct final answer is itself strong evidence a correct method was used. On a "show that" question that evidence is destroyed on purpose: the final answer is printed in the question, so a correct final answer is evidence of nothing — it is equally consistent with genuine working and with simply copying what was already given. The mark scheme's only defence is to refuse to award the A mark unless the M mark — the actual method, shown on the page — is there first. This is also why "the answer is printed on the paper" has its own dedicated symbol in the mark schemes' general abbreviations list: it is common enough, specifically because "show that" questions are common enough, to need its own shorthand. The Statistics-specific marking note reinforces the same point from the examiner's side: "Any correct method should gain credit. If you cannot see how to apply the mark scheme but believe the method to be correct then please send to review" — the examiner is instructed to credit what they can see, not what they can infer might have happened off the page.

Traps — 4

answer-only-no-intermediate-step
The single most literal version of the trap, on a real Normal-distribution "show that" question (Oct 2021, Q6(a)): "many students wrote down the standardisation followed by 0.15 missing out the intermediate step and the accurate answer. Some students simply wrote down 0.1500… with no working at all." Two separate failures are named in that one quote — some students showed the standardisation but skipped the accurate table value in between; others showed nothing at all and just wrote the given probability. Both lose marks, because both leave the examiner unable to distinguish "did the work, wrote it up badly" from "copied the given answer."
reverse-engineered-working
A more subtle version, on a real interquartile-range "show that" question (Oct 2021, Q3(b)): "the main errors, since the answer was given, were making up two ages... which differed by 16 as their 'working'." This is worse than showing nothing, not better — it produces two numbers that are correct only in that they happen to be 16 apart, with no interpolation, no cumulative frequency, no class boundary in sight. A "working" that was clearly reverse-engineered from the given answer, rather than derived from the data, does not read as method at all.
shown-result-not-fully-stated
Confirmed on a "show that there are 3 outliers" question (Oct 2021, Q3(c)): "having gained the correct limits some did not list the 3 outliers in this part in order to show there are 3 outliers." The calculation (the outlier boundaries) was right; what was missing was the last, cheapest step — actually pointing at the three data values that satisfy it. When the question asks you to show a count, a stated boundary is not the same as a demonstrated count, and the gap between them is one line of writing.
insufficient-accuracy-in-a-show-that
A precision-specific version, on a real regression-line "show that" question (Jun 2024, Q4(c) — the actual target line was g=42.3+0.722dg = -42.3 + 0.722d, 3 s.f.): "Many students were able to show the given regression line but too often students lost the final A mark as they failed to give values to the required degree of accuracy. b = 12105.12/16769.78 = 0.722 was not accurate enough to gain the final mark due to it being a 'show that' question. Students should be encouraged in these types of questions to give answers to at least one more decimal place than the given value." The candidate's 0.722 was correct — 12105.12 ÷ 16769.78 really does round to 0.722 — but stopping the working at the same 3 significant figures as the printed line cannot prove the calculation reached it rather than having been rounded to match; the extra digit (0.7218…) is the only thing that can.

Say it out loud

Out loud, from memory, no notes: explain why the mark scheme genuinely cannot pay for the printed answer to someone who has never seen this topic — where does your explanation get vague or hand-wavy? That's the exact spot to re-study, and it only works if you check it: read back over the mechanism above the moment you finish talking and mark precisely where you drifted from it.

Spec cross-cutting

1 lesson

"Show that" answer discipline

A "show that" question already gives you the answer. That is not a hint about how easy the question is — it is a warning about how the mark scheme has to be built. When the final number is printed on the page, an examiner cannot give you credit for reaching it, because copying a printed number proves nothing about whether you can produce it yourself. Every mark in a "show that" question is attached to the *steps between* the start of the question and the number you were already handed — and Pearson has said, in three separate examiner reports spanning three years, that this is the single thing students most often get wrong on this paper, in every topic it touches.

The card

"Show that X" means the answer is given — the marks are for the steps to X, not for X itself.
Never write the target value first. Show the method, and let the target value appear as your last line.
Work to one more d.p./s.f. than the target throughout — matching its precision is how the last mark is lost.
"Show that there are N of something" needs the N things named, not the rule that would find them.
M0 A1 is impossible: no visible method means the accuracy marks have nothing to attach to, whatever the final line says.

Why it works — Why the mark scheme genuinely cannot pay for the printed answer

This follows directly from how M, A and B marks work, not from an arbitrary examiner preference (WST01-verified-facts.md §4; this exact "General Instructions for Marking" wording opens at least 6 of the 14 reviewed mark schemes, byte-for-byte — confirmed directly against the real Jan 2024 mark scheme PDF; Jun 2024's own preamble is shorter but keeps the same M-before-A dependency rule). An M mark is "given for a correct method or an attempt at a correct method" — it has to be visible to be given, because there is nothing else for the examiner to assess it against. An A mark is "dependent... and can only be awarded if the previous M mark has been earned. E.g. M0 A1 is impossible." On an ordinary question this dependency is nearly invisible, because a correct final answer is itself strong evidence a correct method was used. On a "show that" question that evidence is destroyed on purpose: the final answer is printed in the question, so a correct final answer is evidence of nothing — it is equally consistent with genuine working and with simply copying what was already given. The mark scheme's only defence is to refuse to award the A mark unless the M mark — the actual method, shown on the page — is there first. This is also why "the answer is printed on the paper" has its own dedicated symbol in the mark schemes' general abbreviations list: it is common enough, specifically because "show that" questions are common enough, to need its own shorthand. The Statistics-specific marking note reinforces the same point from the examiner's side: "Any correct method should gain credit. If you cannot see how to apply the mark scheme but believe the method to be correct then please send to review" — the examiner is instructed to credit what they can see, not what they can infer might have happened off the page.

Traps — 4

answer-only-no-intermediate-step
The single most literal version of the trap, on a real Normal-distribution "show that" question (Oct 2021, Q6(a)): "many students wrote down the standardisation followed by 0.15 missing out the intermediate step and the accurate answer. Some students simply wrote down 0.1500… with no working at all." Two separate failures are named in that one quote — some students showed the standardisation but skipped the accurate table value in between; others showed nothing at all and just wrote the given probability. Both lose marks, because both leave the examiner unable to distinguish "did the work, wrote it up badly" from "copied the given answer."
reverse-engineered-working
A more subtle version, on a real interquartile-range "show that" question (Oct 2021, Q3(b)): "the main errors, since the answer was given, were making up two ages... which differed by 16 as their 'working'." This is worse than showing nothing, not better — it produces two numbers that are correct only in that they happen to be 16 apart, with no interpolation, no cumulative frequency, no class boundary in sight. A "working" that was clearly reverse-engineered from the given answer, rather than derived from the data, does not read as method at all.
shown-result-not-fully-stated
Confirmed on a "show that there are 3 outliers" question (Oct 2021, Q3(c)): "having gained the correct limits some did not list the 3 outliers in this part in order to show there are 3 outliers." The calculation (the outlier boundaries) was right; what was missing was the last, cheapest step — actually pointing at the three data values that satisfy it. When the question asks you to show a count, a stated boundary is not the same as a demonstrated count, and the gap between them is one line of writing.
insufficient-accuracy-in-a-show-that
A precision-specific version, on a real regression-line "show that" question (Jun 2024, Q4(c) — the actual target line was g=42.3+0.722dg = -42.3 + 0.722d, 3 s.f.): "Many students were able to show the given regression line but too often students lost the final A mark as they failed to give values to the required degree of accuracy. b = 12105.12/16769.78 = 0.722 was not accurate enough to gain the final mark due to it being a 'show that' question. Students should be encouraged in these types of questions to give answers to at least one more decimal place than the given value." The candidate's 0.722 was correct — 12105.12 ÷ 16769.78 really does round to 0.722 — but stopping the working at the same 3 significant figures as the printed line cannot prove the calculation reached it rather than having been rounded to match; the extra digit (0.7218…) is the only thing that can.

Say it out loud

Out loud, from memory, no notes: explain why the mark scheme genuinely cannot pay for the printed answer to someone who has never seen this topic — where does your explanation get vague or hand-wavy? That's the exact spot to re-study, and it only works if you check it: read back over the mechanism above the moment you finish talking and mark precisely where you drifted from it.

Say these out loud before the exam

Every prompt below is answerable from the sheet above. If one stops you, that’s the page to go back to — and the fact that it stopped you is worth more than another read-through of the pages that didn’t.

  1. In one sentence: why does "the die is fair" have to be treated as an assumption you're licensed to use, rather than as a fact you've verified — and what practical difference does that make to how you'd answer a "critique this model" question if one appeared?
  2. What is the "spec-1.1-expected-as-a-standalone-question" trap, and how do you catch it?
  3. What is the "assumption-treated-as-a-provable-fact" trap, and how do you catch it?
  4. What is the "refinement-answer-not-connected-to-the-specific-evidence-given" trap, and how do you catch it?
  5. Without looking: what does this lesson say about what a model actually is: assumption, parameter, refinement?
  6. Without looking: what does this lesson say about why a topic with no past-paper question of its own is still worth 15-20 marks?
  7. In one sentence: why can two histogram bars have exactly the same height and still represent completely different numbers of data points?
  8. What is the "histogram-height-read-as-frequency" trap, and how do you catch it?
  9. What is the "stem-and-leaf-read-in-the-wrong-direction" trap, and how do you catch it?
  10. What is the "comparison-lacks-a-named-statistic-figures-or-the-right-focus" trap, and how do you catch it?
  11. Without looking: what does this lesson say about what a histogram actually draws, and what's actually examined?
  12. Without looking: what does this lesson say about reading a stem-and-leaf diagram — and reading a back-to-back one the same way on both sides?
  13. Without looking: what does this lesson say about comparing two distributions — the discipline that earns the marks?
  14. In one sentence: when you decode a coded variance back to the original scale, why do you divide by b2b^2 rather than by bb?
  15. What is the "mean-denominator-is-class-count-not-total-frequency" trap, and how do you catch it?
  16. What is the "mean-numerator-uses-class-width-instead-of-midpoint" trap, and how do you catch it?
  17. What is the "coding-decode-forgot-to-square-constant" trap, and how do you catch it?
  18. What is the "coding-decode-multiplied-instead-of-divided" trap, and how do you catch it?
  19. What is the "coding-not-decoded-at-all" trap, and how do you catch it?
  20. What is the "malformed-standard-deviation-formula" trap, and how do you catch it?
  21. What is the "mean-rounded-mid-calculation" trap, and how do you catch it?
  22. What is the "forgot-to-square-sd-recovering-sum-of-squares" trap, and how do you catch it?
  23. What is the "interpolation-convention-can-shift-the-class-at-a-boundary" trap, and how do you catch it?
  24. What is the "interpolation-answer-reverse-engineered-to-match-a-given-target" trap, and how do you catch it?
  25. Without looking: what does this lesson say about mean — from a plain list to a grouped table?
  26. Without looking: what does this lesson say about standard deviation — the formula you have to carry in your head?
  27. Without looking: what does this lesson say about coding — trading awkward numbers for easy ones, without changing the real answer?
  28. In one sentence: why does the lower outlier fence have to be built by SUBTRACTING 1.5×IQR1.5 \times IQR from Q1Q_1, never by adding it?
  29. What is the "outlier-lower-fence-direction-reversed" trap, and how do you catch it?
  30. What is the "whisker-drawn-to-fence-not-to-data" trap, and how do you catch it?
  31. What is the "remembered-formula-overrides-the-rule-given-in-the-question" trap, and how do you catch it?
  32. What is the "show-that-outliers-not-listed" trap, and how do you catch it?
  33. What is the "mean-assumed-more-accurate-than-the-median" trap, and how do you catch it?
  34. What is the "comparison-missing-supporting-figures" trap, and how do you catch it?
  35. What is the "comparison-answers-the-wrong-question" trap, and how do you catch it?
  36. What is the "stem-and-leaf-read-in-the-wrong-direction" trap, and how do you catch it?
  37. Without looking: what does this lesson say about the five-number summary, and what a box plot actually draws?
  38. Without looking: what does this lesson say about skewness — read the shape, don't calculate a number?
  39. In one sentence: what has to be true about events AA and BB for the shortcut P(A)/P(B)P(A)/P(B) to actually equal the correct value of P(AB)P(A \mid B)?
  40. What is the "conditional-probability-as-raw-ratio-of-marginals" trap, and how do you catch it?
  41. What is the "conditioning-probability-refolded-into-the-numerator" trap, and how do you catch it?
  42. Without looking: what does this lesson say about elementary probability: sample spaces, and counting equally likely outcomes?
  43. Without looking: what does this lesson say about from words to notation: complement, union, intersection, and the conditioning bar?
  44. In one sentence: why can two events with positive probability never be both mutually exclusive and independent at the same time?
  45. What is the "independence-and-mutual-exclusivity-conflated" trap, and how do you catch it?
  46. What is the "conditional-probability-as-raw-ratio-of-marginals" trap, and how do you catch it?
  47. What is the "independence-assumed-instead-of-tested" trap, and how do you catch it?
  48. What is the "tree-diagram-numerator-missing-a-branch" trap, and how do you catch it?
  49. What is the "extra-probability-folded-into-the-numerator" trap, and how do you catch it?
  50. Without looking: what does this lesson say about sample space, mutually exclusive events, and where the conditional probability formula actually comes from?
  51. Without looking: what does this lesson say about independence — a different kind of question entirely, and one you have to calculate?
  52. In one sentence: why does the region outside both circles on a Venn diagram still need a value written in it, even when that value turns out to be zero?
  53. What is the "original-fraction-reused-after-removal" trap, and how do you catch it?
  54. What is the "second-draw-branches-swapped" trap, and how do you catch it?
  55. What is the "denominator-kept-constant-across-repeated-draws" trap, and how do you catch it?
  56. What is the "venn-blank-region-assumed-zero" trap, and how do you catch it?
  57. What is the "independence-assumed-instead-of-tested" trap, and how do you catch it?
  58. Without looking: what does this lesson say about sampling with and without replacement — what changes between draws?
  59. Without looking: what does this lesson say about venn diagrams: reading regions, and the rule about what "blank" means?
  60. In one sentence: why does rounding a square root to the same number of significant figures the final answer needs — before dividing by it — put that final answer at genuine risk, when rounding it to one extra figure first would not?
  61. What is the "sqrt-omitted-in-pmcc-denominator" trap, and how do you catch it?
  62. What is the "pmcc-recomputed-unnecessarily-after-coding" trap, and how do you catch it?
  63. Without looking: what does this lesson say about recap: sxx and sxy, then the sum regression never needed?
  64. Without looking: what does this lesson say about reading r: use, interpretation, and limitations?
  65. Without looking: what does this lesson say about finishing the derivation — where the sign question actually comes from?
  66. In one sentence: why can two predictions made from the exact same regression line, using the exact same correct arithmetic, have completely different reliability?
  67. What is the "sxx-computed-with-xbar-squared" trap, and how do you catch it?
  68. What is the "gradient-fraction-inverted" trap, and how do you catch it?
  69. What is the "intercept-from-raw-totals-not-means" trap, and how do you catch it?
  70. What is the "stops-after-a-and-b-never-states-the-line" trap, and how do you catch it?
  71. What is the "fraction-left-in-final-line" trap, and how do you catch it?
  72. What is the "gradient-interpreted-at-the-wrong-scale" trap, and how do you catch it?
  73. What is the "interpretation-gives-correlation-not-a-rate" trap, and how do you catch it?
  74. What is the "reliability-judged-without-naming-the-range" trap, and how do you catch it?
  75. Without looking: what does this lesson say about scatter diagrams, and which variable explains which?
  76. Without looking: what does this lesson say about fitting the line: what sxxs_{xx} and sxys_{xy} are, and what the formula booklet already gives you?
  77. In one sentence: why does the fact that XX must take SOME value force p(x)\sum p(x) to equal exactly 1, rather than merely 'roughly' or 'approximately' 1?
  78. What is the "possible-value-left-out-of-domain" trap, and how do you catch it?
  79. What is the "sum-to-1-equation-not-written-down" trap, and how do you catch it?
  80. Without looking: what does this lesson say about the probability function, and the one formula that isn't on the sheet?
  81. Without looking: what does this lesson say about the discrete uniform distribution — thin on this paper, but genuinely simple?
  82. In one sentence: why does adding a constant bb never change Var(X)\text{Var}(X) at all, while multiplying by a constant aa changes it by a factor of a2a^2 rather than just aa?
  83. What is the "coefficient-not-squared-in-variance" trap, and how do you catch it?
  84. What is the "e-x-squared-confused-with-var-x" trap, and how do you catch it?
  85. What is the "formulae-not-known-from-memory" trap, and how do you catch it?
  86. Without looking: what does this lesson say about what's already on the formula sheet, and what you have to know cold?
  87. Without looking: what does this lesson say about the same square, wearing a different costume, in a different part of the paper?
  88. In one sentence: why does P(X > a) always require a subtraction from 1, while P(X < a) never does?
  89. What is the "subtract-from-1-direction-confusion" trap, and how do you catch it?
  90. What is the "rounded-z-value-instead-of-4dp-table-value" trap, and how do you catch it?
  91. What is the "sign-error-reversing-standardisation" trap, and how do you catch it?
  92. What is the "conditional-probability-not-recognised-with-normal" trap, and how do you catch it?
  93. What is the "show-that-standardisation-not-shown" trap, and how do you catch it?
  94. Without looking: what does this lesson say about what standardisation actually does, and what spec 6.1 does not ask for?
  95. Without looking: what does this lesson say about two tables, two different jobs — and the exam expects you to tell them apart?
  96. In one sentence: why does writing "0.1500..." with no working score zero out of three, when the number itself is exactly right?
  97. What is the "answer-only-no-intermediate-step" trap, and how do you catch it?
  98. What is the "reverse-engineered-working" trap, and how do you catch it?
  99. What is the "shown-result-not-fully-stated" trap, and how do you catch it?
  100. What is the "insufficient-accuracy-in-a-show-that" trap, and how do you catch it?
  101. Without looking: what does this lesson say about what "show that" actually asks for — and why it keeps being the same complaint?
  102. Without looking: what does this lesson say about two more places the same discipline bites — briefly, because the mechanism is now familiar?

Beyond the spec

Every item below already lives inside a lesson, labelled the same way there — content the spec doesn’t strictly require, pulled into one place because it’s worth carrying alongside the rest of the sheet, not because it’s tested.

  1. Conditional probability, and independence vs. mutually exclusive

    The conditional probability formula reverses cleanly. Start from the two ways of writing P(AB)P(A \cap B) from the multiplication law: P(AB)=P(B)P(AB)=P(A)P(BA)P(A \cap B) = P(B)P(A \mid B) = P(A)P(B \mid A). Set the middle and right expressions equal — P(B)P(AB)=P(A)P(BA)P(B)P(A \mid B) = P(A)P(B \mid A) — and divide through by P(B)P(B): P(AB)=P(A)P(BA)P(B)P(A \mid B) = \dfrac{P(A)P(B \mid A)}{P(B)}. This is Bayes' theorem: it lets you find P(AB)P(A \mid B) from P(BA)P(B \mid A) — the *reverse* conditional — which matters whenever the reverse direction is the one you can actually measure or are actually given. A concrete shape this takes on a WST01-style question: a factory's three machines produce known proportions of its output, each machine has a different known defect rate — that's P(machine)P(\text{machine}) and P(defectivemachine)P(\text{defective} \mid \text{machine}), both easy to state — but the question asks the reverse, P(a specific machineitem is defective)P(\text{a specific machine} \mid \text{item is defective}), which needs Bayes' theorem to get at directly. The formula booklet's own version of this (verified in WST01-verified-facts.md §2a) is written as a ratio of two products rather than the two-line derivation above, but it is the identical statement — this derivation just shows where that printed ratio actually comes from, the same discipline the mechanism block above applies to the independence rules.

    WST01's own spec content list (WST01-verified-facts.md §1) never names 'Bayes' theorem' anywhere in items 3.1–3.4 — but the real formula booklet does print the full ratio-of-products form of it, in the S1 section, right next to the general multiplication and addition laws (WST01-verified-facts.md §2a). That's worth knowing precisely because it means recognising the shape of a Bayes'-theorem question is useful exam technique even though the name itself is never examined — the tool is sitting on the sheet whether or not you know what to call it.

  2. Correlation coefficient and regression — the calculation mechanics

    Suppose xx and yy are coded as u=xpqu = \dfrac{x-p}{q} and v=ystv = \dfrac{y-s}{t}, for constants p,q,s,tp, q, s, t — exactly the shape spec 4.2's own guidance names ('linear change of variable may be required'). Since uiuˉ=xixˉqu_i - \bar u = \dfrac{x_i - \bar x}{q} for every data point (the shift pp cancels the moment a mean is subtracted, because uˉ=(xˉp)/q\bar u = (\bar x - p)/q too), squaring and summing gives Suu=(uiuˉ)2=1q2(xixˉ)2=Sxxq2S_{uu} = \sum(u_i - \bar u)^2 = \dfrac{1}{q^2}\sum(x_i - \bar x)^2 = \dfrac{S_{xx}}{q^2} — and by identical reasoning, Svv=Syy/t2S_{vv} = S_{yy}/t^2. The cross term picks up both scale factors at once: Suv=(uiuˉ)(vivˉ)=1qt(xixˉ)(yiyˉ)=SxyqtS_{uv} = \sum(u_i - \bar u)(v_i - \bar v) = \dfrac{1}{qt}\sum(x_i - \bar x)(y_i - \bar y) = \dfrac{S_{xy}}{qt}.

    Spec 4.3's own guidance says outright that derivations are 'not required' for this unit — a student can apply 'r is unaffected by (linear) coding,' the exact phrase a real mark scheme credits, without ever seeing why it's true. This is why, and it also surfaces the one genuine subtlety that credited phrase glosses over — a subtlety the reviewed archive gives no sign of the exam actually testing, but one worth understanding rather than trusting blindly.

  3. Regression — gradient interpretation and extrapolation/reliability

    'Least squares' names an actual optimisation: choose aa and bb to make S=(yiabxi)2S = \sum (y_i - a - bx_i)^2 — the total of the squared vertical gaps between each real data point and the line — as small as possible. Calculus finds that minimum by setting both partial derivatives to zero. Sa=2(yiabxi)=0\frac{\partial S}{\partial a} = -2\sum(y_i - a - bx_i) = 0 gives yinabxi=0\sum y_i - na - b\sum x_i = 0, so a=yˉbxˉa = \bar y - b\bar x — the exact intercept formula used throughout this lesson, arrived at without assuming it. Substitute x=xˉx = \bar x into y^=a+bx\hat y = a + bx and this identity forces y^=(yˉbxˉ)+bxˉ=yˉ\hat y = (\bar y - b\bar x) + b\bar x = \bar y: the point (xˉ,yˉ)(\bar x, \bar y) satisfies the line's own equation for ANY value of bb, which is exactly why every least squares line, for every dataset, passes through its own mean point — not a property of this data, a property of the method itself. The second equation, Sb=2xi(yiabxi)=0\frac{\partial S}{\partial b} = -2\sum x_i(y_i - a - bx_i) = 0, gives xiyiaxibxi2=0\sum x_iy_i - a\sum x_i - b\sum x_i^2 = 0; substituting the expression for aa just found and collecting every term containing bb onto one side eventually reduces to b(xi2(xi)2n)=xiyi(xi)(yi)nb\left(\sum x_i^2 - \frac{(\sum x_i)^2}{n}\right) = \sum x_iy_i - \frac{(\sum x_i)(\sum y_i)}{n} — which is exactly bSxx=Sxyb \cdot S_{xx} = S_{xy}, so b=Sxy/Sxxb = S_{xy}/S_{xx}. Nobody chose that formula for its shape; it's the unique answer to 'which line makes the total squared error smallest', and the booklet simply saves every student the calculus needed to arrive at it.

    Spec 4.2's own guidance says outright that derivations are 'not required' for this unit — a student can score full marks using b=Sxy/Sxxb = S_{xy}/S_{xx} and a=yˉbxˉa = \bar y - b\bar x exactly as the formula booklet hands them over, with no idea where either formula comes from. This is where they come from, and it explains, rather than just asserts, a fact used elsewhere in this lesson: why the fitted line always passes through (xˉ,yˉ)(\bar x, \bar y).

  4. The Normal distribution — standardisation, table precision, and conditional probability

    The Normal probability density function is f(x)=1σ2πe(xμ)22σ2f(x) = \dfrac{1}{\sigma\sqrt{2\pi}} e^{-\frac{(x-\mu)^2}{2\sigma^2}} — the formula spec 6.1 explicitly does not ask for. A probability like P(a<X<b)P(a < X < b) is, for any continuous distribution, the area under this curve between aa and bb, which integration computes. Here is the actual reason a table exists instead of a formula: nobody has ever found an elementary closed-form antiderivative for ex2e^{-x^2}. Not because it hasn't been tried — it is one of the most famous non-elementary integrals in mathematics, alongside sinxx\frac{\sin x}{x}. Almost everything else met at this level — polynomials, trig functions, exponentials, and most combinations of them — integrates to another named function. This one provably does not: it can be shown (Liouville's theorem, well beyond this course) that no finite combination of the standard functions differentiates to give it back. The only way to get a number out of it is numerical approximation — compute it once, to high precision, for every value that could plausibly be asked for, and print the results in advance. That is exactly what both Normal-distribution tables in the formula booklet are: a definite integral nobody can write in closed form, evaluated ahead of time so nobody sitting the exam ever has to. It's also why spec 6.1 can honestly say derivation of the mean, variance, and CDF is not required — there is no route to Φ(z)\Phi(z) written on paper that a student could reasonably be asked to reproduce.

    Spec 6.1 explicitly does not require knowledge of the Normal probability density function, or derivation of its mean, variance, or cumulative distribution function (verified guidance, facts bank §1) — full marks on this topic are available while treating Φ(z)\Phi(z) purely as 'the number the table gives you.' This is the one-paragraph answer to a question that a table-only treatment leaves permanently unresolved: why does a table exist for this distribution at all, when every other function met at this level gets a formula instead?

Statistics 1 · condensed sheet · not affiliated with or endorsed by Pearson Edexcel