Exam technique
How marks are actually earned
Every level exemplar, common trap and conditional-judgement drill in this paper, pulled out of the lessons that introduced them and grouped by kind — not held hostage to whichever lesson happened to teach it first.
Common traps — 60
Named failure modes, so you can pattern-match a trap on sight instead of rediscovering it mid-answer.
spec-1.1-expected-as-a-standalone-question
This course's own research pass is direct about what it found (and didn't find) searching all fourteen reviewed WST01 series for the words "model"/"modelling": no standalone question tests spec 1.1 in isolation anywhere in that record — every direct hit turns out to be about a regression model instead (spec 4.1/4.2). This isn't an examiner-report-documented misconception the way most trap-taxonomy items on this course are; it's a genuine, checkable finding about how this content is actually examined, and it produces a real exam-navigation trap in its own right: revising this spec point by looking for a past "define a mathematical model" question to practise on is a search that will come back empty every time, because the AO3 marks this content maps to (15-20 of 75) are distributed as sub-parts riding on top of other questions, not concentrated in a question of their own.
Mathematical modelling in probability and statisticsassumption-treated-as-a-provable-fact
VERIDIAN-original naming for a genuine failure mode this lesson's own material is built to guard against, not a quoted examiner misconception (none exists in the reviewed record for this specific spec point). "The die is fair" and "growth is Normally distributed" are grammatically identical to plain statements of fact, and it's easy to read them that way — as something the question has already established, rather than something it's asking you to accept for the purpose of the calculation that follows. The practical cost: a student who reads an assumption as a fact has nothing to say when a later part of the same question hands them evidence against it (see the marked-solution's part (c) above), because they never registered there was anything provisional about the claim in the first place.
Mathematical modelling in probability and statisticsrefinement-answer-not-connected-to-the-specific-evidence-given
VERIDIAN-original naming, illustrated directly by this lesson's own marked-solution common-wrong-path above. A true-sounding, generically applicable statement — "real spinners are never perfectly fair," "real dice have manufacturing tolerances" — earns nothing on a "critique this model" or "explain what this evidence suggests" part, however scientifically reasonable it sounds, if it isn't tied to the SPECIFIC observation the question actually reported. The credited answer names what was specifically observed (sector 5 over-represented, not the spinner in general) and states the specific consequence for the model (which probability moves, and why the total must still sum to 1) — not a general truth that would apply equally well to any question of this shape.
Mathematical modelling in probability and statisticshistogram-height-read-as-frequency
Confirmed, in the one examiner report reviewed in this research pass that treats a histogram-reading question directly: "the majority of those that were unsuccessful failed to realise that the area of bar represented the frequency and so simply gave frequencies of 25 and 5" (Jan 2023, Q1(a)). Frequency density is a RATE — data points per unit of x-axis — not a count; multiplying by the class width is what turns it into one, and skipping that step is the entire error, not a rounding slip layered on top of an otherwise correct method.
Reading Data Representations, and Comparing Distributions in Wordsstem-and-leaf-read-in-the-wrong-direction
Confirmed: "some students read the stem and leaf diagram the wrong way round and so incorrectly identified Q1 = 31 and Q3 = 51" (Jun 2024, Q1(b)) — swapping which end of the ordered leaves is the smaller quartile and which is the larger. On a back-to-back diagram specifically, this usually comes from reading the LEFT-hand side in plain left-to-right print order instead of recognising it mirrors the right-hand side's own ascending-outward rule (see the teach block above for the mechanism). This same trap is named in the Outliers, Box Plots and Comparing Distributions lesson's own trap-taxonomy, since a misread stem-and-leaf diagram is exactly as capable of feeding a wrong Q1/Q3 into an outlier-fence calculation as it is of feeding one into this lesson's own reading-practice content — it is cited in both places because it is genuinely load-bearing for both skills, not duplicated by accident.
Reading Data Representations, and Comparing Distributions in Wordscomparison-lacks-a-named-statistic-figures-or-the-right-focus
Three separate confirmed patterns, all under the same "compare two distributions in words" guidance point, and all already given full standalone treatment (their own concept-ladder, marked-solution part and trap-taxonomy entries) in the Outliers, Box Plots and Comparing Distributions lesson — restated here only because they are exactly as relevant when the two distributions being compared were read off a histogram or a stem-and-leaf diagram, which is this lesson's own subject. (1) The single most common failure: treating the mean as inherently "more accurate" than the median — "A very common misconception was that the mean is more accurate than the median... showing no appreciation that the mean is just one measure of average and the median is another" (Jun 2022, Q1(e)). (2) Naming no actual figures: "surprisingly too many students failed to give supporting figures... a reference to a named statistic and supporting figures [is required]" (Jun 2024, Q1(e)). (3) Answering a different question from the one asked: "few comments referring to the distribution... others commented on the similarities [when asked for differences]" (Oct 2021, Q3(e)). A comparison that names a statistic, states both figures, and addresses exactly what the question asked for is the only shape of answer that scores against any of the three.
Reading Data Representations, and Comparing Distributions in Wordsmean-denominator-is-class-count-not-total-frequency
Confirmed, and stated by the examiner report as a genuinely disappointing pattern, not a rare one: "it was disappointing to see that some students were unsure on how to calculate a mean from a frequency table. Some added the frequencies and divided by 5 whilst others used the sum of class width multiplied by the frequencies" (Jan 2023, Q1(c)). The denominator of a mean is always Σf — the total number of data values — never the number of rows the table happens to be split across.
Measures of location and dispersion — mean, coding, and standard deviationmean-numerator-uses-class-width-instead-of-midpoint
The second half of the same confirmed finding: some candidates used Σ(class width × frequency) as the numerator instead of Σ(midpoint × frequency). Worth noticing explicitly: for a table where every class has the SAME width (as in the marked-solution example above), this particular wrong shortcut always collapses to exactly the class width itself, no matter what the actual frequencies are — Σ(h×f)/Σf = h when h is constant — which is itself a giveaway that something's gone wrong, since a genuine mean has to depend on where the data actually sits, not just on how the table happens to be split into intervals.
Measures of location and dispersion — mean, coding, and standard deviationcoding-decode-forgot-to-square-constant
Confirmed, real, and directly quoted: candidates "scaling using 0.5 instead of 0.5²" when decoding a coded variance (Jun 2022, Q3(e)). Coding's effect on variance is always squared, because variance is itself built from squared deviations — see the concept-ladder above for the full mechanism, not just the rule.
Measures of location and dispersion — mean, coding, and standard deviationcoding-decode-multiplied-instead-of-divided
A more subtle version of the same real question, quoted directly: candidates "knowing they should use 0.5² but multiplying instead of dividing" (Jun 2022, Q3(e)). Getting the CONSTANT right (squaring it) but the OPERATION wrong moves the variance in the wrong direction entirely — decoding should always make the variance BIGGER than the coded value, since the original data is more spread out on its own (uncoded) scale than the easy numbers coding produced.
Measures of location and dispersion — mean, coding, and standard deviationcoding-not-decoded-at-all
Confirmed, real, and distinct from either error above: the same examiner report records candidates "not decoding" at all — reporting the coded variance, or a value still carrying "the assumed mean of 255" folded in incorrectly, as if it were the answer for the original data. A coded answer is a WORKING TOOL, never the final one; every coded calculation has to be translated back before it answers the question that was actually asked.
Measures of location and dispersion — mean, coding, and standard deviationmalformed-standard-deviation-formula
Confirmed as a real pattern from a genuine examiner report — a formula that isn't shaped like a standard deviation at all, rather than a numerically close slip. WST01-verified-facts.md describes (its own summary, not a Pearson direct quotation) two specific malformed shapes candidates produced in the same series, Jan 2023 Q1(d): one missing the ÷n step before the square root, one missing the square root entirely while also dividing the whole subtraction by n instead of just the first term. Either produces a wildly implausible number — see the marked-solution's commonWrongPath above for exactly how implausible, with real numbers attached. (A related but distinct error — squaring a given standard deviation before using it, rather than misshaping the SD formula itself — is confirmed independently in Jan 2021 Q6(c); see the separate forgot-to-square-sd-recovering-sum-of-squares entry below.)
Measures of location and dispersion — mean, coding, and standard deviationmean-rounded-mid-calculation
Confirmed, and directly quoted: "students should be encouraged to work with exact answers in their working of calculations" (Jan 2023, Q1(d)) — rounding the mean before it feeds into a variance calculation was a documented, real source of inaccuracy on this exact question. Carry the exact value through every intermediate step; round only the final answer, to the paper's own stated 3 significant figures unless told otherwise.
Measures of location and dispersion — mean, coding, and standard deviationforgot-to-square-sd-recovering-sum-of-squares
Confirmed, real, and directly quoted: "forgot to square the standard deviation of 2 when finding Σy² for the second group" (Jan 2021, Q6(c)). Recovering Σx² from a given mean and standard deviation requires rearranging Var = Σx²/n − mean² to Σx² = n(SD² + mean²) — the SD has to be squared into a variance first, the same squaring rule that governs coding above, showing up again in a differently-shaped question.
Measures of location and dispersion — mean, coding, and standard deviationinterpolation-convention-can-shift-the-class-at-a-boundary
Confirmed, real, and directly quoted: "those that worked with n rather than n + 1 were usually more successful" (Jan 2023, Q1(e)). Both conventions for the quartile position — n/4 or (n+1)/4 — are genuinely accepted by the mark scheme; the risk with (n+1)/4 is specifically at a class boundary, where a small extra push past a boundary can land the position inside a DIFFERENT class than n/4 would, changing which numbers the interpolation formula actually uses. See the mechanism block above for exactly this happening with this lesson's own masses data (VERIDIAN's own explanation for why the finding holds, clearly separated there from the verified finding itself).
Measures of location and dispersion — mean, coding, and standard deviationinterpolation-answer-reverse-engineered-to-match-a-given-target
A distinct, real trap on "show that" interpolation questions specifically, confirmed and directly quoted: "the main errors, since the answer was given, were making up two ages... which differed by 16 as their 'working'" (Oct 2021, Q3(b)). This is not a wrong CONVENTION or a wrong FORMULA — it's inventing numbers with the right gap between them instead of genuinely interpolating either quartile. Its full treatment, including why a mark scheme structurally cannot credit this kind of working, lives in this course's "Show that" answer discipline lesson; it's named here because it's specifically an interpolation trap, and this lesson would be incomplete without at least naming it.
Measures of location and dispersion — mean, coding, and standard deviationoutlier-lower-fence-direction-reversed
Confirmed directly, and the research bank's own phrasing makes clear this was not a rare slip: "there were a surprising number of errors... with some multiplying the quartiles by 1.5 and others using Q1 + 1.5×IQR for the lower limit" (Jan 2021, Q2). The lower fence has to SUBTRACT from Q1, moving further below it — the mechanism block above derives why, rather than asking you to remember it as a rule with no reason behind it.
Outliers, Box Plots, and Comparing Distributionswhisker-drawn-to-fence-not-to-data
The research bank's own account of the same Jan 2021 Q2 finding (its documented summary of the examiner report, not a direct quotation) is that box-plot whiskers were commonly drawn out to the outlier boundary itself (98, in that series) rather than to the actual highest non-outlier value in the data (97). A fence is a number used to test values against; it's never itself a point that gets drawn. See the diagram block above for the same error, reproduced with this lesson's own numbers (74.5 versus the real value, 61).
Outliers, Box Plots, and Comparing Distributionsremembered-formula-overrides-the-rule-given-in-the-question
Confirmed: "a small number of candidates opted to apply the more commonly used outlier formula of Q3 + 1.5×(Q3−Q1) rather than the one quoted in the question" (Jun 2022, Q1(b)). This matters because the specification itself states that "any rule to identify outliers will be specified in the question" (spec S1.3, item 2.4, guidance) — there is deliberately no single fixed rule to memorise, and treating 1.5×IQR as a universal constant is itself the error, even on the (common) occasions where the question happens to specify 1.5 anyway.
Outliers, Box Plots, and Comparing Distributionsshow-that-outliers-not-listed
Confirmed, on a "show that there are 3 outliers" question: "having gained the correct limits some did not list the 3 outliers in this part in order to show there are 3 outliers" (Oct 2021, Q3(c)). Finding the correct fences is necessary but not sufficient — a "show that N outliers exist" question needs the N values actually named, not just the machinery that would find them.
Outliers, Box Plots, and Comparing Distributionsmean-assumed-more-accurate-than-the-median
Confirmed, and stated by the examiner report as a genuinely common pattern rather than a rare one: "very few candidates scored this mark... A very common misconception was that the mean is more accurate than the median. A large number of responses simply explained how to calculate a mean or said that it was because the mean is the average, showing no appreciation that the mean is just one measure of average and the median is another" (Jun 2022, Q1(e)). The credited answers were specific and short: "the mean uses all the data" / "the mean includes the outliers." Neither statistic is "more accurate" — they are different measures of the same idea, and the concept-ladder above works through exactly why an outlier moves one and not the other.
Outliers, Box Plots, and Comparing Distributionscomparison-missing-supporting-figures
Confirmed: "surprisingly too many students failed to give supporting figures... a question like this will require some context..., a reference to a named statistic and supporting figures" (Jun 2024, Q1(e)). "Branch A had a higher average" names nothing and gives no numbers; "Branch A had a higher median (34 minutes against 28)" does both, and only the second shape of answer scores.
Outliers, Box Plots, and Comparing Distributionscomparison-answers-the-wrong-question
Confirmed: "few comments referring to the distribution of ages were seen... others commented on the similarities [when asked for differences]" (Oct 2021, Q3(e)). Read what the question is actually asking for — differences, similarities, a specific statistic — before writing the comparison, since a technically-true observation about the wrong aspect of the data scores nothing.
Outliers, Box Plots, and Comparing Distributionsstem-and-leaf-read-in-the-wrong-direction
Confirmed, on a real quartile-from-stem-and-leaf question: "some students read the stem and leaf diagram the wrong way round and so incorrectly identified Q1 = 31 and Q3 = 51" (Jun 2024, Q1(b)) — they swapped which end of the ordered leaves is the lower quartile and which is the upper. Whatever representation a box plot is built from — a stem-and-leaf diagram, a table, a raw list — check which end you're counting from before quoting a quartile out of it.
Outliers, Box Plots, and Comparing Distributionsconditional-probability-as-raw-ratio-of-marginals
Verified directly, quoted from a real WST01 conditional-probability question: "we occasionally still saw P(0.35∩0.4)/0.4 or 0.35/0.4" (Jan 2021, Q1(c)) — dividing two given probabilities directly, as though P(A|B) meant P(A) ÷ P(B) rather than P(A ∩ B) ÷ P(B). The mechanism block above proves this is never a harmless shortcut: it can only match the correct value (when A is entirely contained in B) or overshoot it — never undershoot — which is exactly why it so often lands above 1, or on a plausible-looking but wrong number just below it.
Elementary probability, and conditional-probability notationconditioning-probability-refolded-into-the-numerator
A second, verified instance of conditional-probability notation going wrong, checked directly against the real Pearson examiner-report PDF rather than taken on trust from a summary of it (Jan 2023, Q2(d), a tree-diagram question): candidates were finding P(A|B) from a tree diagram where the correct denominator, P(B) = 61/234, had already been found in an earlier part, and the correct numerator was the single branch product 5/9 × 4/8 × 8/13 = 20/117. The two malformed answers the examiner report actually quotes both use that 61/234 CORRECTLY as the denominator — the error is entirely in the numerator, where an extra copy of the same 61/234 gets folded in: one wrote (5/9 × 4/8 × 8/13 + 61/234)/(61/234), another wrote (5/9 × 4/8 × 8/13 × 61/234)/(61/234). Different arithmetic from the raw-ratio trap above, but the same root confusion: not treating P(A∩B) as one clean, self-contained quantity, separate from whatever's about to divide it — instead letting the denominator's own value leak back into the numerator.
Elementary probability, and conditional-probability notationindependence-and-mutual-exclusivity-conflated
The chronic, named error across the whole archive: "there was the usual confusion between events being independent and events being mutually exclusive highlighted by statements such as 'they are not independent as they overlap'" (Oct 2021, Q1(b)). Pearson's own word "usual" marks this as a recurring pattern, not a single script's slip. The mechanism block above shows exactly where the reasoning goes wrong: NOT overlapping does rule out independence (when both probabilities are positive) — but overlapping on its own proves nothing about independence, which needs the actual P(A∩B) = P(A)P(B) check.
Conditional probability, and independence vs. mutually exclusiveconditional-probability-as-raw-ratio-of-marginals
Verified directly, quoted from a real conditional-probability question: "we occasionally still saw P(0.35∩0.4)/0.4 or 0.35/0.4" (Jan 2021, Q1(c)) — dividing two given probabilities directly, as though P(A|B) meant P(A) ÷ P(B) rather than P(A∩B) ÷ P(B). A fast self-check: if the "conditional probability" you've computed by dividing two marginals exceeds 1, you've made exactly this error — a real probability never can.
Conditional probability, and independence vs. mutually exclusiveindependence-assumed-instead-of-tested
"Candidates often assume independence when an appropriate conditional probability should be used instead" (Jun 2022, general examiner comment — stated as a paper-wide diagnosis, not tied to one question). A second, more specific instance on a Venn-diagram conditional-probability question: "common errors were to assume independence or to do 1 − 1/15 before dividing by 3/8" (Jun 2022, Q4(b)) — independence substituted in as a shortcut for the actual conditioning calculation the question required.
Conditional probability, and independence vs. mutually exclusivetree-diagram-numerator-missing-a-branch
"the most common error was using a product of 2 probabilities rather than 3 in the numerator" (Oct 2021, Q4(c)), on a three-branch tree-diagram conditional probability question. The worked-chain example above is built around this exact trap: forgetting one branch of a multi-stage path produces a wrong but perfectly plausible-looking fraction, with no obvious sign anything went wrong.
Conditional probability, and independence vs. mutually exclusiveextra-probability-folded-into-the-numerator
A second, independently confirmed instance of the numerator/denominator confusion (Jan 2023, Q2(d)) — corrected 2026-08-28 against the real examiner-report PDF, since an earlier reading of this citation had the mechanism backwards: the conditioning denominator itself (61/234, carried from an earlier part) was correct in both wrong answers candidates gave. The actual slip was in the numerator — an extra, spurious copy of that same 61/234 got folded in, either added or multiplied: (5/9 × 4/8 × 8/13 + 61/234)/(61/234) and (5/9 × 4/8 × 8/13 × 61/234)/(61/234), instead of the correct numerator 5/9 × 4/8 × 8/13 = 20/117. The lesson here: once a probability from an earlier part is sitting on the page, it's tempting to reuse it a second time inside a new calculation where it doesn't belong — write P(A|B) = P(A∩B)/P(B) out in full first, work out P(A∩B) as its own clean fraction, and only then substitute, so there's a formula on the page to check the substitution against.
Conditional probability, and independence vs. mutually exclusiveoriginal-fraction-reused-after-removal
Confirmed directly on a real without-replacement counters question: "common errors seen were to repeat the 4/8 (given in the question) following the 1st was green onto the branches following the 1st was blue" (Jan 2023, Q2). The mechanism is not carelessness with arithmetic — the fraction is not wrong in itself, it is simply the wrong bag's fraction, carried onto branches that describe a bag with one fewer counter in it. The fix is procedural, not conceptual: rebuild the composition of what remains before writing a single second-draw fraction, every time.
Sampling with/without replacement, tree diagrams, and Venn diagramssecond-draw-branches-swapped
Confirmed directly on the same real question, immediately after the reused-fraction error above: "on the second branches 5/8 and 3/8 were sometimes given the wrong way round" (Jan 2023, Q2). This is a different mechanism from the reused-fraction error above, not a restatement of it: the bag was correctly reduced to 8 counters, so the right pair of fractions was on the page — but the two fractions were then swapped onto the wrong colour's branch, as if the reducing had been done correctly and then mis-filed. A tree diagram that has been rebuilt with the right numbers can still be wrong if those numbers are not checked against which branch they actually belong to.
Sampling with/without replacement, tree diagrams, and Venn diagramsdenominator-kept-constant-across-repeated-draws
Confirmed on a real without-replacement counting question spanning four draws: "the most common response was scoring 1 mark for the special case using (62/88)⁴" (Jun 2022, Q3(d)) — raising a single fraction to the fourth power, which keeps the denominator fixed at 88 the whole way through, exactly as though every counter were replaced before the next draw. Without replacement, a bag that loses one counter per draw needs a different denominator — and often a different numerator — at every single stage; four genuine draws without replacement need four genuinely different fractions multiplied together, never the same one raised to a power.
Sampling with/without replacement, tree diagrams, and Venn diagramsvenn-blank-region-assumed-zero
Confirmed directly on a real Venn-diagram question: "some candidates left out the 0 in the outside region of N and should be reminded that blank spaces are not assumed to be 0s in Venn diagrams" (Jun 2022, Q4(c)). Even when the correct value for a region genuinely is zero, it has to be written there — an empty space on the diagram reads as unfinished working, not as an answer.
Sampling with/without replacement, tree diagrams, and Venn diagramsindependence-assumed-instead-of-tested
A real examiner report's own general, paper-wide comment states plainly: "Candidates often assume independence when an appropriate conditional probability should be used instead" — named as a pattern across the whole paper, not tied to one question. The same series' report names the concrete version of it on a Venn-diagram question: "common errors were to assume independence" (Jun 2022, Q4(b)). Independence is a conclusion a calculation reaches — comparing with — never a shortcut taken on the way to one; see the common wrong path attached to the marked Venn-diagram solution above for exactly what this costs.
Sampling with/without replacement, tree diagrams, and Venn diagramssqrt-omitted-in-pmcc-denominator
Confirmed on the real question this comes from — Jun 2024 Q4(b), the actual PMCC step of the same question whose part (c) is a 'show that' for the regression line (verified directly against the official Pearson mark scheme, WST01_01_2406_MS, and examiner report, WST01_01_2406_ER, during this lesson's own review, not just against the facts bank's secondhand write-up of it): *'Part (b) was answered well with many students able to calculate a correct value of the product moment correlation coefficient. Common error included the omission of the square root in the denominator.'* The failure isn't an accuracy slip — it's dividing by directly instead of , which doesn't just lose precision, it produces a genuinely different (and, since the one step that keeps bounded between and has been skipped, often an out-of-range) number.
Correlation coefficient and regression — the calculation mechanicspmcc-recomputed-unnecessarily-after-coding
A real mark scheme credits a clean, quotable fact directly: *'r not affected by (linear) coding'* (Jan 2024 MS, Q2(d)). The trap is spending time — and sometimes marks — undoing a coding scheme that never needed undoing: recomputing , and from the original, uncoded and values after already finding from the coded data, on the mistaken assumption that needs 'converting back' the way the regression coefficient genuinely does. scales with the coding constants; doesn't, by construction (see the mechanism and derivation earlier in this lesson) — and a script that recalculates from scratch for the uncoded data isn't doing extra-safe working, it's demonstrating it hasn't understood the property the mark scheme is actually crediting.
Correlation coefficient and regression — the calculation mechanicssxx-computed-with-xbar-squared
Confirmed directly on a real question: *'the most common error made was using x̄² in the calculation of Sxx'* (Oct 2021, Q2(b)). and are not the same number — for this lesson's own dataset they're 36 against 180 — and the two formulas that use them, (wrong) against (correct, and the one printed in the booklet), look similar enough on the page that the substitution can slip past unnoticed.
Regression — gradient interpretation and extrapolation/reliabilitygradient-fraction-inverted
Confirmed on a real regression question: gradient errors came from *'the fraction the wrong way up'* (Oct 2021, Q2(c)) — computing instead of . The fix is naming the formula out loud before substituting: is 'the sum that mixes x and y' over 'the sum that's just x', in that order, every time.
Regression — gradient interpretation and extrapolation/reliabilityintercept-from-raw-totals-not-means
Confirmed on the same question, with the exact numbers the report names: intercept errors came from substituting the raw column totals — the report gives 273 and 93 — straight into in place of the actual means those totals needed to be divided by first to produce (Oct 2021, Q2(c)). and differ by a factor of ; using one where the formula asks for the other doesn't produce a slightly-off answer, it produces one wrong by exactly that factor.
Regression — gradient interpretation and extrapolation/reliabilitystops-after-a-and-b-never-states-the-line
Confirmed on the same question again: candidates who correctly found both and but never wrote them back up as the actual equation of the line lost the final mark of the part for stopping one step early (Oct 2021, Q2(c)). Finding the two numbers is not the same task as answering 'find the equation of the regression line' — the sentence , with the values substituted in, is the actual deliverable.
Regression — gradient interpretation and extrapolation/reliabilityfraction-left-in-final-line
Confirmed directly: *'it is also important to note that fractions are not accepted in a final regression equation; the values of a and b are both estimates and so a fractional answer is not appropriate'* (Jun 2022, Q2(c)). Once has actually been divided out, it stays a decimal (to at least 3 s.f.) for the rest of the question — reverting to the exact fraction it started as, e.g. , in the final line costs the mark even when the value is numerically correct.
Regression — gradient interpretation and extrapolation/reliabilitygradient-interpreted-at-the-wrong-scale
The richest single trap in this topic, confirmed independently in two different subjects' worth of context. One report gives the cost in dollars: the main error 'was not recognising that a single rise in the number of employees led to a rise in the amount spent on paper of $156 and not $1.56' (Oct 2021, Q2(d)) — a hundred-fold misreading of the gradient's own scale. Another shows the same failure from the other direction: 'only the most able candidates successfully managed to write that the GDP increases by 31.2 billion dollars for every 1 million increase in the population' (Jun 2022, Q2(d)) — most answers gave the direction of change and stopped, never converting the raw coefficient into a sentence about what one real unit of the explanatory variable actually buys. Whatever is measured in — dollars, hundreds of dollars, millions of people — 'one unit of x' in the interpretation sentence has to mean that, not '1' read off the page with no units attached.
Regression — gradient interpretation and extrapolation/reliabilityinterpretation-gives-correlation-not-a-rate
Confirmed twice: *'too many students referred to positive correlation or "as x increases then y increases"... some students mixed up the units (grams and °C)'* (Jan 2023, Q6(a)); and separately, *'a few lacked the context required, whilst others gave the interpretation the wrong way round'* (Jun 2024, Q4(d)). Stating that and move together restates something the scatter diagram already showed, and earns nothing on its own — a full-credit sentence names both variables, states the correct-scale rate of change, and gets the direction right, all three, in one sentence.
Regression — gradient interpretation and extrapolation/reliabilityreliability-judged-without-naming-the-range
Confirmed directly: *'it was rare to see responses which accurately assessed the reliability of the estimate found... many stated the estimate was unreliable, but they were unable to refer to the correct variable or value that is not in the range... it was also no surprise to find responses saying that the estimate was reliable, even with correct working earlier in the question [that showed it wasn't]'* (Jun 2022, Q2(e)). The credited form is specific: a separate report records that successful answers 'made reference to 90 being outside of the range and therefore unreliable' (Jan 2023, Q6(c)) — name the value, name the range, connect the two, or the comment doesn't score even when the underlying judgement (reliable / unreliable) happens to be right.
Regression — gradient interpretation and extrapolation/reliabilitypossible-value-left-out-of-domain
The headline trap of this whole spec point, in Pearson's own words: "the most common error was to miss out the value x = 0... had students checked whether the sum of their probabilities equalled 1 they may have realised that they had missed zero out" (Oct 2021, Q4(d)). The defence is the mechanism block above, not repetition: Σp(x) = 1 because X is certain to take SOME value, and a total that falls short of 1 is a direct signal that a genuinely possible value never made it into the working at all — not just an arithmetic slip to hunt for.
Discrete random variables — the probability function and the discrete uniform distributionsum-to-1-equation-not-written-down
The construction-side version of the same idea, confirmed directly: "many candidates scored the 2nd M mark but failed to write down the equation for the sum of probabilities = 1 for the first M mark" (Jun 2022, Q5(d)). This matters specifically on a "show that" item, where the target value is already printed on the page: reaching the correct final number without ever writing the equation that forces it isn't a shortcut, it's a derivation that never actually happened — see the marked-solution's WarrantCheck above for exactly this failure, built around this same real quote.
Discrete random variables — the probability function and the discrete uniform distributioncoefficient-not-squared-in-variance
The headline trap of this whole spec point, confirmed independently at least three times across three years — the same underlying mechanism, never squaring the multiplicative constant when finding Var(aX+b): "others forgot to square [a constant] when subtracting" (Jan 2021, Q4(b–d)); "the most common error seen was a × 4.14 = 66.24 rather than a² × 4.14 = 66.24" (Jun 2024, Q2(c)); "common errors included incorrect use of variance expressions and failure to realise that Var(aX) = a²Var(X)" (Jun 2024, Q3(d) — the same paper, a second independent instance). The defence is the derivation above, not repetition: the square is there because variance is defined as an average of a squared deviation, and a pure scaling of X scales that deviation by a, which the squaring then squares again.
E(aX+b) and Var(aX+b) for discrete random variablese-x-squared-confused-with-var-x
Confirmed on the same question that produced the "forgot to square" trap above: "some still think E(X²) = Var(X)" (Jan 2021, Q4(b–d)). These are two different objects computed from two different formulae — E(X²) = Σx²P(X=x) is a raw second moment; Var(X) = E(X²) − [E(X)]² is that moment with the SQUARE OF THE MEAN subtracted back off. Skipping the subtraction step doesn't just lose a mark on a routine question — reused inside an aX+b question, it feeds the wrong number into everything that follows it.
E(aX+b) and Var(aX+b) for discrete random variablesformulae-not-known-from-memory
"others did not know the formulae for E(aX+b) and Var(aX+b)" (Jan 2021, Q4(b–d)) — and this is worth taking literally rather than as a synonym for the other two traps above. E(aX+b) and Var(aX+b) are explicitly NOT printed anywhere in the Mathematical Formulae and Statistical Tables booklet (verified in WST01-verified-facts.md §2a, against the spec's own statement that formulae "students are expected to know... will not appear in the booklet"). A student who forgets the base formulae for E(X) and Var(X) can look them up; a student who forgets the aX+b shortcuts has nowhere in the exam room to check.
E(aX+b) and Var(aX+b) for discrete random variablessubtract-from-1-direction-confusion
Confirmed in three separate series, failing in BOTH directions — the same underlying gap producing opposite mistakes depending on the question. Under-subtracting: "a significant number of students lost 2 marks as they failed to subtract from one the value obtained from the normal tables. A simple diagram would have helped many to avoid this error" (Jan 2023, Q5(a)). Over-subtracting: "a few lost the final mark as they went on to subtract the correct answer from 1, which of course is P(X > 18) and not P(X < 18) which is what was required" (Jun 2024, Q5(a)). A third series confirms the skill is fragile even when it goes right: "most students standardising correctly and the majority realising that they then needed to subtract the value found in the tables from 1" (Jan 2021, Q3(a)) — implying, in Pearson's own words, that a real minority did not. The fix in every case is the same: sketch which side of z the question describes before opening the table at all.
The Normal distribution — standardisation, table precision, and conditional probabilityrounded-z-value-instead-of-4dp-table-value
The exact same mechanism, confirmed in three independent series, each with its own quoted wrong value: "many students were using a z value of 1.03 or 1.04 rather than the value 1.0364 from the 'Percentage Points of the Normal Distribution' table" (Jan 2021, Q3(b)); "not using an inaccurate value such as 1.64 instead of 1.6449" (Oct 2021, Q6(b)); "the most common error included the use of an inaccurate z value... students should be reminded that when values are required from the tables, they need to be 4 decimal places. A common error was to use z value = 0.25" (Jun 2024, Q5(b)). This is not carelessness in isolation — it is exam technique stated verbatim on every WST01 question-paper front page: "Values from the statistical tables should be quoted in full" (verified, facts bank §2).
The Normal distribution — standardisation, table precision, and conditional probabilitysign-error-reversing-standardisation
Confirmed in two series, once in each direction of the sign: "others gained this mark but were using −1.0364, an error that could probably have been avoided if they had drawn a suitable diagram" (Jan 2021, Q3(b)) — where the value needed was positive; "using the wrong sign for 1.6449 appropriate to their standardisation giving 34.1, a value higher than the upper limit" (Oct 2021, Q6(b)) — where an unflipped sign produced an answer that was, on inspection, impossible. That second detail is the real lesson: the wrong answer was checkable as wrong using nothing but the question's own numbers, and the check was skipped anyway.
The Normal distribution — standardisation, table precision, and conditional probabilityconditional-probability-not-recognised-with-normal
The most consistently documented trap in the whole facts bank — present, in a different concrete form, in every one of the 5 examiner reports reviewed. "Many did not realise that a conditional probability was required... a common error P(W<18)/0.85... but there were a number of correct attempts of the form (0.85−0.5)/0.85 which usually led to the correct answer" (Jan 2021, Q3(c)). "Many students did not realise that the ratios only applied to the middle 80% of the data" (Oct 2021, Q6(c)). "Like question 2, a common error was that students failed to realise that a conditional probability was required. A common error was to find P(L≤5) and go no further" (Jan 2023, Q5(e)). Stated as a paper-wide diagnosis, not a topic-specific aside: "Candidates often assume independence when an appropriate conditional probability should be used instead" (Jun 2022, general comment). One series' report names this sub-skill, in these words, as the hardest thing on the entire paper: a Jun 2022 question is flagged as "the final part of the paper... also the most discriminating part."
The Normal distribution — standardisation, table precision, and conditional probabilityshow-that-standardisation-not-shown
A cross-cutting instruction-following failure, not a maths error, confirmed directly on a Normal-distribution question: "many students wrote down the standardisation followed by 0.15 missing out the intermediate step and the accurate answer. Some students simply wrote down 0.1500… with no working at all" (Oct 2021, Q6(a)). Stated as a general rule in a later series: "if asked to use standardisation then the standardisation should be shown" (Jun 2024, general introduction). This costs marks even when the final decimal is completely correct — a 'show that' mark scheme has nothing to attach a mark to if the standardisation step itself never appears on the page.
The Normal distribution — standardisation, table precision, and conditional probabilityanswer-only-no-intermediate-step
The single most literal version of the trap, on a real Normal-distribution "show that" question (Oct 2021, Q6(a)): "many students wrote down the standardisation followed by 0.15 missing out the intermediate step and the accurate answer. Some students simply wrote down 0.1500… with no working at all." Two separate failures are named in that one quote — some students showed the standardisation but skipped the accurate table value in between; others showed nothing at all and just wrote the given probability. Both lose marks, because both leave the examiner unable to distinguish "did the work, wrote it up badly" from "copied the given answer."
"Show that" answer disciplinereverse-engineered-working
A more subtle version, on a real interquartile-range "show that" question (Oct 2021, Q3(b)): "the main errors, since the answer was given, were making up two ages... which differed by 16 as their 'working'." This is worse than showing nothing, not better — it produces two numbers that are correct only in that they happen to be 16 apart, with no interpolation, no cumulative frequency, no class boundary in sight. A "working" that was clearly reverse-engineered from the given answer, rather than derived from the data, does not read as method at all.
"Show that" answer disciplineshown-result-not-fully-stated
Confirmed on a "show that there are 3 outliers" question (Oct 2021, Q3(c)): "having gained the correct limits some did not list the 3 outliers in this part in order to show there are 3 outliers." The calculation (the outlier boundaries) was right; what was missing was the last, cheapest step — actually pointing at the three data values that satisfy it. When the question asks you to show a count, a stated boundary is not the same as a demonstrated count, and the gap between them is one line of writing.
"Show that" answer disciplineinsufficient-accuracy-in-a-show-that
A precision-specific version, on a real regression-line "show that" question (Jun 2024, Q4(c) — the actual target line was , 3 s.f.): "Many students were able to show the given regression line but too often students lost the final A mark as they failed to give values to the required degree of accuracy. b = 12105.12/16769.78 = 0.722 was not accurate enough to gain the final mark due to it being a 'show that' question. Students should be encouraged in these types of questions to give answers to at least one more decimal place than the given value." The candidate's 0.722 was correct — 12105.12 ÷ 16769.78 really does round to 0.722 — but stopping the working at the same 3 significant figures as the printed line cannot prove the calculation reached it rather than having been rounded to match; the extra digit (0.7218…) is the only thing that can.
"Show that" answer discipline