"Show that" answer discipline

~30 min · WST01 · cross-cutting

WST01 · cross-cutting · 30 min

A "show that" question already gives you the answer. That is not a hint about how easy the question is — it is a warning about how the mark scheme has to be built. When the final number is printed on the page, an examiner cannot give you credit for reaching it, because copying a printed number proves nothing about whether you can produce it yourself. Every mark in a "show that" question is attached to the *steps between* the start of the question and the number you were already handed — and Pearson has said, in three separate examiner reports spanning three years, that this is the single thing students most often get wrong on this paper, in every topic it touches.

What "show that" actually asks for — and why it keeps being the same complaint

A "show that" question states the answer up front — "show that the interquartile range is 16.25", "show that P(X > 196) = 0.15" — and asks you to produce it. That phrasing changes what is being marked. On a normal question, the mark scheme rewards a correct final value (an A mark, "correct answer only"). On a "show that" question, the final value is already correct by construction — you were told what it is — so the marks that would normally sit on the final answer have nowhere to attach except the working that gets you there. A "show that" question is not an easier version of the same question; it is a request for the method, with the answer supplied only so you can check your own working against it.

This is not a one-off marking quirk. It is Pearson's own stated diagnosis of the paper's single biggest non-content problem, repeated in near-identical language across three examiner reports that are not consecutive series — meaning the pattern was still present roughly three years after it was first flagged, not a passing issue that got fixed:

> "A general comment, which applies to other questions as well, is that many students do not really understand that evidence has to be given for 'show that' questions." (Oct 2021, general comment)

> "Students need to understand that a full method should be given for 'show that' questions." (Jan 2023, general comment)

> "Students should be advised to read questions carefully as marks were dropped carelessly for not meeting the required demands. In particular questions that ask a student to show something is true require all the steps in the working to be shown." (Jun 2024, general comment)

These are general, opening-paragraph comments in each report — meaning Pearson chose to say this once, about the whole paper, rather than only under the individual questions where it happened to bite. The rest of this lesson is the specific, quoted evidence for why: real instances of the same underlying failure recurring in the summary statistics section, the correlation section, the discrete random variables section, and the Normal distribution section — four different topics, one shared mistake.

Mechanism

Why the mark scheme genuinely cannot pay for the printed answer

This follows directly from how M, A and B marks work, not from an arbitrary examiner preference (WST01-verified-facts.md §4; this exact "General Instructions for Marking" wording opens at least 6 of the 14 reviewed mark schemes, byte-for-byte — confirmed directly against the real Jan 2024 mark scheme PDF; Jun 2024's own preamble is shorter but keeps the same M-before-A dependency rule). An M mark is "given for a correct method or an attempt at a correct method" — it has to be visible to be given, because there is nothing else for the examiner to assess it against. An A mark is "dependent... and can only be awarded if the previous M mark has been earned. E.g. M0 A1 is impossible." On an ordinary question this dependency is nearly invisible, because a correct final answer is itself strong evidence a correct method was used. On a "show that" question that evidence is destroyed on purpose: the final answer is printed in the question, so a correct final answer is evidence of nothing — it is equally consistent with genuine working and with simply copying what was already given. The mark scheme's only defence is to refuse to award the A mark unless the M mark — the actual method, shown on the page — is there first. This is also why "the answer is printed on the paper" has its own dedicated symbol in the mark schemes' general abbreviations list: it is common enough, specifically because "show that" questions are common enough, to need its own shorthand. The Statistics-specific marking note reinforces the same point from the examiner's side: "Any correct method should gain credit. If you cannot see how to apply the mark scheme but believe the method to be correct then please send to review" — the examiner is instructed to credit what they can see, not what they can infer might have happened off the page.

Named traps

answer-only-no-intermediate-step
The single most literal version of the trap, on a real Normal-distribution "show that" question (Oct 2021, Q6(a)): "many students wrote down the standardisation followed by 0.15 missing out the intermediate step and the accurate answer. Some students simply wrote down 0.1500… with no working at all." Two separate failures are named in that one quote — some students showed the standardisation but skipped the accurate table value in between; others showed nothing at all and just wrote the given probability. Both lose marks, because both leave the examiner unable to distinguish "did the work, wrote it up badly" from "copied the given answer."
reverse-engineered-working
A more subtle version, on a real interquartile-range "show that" question (Oct 2021, Q3(b)): "the main errors, since the answer was given, were making up two ages... which differed by 16 as their 'working'." This is worse than showing nothing, not better — it produces two numbers that are correct only in that they happen to be 16 apart, with no interpolation, no cumulative frequency, no class boundary in sight. A "working" that was clearly reverse-engineered from the given answer, rather than derived from the data, does not read as method at all.
shown-result-not-fully-stated
Confirmed on a "show that there are 3 outliers" question (Oct 2021, Q3(c)): "having gained the correct limits some did not list the 3 outliers in this part in order to show there are 3 outliers." The calculation (the outlier boundaries) was right; what was missing was the last, cheapest step — actually pointing at the three data values that satisfy it. When the question asks you to show a count, a stated boundary is not the same as a demonstrated count, and the gap between them is one line of writing.
insufficient-accuracy-in-a-show-that
A precision-specific version, on a real regression-line "show that" question (Jun 2024, Q4(c) — the actual target line was g=42.3+0.722dg = -42.3 + 0.722d, 3 s.f.): "Many students were able to show the given regression line but too often students lost the final A mark as they failed to give values to the required degree of accuracy. b = 12105.12/16769.78 = 0.722 was not accurate enough to gain the final mark due to it being a 'show that' question. Students should be encouraged in these types of questions to give answers to at least one more decimal place than the given value." The candidate's 0.722 was correct — 12105.12 ÷ 16769.78 really does round to 0.722 — but stopping the working at the same 3 significant figures as the printed line cannot prove the calculation reached it rather than having been rounded to match; the extra digit (0.7218…) is the only thing that can.

Marked, line by line

The heights, X cm, of adult males in a population are modelled by X ~ N(170, 625). Show that P(X > 196) = 0.15, correct to 2 significant figures. (3) — VERIDIAN-original question and dataset (μ = 170, σ = 25 built specifically to give a clean standardised value of z = 1.04). Built to demonstrate the exact trap named in the Oct 2021 examiner report for this question type (Q6(a), quoted above) — not a reproduction of any past-paper question or dataset.

3 marks available

  1. 01

    Z=X17025Z = \dfrac{X - 170}{25}, so P(X>196)=P(Z>19617025)=P(Z>1.04)P(X > 196) = P\left(Z > \dfrac{196 - 170}{25}\right) = P(Z > 1.04).

    Method mark for standardising the given value using Z = (X − μ)/σ. This is the exact line Pearson names as going missing on this question type — its own general comment on a later series states plainly: "If asked to use standardisation then the standardisation should be shown" (Jun 2024, general comment). On a "show that" question the standardised value 1.04 has to be visible on the page even though the final probability is already given to you.

    M1
  2. 02

    Using the Normal Distribution Function table, Φ(1.04)=0.8508\Phi(1.04) = 0.8508, so P(Z>1.04)=10.8508=0.1492P(Z > 1.04) = 1 - 0.8508 = 0.1492.

    Accuracy mark for the table value quoted in full (4 decimal places, per the front-page rounding instruction: "Values from the statistical tables should be quoted in full") and the correct 1 − Φ step. This is the "intermediate step and the accurate answer" the Oct 2021 examiner report specifically names as missing when a student skips straight to the given probability instead of showing where it came from.

    A1
  3. 03

    0.1492=0.150.1492 = 0.15 (2 s.f.), as required.

    cso — correct solution only. Because 0.15 is already printed in the question, this mark is not for writing down 0.15 — it is for showing that your own calculation, carried through to 0.1492, is what that printed value rounds from. Without the 0.1492 above, this line has nothing to attach to.

    A1

In your own words

In one sentence: why does writing "0.1500..." with no working score zero out of three, when the number itself is exactly right?

Worked, in full

Show that the interquartile range of the volunteers' ages is 16.25 — and why two ages that happen to be 16.25 apart isn't the same thing

  1. 01

    The ages of 50 volunteers are grouped as: 10–20 years, 5 people; 20–30 years, 12 people; 30–40 years, 18 people; 40–50 years, 10 people; 50–60 years, 5 people. (VERIDIAN-original data, built specifically so both quartiles fall mid-class and interpolate to clean values — see the flag block at the end of this lesson.) Before touching a quartile formula, build the cumulative frequencies: 5, 17, 35, 45, 50. Every interpolation step below depends on having these right first.

    Earns: B1 — correct cumulative frequency table. This line has no algebra in it at all, but skipping it in your working is exactly the kind of missing intermediate step Pearson's examiner reports repeatedly flag: the class each quartile falls in cannot be checked without it being shown.

  2. 02

    Locate and interpolate Q1Q_1. Position n/4=50/4=12.5n/4 = 50/4 = 12.5. The cumulative frequency reaches 5 at the end of the 10–20 class and 17 at the end of the 20–30 class, so position 12.5 falls inside the 20–30 class. Interpolating: Q1=20+12.5512×10=20+6.25=26.25Q_1 = 20 + \dfrac{12.5 - 5}{12} \times 10 = 20 + 6.25 = 26.25.

    Earns: M1 A1 — correct method (identify the class from the cumulative frequencies, then interpolate across it) and correct value. This is the step a real examiner report names as the actual skill being tested on this question type, separately from the final subtraction: reading which class a quartile position falls into before applying the interpolation formula to it.

  3. 03

    Locate and interpolate Q3Q_3 the same way. Position 3n/4=37.53n/4 = 37.5. The cumulative frequency reaches 35 at the end of the 30–40 class and 45 at the end of the 40–50 class, so position 37.5 falls inside the 40–50 class. Interpolating: Q3=40+37.53510×10=40+2.5=42.5Q_3 = 40 + \dfrac{37.5 - 35}{10} \times 10 = 40 + 2.5 = 42.5.

    Earns: M1 A1 — the same method applied to the upper quartile. Both quartiles need their own full interpolation line; one correct interpolation does not imply the other, and a script that shows only one of the two has only demonstrated half the required method.

  4. 04

    IQR=Q3Q1=42.526.25=16.25IQR = Q_3 - Q_1 = 42.5 - 26.25 = 16.25, as given.

    Earns: A1 cso — correct solution only. Because 16.25 is printed in the question, this mark is for reaching it via the two interpolated values shown above, not for writing 16.25 itself. A script that wrote any two numbers 16.25 apart without stages 2–3 behind them would not earn this mark, however confidently the subtraction is presented.

Source — Examiner report, Oct 2021

"the main errors, since the answer was given, were making up two ages... which differed by 16 as their 'working'"

Two more places the same discipline bites — briefly, because the mechanism is now familiar

The same failure recurs wherever a question supplies the final value in advance, not just in the two topics worked through above. On a regression "show that" question the printed coefficients are usually stated to 3 significant figures — the real Jun 2024 Q4(c) target was g=42.3+0.722dg = -42.3 + 0.722d — and the same examiner report is explicit that carrying your own working to only that same precision is not safe practice: writing b=Sxy/Sxx=12105.12÷16769.78=0.722b = S_{xy}/S_{xx} = 12105.12 \div 16769.78 = 0.722 as your line of working was judged not accurate enough to earn the final mark, precisely because it stops at the same precision as the printed target rather than the one extra decimal place (b=0.7218b = 0.7218) that would prove the calculation, not just the given line, produced it.

The same discipline extends past questions literally phrased as "show that." A real, fully-worked discrete random variable question (WST01-verified-facts.md §5, Jan 2025 Q1(d)) instructs a candidate to find Var(B)\text{Var}(B) "showing your working" — the mark scheme's own wording, not a paraphrase — even though nothing in that part is printed in advance for the candidate to reverse-engineer toward. The underlying expectation is the same one this whole lesson has been building: a numeric answer with no visible method behind it is worth less than the mark scheme allocates to it, whether or not that answer happens to be correct, and whether or not the question used the exact words "show that."

Retrieval — with feedback on every choice

Question 1
3 marks

X ~ N(170, 625). Show that P(X > 196) = 0.15, correct to 2 significant figures. (3 marks)

Which response below is certain to score all 3 marks?

Question 2
2 marks

Show that the interquartile range of the volunteers' ages (grouped data, n = 50) is 16.25.

A script states Q1=26Q_1 = 26 and Q3=42.25Q_3 = 42.25, giving IQR =16.25= 16.25, with no working shown for either quartile. What is the safest description of what this earns?

Question 3
2 marks

Having correctly calculated the outlier boundaries as 8 and 52, a question asks you to "show that there are 3 outliers" in a dataset. What must your answer additionally include to be safe?

Question 4
1 mark

A real WST01 question gives Sxy=12105.12S_{xy} = 12105.12 and Sxx=16769.78S_{xx} = 16769.78, and asks you to show that the regression line is g=42.3+0.722dg = -42.3 + 0.722d (3 s.f.) — the actual Jun 2024 Q4(c).

What is the safest working practice for how many decimal places to carry the gradient b=Sxy/Sxxb = S_{xy}/S_{xx} to, before rounding down to the printed line?

Reference — not a study method, a lookup
  • "Show that X" means the answer is given — the marks are for the steps to X, not for X itself.
  • Never write the target value first. Show the method, and let the target value appear as your last line.
  • Work to one more d.p./s.f. than the target throughout — matching its precision is how the last mark is lost.
  • "Show that there are N of something" needs the N things named, not the rule that would find them.
  • M0 A1 is impossible: no visible method means the accuracy marks have nothing to attach to, whatever the final line says.

Not affiliated with or endorsed by Pearson Edexcel. Every quotation attributed to an examiner report in this lesson (the three general comments; the Oct 2021 Q6(a), Q3(b) and Q3(c) quotes; the Jun 2024 Q4(c) quote; the Jun 2024 general comment; the "showing your working" mark-scheme wording from Jan 2025 Q1(d)) is transcribed verbatim from WST01-verified-facts.md, itself independently checked against the primary Pearson PDFs — none of it is carried over from prior course material, since no prior WST01 material exists in this repo. The regression-line numbers built into the trap taxonomy, the teach block and the fourth MCQ (Sxy = 12105.12, Sxx = 16769.78, gradient 0.7218… rounding to the printed 0.722, intercept −42.3) are the real values from that same Jun 2024 Q4(c) — not VERIDIAN-original — re-verified directly against the official Pearson mark scheme (WST01_01_2406_MS) and examiner report (WST01_01_2406_ER) PDFs during this lesson's own review, not only against WST01-verified-facts.md's secondhand transcription of them. The Normal-distribution scenario (μ = 170, σ = 25, threshold 196) and the ages/IQR grouped-frequency dataset (5, 12, 18, 10, 5 across five ten-year classes) are both VERIDIAN-original — built to demonstrate the exact documented trap for each topic, not reproductions of any real Pearson dataset, and independently checked by direct calculation (Φ(1.04) = 0.8508 via the error function; the two interpolations and their subtraction) before being written into this lesson. The per-line M/A/B mark allocations attached to both worked examples are modelled on the verified mark-scheme conventions in WST01-verified-facts.md §4 (what M, A, B and cso mean; that M0 A1 is impossible), not transcribed from a real mark scheme, which for an original question does not exist.

Question 13 marks

X ~ N(170, 625). Show that P(X > 196) = 0.15, correct to 2 significant figures. (3 marks)

Which response below is certain to score all 3 marks?

  • Standardise to get z=1.04z = 1.04, look up Φ(1.04)=0.8508\Phi(1.04) = 0.8508 from the tables, subtract from 1 to get 0.14920.1492, then state 0.1492=0.150.1492 = 0.15 (2 s.f.)

    Correct. All three mark-scheme steps are visible in order: the standardisation (M1), the accurate table value with the 1 − Φ step (A1), and the connection between that accurate value and the printed target (A1 cso). Nothing here relies on the examiner inferring work that isn't shown.

  • B"P(X>196)=0.15P(X > 196) = 0.15, as given in the question."

    This is the exact response a real examiner report records candidates giving on this question type, verbatim: some students "simply wrote down 0.1500… with no working at all." It scores zero, not partial credit, because there is no method mark to attach the dependent accuracy marks to — the answer being numerically correct is irrelevant when nothing demonstrates how it was reached.

  • CStandardise to z=1.04z = 1.04 correctly, then write "P(Z>1.04)=0.15P(Z > 1.04) = 0.15" straight away with no table value shown

    This earns the method mark for the standardisation, but not the accuracy mark that follows it — the accurate table value (0.8508, or the resulting 0.1492) never appears, only the given target repeated. This is precisely the 'missing out the intermediate step and the accurate answer' failure the Oct 2021 report names, distinct from writing nothing at all.

  • DUse a calculator to find P(X>196)0.149P(X > 196) \approx 0.149 directly, and write only that decimal

    A calculator route can be entirely correct mathematically and still fail a 'show that' question, because the mark scheme is built around the standardisation and table-lookup steps specifically, not just around producing a numerically close decimal. Without the standardised value z=1.04z = 1.04 shown, there is no method visible for the examiner to credit.

Traps tested: Answer only no intermediate step

Question 22 marks

Show that the interquartile range of the volunteers' ages (grouped data, n = 50) is 16.25.

A script states Q1=26Q_1 = 26 and Q3=42.25Q_3 = 42.25, giving IQR =16.25= 16.25, with no working shown for either quartile. What is the safest description of what this earns?

  • Little or nothing — the subtraction is correct but nothing demonstrates the two quartile values were actually interpolated rather than chosen to be 16.25 apart

    Correct. This is the reverse-engineering trap a real examiner report names directly: "the main errors, since the answer was given, were making up two ages... which differed by 16 as their 'working'." Both quartiles here are also individually wrong (26.25 and 42.5 are the correct values) in a way genuine interpolation working would have caught — which is itself evidence the numbers were picked to match the gap, not derived from the data.

  • BFull marks — the final answer matches the value given in the question

    This is exactly the assumption the whole lesson has been arguing against: on a 'show that' question, matching the given value is necessary but proves nothing on its own, because the value was never in doubt. The marks are attached to the interpolation method, which is entirely absent here.

  • CFull marks for Q1 and Q3 individually, since the method for finding a quartile doesn't need to be shown, only the final IQR

    The interpolation method is exactly what the mark scheme is testing at this stage of the paper (2.3, 'simple interpolation may be required') — a quartile value with no cumulative frequency and no interpolation line behind it is an unsupported assertion, not a scored calculation.

  • DIt cannot be assessed without knowing whether the student used the n or the n+1 convention for the quartile positions

    That convention choice matters for a genuinely worked interpolation (and the mark scheme tolerates either), but it is irrelevant here — there is no interpolation shown to apply either convention to. The problem with this script isn't which formula was used; it's that no formula was used at all.

Traps tested: Reverse engineered working · Answer only no intermediate step · Irrelevant technicality substituted for the real issue

Question 32 marks

Having correctly calculated the outlier boundaries as 8 and 52, a question asks you to "show that there are 3 outliers" in a dataset. What must your answer additionally include to be safe?

  • The three specific data values that lie outside 8 and 52, named explicitly

    Correct. A real examiner report on exactly this question type records the gap directly: "having gained the correct limits some did not list the 3 outliers in this part in order to show there are 3 outliers." The boundaries alone show where an outlier WOULD be; they do not show that exactly three data points actually are.

  • BNothing further — stating the correct boundaries already proves there are exactly 3 outliers

    The boundaries define the rule, not the count. Demonstrating that exactly three data points satisfy that rule requires checking the data against it, which is a separate, required step — and it's specifically the step the examiner report records candidates skipping.

  • CA recalculation of the boundaries using a different outlier rule, to double-check the count

    The spec is explicit that any outlier rule used will be given in the question — inventing or substituting a second rule is not what's being asked for, and risks the separate, also-documented trap of applying a remembered generic rule instead of the one the question specified.

  • DA box plot showing the outliers as individual points beyond the whiskers

    A box plot can illustrate outliers but this question asks you to 'show that there are 3' as a written justification, not to draw one. Drawing a diagram doesn't substitute for stating which three values were checked against the boundaries and found to lie outside them.

Traps tested: Shown result not fully stated · Unrequested second method substituted · Diagram substituted for required written justification

Question 41 mark

A real WST01 question gives Sxy=12105.12S_{xy} = 12105.12 and Sxx=16769.78S_{xx} = 16769.78, and asks you to show that the regression line is g=42.3+0.722dg = -42.3 + 0.722d (3 s.f.) — the actual Jun 2024 Q4(c).

What is the safest working practice for how many decimal places to carry the gradient b=Sxy/Sxxb = S_{xy}/S_{xx} to, before rounding down to the printed line?

  • Carry bb to at least one more decimal place than the 3 s.f. printed in the line — work with b=0.7218b = 0.7218, not 0.7220.722 — then round only at the very last step

    Correct, and it is Pearson's own stated advice on this real question: "Students should be encouraged in these types of questions to give answers to at least one more decimal place than the given value." Rounding the gradient to 3 s.f. before it feeds into the intercept calculation is exactly what cost marks here.

  • BStop at b=12105.12÷16769.78=0.722b = 12105.12 \div 16769.78 = 0.722, matching the 3 s.f. already printed in the line

    This is the exact working the examiner report records as insufficient on this real question: "b = 12105.12/16769.78 = 0.722 was not accurate enough to gain the final mark due to it being a 'show that' question." The 0.722 itself is correct — the working simply doesn't carry enough precision to prove it was calculated rather than copied from the given line.

  • CIt doesn't matter, since the mark scheme only checks the final printed line against the question

    It matters specifically because the final line IS being checked against a printed target — the whole reason a 'show that' question is harder to get marks on than an ordinary one is that the examiner can see exactly how close your working came, not just whether you wrote the right final digits.

  • DAs few decimal places as possible, to keep the working short and reduce the chance of a transcription error

    This trades one risk for a larger one. A transcription error in a longer decimal is a real but recoverable risk; rounding too early on a 'show that' question is the documented, repeated way marks are actually lost on this question type.

Traps tested: Insufficient accuracy in a show that · Answer only no intermediate step

Practice this for real

This site teaches the mechanism; the exam is sat on Pearson's own real questions. Go find and attempt these yourself — nothing here substitutes for actually sitting a timed paper.

Examiner report
Oct 2021 · Q3(b) — cited directly in this lesson
Pearson's official past-papers portal

Select International Advanced Level → Mathematics → any series, then look for WST01.

Statistics 1 · progress saved in this browser · sign in to sync across devices

That’s the end of Statistics 1.

You've finished the reading order. 13 lessons left unmarked — worth a pass before you call it done.

Back to the contents