Reading Data Representations, and Comparing Distributions in Words

~50 min · WST01 · 2.1

WST01 · 2.1 · 50 min

A histogram's bars can lie to you at a glance. The vertical axis reads frequency DENSITY, not frequency — and the two are only the same number when every class happens to be the same width, which a real exam question is under no obligation to give you. The single most repeated error the mark schemes for this section record is exactly this: reading a bar's height straight off the axis and writing it down as the answer, skipping the one multiplication — by class width — that turns a density into an actual count of people. A stem-and-leaf diagram hides a quieter version of the same problem: read it from the wrong end, and Q1 and Q3 swap places without anything on the page telling you they have. And once you've read the real figures off either diagram, a "compare the two distributions" question is asking you to use them — name a statistic, give both figures, answer what was actually asked — not just to have found them.

Before you read on

Two or three questions on exactly what this lesson teaches. Being wrong here is fine — it's the fastest way to find out what to pay attention to next.

What a histogram actually draws, and what's actually examined

Spec 2.1 groups histograms with stem-and-leaf diagrams and box plots under "representation and summary of data," and the guidance attached to it draws a firm line: "drawing of histograms/stem-and-leaf/box plots will not be the direct focus of exam questions — reading/using them will be." A histogram's vertical axis is labelled frequency DENSITY, not frequency — and the two coincide only when every class happens to share the same width. A real exam question is never obliged to give you equal widths, and reading the diagram correctly, not drawing one, is the actual skill under test.

The same guidance point adds one more instruction worth taking literally: "back-to-back stem and leaf diagrams may be required" — meaning a single diagram can present two data sets side by side, sharing one column of stems, with each data set's leaves running outward from the stem on its own side. Reading a histogram correctly and reading a back-to-back stem-and-leaf diagram correctly turn out to be the same underlying skill: knowing precisely what a diagram's shape is standing for, rather than reading a number straight off an axis or a printed row and assuming it's already the answer.

Why area, not height, is frequency

In plain terms

Imagine four tubs of different widths standing out in the rain, and after the storm you check how DEEP the water sits in each one — not how much water it actually holds. A wide, shallow tub can hold just as much water as a narrow, deep one; the depth alone doesn't tell you the volume, because the tub's width matters too. A histogram bar works the same way. Its height only tells you how "crowded" that stretch of the x-axis is, per unit of width — not how many data points actually fall inside it. To get the real count, you have to account for how WIDE the bar is, the same way you'd need a tub's width before its water depth told you anything about the volume inside it.

The height of a histogram bar is a RATE, not a count — it's called the frequency density, and it measures "how many data points per unit of x-axis, in this stretch." Two bars can have the same height and still represent wildly different numbers of data points, if their widths differ: a bar of height 4 and width 2 represents 8 data points; a bar of height 4 and width 10 represents 40. The quantity that actually equals the frequency is the AREA of the bar — height multiplied by width — which is exactly why every class on a histogram is drawn so that its area, not its height, stands for the count of data inside it.

Formally

Formally, frequency density is defined as fd=fwfd = \dfrac{f}{w}, where ff is the class frequency and ww is the class width — rearranged, f=fd×wf = fd \times w, the area of the bar. This definition exists specifically so that classes of DIFFERENT widths remain visually comparable: without dividing by width first, a wide class would always look artificially "busier" than a narrow one purely because it swept up more of the x-axis, even at an identical underlying rate. Reading a bar's height directly as its frequency silently assumes every class has width 1 — true only by coincidence, never as a general rule — and is exactly the substitution the confirmed error record shows happening in practice: "the majority of those that were unsuccessful failed to realise that the area of bar represented the frequency and so simply gave frequencies of 25 and 5" (Jan 2023, Q1(a)).

Worked, in full

Reading a histogram — 60 cyclists' distances, four unequal-width classes

  1. 01

    60 participants in a charity cycling event had their distance travelled (in km) recorded, grouped into four classes of DIFFERENT widths: 0d<100 \leq d < 10 (width 10), 10d<1510 \leq d < 15 (width 5), 15d<2515 \leq d < 25 (width 10), and 25d<4025 \leq d < 40 (width 15). The histogram's vertical axis is labelled "frequency density (people per km)," and the heights read off it are 1.2, 4.8, 1.8, and 0.4 respectively. Multiplying each height by its own class's width: 1.2×10=121.2 \times 10 = 12; 4.8×5=244.8 \times 5 = 24; 1.8×10=181.8 \times 10 = 18; 0.4×15=60.4 \times 15 = 6.

    Earns: M1 — attempts frequency = frequency density × class width for at least one class, using that class's own width, not a shared or assumed one.

  2. 02

    Check the total: 12+24+18+6=6012 + 24 + 18 + 6 = 60, matching the 60 cyclists the question states — a check worth running every time a histogram question gives you the overall total, since it catches a wrong method immediately. Contrast this with what happens if the heights are used directly instead of being multiplied by width: 1.2+4.8+1.8+0.4=8.21.2 + 4.8 + 1.8 + 0.4 = 8.2, nowhere near 60 — exactly the substitution the confirmed record shows happening for real: "the majority of those that were unsuccessful failed to realise that the area of bar represented the frequency and so simply gave frequencies of 25 and 5" (Jan 2023, Q1(a)) — different numbers, the same missing multiplication.

    Earns: A1 — correct frequencies for all four classes (12, 24, 18, 6), with the total checked against the 60 stated in the question.

  3. 03

    The finished frequency table: 0d<100 \leq d < 10: 12; 10d<1510 \leq d < 15: 24; 15d<2515 \leq d < 25: 18; 25d<4025 \leq d < 40: 6.

    Earns: B1 — table correctly laid out against its class boundaries, not left as a bare list of four numbers.

  4. 04

    Which class contains the most cyclists? By frequency (area), it's the second class, 10d<1510 \leq d < 15, with 24 — even though the third class, 15d<2515 \leq d < 25, covers TWICE the distance range. Height alone happens to give the same ranking here too (4.8 is still the largest density), but that agreement is a coincidence of this particular data set, not something to rely on in general — the MCQs later in this lesson include a case where reading by height instead of area actually changes the answer.

    Earns: B1 — correct class identified with its frequency, and the height-vs-area distinction stated explicitly rather than assumed.

Source — Examiner report, Jan 2023

"the majority of those that were unsuccessful failed to realise that the area of bar represented the frequency and so simply gave frequencies of 25 and 5"

In your own words

In one sentence: why can two histogram bars have exactly the same height and still represent completely different numbers of data points?

Diagram — The finished histogram for the 60-cyclist data — area, not height, giving each class's frequency
Distance cycled (km)Frequency density (people per km)Class 1 heightClass 2 heightClass 3 heightClass 4 heightClass 1: frequency 12Class 2: frequency 24Class 3: frequency 18Class 4: frequency 6

x-axis: Distance cycled (km) · y-axis: Frequency density (people per km)

Class 1 height
Bar height (density) 1.2 — the class's RATE, not its frequency.
Class 2 height
Bar height (density) 4.8 — the tallest bar, but height alone still isn't the frequency.
Class 3 height
Bar height (density) 1.8.
Class 4 height
Bar height (density) 0.4 — the shortest bar, and (correctly, once width is included) the class with fewest people.
Class 1: frequency 12
Width 10, height (density) 1.2 — area 1.2 × 10 = 12.
Class 2: frequency 24
Width 5, height (density) 4.8 — the TALLEST bar on the diagram, but not the widest, and not necessarily the one with the most people just because it's tallest.
Class 3: frequency 18
Width 10, height (density) 1.8 — area 1.8 × 10 = 18.
Class 4: frequency 6
Width 15 — the WIDEST bar on the diagram, but the shortest, and (correctly) the class with the fewest people: area 0.4 × 15 = 6.

Common error: Reading each bar's height value and writing it straight down as that class's frequency — here, 1.2, 4.8, 1.8 and 0.4. Adding those four numbers gives 8.2, wildly short of the 60 cyclists actually recorded — a warning sign the height-only reading has no way to notice on its own.

Correct: Frequency = frequency density × class width — the AREA of the bar. Here: 1.2 × 10 = 12, 4.8 × 5 = 24, 1.8 × 10 = 18, 0.4 × 15 = 6, summing to exactly 60.

examiner-report · Jan 2023 · Q1(a)

Reading a stem-and-leaf diagram — and reading a back-to-back one the same way on both sides

A stem-and-leaf diagram splits each value into a "stem" (every digit except the last) and a "leaf" (the last digit), then lists every leaf for a given stem in ASCENDING order, read outward from the stem. For an ordinary (one-sided) diagram this is straightforward: stem 2 with leaves 1 3 6 8 9 represents 21, 23, 26, 28, 29 — smallest nearest the stem, largest furthest away — so reading the leaves left to right is the same as reading the values in ascending order.

A back-to-back diagram puts two data sets either side of one shared stem column — exactly the arrangement the spec's own guidance names directly: "back-to-back stem and leaf diagrams may be required." The RIGHT-hand side keeps the ordinary convention: leaves ascend as you move away from the stem, so reading left to right (stem outward) gives ascending values. The LEFT-hand side is a MIRROR IMAGE of this, not a reversal of the rule itself: its leaves still ascend as you move away from the stem — which on this side means moving further left, away from the centre. Printed left to right, a left-hand row therefore reads in DESCENDING order; the smallest value on that side is always the leaf closest to the stem, never the one printed first.

This is exactly the distinction the confirmed error record shows going wrong in practice: "some students read the stem and leaf diagram the wrong way round and so incorrectly identified Q1 = 31 and Q3 = 51" (Jun 2024, Q1(b)) — reading the left-hand side in plain left-to-right print order, rather than recognising it mirrors the right-hand side's own ascending-outward rule, silently swaps which end of the data is which. The fix isn't a new rule to memorise for the left side specifically — it's remembering that BOTH sides use the same rule (ascend outward from the stem), even though that makes them look printed in opposite directions.

Tree growth over one growing season (cm) was recorded for two groups of saplings — Group A (Oak, on the left) and Group B (Pine, on the right) — in the back-to-back diagram used for the rest of this lesson. Key: 2 | 1 | 4 means 12 cm for Group A and 14 cm for Group B — same shared stem (1) combined with each side's own leaf, exactly as an ordinary one-sided diagram would combine them. Stem 1: 9 5 2 | 1 | 4 6 7 8 9. Stem 2: 9 8 6 3 1 | 2 | 0 2 4 5 6 8 9. Stem 3: 9 7 4 | 3 | 1 3 5. Reading Group A's side in ITS OWN ascending-outward order (right to left, stem outward, since Group A sits on the left of the diagram), the 11 sorted values are 12, 15, 19, 21, 23, 26, 28, 29, 34, 37, 39.

Complete it yourself

Complete the chain — reading Q1, the median and Q3 for Group B (Pine saplings) from the diagram above

  1. 01

    Reading Group B's side in its normal ascending-outward order (left to right, stem outward — the RIGHT-hand convention), the 15 sorted values are 14, 16, 17, 18, 19, 20, 22, 24, 25, 26, 28, 29, 31, 33, 35.

  2. 02

    n=15n = 15 for Group B. The median position is n+12=8\dfrac{n+1}{2} = 8th value; Q1Q_1 position is n+14=4\dfrac{n+1}{4} = 4th value; Q3Q_3 position is 3(n+1)4=12\dfrac{3(n+1)}{4} = 12th value — all three land on whole positions here, so no interpolation is needed for this data set.

Comparing two distributions — the discipline that earns the marks

Once you've read the real figures off a diagram — a median from a stem-and-leaf, frequencies from a histogram — a "compare" question is asking you to USE them, not just to have found them. The shape of answer that scores is short and formulaic: name a statistic, give BOTH groups' figures for it, and state the direction of the difference. "Group A had a higher median (26 cm against 24 cm)" does all three; "Group A's saplings grew more" does none of them, and scores nothing no matter how confidently it's written.

Two things sink a comparison beyond simply forgetting the figures. First, treating the mean as somehow "more accurate" than the median — they are different measures of the same idea, not competing estimates of one, and the full mechanism for why an outlier moves one and not the other is worked through in the Outliers, Box Plots and Comparing Distributions lesson rather than repeated here. Second, answering a different question from the one actually asked — commenting on how SIMILAR two distributions are when the question asked for their differences, for instance. Read the question's own wording before writing the comparison; a technically fluent answer to the wrong prompt still scores nothing.

Marked, line by line

A clinic's patient waiting-time histogram has one class, 20t<3020 \leq t < 30 (width 10), with a known frequency of 18, drawn with bar height 3.6 cm on the diagram. A second class, 30t<4530 \leq t < 45 (width 15), has bar height 1.2 cm on the same diagram, drawn to the same scale. (a) Find the frequency for the second class. (3) Tree growth over one growing season (cm) was recorded for two groups of saplings, Group A (Oak) and Group B (Pine), in the back-to-back stem-and-leaf diagram shown earlier in this lesson (Group A: n=11n = 11; Group B: n=15n = 15). (b) Find Q1Q_1 and Q3Q_3 for Group A. (3) (c) Group B's median growth is 24 cm and its interquartile range is 11 cm. Using this and your answer to (b), compare the typical growth and the spread of growth for the two groups. (2) — VERIDIAN-original question; not a reproduction of any past-paper question. The scale-calibration method in part (a) and the two-sample back-to-back comparison in (b)/(c) both match how real WST01 questions in this area of the spec are actually posed, but every figure was chosen and checked for this lesson.

8 marks available

(a)3 marks

  1. 01

    density1=1810=1.8\text{density}_1 = \dfrac{18}{10} = 1.8

    Method mark for finding the known class's frequency density from its given frequency and width.

    M1
  2. 02

    scale=1.83.6=0.5\text{scale} = \dfrac{1.8}{3.6} = 0.5 (density units per cm of height)

    Method mark for converting the known bar's height into a height-to-density scale, calibrated from the class actually given (frequency 18, width 10), not an assumed one.

    M1
  3. 03

    density2=1.2×0.5=0.6\text{density}_2 = 1.2 \times 0.5 = 0.6; frequency =0.6×15=9= 0.6 \times 15 = 9

    Accuracy mark for the correct frequency (9), from scaling the SECOND bar's own height and then multiplying by its OWN width (15), not the first class's width.

    A1

(b)3 marks

  1. 101

    For Group A, n=11n = 11. Q1Q_1 position =n+14=3= \dfrac{n+1}{4} = 3rd value.

    Method mark for the correct position formula, applied to the correct sample size.

    M1
  2. 102

    Reading Group A's side in its own ascending-outward order (right to left, stem outward, since Group A sits on the LEFT of the diagram): the 3rd value is 19, so Q1=19Q_1 = 19.

    Accuracy mark for the correct Q1 value, read from the correct end of the left-hand side.

    A1
  3. 103

    Q3Q_3 position =3(n+1)4=9= \dfrac{3(n+1)}{4} = 9th value =34= 34, so Q3=34Q_3 = 34.

    Accuracy mark for the correct Q3 value.

    A1

(c)2 marks

  1. 201

    Median: Group A =26= 26 cm (from the middle value of the 11 sorted values in part (b)'s data); Group B =24= 24 cm (given). Group A's typical growth is slightly higher, by 2 cm.

    Independent mark for a location comparison that names the statistic (median) and gives both figures.

    B1
  2. 202

    IQR: Group A =3419=15= 34 - 19 = 15 cm; Group B =11= 11 cm (given). Group A's growth is more spread out — its IQR is 4 cm larger than Group B's.

    Independent mark for a spread comparison that names the statistic (IQR) and gives both figures — addressing a DIFFERENT aspect of the comparison from the first mark, not repeating it.

    B1

Named traps

histogram-height-read-as-frequency
Confirmed, in the one examiner report reviewed in this research pass that treats a histogram-reading question directly: "the majority of those that were unsuccessful failed to realise that the area of bar represented the frequency and so simply gave frequencies of 25 and 5" (Jan 2023, Q1(a)). Frequency density is a RATE — data points per unit of x-axis — not a count; multiplying by the class width is what turns it into one, and skipping that step is the entire error, not a rounding slip layered on top of an otherwise correct method.
stem-and-leaf-read-in-the-wrong-direction
Confirmed: "some students read the stem and leaf diagram the wrong way round and so incorrectly identified Q1 = 31 and Q3 = 51" (Jun 2024, Q1(b)) — swapping which end of the ordered leaves is the smaller quartile and which is the larger. On a back-to-back diagram specifically, this usually comes from reading the LEFT-hand side in plain left-to-right print order instead of recognising it mirrors the right-hand side's own ascending-outward rule (see the teach block above for the mechanism). This same trap is named in the Outliers, Box Plots and Comparing Distributions lesson's own trap-taxonomy, since a misread stem-and-leaf diagram is exactly as capable of feeding a wrong Q1/Q3 into an outlier-fence calculation as it is of feeding one into this lesson's own reading-practice content — it is cited in both places because it is genuinely load-bearing for both skills, not duplicated by accident.
comparison-lacks-a-named-statistic-figures-or-the-right-focus
Three separate confirmed patterns, all under the same "compare two distributions in words" guidance point, and all already given full standalone treatment (their own concept-ladder, marked-solution part and trap-taxonomy entries) in the Outliers, Box Plots and Comparing Distributions lesson — restated here only because they are exactly as relevant when the two distributions being compared were read off a histogram or a stem-and-leaf diagram, which is this lesson's own subject. (1) The single most common failure: treating the mean as inherently "more accurate" than the median — "A very common misconception was that the mean is more accurate than the median... showing no appreciation that the mean is just one measure of average and the median is another" (Jun 2022, Q1(e)). (2) Naming no actual figures: "surprisingly too many students failed to give supporting figures... a reference to a named statistic and supporting figures [is required]" (Jun 2024, Q1(e)). (3) Answering a different question from the one asked: "few comments referring to the distribution... others commented on the similarities [when asked for differences]" (Oct 2021, Q3(e)). A comparison that names a statistic, states both figures, and addresses exactly what the question asked for is the only shape of answer that scores against any of the three.

Retrieval — with feedback on every choice

Question 1
2 marks

A histogram bar for the class 8x<128 \leq x < 12 (width 4) has height (frequency density) 6.5. What is the frequency for this class?

Question 2
3 marks

Class P: width 10, frequency density 2. Class Q: width 3, frequency density 5.

Which class contains MORE data values?

Question 3
2 marks

Group A (Oak saplings, n=11n = 11), sorted: 12, 15, 19, 21, 23, 26, 28, 29, 34, 37, 39.

What is Q1Q_1 for Group A?

Question 4
3 marks

Group A: median growth 26 cm, IQR 15 cm. Group B: median growth 24 cm, IQR 11 cm.

Which response to "compare the growth of the two groups" would be awarded full marks?

Reference — not a study method, a lookup
  • Histogram: frequency = frequency density × class width — AREA, not height, is frequency.
  • Different-width bars can only be compared by area. A taller bar can still represent FEWER data points.
  • Self-check: your frequencies should sum to the total the question states.
  • Back-to-back stem-and-leaf: on BOTH sides, leaves ascend outward from the stem — the left side just prints that in reverse.
  • Quartile position: Q1 = (n+1)/4th value, median = (n+1)/2th, Q3 = 3(n+1)/4th — count from the correct end.
  • Comparing distributions: name a statistic, give both figures, state the direction, answer what's actually asked.

Not affiliated with or endorsed by Pearson Edexcel. Every quotation attributed to a mark scheme, examiner report or the specification in this lesson was independently verified against the primary Pearson document by the research pass this lesson was written from. Two pairs of real figures appear directly in this lesson, both as evidence of a genuine documented error rather than as a full dataset: "25 and 5," the wrong frequencies a real cohort gave for a histogram question (Jan 2023, Q1(a)), and "Q1 = 31 and Q3 = 51," a real cohort's wrong-way-round quartile misread (Jun 2024, Q1(b)) — the research bank does not supply the correct figures or the original diagrams for either question, only these examiner-quoted wrong answers, so no attempt is made here to reconstruct either past paper's actual data. Every other number in this lesson — the 60-cyclist histogram, the clinic waiting-time scale-calibration question, and the two-group (Oak/Pine sapling) back-to-back stem-and-leaf data used in the teach content, the chain drill and the marked solution — is VERIDIAN-original, computed and independently checked before being written into this file, not carried over from any real series. Every question in this lesson — prequestion, worked chain, chain drill, marked solution and MCQ alike — is VERIDIAN-original, inspired by confirmed real question types and traps but never a reproduction of a real Pearson question. Two trap-taxonomy items (the stem-and-leaf misread direction, and the comparing-distributions discipline) are deliberately also named in the Outliers, Box Plots and Comparing Distributions lesson, for the honest reason given at each entry above — not by accident, and not as a shortcut around writing this lesson's own content.

Question 12 marks

A histogram bar for the class 8x<128 \leq x < 12 (width 4) has height (frequency density) 6.5. What is the frequency for this class?

  • 6.5×4=266.5 \times 4 = 26

    Correct. Frequency = frequency density × class width =6.5×4=26= 6.5 \times 4 = 26.

  • B6.5

    This reports the bar's height directly as the frequency, skipping the multiplication by class width — the confirmed real error, restated with fresh numbers: "failed to realise that the area of bar represented the frequency" (Jan 2023, Q1(a)).

  • C6.5÷4=1.6256.5 \div 4 = 1.625

    This divides instead of multiplies. Frequency density is already frequency divided by width — recovering the frequency means multiplying back by the width, not dividing by it again.

  • D6.54=2.56.5 - 4 = 2.5

    Frequency density and class width combine by multiplication (they represent an area), not by subtraction — there's no meaningful operation here that involves subtracting a length from a rate.

Traps tested: Histogram height read as frequency · Density width relationship inverted · Density and width combined by subtraction

Question 23 marks

Class P: width 10, frequency density 2. Class Q: width 3, frequency density 5.

Which class contains MORE data values?

  • Class P — its area (10×2=2010 \times 2 = 20) is larger than Class Q's (3×5=153 \times 5 = 15), even though Q's bar is taller

    Correct. Class P's frequency is 20, Class Q's is 15 — despite Q having the taller bar (density 5 against 2), P is the wider class and its AREA, not its height, is what decides which contains more data.

  • BClass Q — its bar is taller, since its frequency density (5) is greater than Class P's (2)

    This compares the two bars by height alone. Height only tells you the RATE for that class, not the total count — and here, comparing by height gives exactly the wrong answer: Class P actually has more data (20 against 15), even though its bar is shorter.

  • CThey contain the same number of values, since one has a wider class and the other a taller bar and the two effects cancel out

    There's no general reason width and height differences would cancel out — here they don't: 10×2=2010 \times 2 = 20 is not equal to 3×5=153 \times 5 = 15. Each class's actual area has to be calculated, not assumed to balance.

  • DCannot be determined without knowing the exact shape of each bar

    A histogram bar's shape is exactly what frequency density and class width already specify — height × width gives the area (frequency) directly, with nothing further needed.

Traps tested: Histogram height read as frequency · Width and height differences assumed to cancel · Overclaims uncertainty

Question 32 marks

Group A (Oak saplings, n=11n = 11), sorted: 12, 15, 19, 21, 23, 26, 28, 29, 34, 37, 39.

What is Q1Q_1 for Group A?

  • 19 — the 3rd value, using position n+14=3\dfrac{n+1}{4} = 3

    Correct. Q1Q_1 position is the 3rd of the 11 ordered values, which is 19.

  • B34 — the 9th value

    This is actually Q3Q_3 (position 3(n+1)4=9\dfrac{3(n+1)}{4} = 9), not Q1Q_1 — swapping which end of the ordered data each quartile comes from is exactly the confirmed real error: "some students read the stem and leaf diagram the wrong way round and so incorrectly identified Q1 = 31 and Q3 = 51" (Jun 2024, Q1(b)).

  • C26 — the middle value

    26 is the MEDIAN (position n+12=6\dfrac{n+1}{2} = 6), not Q1Q_1 — the median, Q1Q_1 and Q3Q_3 are three different statistics at three different positions, and this answer uses the wrong one of the three.

  • D12 — the smallest value

    12 is the MINIMUM, the very first value, not Q1Q_1Q1Q_1 is a quarter of the way through the ordered data, not the starting point of it.

Traps tested: Stem and leaf read in the wrong direction · Wrong statistic supplied for the one asked · Minimum confused with lower quartile

Question 43 marks

Group A: median growth 26 cm, IQR 15 cm. Group B: median growth 24 cm, IQR 11 cm.

Which response to "compare the growth of the two groups" would be awarded full marks?

  • "Group A had a higher median (26 cm against 24 cm), so typical growth was slightly greater for Group A; Group A also had a higher IQR (15 cm against 11 cm), so growth was more variable for Group A."

    Correct — it names two statistics (median and IQR), gives both groups' figures for each, and states the direction each time. This is exactly the shape of answer a real mark scheme requires: "a reference to a named statistic and supporting figures" (Jun 2024, Q1(e)).

  • B"Group A grew more, and the mean would confirm this more accurately than the median does."

    This is the named misconception itself, close to verbatim: "A very common misconception was that the mean is more accurate than the median... showing no appreciation that the mean is just one measure of average and the median is another" (Jun 2022, Q1(e)).

  • C"Group A had more growth overall."

    No named statistic and no figures for either group. A confirmed real error record states: "surprisingly too many students failed to give supporting figures" (Jun 2024, Q1(e)) — this is exactly that failure.

  • D"Both groups show a similar spread of growth, based on the shape of the stem-and-leaf diagram."

    This both gives no figures and gets the actual comparison backwards: the IQRs (15 against 11) are not close, and the true difference in spread — the more striking of the two real differences between these groups — goes unaddressed. A confirmed error record names exactly this shape of failure on a differences-focused question: "others commented on the similarities [when asked for differences]" (Oct 2021, Q3(e)).

Traps tested: Mean assumed more accurate than the median · Comparison missing supporting figures · Comparison answers the wrong question

Practice this for real

This site teaches the mechanism; the exam is sat on Pearson's own real questions. Go find and attempt these yourself — nothing here substitutes for actually sitting a timed paper.

Examiner report
Jan 2023 · Q1(a) — cited directly in this lesson
Pearson's official past-papers portal

Select International Advanced Level → Mathematics → any series, then look for WST01.

Statistics 1 · progress saved in this browser · sign in to sync across devices

Up next

Measures of location and dispersion — mean, coding, and standard deviation

Coding turns an awkward mean — 255.2 — into an easy one, 0.1. That saving isn't free: add a constant and the mean slides by exactly that constant, but multiply by one and the mean scales by it once while the *variance* scales by its square — and has to be undone by squaring again on the way back. One real WST01 examiner's report records three separate ways candidates got that one decoding step wrong, in a single sitting. The standard deviation formula fails in an even more literal way: two different exam series produced malformed versions of it — missing a division, or missing a square root in the place it actually belongs — and both examiner reports trace the damage back to the same habit, rounding a value mid-calculation instead of carrying it through exactly. None of this is really about arithmetic. It's about which few numbers this exam expects you to reconstruct from memory, because the formula booklet — deliberately — will not hand them to you.

55 min