Outliers, Box Plots, and Comparing Distributions
~55 min · WST01 · 2.4
WST01 · 2.4 · 55 min
A box plot's whiskers only ever touch real data. The fence that decides who counts as an outlier is arithmetic you compute and then throw away — it never appears as a mark on the finished plot — and one of the most consistently documented box-plot errors on this paper is drawing a whisker straight to that invisible number instead of stopping at the last real value before it. The formula for the fence isn't fixed, either: the specification insists every question state its own rule, precisely so a memorised "1.5 × IQR" can't quietly become the wrong number for the one sitting in front of you.
Before you read on
Two or three questions on exactly what this lesson teaches. Being wrong here is fine — it's the fastest way to find out what to pay attention to next.
The five-number summary, and what a box plot actually draws
A box plot is built from exactly five numbers, in order: the minimum, the lower quartile (), the median (), the upper quartile (), and the maximum. The box itself runs from to — the middle 50% of the data, the interquartile range (IQR) — with a line drawn inside it at the median; whiskers then extend outward from the box toward the two most extreme values. Spec 2.1 groups box plots with histograms and stem-and-leaf diagrams under "representation," and the guidance attached to that spec item is explicit about which half of the skill is actually examined: "drawing of histograms/stem-and-leaf/box plots will not be the direct focus of exam questions — reading/using them will be." Almost everything below is therefore about reading a box plot correctly, or using given figures to work out where its features belong — not about ruling one out neatly with a pencil.
The two whiskers are not simply "lines to the minimum and maximum." That description is only true when neither extreme is an outlier. The moment a value is flagged as an outlier by whatever rule the question has given you, its own whisker stops short: it runs only as far as the most extreme value that is NOT an outlier, and the outlier itself is marked separately — usually as an isolated point or cross beyond the end of the whisker. Get this one distinction right and most of the outlier-and-box-plot content on this paper falls into place; get it wrong and a whisker drawn to the wrong place is one of the most consistently documented box-plot errors in the record this lesson is built from.
The interquartile range itself, , is one of the formulae you're expected to know without looking it up — the specification's own "Notation and formulae" section states outright that formulae in that category "will not appear in the booklet." Mean and standard deviation sit on the same short list. The five-number summary and the box-plot picture, by contrast, are things you USE, not derive from memory — you will very often be handed , the median and directly, since "calculation of mean/mode/median/range/IQR will not be the direct focus of exam questions" either (spec 2.2 guidance).
Skewness — read the shape, don't calculate a number
Skewness sits under the same spec item as outliers (2.4), and at this level it's a reading skill, not a calculation: no skewness formula appears anywhere in the WST01 formula booklet, and none is on the short list of formulae you're expected to memorise either. What you're actually being asked to do is compare the two halves of a box plot and say which one is more stretched out.
Two comparisons do almost all of the work, and they usually agree with each other. First, compare the two halves of the box itself: the distance from to the median against the distance from the median to . If the upper half ( minus the median) is the bigger gap, the middle 50% of the data is more spread out on the high side — a sign of positive (right) skew, a distribution with a longer tail stretching toward large values. If the lower half (the median minus ) is the bigger gap, that's negative (left) skew instead. Second, compare the two whiskers the same way, when neither is cut short by an outlier: a longer upper whisker corroborates positive skew, a longer lower whisker corroborates negative skew.
Both comparisons are about the actual numbers, not the picture. A box plot sketched by hand on an exam script is not drawn to scale, and a box that LOOKS lopsided can mislead in either direction. Subtract the real values — minus the median, the median minus — before deciding, the same "show the comparison, don't just assert the conclusion" discipline that governs outlier questions.
Skew direction also predicts, loosely, which way the mean sits relative to the median: a long right tail tends to drag the mean above the median, and a long left tail tends to drag it below. "Tends to" is doing real work in that sentence — it's a common pattern, not a guarantee — but it's the same underlying mechanism the next section makes precise: the mean is pulled toward wherever the extreme values sit, and the median mostly isn't.
Mechanism
Where 1.5 × IQR fences come from, and why the direction is not optional
The IQR measures the width of the middle 50% of the data — how spread out the 'ordinary' half of it is. An outlier rule takes that width as its own unit of measurement and asks: how many IQRs outside the box does a value have to sit before it stops looking like a natural continuation of the data and starts looking like something genuinely unusual? A rule like answers that by marking a fence at 1.5 IQRs beyond EACH quartile, extending outward, away from the box. That direction isn't a convention to memorise; it's forced by what a fence is for. The lower fence has to sit below , so it's minus the distance — subtracting. The upper fence has to sit above , so it's plus the distance — adding. A fence built the other way round, added to or subtracted from , moves INTO the box rather than away from it, and for a large enough multiplier can end up on the wrong side of the median entirely — at which point it would start flagging perfectly ordinary values near the middle of the data as outliers, exactly the opposite of what the rule exists to do. The confirmed error record shows precisely this: "there were a surprising number of errors... with some multiplying the quartiles by 1.5 and others using Q1 + 1.5×IQR for the lower limit" (Jan 2021, Q2) — both are the same underlying mistake, arithmetic that produces a real number but not the number the fence is defined to be. There's a second reason not to treat 1.5 as a fixed fact at all: the specification's own guidance on this spec item states plainly that "any rule to identify outliers will be specified in the question" — 1.5 is the multiplier used throughout this lesson because it's the one seen most often in the reviewed record, not because it's fixed. A real question can hand you a different multiplier, and when it does, that stated rule overrides whatever you remember from practice. The confirmed record shows exactly what happens when a student substitutes the remembered version anyway: "a small number of candidates opted to apply the more commonly used outlier formula of Q3 + 1.5×(Q3−Q1) rather than the one quoted in the question" (Jun 2022, Q1(b)). What has to transfer between questions is the outward-facing DIRECTION of the rule, not a specific number attached to it.
Why the mean moves and the median doesn't
In plain terms
Line up 15 runners by their finishing time, fastest to slowest, and ask 'what's a typical time?' Pick the runner standing exactly in the middle: that's the median, and it doesn't care how fast the fastest runner was or how slow the slowest was — only that they ARE the fastest and the slowest, standing at the two ends of the line. Now add up every single runner's time and divide by 15: that's the mean. If the slowest runner took 92 minutes instead of a more ordinary 61, the mean has to answer for that extra half hour, because it's built from every time in the line — while the runner standing in the middle hasn't moved an inch.
In numbers: for the 15 runner-times used later in this lesson, replacing the outlier (92) with a value 30 minutes smaller changes the mean by exactly minutes — always exactly (the change) ÷ (the number of values), because the mean is , a sum divided by a count, and every term in that sum matters equally. The median, by contrast, is found by RANK, not by size: it's whichever value sits in the middle position once the data is ordered. Change the outlier's exact size without changing which position it holds — still the largest — and the median (the 8th of 15 values, unaffected by anything happening at position 15) doesn't move at all.
Formally
Formally: the mean is a linear function of every data value, , so a change of in any single changes by exactly — a small, predictable, but NEVER-zero shift, however extreme that one value becomes. The median is an order statistic: for an ordered data set of size , it's (for odd ) the value at position , determined entirely by rank. Any change to a value that doesn't alter its rank relative to the others leaves the median completely unchanged, no matter how large that change is. This is exactly the property the credited exam answers name directly — 'the mean uses all the data' and 'the mean includes the outliers' (Jun 2022, Q1(e)) — and exactly what the wrong, uncredited answer misses when it calls the mean 'more accurate': accuracy isn't the axis these two statistics differ on. They're both legitimate measures of 'a typical value,' and they differ only in how much an extreme value is allowed to move them.
In your own words
In one sentence: why does the lower outlier fence have to be built by SUBTRACTING from , never by adding it?
Worked, in full
Fences and outliers for the runner-times data — Q1 = 42, median = 47, Q3 = 55 (n = 15)
- 01
The sorted finishing times, in minutes, of 15 runners in an obstacle-course race are 25, 35, 39, 42, 44, 45, 46, 47, 49, 51, 53, 55, 58, 61, 92. From these, (the 4th value), the median (the 8th), and (the 12th) — quartile positions read off directly rather than derived here, matching how this content is actually tested (the spec keeps quartile calculation itself off the direct focus of exam questions). The rule for this question: a value is an outlier if it lies more than beyond the nearer quartile. .
Earns: M1 — attempts , using the given quartiles.
- 02
Scale the IQR by the multiplier the question actually gives you — here, 1.5. . This is a number to compute fresh every time, not a value to remember: a different question can give a different multiplier, and reaching for 1.5 out of habit when a question specifies something else is a confirmed, named error in its own right (see the trap taxonomy below).
Earns: M1 — attempts using their own IQR.
- 03
Build both fences, and note the direction is forced, not chosen: the lower fence sits BELOW , so it subtracts — ; the upper fence sits ABOVE , so it adds — . A fence built the other way round points back into the box, not away from it.
Earns: A1 — both fences correct, 22.5 and 74.5. The two commonest wrong versions of this exact line are quoted directly in the embedded evidence below — losing this mark is almost never an arithmetic slip; it's the direction, or the ×IQR step being skipped.
- 04
Compare every extreme value against its fence, and show the comparison rather than just stating the verdict — the same 'show that' discipline this exam applies across every topic on this paper, not just this one. Minimum: is ? No — , so the minimum is NOT an outlier. Maximum: is ? Yes, so the maximum IS an outlier.
Earns: A1 — correct conclusion for both extremes, with the comparison actually shown. A bare 'not an outlier' / 'is an outlier' with no numbers next to it is exactly the incompleteness an examiner report elsewhere in this record flags on a 'show there are 3 outliers' question — see the trap taxonomy.
- 05
Translate that into the box plot itself. Since 92 is the only outlier, the upper whisker cannot reach it — it stops at the highest value in the data that is NOT an outlier, which is 61, and 92 is plotted as a separate point beyond the end of the whisker. The lower whisker is unaffected and runs all the way to the true minimum, 25, since nothing was excluded at that end.
Earns: B1 — correct whisker endpoints (25 and 61), with the outlier (92) shown as a separate point. This is the step where marks are lost even after the fences themselves are found correctly — it's a different skill from finding the fence, and it's tested separately.
Source — Examiner report, Jan 2021
"there were a surprising number of errors... with some multiplying the quartiles by 1.5 and others using Q1 + 1.5×IQR for the lower limit"
Complete it yourself
Complete the chain — fences and outliers for 11 call-centre phone calls (Q1 = 6, median = 9, Q3 = 14 minutes)
- 01
.
- 02
The rule given in the question: a value is an outlier if it lies more than beyond the nearer quartile. .
x-axis: Finishing time (minutes) · y-axis: (box-plot lane — not a data axis)
- Lower whisker
- A straight line from the true minimum (25) to Q1 (42) — the minimum is not an outlier, so the whisker reaches it exactly.
- Median line
- A single line inside the box at the median (47), splitting the box — not necessarily in half, since the two quartile-to-median gaps needn't be equal (unequal gaps are exactly the skewness signal from the teach block above).
- Upper whisker
- A straight line from Q3 (55) to 61 — the highest value in the data that is NOT an outlier. It stops well short of the fence (74.5), and further still short of the outlier itself (92).
- Q1 = 42
- Lower edge of the box — 25% of the data lies at or below this value.
- Q3 = 55
- Upper edge of the box — 75% of the data lies at or below this value.
- Outlier = 92
- Plotted as an isolated point, not joined to the whisker by any line — 92 > 74.5, the upper fence, confirmed in stage 4 of the worked chain above.
- Upper fence = 74.5 (not drawn)
- The boundary used to TEST the maximum value against — it's arithmetic, not a feature of the finished plot. Nothing is ever drawn AT a fence itself, which is exactly the mistake the commonError below documents.
Common error: Drawing the upper whisker out to the fence, 74.5 — or further still, out to the outlier itself, 92 — instead of stopping at 61, the highest value in the data that is genuinely not an outlier. The research bank records this exact shape of error (its own summary of the Jan 2021 Q2 examiner report, not a direct quotation): in that series the fence sat at 98 and the actual highest non-outlier value was 97, and whiskers were drawn to the boundary rather than to the real data point beneath it.
Correct: A whisker only ever ends at an actual value present in the data. A fence is a number used to TEST values against — it decides who counts as an outlier — but it's never itself a point on the finished plot, and neither is an outlier joined to the rest of the box by a line.
Marked, line by line
Two branches of a gym, P and Q, each recorded the length of a sample of 40 visits, in minutes. Summary statistics: Branch P — minimum 10, Q1 = 28, median 34, Q3 = 45, maximum 70, mean 35.2. Branch Q — minimum 5, Q1 = 22, median 34, Q3 = 38, maximum 95, mean 33.8. (a) Using the rule "a value is an outlier if it lies more than 1.5 × IQR beyond the nearer quartile," determine whether the maximum visit length at Branch P is an outlier. (3) (b) State, with a reason based on the figures given, whether the distribution of visit lengths at Branch P is more likely to show positive skew, negative skew, or no skew. (2) (c) Branch P has the higher mean, even though Branch Q has the higher maximum. Explain why this difference in maximum values affects the two branches' means differently from how it affects their medians. (2) (d) Compare the typical visit length at the two branches. (2) — VERIDIAN-original question and dataset; not a reproduction of any past-paper question. The context (two comparable groups, given as summary statistics rather than raw data) matches how real WST01 comparison questions are commonly posed, but every figure below was chosen and checked for this lesson, not carried over from a real series.
9 marks available
(a) — 3 marks
- 01M1
Method mark for attempting the IQR from the two given quartiles.
- 02M1
Upper fence
Method mark for the upper fence, built in the correct direction — added to Q3, not subtracted, and not applied to Q1 instead.
- 03A1
, so the maximum visit length at Branch P is NOT an outlier.
Accuracy mark for the correct conclusion WITH the comparison shown. The two values are close enough (70 against a fence of 70.5) that asserting the answer without the comparison would not be credited under the 'show that' discipline this exam applies throughout — the closeness is deliberate: this question can't be answered by eye.
(b) — 2 marks
- 101B1
Median Q1 ; Q3 median .
Independent mark for comparing the correct two gaps — the upper half of the box against the lower half.
- 102B1
, so the upper half of the middle 50% is more spread out than the lower half: the distribution shows positive (right) skew.
Independent mark for the correct direction, stated with the reasoning that produced it. (The whisker check corroborates it: Q1 − min = 18 against max − Q3 = 25 — the upper whisker is longer too.)
(c) — 2 marks
- 201B1
The mean is calculated from EVERY value in the data (), so a branch with a more extreme maximum has that value baked directly into its mean — a bigger maximum pulls the mean up, in proportion to how much bigger it is.
Independent mark for correctly identifying that the mean is affected because it's calculated using every value, including the maximum.
- 202B1
The median only depends on the RANK of the middle value(s), not the size of the extremes — Branch Q's maximum could be 95 or 195 and its median would be unaffected either way, since neither changes which value sits in the middle of the ordered list.
Independent mark for correctly identifying that the median is a positional measure, unaffected by how extreme the largest value becomes.
(d) — 2 marks
- 301B1
Both branches have the same median visit length (34 minutes), so 'typical' visits are similar in length at both branches.
Independent mark for a comparative statement that names a specific statistic and gives both figures — not just 'similar', a number for each branch.
- 302B1
But Branch Q's visits are more spread out overall: its range (90 minutes) is far larger than Branch P's (60 minutes), driven by Branch Q's more extreme maximum (95 against 70) — even though the two branches' interquartile ranges are close (16 against 17).
Independent mark for a second comparative statement, again with named statistics and figures on both sides, addressing spread rather than repeating the first mark's point about location.
Named traps
- outlier-lower-fence-direction-reversed
- Confirmed directly, and the research bank's own phrasing makes clear this was not a rare slip: "there were a surprising number of errors... with some multiplying the quartiles by 1.5 and others using Q1 + 1.5×IQR for the lower limit" (Jan 2021, Q2). The lower fence has to SUBTRACT from Q1, moving further below it — the mechanism block above derives why, rather than asking you to remember it as a rule with no reason behind it.
- whisker-drawn-to-fence-not-to-data
- The research bank's own account of the same Jan 2021 Q2 finding (its documented summary of the examiner report, not a direct quotation) is that box-plot whiskers were commonly drawn out to the outlier boundary itself (98, in that series) rather than to the actual highest non-outlier value in the data (97). A fence is a number used to test values against; it's never itself a point that gets drawn. See the diagram block above for the same error, reproduced with this lesson's own numbers (74.5 versus the real value, 61).
- remembered-formula-overrides-the-rule-given-in-the-question
- Confirmed: "a small number of candidates opted to apply the more commonly used outlier formula of Q3 + 1.5×(Q3−Q1) rather than the one quoted in the question" (Jun 2022, Q1(b)). This matters because the specification itself states that "any rule to identify outliers will be specified in the question" (spec S1.3, item 2.4, guidance) — there is deliberately no single fixed rule to memorise, and treating 1.5×IQR as a universal constant is itself the error, even on the (common) occasions where the question happens to specify 1.5 anyway.
- show-that-outliers-not-listed
- Confirmed, on a "show that there are 3 outliers" question: "having gained the correct limits some did not list the 3 outliers in this part in order to show there are 3 outliers" (Oct 2021, Q3(c)). Finding the correct fences is necessary but not sufficient — a "show that N outliers exist" question needs the N values actually named, not just the machinery that would find them.
- mean-assumed-more-accurate-than-the-median
- Confirmed, and stated by the examiner report as a genuinely common pattern rather than a rare one: "very few candidates scored this mark... A very common misconception was that the mean is more accurate than the median. A large number of responses simply explained how to calculate a mean or said that it was because the mean is the average, showing no appreciation that the mean is just one measure of average and the median is another" (Jun 2022, Q1(e)). The credited answers were specific and short: "the mean uses all the data" / "the mean includes the outliers." Neither statistic is "more accurate" — they are different measures of the same idea, and the concept-ladder above works through exactly why an outlier moves one and not the other.
- comparison-missing-supporting-figures
- Confirmed: "surprisingly too many students failed to give supporting figures... a question like this will require some context..., a reference to a named statistic and supporting figures" (Jun 2024, Q1(e)). "Branch A had a higher average" names nothing and gives no numbers; "Branch A had a higher median (34 minutes against 28)" does both, and only the second shape of answer scores.
- comparison-answers-the-wrong-question
- Confirmed: "few comments referring to the distribution of ages were seen... others commented on the similarities [when asked for differences]" (Oct 2021, Q3(e)). Read what the question is actually asking for — differences, similarities, a specific statistic — before writing the comparison, since a technically-true observation about the wrong aspect of the data scores nothing.
- stem-and-leaf-read-in-the-wrong-direction
- Confirmed, on a real quartile-from-stem-and-leaf question: "some students read the stem and leaf diagram the wrong way round and so incorrectly identified Q1 = 31 and Q3 = 51" (Jun 2024, Q1(b)) — they swapped which end of the ordered leaves is the lower quartile and which is the upper. Whatever representation a box plot is built from — a stem-and-leaf diagram, a table, a raw list — check which end you're counting from before quoting a quartile out of it.
Retrieval — with feedback on every choice
A data set has and . Using the rule "an outlier is a value more than beyond the nearer quartile," is the value 47 an outlier?
For this data (Q1 = 40, Q3 = 52), a value is defined as an outlier if it lies more than 2 × IQR beyond the nearer quartile.
Is the value 73 an outlier, under the rule as stated in THIS question?
A box plot has , median , , with whiskers reaching 5 (minimum) and 24 (maximum, not an outlier). What does this suggest about the skew of the distribution?
A data set has , . Using , the upper fence is 47.5. The highest value in the data that is NOT an outlier is 40; one value, 55, is an outlier. Where does the upper whisker end, and where is 55 shown?
Data set A has a higher median (42 minutes) than data set B (35 minutes). Which response to "compare the typical time recorded by A and B" would be awarded full marks?
- Five-number summary: min, Q1, median, Q3, max. Box = Q1 to Q3 (the IQR); line inside = median.
- IQR = Q3 − Q1 — not in the booklet, memorise it.
- The outlier rule is ALWAYS given in the question. Use exactly what's stated, never a remembered formula.
- Fences point outward: lower = Q1 − (k × IQR); upper = Q3 + (k × IQR). Never the reverse.
- A whisker stops at the most extreme NON-outlier value — never at the fence, never at the outlier.
- Mean uses every value (outliers pull it); median uses only position (outliers barely move it).
- Comparing distributions: name a statistic, give both figures, state the direction. No numbers, no marks.
Not affiliated with or endorsed by Pearson Edexcel. Every quotation and figure attributed to a mark scheme, examiner report or the specification in this lesson was independently verified against the primary Pearson document by the research pass this lesson was written from, not invented or carried over from other course material. Every question in this lesson — prequestion, worked chain, chain drill, marked solution and MCQ alike — is VERIDIAN-original, inspired by confirmed real question types and traps but never a reproduction of a real Pearson question. Three complete numeric scenarios are original to this lesson and do not appear in the research bank: the 15-runner race-time data used in the worked chain and diagram, the 11-call phone data used in the chain drill, and the two-branch gym data used in the marked solution — every figure in all three was computed and checked before being written into this file. Because the questions are original, the per-line mark allocations attached to them are modelled on verified WST01 mark-scheme conventions (what M, A and B marks mean, when a mark is independent versus dependent, the paper-wide 'show that' discipline) rather than transcribed from a real mark scheme, which for an original question does not exist.
A data set has and . Using the rule "an outlier is a value more than beyond the nearer quartile," is the value 47 an outlier?
- No — the upper fence is , and
Correct. , and , so the upper fence is . 47 sits just inside it. This is deliberately a close call: the fence and the value differ by only 1, exactly why the comparison has to be shown in full rather than judged by eye.
- BYes — 47 is greater than (30)
Every value above is, by definition, in the upper quarter of the data — that alone doesn't make it unusual. An outlier has to clear the FENCE, a full beyond , not just clear itself.
- CYes — using Q1's fence instead: , and
This tests 47 against a fence built from , but 47 sits in the UPPER part of the data, so it has to be tested against the fence built from — the nearer quartile — not .
- DNo — because 47 isn't the maximum value in the data
Whether a value is the overall maximum is irrelevant to whether it's an outlier — an outlier is decided purely by its distance from the nearer quartile, measured against the fence, whatever else is in the data set.
Traps tested: Any value past q3 treated as outlier · Wrong quartile fence used · Outlier status confused with being the extreme value
For this data (Q1 = 40, Q3 = 52), a value is defined as an outlier if it lies more than 2 × IQR beyond the nearer quartile.
Is the value 73 an outlier, under the rule as stated in THIS question?
- No — the upper fence is , and
Correct. , and the question's own rule uses a multiplier of 2, not 1.5: , and 73 falls short of it.
- BYes — the upper fence is , and
This applies the standard rule from memory instead of the rule this specific question states. It's a confirmed, real error: "a small number of candidates opted to apply the more commonly used outlier formula... rather than the one quoted in the question" (Jun 2022, Q1(b)) — and here, unlike in that example, using the wrong multiplier changes the actual conclusion, not just the working.
- CYes — any value more than plus the IQR itself () counts as an outlier
This is a third formula, matching neither the rule this question gives nor the more familiar default — it isn't derived from anything stated in the question.
- DCannot be determined without seeing the full data set
The rule plus and is everything the calculation needs — the individual raw values play no part in deciding whether 73 clears the fence.
Traps tested: Remembered formula overrides the rule given in the question · Invented rule not matching the question · Overclaims uncertainty
A box plot has , median , , with whiskers reaching 5 (minimum) and 24 (maximum, not an outlier). What does this suggest about the skew of the distribution?
- Negative (left) skew — both the box (median Q1 , against Q3 median ) and the whiskers (Q1 min , against max Q3 ) show more spread on the LOW side
Correct. Both comparisons agree: the lower half of the box is wider, and the lower whisker is longer, so more of the spread sits below the median than above it — a longer left tail, i.e. negative skew.
- BPositive (right) skew
The upper half of the box (median to ) is the SMALLER of the two gaps here, and the upper whisker is the SHORTER of the two — both point toward more spread on the low side, not the high side. Positive skew would need the opposite pattern.
- CRoughly symmetric, since the median sits inside the box, not at either edge
The median sitting somewhere inside the box is true of almost every box plot and says nothing about symmetry on its own — comparing the actual gaps (4 against 2, and 10 against 3) is what decides it, and neither pair is close to equal.
- DNothing can be concluded about skew from a box plot — only from the raw data
Reading skew direction from a box plot's proportions is exactly the spec 2.4 skill being tested here — a legitimate, examinable comparison, not a shortcut that needs the raw data to back it up.
Traps tested: Skew direction reversed · Skew read from visual impression not figures · Box plot skew reading dismissed
A data set has , . Using , the upper fence is 47.5. The highest value in the data that is NOT an outlier is 40; one value, 55, is an outlier. Where does the upper whisker end, and where is 55 shown?
- The whisker ends at 40; 55 is plotted separately, beyond the end of the whisker
Correct. The whisker only ever reaches an actual value in the data that survives the outlier rule; 55 fails that rule, so it is shown as an isolated point instead.
- BThe whisker ends at 47.5 (the fence); 55 is plotted separately
47.5 is a boundary used to TEST values, not a value that exists in the data — a whisker only ever ends at a real data point, and the confirmed record shows exactly this error happening in practice.
- CThe whisker ends at 55, and no point is plotted separately
This folds the outlier back into the ordinary range, defeating the entire purpose of identifying it — an outlier is shown SEPARATELY, precisely because it isn't a natural continuation of the rest of the data.
- DThe whisker ends at (25); there is no upper whisker at all once an outlier exists
An outlier shortens the whisker to the nearest genuine non-outlier value — it doesn't remove the whisker. Here that value is 40, well above .
Traps tested: Whisker drawn to fence not to data · Outlier included in the whisker · Whisker omitted entirely once an outlier exists
Data set A has a higher median (42 minutes) than data set B (35 minutes). Which response to "compare the typical time recorded by A and B" would be awarded full marks?
- "Data set A had the higher median (42 minutes, against 35 minutes for B), so on average the times recorded in A were longer."
Correct — it names the statistic being compared, gives both figures, and states the direction. This is the exact shape of answer a real mark scheme requires: "a reference to a named statistic and supporting figures" (Jun 2024, Q1(e)).
- B"Data set A has a higher average."
No figures, and 'average' doesn't say which one. A confirmed real error record puts it plainly: "surprisingly too many students failed to give supporting figures" (Jun 2024, Q1(e)) — this is exactly that failure.
- C"The two data sets are broadly similar."
This doesn't even attempt the comparison the question asked for (typical time), let alone give a direction or a figure — a technically-vague observation like this scores nothing.
- D"Data set A shows a positive correlation with data set B."
Correlation measures the relationship between two paired VARIABLES measured on the same individuals — it doesn't apply to comparing two separate groups' typical values at all, so this answer is off-topic rather than merely imprecise.
Traps tested: Comparison missing supporting figures · Comparison too vague no direction · Irrelevant statistical concept invoked
Practice this for real
This site teaches the mechanism; the exam is sat on Pearson's own real questions. Go find and attempt these yourself — nothing here substitutes for actually sitting a timed paper.
- Examiner report
- Jan 2021 · Q2 — cited directly in this lesson
Select International Advanced Level → Mathematics → any series, then look for WST01.
Up next
Elementary probability, and conditional-probability notation
P(A \mid B) is not P(A) divided by P(B) — the bar means something has already happened, and 'something has already happened' means the sample space itself just got smaller. Every genuine sample-space question, at bottom, is a counting question: list the possibilities carefully, count the ones you want, divide by the ones that are possible. Conditional notation asks you to do exactly that counting inside a SHRUNK sample space — restricted to whatever's on the right of the bar — and the single most repeatable way to get it wrong, confirmed on real WST01 papers, is skipping that restriction and dividing two numbers that were never counted from the same space to begin with.
50 min