Mathematical modelling in probability and statistics
~30 min · WST01 · 1.1
WST01 · 1.1 · 30 min
"Assume the die is fair" is not a fact about the die — it's a decision about which maths you're allowed to do. Spec item 1.1 is genuinely the thinnest, quietest topic on this paper: across all fourteen WST01 series reviewed for this course, not one question tests "the basic ideas of mathematical modelling" on its own — every direct hit for the word "model" turns out to be about a regression line, a different spec point entirely. And yet the assessment objective this content maps to, AO3, is worth 15 to 20 of this paper's 75 marks — nearly a fifth of it, sitting quietly inside almost every other question on the script. This lesson exists to make that content visible on its own terms, since no single past-paper question ever will.
Before you read on
Two or three questions on exactly what this lesson teaches. Being wrong here is fine — it's the fastest way to find out what to pay attention to next.
What a model actually is: assumption, parameter, refinement
Spec item 1.1's own text is a single sentence: "The basic ideas of mathematical modelling as applied in probability and statistics." It sits first, unit-wide, before representation of data, before probability itself — and the unit's own description (spec S1.1) opens with it too: "Mathematical models in probability and statistics; representation and summary of data; probability; correlation and regression; discrete random variables; discrete distributions; the Normal distribution." Read plainly, everything listed after the semicolon is an APPLICATION of the first phrase, not a separate concern from it. Worth being upfront about, in the same spirit this course insists on for every thin topic: not one of the fourteen WST01 series reviewed for this course's own research pass anchors a full question on "the basic ideas of mathematical modelling" by itself — every direct hit for the word "model" or "modelling" inside those fourteen papers turned out to be about a REGRESSION model (spec 4.1/4.2), a different, better-anchored spec point with its own lesson (regression-gradient-interpretation-extrapolation.ts). This lesson exists anyway, on spec 1.1's own terms, because — as the section below on AO3 makes concrete — the marks for this content don't disappear just because no single question is labelled with them.
Three pieces of vocabulary do almost all the work, and none of them is a formula to memorise: a model is a simplified description of a real situation, built by deliberately choosing what to ignore so the rest becomes calculable — "the die is fair," "journey times are Normally distributed," "the two events are independent" are all models in exactly this sense, even though they sound like plain factual statements. An assumption is the specific, statable choice a model rests on — usually one sentence, sometimes implicit in how a question is set up rather than spelled out. A parameter is what pins down exactly WHICH member of a family of models is in play, once the family itself has been chosen by assumption: assume Normality, and and are the parameters that say which particular Normal curve; assume a die is fair with faces, and itself (plus which values are printed on those faces) is what pins down which discrete uniform distribution applies.
"Refinement" is the fourth piece, and it's explicitly named in the assessment objective this content maps to, not invented for this lesson: AO3 credits candidates who "recall/select/use knowledge of standard mathematical models; interpret results from models in terms of the original situation, INCLUDING ASSUMPTIONS AND REFINEMENT" (spec pp. 68-69, AO3 row for S1, emphasis added). Refining a model means relaxing or replacing an assumption that turns out not to hold well enough, in favour of one closer to reality — usually at the cost of being harder, or impossible, to calculate with cleanly. A biased die can no longer be described by a single number ( for every face); it needs a full table of possibly-unequal probabilities instead, one for each face, still constrained to sum to 1 (spec 5.2's own requirement for any discrete probability function) but no longer forced to be equal. That's a genuinely worse trade for calculation, and a genuinely better one for realism — which is exactly the trade every refinement in this unit makes.
Why a topic with no past-paper question of its own is still worth 15-20 marks
Here is the honest shape of the evidence this lesson is built from, stated plainly rather than dressed up: this course's own research pass searched every question paper across fourteen WST01 series for the words "model" and "modelling," and found no standalone question testing spec 1.1 in isolation anywhere in that record. Its own conclusion, worth quoting directly because it's the clearest single sentence this facts bank has for this exact topic: spec item 1.1 "appears to function as context framing embedded throughout the paper... rather than as its own question." That's not a Pearson statement — it's this course's own research finding, and it matters practically, not just as a piece of trivia: it means revising spec 1.1 by hunting for a past "define a mathematical model" question is a search that will come back empty, every time.
What DOES show up, checkably, is the assessment-objective weighting behind this content. AO3 — "recall/select/use knowledge of standard mathematical models; interpret results from models in terms of the original situation, including assumptions and refinement" — carries 15 to 20 of this paper's 75 marks, roughly a fifth of the whole exam (spec pp. 68-69). That range is unusually high for this qualification: the Pure Mathematics units cap AO3 far lower on the same spec table — P1 at 5-15, P2 through P4 at just 5-10 — while Statistics 1 sits at 15-20, a range shared by only two other units on the whole qualification, Mechanics 1 and Decision 1, where interpreting a model in context is weighted almost as heavily as raw calculation technique. This course's own research pass draws the connection explicitly, and is careful to flag it as its own analysis rather than a claim Pearson states outright: the AO3/AO4 marks scattered through worded-interpretation parts — regression-gradient meaning, comparing distributions, judging whether a prediction is reliable — are "consistently flagged across every series reviewed as the parts students drop marks on, precisely because AO3/AO4 content is being under-taught relative to its real weight on the paper." A modelling-assumption sub-part rarely announces itself; it arrives disguised as the last two marks of a probability, discrete-random-variable, or Normal-distribution question that looked, up to that point, like pure calculation.
Mechanism
The licensing chain: assumption → distributional family → parameters → calculation
Every calculation in this unit that starts from a modelling assumption follows the same four-step chain, whether the question ever spells the steps out or not. Step 1, the ASSUMPTION: a specific, statable claim about the situation — "the die is fair," "the events are independent," "the underlying variable is Normally distributed." This is a choice, made either explicitly by the question or as a standard convention this unit works under unless told otherwise; it is never something the arithmetic itself can prove. Step 2, the FAMILY the assumption licenses: fairness (every outcome equally likely) licenses the discrete uniform distribution; Normality licenses the Normal distribution; independence licenses in place of the more general (this exact substitution — reaching for independence without it being licensed — is part of what this course's own research records as the richest, most consistently-tested sub-topic in the whole unit, appearing in a different guise in all five examiner reports reviewed, and has its own full lesson: conditional-probability-independence-vs-mutually-exclusive.ts). Step 3, the PARAMETERS that pin the family down to one specific member: for a fair -sided die, that's just itself and the values printed on its faces; for a Normal model, it's and . Step 4, the CALCULATION: everything this unit actually asks you to compute — a probability, an expectation, a standardised z-value — is arithmetic performed on the family and parameters step 1 licensed, not on the real situation directly. The chain only runs forward from a genuine assumption; it never runs the other way. A model can be undermined by evidence found AFTER a calculation is done — a box plot showing skew where Normality was assumed (spec 2.4, outliers-boxplots-comparing-distributions.ts's own territory), a regression line asked to predict miles outside the range of it was fitted to (spec 4.2, regression-gradient-interpretation-extrapolation.ts's own territory) — but that evidence never rewrites the arithmetic itself; it only tells you the assumption that licensed the arithmetic in the first place needs revisiting. Confusing "the calculation was done correctly" with "the model was appropriate" is exactly the gap this whole spec point exists to close, and it's a distinct question every single time: was the METHOD right, and separately, was the ASSUMPTION the method rested on actually reasonable for this situation?
Worked, in full
Reading the modelling step underneath a real WST01 question — the Jan 2025 dice (spec 1.1, working silently beneath spec 5.4)
- 01
A real WST01 question (Jan 2025, Q1) begins: a four-sided die, with faces numbered 1, 2, 3, 4, is rolled once, and its score is modelled by the discrete random variable . Before a single probability can be written down, a modelling decision has to be made, even though the question never asks for it explicitly: is the die assumed fair? Every real die has manufacturing tolerances, wear, and an asymmetric hole where the numbers are printed — none of that is checked here. "Fair" is the standard convention this whole unit works under for any die or coin unless a question states otherwise, and this question relies on it silently from its very first sentence.
Earns: Nothing on the real mark scheme — this step is never asked for directly, which is exactly the point spec 1.1 is making: the assumption is doing real work while remaining invisible on the page.
- 02
Because each of the four faces is assumed equally likely, takes each value in with the same probability, — precisely the definition of the discrete uniform distribution (spec 5.4). On the real question this is taken from, naming this distribution is itself a credited step.
Earns: B1 — for correctly naming the distribution, credited on the real mark scheme as the answer "Discrete uniform" (Jan 2025, Q1(a)).
- 03
Use the model to compute: — matching the real mark scheme's own stated answer for this same question's very next part (Jan 2025, Q1(b)).
Earns: The marks attached to Q1(b) reward exactly this substitution. Once the model (stage 2) is accepted, everything from here is arithmetic on numbers the model itself supplied.
- 04
Make the dependency explicit — the actual spec-1.1 payoff hiding underneath a question that's officially about spec 5.4 and 3.1. Every number in stage 3 is downstream of the single assumption made in stage 1. If red-die fairness failed — say a manufacturing flaw made "4" land noticeably more often than the other three faces — would no longer be discrete uniform, through would no longer each be exactly , and would not be , even though the ARITHMETIC MOVE in stage 3 (add the two lowest scores' probabilities) would look identical on the page. Nothing about how you calculate changes; what changes is whether the numbers you're calculating with are correct in the first place.
Earns: Nothing on the real mark scheme — this is the step spec 1.1 tests by proxy, through the credit already given at stage 2, not through a separate question anywhere in this unit's own past-paper record.
Source — Mark scheme, Jan 2025
"Discrete uniform"
In your own words
In one sentence: why does "the die is fair" have to be treated as an assumption you're licensed to use, rather than as a fact you've verified — and what practical difference does that make to how you'd answer a "critique this model" question if one appeared?
Complete it yourself
Complete the chain — what changes, and what can't be said for certain, when a fairground spinner's own fairness is called into question
- 01
A fairground spinner has 8 equal sectors, numbered 1 to 8. It is spun once, and its score is modelled by the discrete random variable . Because the sectors are assumed equal (the modelling assumption made here), is modelled as discrete uniform on : for each .
- 02
Using this model, .
Marked, line by line
A game at a school fête uses a spinner with five equal sectors, labelled 1, 2, 3, 4, 5. A player spins once; the score is modelled by the discrete random variable . (a) State the modelling assumption needed to treat as discrete uniform on , and hence write down for each value of . (2) (b) Using this model, find . (2) (c) Over many fête days, the organiser keeps a tally and notices sector 5 is landed on noticeably more often than the model in part (a) predicts, while the other four sectors land close to equally often among themselves. Explain what this evidence suggests about the assumption made in part (a), and state what would need to change in the model as a result. (2) (d) Without carrying out any further calculation, state whether the answer to part (b) can be said with confidence to become larger, smaller, or whether it's impossible to say for certain, once the model is refined as described in part (c). (1) — VERIDIAN-original question; not a reproduction of any past-paper question. Built specifically because WST01-verified-facts.md records no real spec-1.1 question this course could anchor a marked-solution to (see this lesson's own header note) — every figure below was computed and checked before being written into this file.
7 marks available
(a) — 2 marks
- 01B1
"The spinner is fair (unbiased), so each of the five sectors is equally likely to be the one it stops on."
Independent mark for stating the specific assumption being made — equally likely sectors — not a vaguer claim like "the spinner works normally."
- 02B1
is discrete uniform on : for .
Independent mark for correctly translating the stated assumption into the distribution it licenses, with all five probabilities given.
(b) — 2 marks
- 101M1
Method mark for identifying and summing the correct two outcomes.
- 102A1
Accuracy mark, correct answer only.
(c) — 2 marks
- 201B1
"The evidence suggests the 'each sector equally likely' assumption specifically fails for sector 5 — not for the spinner as a whole, since the other four sectors are reported as close to equal among themselves. The spinner appears biased toward landing on 5, so treating as discrete uniform across all five values is no longer an appropriate model."
Independent mark for connecting the conclusion to what was SPECIFICALLY reported (sector 5 alone, others roughly equal among themselves) rather than a generic "the spinner might be biased" statement.
- 202B1
"For the model to be refined, would need to be modelled as GREATER than , with the other four probabilities adjusted so that all five still sum to 1 (as any discrete probability function must, spec 5.2) — even though the exact new values can't be pinned down from 'noticeably more often' alone."
Independent mark for stating what the refined model would need to satisfy (P(S=5) increased, sum-to-1 preserved) without overclaiming specific numbers the evidence doesn't support.
(d) — 1 mark
- 301B1
"Impossible to say for certain from the evidence given. increasing would, by itself, push up — but the same evidence says nothing about specifically (the other four sectors were only reported as landing close to equally often AMONG THEMSELVES, not as each unchanged from ). Since depends on both terms, a stated change in one without information about the other doesn't settle the total."
Independent mark for correctly declining to over-conclude — recognising that a change to one term of a sum doesn't determine the sum's direction without information about the other term.
Named traps
- spec-1.1-expected-as-a-standalone-question
- This course's own research pass is direct about what it found (and didn't find) searching all fourteen reviewed WST01 series for the words "model"/"modelling": no standalone question tests spec 1.1 in isolation anywhere in that record — every direct hit turns out to be about a regression model instead (spec 4.1/4.2). This isn't an examiner-report-documented misconception the way most trap-taxonomy items on this course are; it's a genuine, checkable finding about how this content is actually examined, and it produces a real exam-navigation trap in its own right: revising this spec point by looking for a past "define a mathematical model" question to practise on is a search that will come back empty every time, because the AO3 marks this content maps to (15-20 of 75) are distributed as sub-parts riding on top of other questions, not concentrated in a question of their own.
- assumption-treated-as-a-provable-fact
- VERIDIAN-original naming for a genuine failure mode this lesson's own material is built to guard against, not a quoted examiner misconception (none exists in the reviewed record for this specific spec point). "The die is fair" and "growth is Normally distributed" are grammatically identical to plain statements of fact, and it's easy to read them that way — as something the question has already established, rather than something it's asking you to accept for the purpose of the calculation that follows. The practical cost: a student who reads an assumption as a fact has nothing to say when a later part of the same question hands them evidence against it (see the marked-solution's part (c) above), because they never registered there was anything provisional about the claim in the first place.
- refinement-answer-not-connected-to-the-specific-evidence-given
- VERIDIAN-original naming, illustrated directly by this lesson's own marked-solution common-wrong-path above. A true-sounding, generically applicable statement — "real spinners are never perfectly fair," "real dice have manufacturing tolerances" — earns nothing on a "critique this model" or "explain what this evidence suggests" part, however scientifically reasonable it sounds, if it isn't tied to the SPECIFIC observation the question actually reported. The credited answer names what was specifically observed (sector 5 over-represented, not the spinner in general) and states the specific consequence for the model (which probability moves, and why the total must still sum to 1) — not a general truth that would apply equally well to any question of this shape.
Retrieval — with feedback on every choice
In the real WST01 question this lesson's worked chain is built from (Jan 2025, Q1), a second die — blue, with faces numbered 1, 3, 5, 7 — is also assumed fair, and its score is modelled by . Which of these correctly describes 's distribution, and why?
A question models daily rainfall as Normally distributed, then later states that real data shows a small number of days with far higher rainfall than the model predicts as at all likely, while ordinary days fit the model well. Which response best identifies what's happening, in the terms this lesson has used?
Spec 1.1 is examined through AO3, worth 15-20 of this paper's 75 marks — and the reviewed past-paper record shows this content is embedded inside other questions rather than tested as a standalone question of its own. What does this mean practically for how to revise it?
- A model = a simplified, assumption-based description used because it's tractable — never a proven fact.
- Assumption → distributional family → parameters → calculation. Wrong assumption = wrong numbers, even with perfect arithmetic.
- "Fair die/coin/spinner" = each outcome assumed equally likely → discrete uniform. Values needn't be consecutive integers.
- Spec 1.1 has no standalone past-paper question in the reviewed record — its AO3 marks (15-20/75) sit inside other questions.
- A "critique this model" or refinement answer must name the SPECIFIC evidence given, not a generic "nothing's perfect" statement.
- One probability changing forces others to change too (sum to 1, spec 5.2) — without more data, you can't say which, or by how much.
Not affiliated with or endorsed by Pearson Edexcel. The spec-1.1 wording, the unit description, and the AO3 table row are transcribed verbatim from WST01-verified-facts.md §1 and §2b, themselves checked against the official spec PDF. The finding that no standalone spec-1.1 question exists across the fourteen WST01 series reviewed, and the connection drawn between AO3's weighting and worded-interpretation marks being dropped paper-wide, are this research pass's own documented conclusions — quoted and paraphrased honestly as such, not presented as a Pearson statement. The ONE real numeric anchor in this lesson is the Jan 2025 Q1 dice question: "Discrete uniform" (the credited answer for naming R's distribution, Q1(a)) is a genuine quoted mark-scheme fragment, and P(R<3) = 1/2 (Q1(b)) is the real recorded answer, independently re-derived here from first principles to confirm consistency; the blue die's face values (1, 3, 5, 7), reused in mcq-1 above, are also real, taken from the same question. This lesson deliberately does not use or invent figures for Jan 2025 Q1's other sub-parts (c)-(g), whose existence but not their answers is recorded in the research bank. Every other numeric scenario in this lesson — the marked-solution's five-sector fête spinner and its P(S≥4) = 2/5, the chain-drill's eight-sector fairground spinner and its P(T>6) = 1/4, and every prequestion/MCQ scenario not built on the real dice fact — is VERIDIAN-original, computed and hand-checked before being written into this file, not a reproduction of any real Pearson question. No WarrantCheck gate is used anywhere in this lesson: that gate requires a real, single-series, examiner-report-sourced citation of a candidate faking a "show that" derivation, which the research bank does not supply for this spec point, and none of this lesson's marked-solution parts are printed-answer ("show that") items in the first place. The B1/M1/A1 mark allocations attached to the VERIDIAN-original questions are modelled on the general marking conventions verified in WST01-verified-facts.md §4, not transcribed from a real mark scheme, which for an original qualitative question does not exist.
In the real WST01 question this lesson's worked chain is built from (Jan 2025, Q1), a second die — blue, with faces numbered 1, 3, 5, 7 — is also assumed fair, and its score is modelled by . Which of these correctly describes 's distribution, and why?
- is discrete uniform on , since the fair-die assumption makes each of the four printed face-values equally likely — the four values themselves don't need to be consecutive integers for the discrete uniform model to apply
Correct. The discrete uniform distribution is defined by "every outcome in the set is equally likely," not by the outcomes being 1, 2, 3, 4,... in order. The fair-die assumption licenses equal probability across whatever four values are actually printed on the faces — here, 1, 3, 5 and 7.
- B cannot be discrete uniform, because its values (1, 3, 5, 7) aren't consecutive integers
This invents a requirement the discrete uniform distribution doesn't have. "Uniform" describes the PROBABILITIES (all equal), not the spacing of the values themselves — a fair die numbered 1, 3, 5, 7 is exactly as much a discrete uniform random variable as one numbered 1, 2, 3, 4.
- C is discrete uniform only once it's independently proven that the blue die is fair, which the question hasn't done
This is the same error prequestion 2 above already named: a modelling assumption doesn't need independent verification to be used — the question states it, and that statement is exactly what licenses the model. Requiring proof before using a stated assumption misunderstands what an assumption is for.
- D should be modelled as Normal, since with more than two possible outcomes the distribution starts to approximate a bell curve
A discrete random variable taking four specific, equally likely values is not well modelled by a continuous, symmetric bell curve — it's four spikes of equal height, not a smooth curve, whatever the number of outcomes. The number of outcomes alone never determines which family is appropriate; the assumption made about them does.
Traps tested: Discrete uniform assumed to require consecutive values · Assumption treated as requiring independent verification · Wrong distribution family invented
A question models daily rainfall as Normally distributed, then later states that real data shows a small number of days with far higher rainfall than the model predicts as at all likely, while ordinary days fit the model well. Which response best identifies what's happening, in the terms this lesson has used?
- The extreme days are evidence against the "Normally distributed" assumption specifically for the tail of the distribution — ordinary days still fit, so the issue is with how the model handles rare, extreme values, not a general failure of the whole model
Correct. This names exactly what was reported (ordinary days fit; extreme days don't) rather than a vaguer verdict on the model as a whole — the same connected-to-the-evidence discipline the marked-solution above is built around.
- BNothing is wrong with the model — a Normal-distribution assumption can never be disproved by a small number of days
This treats the assumption as unfalsifiable, the exact error prequestion 2 above names directly. A stated modelling assumption can absolutely be shown to sit poorly with real evidence, however few data points that evidence involves — a handful of days far outside what a Normal model predicts as plausible is real evidence, not nothing.
- C"No real-world data is ever perfectly Normal, so this is unremarkable and doesn't need commenting on"
True of essentially any Normal-distribution question that could ever be set, which is exactly why it isn't a useful answer here — it doesn't engage with the SPECIFIC pattern reported (ordinary days fitting, extreme days not), the same disconnected-generic-statement trap this lesson's marked-solution common-wrong-path shows costing marks.
- DThe extreme days should be removed from the data set so the Normal model fits again
Deleting inconvenient data to make a model fit is not refining the model — it's hiding evidence the model doesn't cover well. This course's own treatment of outliers elsewhere makes the same point from a different angle: an unusual value is reported and reasoned about, never simply discarded.
Traps tested: Model assumption treated as unfalsifiable · Disconnected generic statement not tied to the evidence · Evidence against model handled by deleting data
Spec 1.1 is examined through AO3, worth 15-20 of this paper's 75 marks — and the reviewed past-paper record shows this content is embedded inside other questions rather than tested as a standalone question of its own. What does this mean practically for how to revise it?
- This spec point isn't likely to arrive as an obvious, labelled "define a model" question — the marks show up as sub-parts riding on probability, discrete-RV, or Normal-distribution questions, so recognising a modelling-assumption sub-part when it appears (rather than waiting for one clearly flagged as such) is the actual skill worth practising
Correct. This is the practical consequence of the finding this whole lesson is built around: the content is real and worth real marks, but it doesn't announce itself the way most spec points do on this paper.
- BSince no standalone question has ever tested it, it's safe to skip this content in revision entirely
This confuses "no question is labelled with this content" with "no marks depend on this content" — 15-20 of 75 marks is a large share of the paper to write off based on question FREQUENCY rather than mark VALUE.
- CIt means a standalone question on this exact spec point is now overdue and likely to appear on a future paper
There's no basis in the reviewed evidence for predicting a specific future question from an absence in fourteen past series — and the evidence actually points the other way: this content's own AO3 marks are structurally distributed across other question types, not saved up for a standalone question that simply hasn't happened yet.
- DAO3's 15-20 marks are entirely accounted for by Normal-distribution questions alone, since that's the richest single section in this course's research record
The Normal distribution is genuinely the richest single sub-topic in this unit's own examiner-report record, but AO3's marks are described as scattered across the whole paper (regression interpretation, comparing distributions, and more, alongside the Normal distribution) — attributing an entire assessment objective to one topic overstates what one topic's own strong evidence actually supports.
Traps tested: Low question frequency mistaken for low mark value · Absence in sample mistaken for overdue appearance · Single topic assumed to account for a whole ao
Practice this for real
This site teaches the mechanism; the exam is sat on Pearson's own real questions. Go find and attempt these yourself — nothing here substitutes for actually sitting a timed paper.
- Mark scheme
- Jan 2025 · Q1(a) — cited directly in this lesson
Select International Advanced Level → Mathematics → any series, then look for WST01.
Up next
Reading Data Representations, and Comparing Distributions in Words
A histogram's bars can lie to you at a glance. The vertical axis reads frequency DENSITY, not frequency — and the two are only the same number when every class happens to be the same width, which a real exam question is under no obligation to give you. The single most repeated error the mark schemes for this section record is exactly this: reading a bar's height straight off the axis and writing it down as the answer, skipping the one multiplication — by class width — that turns a density into an actual count of people. A stem-and-leaf diagram hides a quieter version of the same problem: read it from the wrong end, and Q1 and Q3 swap places without anything on the page telling you they have. And once you've read the real figures off either diagram, a "compare the two distributions" question is asking you to use them — name a statistic, give both figures, answer what was actually asked — not just to have found them.
50 min