Mathematical modelling in probability and statistics
"Assume the die is fair" is not a fact about the die — it's a decision about which maths you're allowed to do. Spec item 1.1 is genuinely the thinnest, quietest topic on this paper: across all fourteen WST01 series reviewed for this course, not one question tests "the basic ideas of mathematical modelling" on its own — every direct hit for the word "model" turns out to be about a regression line, a different spec point entirely. And yet the assessment objective this content maps to, AO3, is worth 15 to 20 of this paper's 75 marks — nearly a fifth of it, sitting quietly inside almost every other question on the script. This lesson exists to make that content visible on its own terms, since no single past-paper question ever will.
The card
A model = a simplified, assumption-based description used because it's tractable — never a proven fact. Assumption → distributional family → parameters → calculation. Wrong assumption = wrong numbers, even with perfect arithmetic. "Fair die/coin/spinner" = each outcome assumed equally likely → discrete uniform. Values needn't be consecutive integers. Spec 1.1 has no standalone past-paper question in the reviewed record — its AO3 marks (15-20/75) sit inside other questions. A "critique this model" or refinement answer must name the SPECIFIC evidence given, not a generic "nothing's perfect" statement. One probability changing forces others to change too (sum to 1, spec 5.2) — without more data, you can't say which, or by how much.
Why it works — The licensing chain: assumption → distributional family → parameters → calculation
Every calculation in this unit that starts from a modelling assumption follows the same four-step chain, whether the question ever spells the steps out or not. Step 1, the ASSUMPTION: a specific, statable claim about the situation — "the die is fair," "the events are independent," "the underlying variable is Normally distributed." This is a choice, made either explicitly by the question or as a standard convention this unit works under unless told otherwise; it is never something the arithmetic itself can prove. Step 2, the FAMILY the assumption licenses: fairness (every outcome equally likely) licenses the discrete uniform distribution; Normality licenses the Normal distribution; independence licenses in place of the more general (this exact substitution — reaching for independence without it being licensed — is part of what this course's own research records as the richest, most consistently-tested sub-topic in the whole unit, appearing in a different guise in all five examiner reports reviewed, and has its own full lesson: conditional-probability-independence-vs-mutually-exclusive.ts). Step 3, the PARAMETERS that pin the family down to one specific member: for a fair -sided die, that's just itself and the values printed on its faces; for a Normal model, it's and . Step 4, the CALCULATION: everything this unit actually asks you to compute — a probability, an expectation, a standardised z-value — is arithmetic performed on the family and parameters step 1 licensed, not on the real situation directly. The chain only runs forward from a genuine assumption; it never runs the other way. A model can be undermined by evidence found AFTER a calculation is done — a box plot showing skew where Normality was assumed (spec 2.4, outliers-boxplots-comparing-distributions.ts's own territory), a regression line asked to predict miles outside the range of it was fitted to (spec 4.2, regression-gradient-interpretation-extrapolation.ts's own territory) — but that evidence never rewrites the arithmetic itself; it only tells you the assumption that licensed the arithmetic in the first place needs revisiting. Confusing "the calculation was done correctly" with "the model was appropriate" is exactly the gap this whole spec point exists to close, and it's a distinct question every single time: was the METHOD right, and separately, was the ASSUMPTION the method rested on actually reasonable for this situation?
Traps — 3
- spec-1.1-expected-as-a-standalone-question
- This course's own research pass is direct about what it found (and didn't find) searching all fourteen reviewed WST01 series for the words "model"/"modelling": no standalone question tests spec 1.1 in isolation anywhere in that record — every direct hit turns out to be about a regression model instead (spec 4.1/4.2). This isn't an examiner-report-documented misconception the way most trap-taxonomy items on this course are; it's a genuine, checkable finding about how this content is actually examined, and it produces a real exam-navigation trap in its own right: revising this spec point by looking for a past "define a mathematical model" question to practise on is a search that will come back empty every time, because the AO3 marks this content maps to (15-20 of 75) are distributed as sub-parts riding on top of other questions, not concentrated in a question of their own.
- assumption-treated-as-a-provable-fact
- VERIDIAN-original naming for a genuine failure mode this lesson's own material is built to guard against, not a quoted examiner misconception (none exists in the reviewed record for this specific spec point). "The die is fair" and "growth is Normally distributed" are grammatically identical to plain statements of fact, and it's easy to read them that way — as something the question has already established, rather than something it's asking you to accept for the purpose of the calculation that follows. The practical cost: a student who reads an assumption as a fact has nothing to say when a later part of the same question hands them evidence against it (see the marked-solution's part (c) above), because they never registered there was anything provisional about the claim in the first place.
- refinement-answer-not-connected-to-the-specific-evidence-given
- VERIDIAN-original naming, illustrated directly by this lesson's own marked-solution common-wrong-path above. A true-sounding, generically applicable statement — "real spinners are never perfectly fair," "real dice have manufacturing tolerances" — earns nothing on a "critique this model" or "explain what this evidence suggests" part, however scientifically reasonable it sounds, if it isn't tied to the SPECIFIC observation the question actually reported. The credited answer names what was specifically observed (sector 5 over-represented, not the spinner in general) and states the specific consequence for the model (which probability moves, and why the total must still sum to 1) — not a general truth that would apply equally well to any question of this shape.
Say it out loud
Out loud, from memory, no notes: explain the licensing chain: assumption → distributional family → parameters → calculation to someone who has never seen this topic — where does your explanation get vague or hand-wavy? That's the exact spot to re-study, and it only works if you check it: read back over the mechanism above the moment you finish talking and mark precisely where you drifted from it.