Elementary probability, and conditional-probability notation

~50 min · WST01 · 3.1

WST01 · 3.1 · 50 min

P(AB)P(A \mid B) is not P(A)P(A) divided by P(B)P(B) — the bar means something has already happened, and 'something has already happened' means the sample space itself just got smaller. Every genuine sample-space question, at bottom, is a counting question: list the possibilities carefully, count the ones you want, divide by the ones that are possible. Conditional notation asks you to do exactly that counting inside a SHRUNK sample space — restricted to whatever's on the right of the bar — and the single most repeatable way to get it wrong, confirmed on real WST01 papers, is skipping that restriction and dividing two numbers that were never counted from the same space to begin with.

Before you read on

Two or three questions on exactly what this lesson teaches. Being wrong here is fine — it's the fastest way to find out what to pay attention to next.

Elementary probability: sample spaces, and counting equally likely outcomes

An experiment is any process with an uncertain result — rolling a die, drawing a card, surveying a customer. Every possible result is an outcome, and the full set of outcomes is the sample space, written Ω\Omega. An event is simply a subset of Ω\Omega: 'the die shows an even number' picks out {2,4,6}\{2, 4, 6\} from Ω={1,2,3,4,5,6}\Omega = \{1,2,3,4,5,6\}. This is genuinely all spec 3.1 asks for as a starting definition — the specification's own unit description names 'the basic ideas of mathematical modelling as applied in probability and statistics' as the whole of this spec point, and the modelling idea itself is usually just as basic: calling a die 'fair,' or a coin 'unbiased,' is an ASSUMPTION about the real object being modelled, not a fact that could be proven from the die itself — a die could always, in principle, be very slightly weighted. The research bank behind this lesson finds this assumption showing up as ordinary phrasing embedded inside real questions ('assume the die is fair') rather than as a topic tested on its own — see this lesson's own header for the honest scope of that.

When every outcome in Ω\Omega is equally likely — a fair die, a well-shuffled deck, a random draw — probability becomes pure counting: P(A)=n(A)n(Ω)P(A) = \dfrac{n(A)}{n(\Omega)}, the number of outcomes in the event divided by the total number of outcomes. This is often called classical probability, and almost every skill in this lesson, right through to conditional notation, is a variation on getting this one count right: how many outcomes are in Ω\Omega, and how many of those are in the event actually being asked about.

Getting n(Ω)n(\Omega) right depends on listing SYSTEMATICALLY, not by eye. Two genuinely different traps live here. First: when an experiment combines several independent choices — two dice, a coin and a die, three coin tosses — the sizes MULTIPLY, not add: two 4-sided dice give 4×4=164 \times 4 = 16 outcomes, not 4+4=84 + 4 = 8, because every one of the 4 first-roll results can be paired with every one of the 4 second-roll results. Second: outcomes built from more than one stage are usually ORDERED — rolling a 2 then a 6 is a different outcome from rolling a 6 then a 2, even though they'd be 'the same pair of numbers' if you stopped caring about order. Treating (2,6)(2,6) and (6,2)(6,2) as one outcome instead of two is exactly the kind of undercounting a systematic list — a table with one row per first result and one column per second — is built to prevent, because a table forces every combination to get its own cell.

From words to notation: complement, union, intersection, and the conditioning bar

Spec 3.2 gives events a small, precise vocabulary, verified directly against the real WST01 formula booklet for which parts of it are printed there and which aren't (WST01-verified-facts.md §2a). The complement AA' is 'everything in Ω\Omega that isn't in AA' — AA and AA' between them always account for the whole sample space, so P(A)=1P(A)P(A') = 1 - P(A). This rule is NOT printed on the formula sheet — the spec's own notation box states plainly that formulae 'expected to know' are left out of the booklet deliberately, and this is one of them; it costs nothing to memorise, but nothing will remind you of it in the exam either. Two events are mutually exclusive when they share no outcomes at all — AB=A \cap B = \emptyset, so P(AB)=0P(A \cap B) = 0 — a statement about whether the two SETS overlap, decided by what the events mean, before any probability is attached.

The general addition law, P(AB)=P(A)+P(B)P(AB)P(A \cup B) = P(A) + P(B) - P(A \cap B), IS printed on the formula sheet, and it's worth reading as a translation rule as much as a formula: 'or' (in the everyday, inclusive sense — 'A or B or both') means \cup, and simply adding P(A)P(A) and P(B)P(B) double-counts whatever's shared, which the P(AB)-P(A \cap B) term exists purely to correct. For mutually exclusive events specifically, that correction term is already zero, so the law simplifies to P(AB)=P(A)+P(B)P(A \cup B) = P(A) + P(B) — a genuine shortcut, but only in that one special case; using the short version when events actually overlap is exactly the double-counting trap above.

Reading a word problem into the right symbol is the whole skill spec 3.2 is testing at this level, and it comes down to a short, fixed dictionary: 'or' (inclusive) \to \cup; 'and'/'both' \to \cap; 'not' \to {}'; and — the one that trips up notation more than any other — 'given that' \to the conditioning bar, \mid. The event named IMMEDIATELY after 'given that' always goes on the RIGHT of the bar; the event actually being asked about goes on the LEFT. 'The probability a customer buys a magazine, given they've bought a newspaper' names the newspaper LAST in the sentence but writes it SECOND in the notation: P(magazinenewspaper)P(\text{magazine} \mid \text{newspaper}) — the grammar of the sentence and the order of the symbols don't have to match, and assuming they do is exactly how the two events end up swapped.

Conditional probability itself, P(AB)P(A \mid B), is defined by restricting the sample space: once you know BB has happened, every outcome outside BB is no longer possible, so the relevant 'whole' shrinks from Ω\Omega down to BB itself. Read as counting — spec 3.1's classical-probability idea, carried straight into spec 3.2's notation — this gives the cleanest possible route to the formula: P(AB)=n(AB)n(B)P(A \mid B) = \dfrac{n(A \cap B)}{n(B)}, the outcomes that are in BOTH AA and the shrunk space BB, divided by the size of BB itself. Dividing top and bottom by n(Ω)n(\Omega) turns that straight into the version on the formula sheet, P(AB)=P(AB)P(B)P(A \mid B) = \dfrac{P(A \cap B)}{P(B)} (requiring P(B)>0P(B) > 0) — the same statement, in probabilities instead of raw counts. Whenever a question hands you equally likely outcomes you can list, the counting version is usually the safer route: there's no formula to misremember, only a smaller list to count correctly.

Diagram — The sample space for two fair dice — 36 equally likely outcomes, and the region where the total is 8 or more
Red dieBlue dieTotal = 8 boundaryRed die shows 6(6, 6)(6, 2)(2, 6)

x-axis: Red die · y-axis: Blue die

Total = 8 boundary
The line separating pairs whose scores sum to fewer than 8 from pairs whose scores sum to 8 or more — the conditioning event used in the worked chain below.
Red die shows 6
Every pair with red = 6 — the event tested against the shaded region below, in the worked chain that follows this diagram.
(6, 6)
Red 6, blue 6 — total 12, the single largest-total outcome.
(6, 2)
Red 6, blue 2 — total 8, the smallest blue value that still keeps red-6 inside the shaded region.
(2, 6)
Red 2, blue 6 — total 8 — the same total as (6,2), but a DIFFERENT outcome, since the dice are ordered.

Common error: Treating (2, 6) and (6, 2) as the same outcome because they're 'the same two numbers' — undercounting the sample space at 21 unordered combinations instead of 36 ordered ones.

Correct: The two dice are distinguishable (red and blue, here) — (2, 6) means red shows 2 and blue shows 6, a genuinely different outcome from (6, 2). Every ordered pair is its own equally likely outcome, which is exactly why a grid — one axis per die — is the systematic way to list this sample space without losing any of them.

Mechanism

Why P(A|B) is never P(A) divided by P(B) — and the one case where they happen to agree

Write both calculations out in raw counts and the difference stops being a rule to memorise and becomes something you can see directly. P(AB)=n(AB)n(B)P(A \mid B) = \dfrac{n(A \cap B)}{n(B)} — the outcomes shared by AA and BB, out of BB's own count. The raw-ratio shortcut, by contrast, is P(A)P(B)=n(A)/n(Ω)n(B)/n(Ω)=n(A)n(B)\dfrac{P(A)}{P(B)} = \dfrac{n(A)/n(\Omega)}{n(B)/n(\Omega)} = \dfrac{n(A)}{n(B)}AA's ENTIRE count, out of BB's own count, with no reference anywhere to whether AA and BB actually overlap. These agree only when AB=AA \cap B = A — that is, when every single outcome in AA also happens to lie inside BB (formally, ABA \subseteq B). That's a real but narrow special case; the moment any part of AA sits OUTSIDE BB, n(AB)<n(A)n(A \cap B) < n(A), and the raw ratio overshoots. In fact this overshoot direction is guaranteed, not just typical: since ABA \cap B is always a subset of AA itself, n(AB)n(A)n(A \cap B) \le n(A) always holds, which means P(AB)P(A)P(B)P(A \mid B) \le \dfrac{P(A)}{P(B)} is true for EVERY pair of events, with equality exactly in that one subset case. The raw-ratio shortcut can never underestimate the true conditional probability — it can only match it, in that one narrow case, or inflate it, which is exactly why it so often produces a value that's suspiciously large, or that exceeds 1 outright and announces itself as impossible.

In your own words

In one sentence: what has to be true about events AA and BB for the shortcut P(A)/P(B)P(A)/P(B) to actually equal the correct value of P(AB)P(A \mid B)?

Worked, in full

Two fair six-sided dice, one red and one blue, are rolled and the scores added. Given that the total is 8 or more, find the probability the red die shows a 6.

  1. 01

    List the sample space systematically as ordered pairs (red, blue) — 6×6=366 \times 6 = 36 equally likely outcomes, since each of the 6 red faces pairs with each of the 6 blue faces. Let CC = 'the total is 8 or more' (the given, conditioning event) and AA = 'the red die shows a 6'. The target is P(AC)=n(AC)n(C)P(A \mid C) = \dfrac{n(A \cap C)}{n(C)}, found by counting directly — every outcome is equally likely, so this is exactly spec 3.1's classical probability, applied to a conditioned event.

    Earns: M1 — correctly identifies the sample space size (36) and sets up the target as a conditional probability via counting, not via a memorised probability formula alone.

  2. 02

    Count n(C)n(C) by listing every pair for every qualifying total. Total 8: (2,6),(3,5),(4,4),(5,3),(6,2)(2,6),(3,5),(4,4),(5,3),(6,2) — 5 pairs. Total 9: (3,6),(4,5),(5,4),(6,3)(3,6),(4,5),(5,4),(6,3) — 4 pairs. Total 10: (4,6),(5,5),(6,4)(4,6),(5,5),(6,4) — 3 pairs. Total 11: (5,6),(6,5)(5,6),(6,5) — 2 pairs. Total 12: (6,6)(6,6) — 1 pair. n(C)=5+4+3+2+1=15n(C) = 5+4+3+2+1 = 15.

    Earns: M1 — systematically lists pairs for every qualifying total, not just one. A1 — correct count, 15.

  3. 03

    Count n(AC)n(A \cap C) by restricting the stage-2 list to red =6= 6: (6,2),(6,3),(6,4),(6,5),(6,6)(6,2),(6,3),(6,4),(6,5),(6,6) — all five already appear in the stage-2 list, since a red 6 guarantees a total of at least 7, and every blue value of 2 or more then pushes it to 8 or above. n(AC)=5n(A \cap C) = 5.

    Earns: M1 — correctly restricts the existing list rather than recounting from scratch. A1 — correct count, 5.

  4. 04

    P(AC)=n(AC)n(C)=515=13P(A \mid C) = \dfrac{n(A \cap C)}{n(C)} = \dfrac{5}{15} = \dfrac{1}{3}.

    Earns: A1 — correct final value, 13\frac{1}{3}, reached entirely by counting, with no formula to misremember.

  5. 05

    Compare this against the raw-ratio shortcut, to see exactly how it fails here: P(A)P(C)=6/3615/36=615=25=0.4\dfrac{P(A)}{P(C)} = \dfrac{6/36}{15/36} = \dfrac{6}{15} = \dfrac{2}{5} = 0.4 — a value that looks like a perfectly ordinary probability, not an obviously impossible one. That's what makes this version of the error more dangerous than one that lands above 1: nothing about 0.40.4 announces itself as wrong. AA (red shows 6) is NOT a subset of CC (total 8\ge 8) here — a red 6 with a blue 1 gives a total of only 7 — and the mechanism block above proves that's exactly when the shortcut overshoots the true value, which it does: 0.4>130.4 > \frac{1}{3}.

    Earns: Nothing further on the mark scheme — full marks were already earned by stage 4. Included because this exact failure — a raw-ratio answer that happens to look valid — is the harder-to-catch cousin of a mistake a real WST01 examiner report documents in its more obviously-wrong form: "we occasionally still saw P(0.35∩0.4)/0.4 or 0.35/0.4" (Jan 2021, Q1(c)).

Source — Examiner report, Jan 2021

"we occasionally still saw P(0.35∩0.4)/0.4 or 0.35/0.4"

Complete it yourself

Complete the chain — a fair 8-sided die (faces 1-8) is rolled and a fair coin is tossed. Given that the coin shows Heads, find the probability the die shows a number greater than 5.

  1. 01

    List the sample space systematically: each outcome pairs one of 8 die values with one of 2 coin results, giving 8×2=168 \times 2 = 16 equally likely outcomes — (1,H),(1,T),(2,H),,(8,T)(1,H), (1,T), (2,H), \ldots, (8,T).

  2. 02

    Let BB = 'the coin shows Heads' (the given, conditioning event) and AA = 'the die shows a number greater than 5'. The target is P(AB)=n(AB)n(B)P(A \mid B) = \dfrac{n(A \cap B)}{n(B)}, found by counting inside the systematic list from stage 1.

Marked, line by line

A box contains 20 cards, numbered 1 to 20. A card is drawn at random. Let AA = 'the number is a multiple of 4' and BB = 'the number is a multiple of 5'. (a) By listing, find P(A)P(A). (2) (b) Find P(A)P(A'). (1) (c) Find P(AB)P(A \cup B), and check your answer by listing the numbers in ABA \cup B directly. (3) (d) Given that the card drawn shows a multiple of 5, find the probability it also shows a multiple of 4. (2) — VERIDIAN-original question and dataset; not a reproduction of any past-paper question. The 1-to-20 card context is built specifically to combine spec 3.1's systematic-listing skill with spec 3.2's notation (complement, union, conditional) inside one multi-part item, the way a real WST01 Section 3 question routinely combines several sub-skills at once.

8 marks available

(a)2 marks

  1. 01

    The multiples of 4 from 1 to 20: 4, 8, 12, 16, 20 — so n(A)=5n(A) = 5 out of n(Ω)=20n(\Omega) = 20.

    Method mark for listing the event systematically rather than estimating a count.

    M1
  2. 02

    P(A)=520=14P(A) = \dfrac{5}{20} = \dfrac{1}{4}

    Accuracy mark, correct answer only, in simplest form.

    A1

(b)1 mark

  1. 101

    P(A)=1P(A)=114=34P(A') = 1 - P(A) = 1 - \dfrac{1}{4} = \dfrac{3}{4}

    Independent mark — the complement rule applied directly to the value already found in (a); no separate listing is needed, since 'not a multiple of 4' is just everything else.

    B1

(c)3 marks

  1. 201

    Multiples of 5 from 1 to 20: 5, 10, 15, 20 — n(B)=4n(B) = 4, so P(B)=420=15P(B) = \frac{4}{20} = \frac{1}{5}. The only number that's a multiple of both 4 and 5 is 20 itself, so n(AB)=1n(A \cap B) = 1 and P(AB)=120P(A \cap B) = \frac{1}{20}.

    Method mark for finding P(B) and P(A ∩ B) by listing, both needed before the addition law can be applied.

    M1
  2. 202

    P(AB)=P(A)+P(B)P(AB)=14+15120=520+420120=820=25P(A \cup B) = P(A) + P(B) - P(A \cap B) = \dfrac{1}{4} + \dfrac{1}{5} - \dfrac{1}{20} = \dfrac{5}{20} + \dfrac{4}{20} - \dfrac{1}{20} = \dfrac{8}{20} = \dfrac{2}{5}

    Accuracy mark for the correct value via the addition law, in simplest form.

    A1
  3. 203

    Checking directly by listing ABA \cup B: 4, 5, 8, 10, 12, 15, 16, 20 — exactly 8 numbers out of 20, confirming 820=25\frac{8}{20} = \frac{2}{5}.

    Independent mark for the direct listing check the question specifically asks for — a genuinely useful habit for spotting a wrong addition-law substitution, not just a formality.

    B1

(d)2 marks

  1. 301

    Given a multiple of 5, restrict attention to B={5,10,15,20}B = \{5, 10, 15, 20\} (n(B)=4n(B) = 4). Of those, only 20 is also a multiple of 4, so n(AB)=1n(A \cap B) = 1.

    Method mark for correctly restricting to the conditioning event B before counting the favourable outcomes within it.

    M1
  2. 302

    P(AB)=n(AB)n(B)=14P(A \mid B) = \dfrac{n(A \cap B)}{n(B)} = \dfrac{1}{4}

    Accuracy mark, correct answer only.

    A1

Named traps

conditional-probability-as-raw-ratio-of-marginals
Verified directly, quoted from a real WST01 conditional-probability question: "we occasionally still saw P(0.35∩0.4)/0.4 or 0.35/0.4" (Jan 2021, Q1(c)) — dividing two given probabilities directly, as though P(A|B) meant P(A) ÷ P(B) rather than P(A ∩ B) ÷ P(B). The mechanism block above proves this is never a harmless shortcut: it can only match the correct value (when A is entirely contained in B) or overshoot it — never undershoot — which is exactly why it so often lands above 1, or on a plausible-looking but wrong number just below it.
conditioning-probability-refolded-into-the-numerator
A second, verified instance of conditional-probability notation going wrong, checked directly against the real Pearson examiner-report PDF rather than taken on trust from a summary of it (Jan 2023, Q2(d), a tree-diagram question): candidates were finding P(A|B) from a tree diagram where the correct denominator, P(B) = 61/234, had already been found in an earlier part, and the correct numerator was the single branch product 5/9 × 4/8 × 8/13 = 20/117. The two malformed answers the examiner report actually quotes both use that 61/234 CORRECTLY as the denominator — the error is entirely in the numerator, where an extra copy of the same 61/234 gets folded in: one wrote (5/9 × 4/8 × 8/13 + 61/234)/(61/234), another wrote (5/9 × 4/8 × 8/13 × 61/234)/(61/234). Different arithmetic from the raw-ratio trap above, but the same root confusion: not treating P(A∩B) as one clean, self-contained quantity, separate from whatever's about to divide it — instead letting the denominator's own value leak back into the numerator.

Retrieval — with feedback on every choice

Question 1
2 marks

A weather forecaster writes P(rain tomorrowcloudy today)=0.6P(\text{rain tomorrow} \mid \text{cloudy today}) = 0.6. What does this number represent?

Question 2
2 marks

A card is drawn from a standard 52-card deck (4 Kings, 4 Queens, 44 others). Given that the card is a King or a Queen, find the probability it is a King.

Question 3
3 marks

Three fair coins are tossed. By listing the sample space systematically, find P(exactly two heads)P(\text{exactly two heads}).

Reference — not a study method, a lookup
  • Ω = the full sample space; an event is a subset of it. Classical probability: P(A) = n(A)/n(Ω), for equally likely outcomes.
  • Combining independent stages: sizes MULTIPLY (4×4=16), not add. Multi-stage outcomes are usually ORDERED — (2,6) ≠ (6,2).
  • Words → notation: 'or' → ∪, 'and'/'both' → ∩, 'not' → ′, 'given that' → | (the event after 'given' goes on the RIGHT of the bar).
  • On the formula sheet: P(A∪B) = P(A)+P(B)−P(A∩B). Not on it: P(A′) = 1−P(A) — memorise it, one line.
  • P(A|B) = P(A∩B)/P(B) = n(A∩B)/n(B) for equally likely outcomes. NEVER P(A)/P(B) — that drops the intersection entirely.
  • P(A)/P(B) only ever equals P(A|B) when A is entirely contained in B. Otherwise it strictly overshoots — check: is the answer > 1?

Not affiliated with or endorsed by Pearson Edexcel. Both examiner-report citations in this lesson were checked directly against the real Pearson PDFs during a review pass, not just against WST01-verified-facts.md's own summaries of them. Jan 2021 Q1(c) (wst01-01-pef-20210304) matches the research bank's transcription exactly: "we occasionally still saw... 0.35/0.4." Jan 2023 Q2(d) (wst01-01-pef-20230302) does NOT match the research bank's summary — the research bank describes its two malformed expressions as 'both missing the correct conditioning denominator,' but the actual PDF shows both expressions using the CORRECT denominator (P(B) = 61/234, from part (c)); the real error is an extra copy of that same 61/234 wrongly folded into the numerator instead. The trap-taxonomy entry above states the corrected, verified version, not the research bank's original description. Of the five examiner-report citations the research bank groups under 'conditional probability confusion' across WST01's whole record, this lesson draws on only these two. The other three (Oct 2021 Q4(c)'s tree-diagram missing branch, Jun 2022 Q4(b)'s assumed-independence-on-a-Venn-diagram, and the Jun 2022 general comment about assuming independence) are already the evidentiary basis of the existing lesson content/veridian/curriculum/wst01/conditional-probability-independence-vs-mutually-exclusive.ts, which covers that ground — tree diagrams, Venn-diagram independence testing, and independence vs. mutual exclusivity — in real depth; reproducing those same three citations here would duplicate that lesson rather than add to it. Jan 2023 Q2(d) is also set inside a tree diagram, but its verified error — a correct denominator with its own value wrongly re-added into the numerator — is a distinct third mechanism from either of the two tree-diagram citations the sibling lesson already owns, which is why it earns a place here instead. Spec item 3.1 itself has no standalone past-paper question anywhere in the reviewed record (see this lesson's own header note); its content here is built from the specification's own wording and general mathematical principle, not from a cited exam question. Every numeric scenario in this lesson — both dice examples, the 8-sided-die-and-coin scenario, the 1-to-20 card question, the college French/Spanish figures, the three-coins and standard-deck MCQs — is VERIDIAN-original, checked with exact fraction arithmetic (Python's fractions.Fraction) before being written in, including the deliberately wrong answers, so the wrong values shown are the actual numbers those specific errors produce.

Question 12 marks

A weather forecaster writes P(rain tomorrowcloudy today)=0.6P(\text{rain tomorrow} \mid \text{cloudy today}) = 0.6. What does this number represent?

  • Among days that were cloudy today, 60% of them are followed by rain tomorrow

    Correct. The notation restricts attention to the world where the event on the RIGHT of the bar (cloudy today) has already happened, and states what fraction of THAT restricted world also has the event on the left (rain tomorrow).

  • BOn 60% of all days, it's both cloudy today and rainy tomorrow

    This describes P(cloudyrain)P(\text{cloudy} \cap \text{rain}) — a probability over ALL days, with no restriction to cloudy ones specifically — not the conditional statement actually written, which is about days that were already known to be cloudy.

  • CAmong days it rains tomorrow, 60% of them were cloudy today

    This reads the notation with the two events swapped — it describes P(cloudyrain)P(\text{cloudy} \mid \text{rain}), not P(raincloudy)P(\text{rain} \mid \text{cloudy}). The event on the right of the bar is always the one already known to have happened; here that's 'cloudy today,' not 'rain tomorrow.'

  • DIt will rain tomorrow with 60% certainty, regardless of today's weather

    This drops the conditioning entirely, treating the figure as an unconditional forecast. The whole reason the forecaster wrote a conditional probability, rather than a plain P(rain)P(\text{rain}), is that today's cloud cover changes the figure — an unconditional statement would need a different, separately-stated number.

Traps tested: Conditional notation confused with intersection · Conditioning event reversed · Conditioning dropped from reading

Question 22 marks

A card is drawn from a standard 52-card deck (4 Kings, 4 Queens, 44 others). Given that the card is a King or a Queen, find the probability it is a King.

  • 12\dfrac{1}{2} — restrict to the 8 Kings-or-Queens; 4 of those 8 are Kings

    Correct. Conditioning on 'King or Queen' shrinks the relevant sample space from all 52 cards down to just those 8; among that 8, exactly 4 are Kings, so P(KingKing or Queen)=48=12P(\text{King} \mid \text{King or Queen}) = \frac{4}{8} = \frac{1}{2}.

  • B452\dfrac{4}{52} — the ordinary probability of drawing a King from the full deck

    This is P(King)P(\text{King}), unconditioned — it ignores the given information entirely and uses the full 52-card deck as the sample space, rather than restricting to the 8 cards the conditioning event allows.

  • C456\dfrac{4}{56} — from 4+4=84 + 4 = 8 Kings-or-Queens, but dividing by 52+4=5652 + 4 = 56

    This correctly counts 4 Kings in the numerator but builds the wrong denominator — conditioning restricts to a SMALLER space (the 8 Kings-or-Queens), never a larger one built by adding extra cards to the original 52.

  • D44=1\dfrac{4}{4} = 1 — since every King is automatically a King-or-Queen

    True that every King is a King-or-Queen, but this only accounts for the Kings themselves as the whole restricted space — it leaves out the 4 Queens, who are equally part of the 'King or Queen' condition and equally belong in the denominator.

Traps tested: Conditioning ignored uses full sample space · Conditioning denominator built incorrectly · Conditioning denominator missing part of the given event

Question 33 marks

Three fair coins are tossed. By listing the sample space systematically, find P(exactly two heads)P(\text{exactly two heads}).

  • 38\dfrac{3}{8} — from 8 equally likely outcomes (HHH, HHT, HTH, HTT, THH, THT, TTH, TTT), of which 3 (HHT, HTH, THH) show exactly two heads

    Correct. 23=82^3 = 8 equally likely ordered outcomes; listing them all makes 'exactly two heads' easy to pick out precisely — three outcomes, one for each coin that could be the single tail.

  • B14\dfrac{1}{4} — treating '0 heads,' '1 head,' '2 heads,' '3 heads' as the 4 possible results, each equally likely

    The 4 possible NUMBERS of heads are not equally likely outcomes themselves — 1 head can happen 3 different ways (HTT, THT, TTH) while 0 heads and 3 heads can each only happen 1 way. Treating the 4 counts as a uniform sample space skips the actual listing of the 8 genuinely equally likely outcomes.

  • C12\dfrac{1}{2} — from P(at least two heads)=48P(\text{at least two heads}) = \frac{4}{8} (HHH, HHT, HTH, THH)

    This correctly counts 4 outcomes with two-or-more heads, but answers 'at least two,' which includes HHH (three heads) — a different event from 'exactly two,' which the question actually asked for.

  • D37\dfrac{3}{7} — from a list that's missing one outcome (7 instead of 8), with the 3 exactly-two-heads outcomes correctly found within it

    The numerator (3) is right, but the sample space itself is short one outcome — most commonly TTT gets left off a hand-written list, since it's easy to stop listing once no more heads seem to be coming. A systematic list (build every combination methodically, e.g. by counting in binary) is exactly what catches a missing outcome like this before it costs a mark.

Traps tested: Counts that arent equally likely treated as the sample space · At least confused with exactly · Sample space listed incompletely

Practice this for real

This site teaches the mechanism; the exam is sat on Pearson's own real questions. Go find and attempt these yourself — nothing here substitutes for actually sitting a timed paper.

Examiner report
Jan 2021 · Q1(c) — cited directly in this lesson
Pearson's official past-papers portal

Select International Advanced Level → Mathematics → any series, then look for WST01.

Statistics 1 · progress saved in this browser · sign in to sync across devices

Up next

Conditional probability, and independence vs. mutually exclusive

Two events overlapping tells you nothing about whether they are independent — and two events being independent tells you they cannot be mutually exclusive. All five WST01 examiner reports read for this course flag a version of the same confusion, in a different disguise each time: candidates assume independence instead of computing P(A \mid B) = P(A \cap B) / P(B), or they judge independence by eye — *"they are not independent as they overlap"* — instead of by the one calculation that actually settles it. Both errors come from treating two genuinely different questions (do these events share any outcomes? does knowing one change the probability of the other?) as though they were the same question.

65 min