Correlation coefficient and regression — the calculation mechanics
~40 min · WST01 · 4.3
WST01 · 4.3 · 40 min
This is the thinnest evidence base of any WST01 lesson in this batch, and it's worth saying plainly rather than dressing it up. Where the regression-interpretation lesson draws on five separate examiner reports, spec item 4.3 — the product moment correlation coefficient itself — is anchored by exactly two sources: one examiner report and one mark scheme. What those two sources hand over is precise and genuinely useful all the same: a real, evidenced trap on the r-calculation itself — an omitted square root in the denominator — a well-documented Pearson accuracy discipline for 'show that' questions this lesson extends to r's own square-root division, and a clean, quotable coding-invariance property Pearson credits outright. This lesson assumes the Sxx/Sxy machinery already taught elsewhere and adds exactly what that machinery didn't need — a third summary sum, and the coefficient it unlocks.
Before you read on
Two or three questions on exactly what this lesson teaches. Being wrong here is fine — it's the fastest way to find out what to pay attention to next.
Recap: Sxx and Sxy, then the sum regression never needed
The regression line needed exactly two summary sums, and — both covered in depth already, including the exact documented trap of substituting for (see the regression-gradient-interpretation-extrapolation lesson, a prerequisite for this one). The product moment correlation coefficient, spec item 4.3, needs one more: — exactly the same computational shape as , with standing in for throughout. It's printed in full in the formula booklet's Statistics S1 section alongside and , so nothing here needs memorising; what needs practising is not repeating, on a formula that looks almost identical, the one substitution error the archive already documents for — confusing for inside it.
With all three sums available, the product moment correlation coefficient is — also given in full in the booklet, so this is a lookup, not a memorisation task. Where answers 'how much does change, on average, per unit of ', answers a genuinely different question: 'how tightly do the actual data points cluster around a straight line, in both directions at once?' carries real units — whatever is measured in, per whatever unit is measured in — and can be any size at all; is built to carry none, which is exactly why it always lands somewhere in the fixed range , whatever units and happen to be measured in.
The sign of always matches the sign of — both come from the same numerator, — so a positive gradient and a positive aren't two separate facts to check on a scatter diagram, they're the same fact read off two different formulas. close to or means the data sits close to a straight line; close to means it doesn't — though, as the next section covers, 'doesn't sit close to a straight LINE' and 'has no relationship at all' are two different claims, and treating them as the same one is 's single most consequential limitation.
Reading r: use, interpretation, and limitations
Spec item 4.3 names 'its use, interpretation and limitations' as content in its own right, not just the calculation — and it's worth being honest about what backs the paragraphs below. Unlike the accuracy trap and the coding property covered later in this lesson, the reviewed archive doesn't contain a quoted examiner-report example testing the limitations of specifically. What follows is standard subject-matter content for this spec point, not a documented exam trap — flagged here plainly rather than dressed up as evidence it isn't.
specifically measures LINEAR association — how well a straight line fits the data — and nothing else. A dataset can show an extremely strong, entirely genuine relationship between and that curves rather than follows a straight line, and will still report a value close to for it, because a curve, however clean, is exactly the shape a straight-line measure is built to miss. A scatter diagram sitting next to a computed value isn't decoration — it's the one check that confirms is even measuring the right kind of pattern in the first place.
A high is evidence of association, not evidence of cause. Two variables can move together closely without one causing the other — both can instead be driven by some third factor that has no way to detect or rule out. Reading a strong correlation as proof of a causal mechanism is a step itself never licenses; it reports how closely two things happened to move together in this data, and nothing about why.
And like the regression line it's built from the same summary sums as, describes the data it was actually calculated from — a strong computed over one range of says nothing certain about whether the same tight linear pattern continues once moves well outside that range. That's the same caution the regression-gradient-interpretation-extrapolation lesson covers in depth for predictions made from the line itself; it applies with identical force to the correlation coefficient describing that line's fit.
Mechanism
Why r has no units when b does
Set and side by side and the difference is exactly one factor: . Every quantity on the right carries units — is measured in (units of ) per (unit of ); is measured in (units of ); is measured in (units of ) — and works out to (units of )/(units of ), the exact reciprocal of 's own units. Multiply the two together and every unit cancels, leaving a pure number. That isn't a coincidence baked into the formula by choice; it follows directly from being built out of a ratio of two MATCHING kinds of spread ( against , both 'sum of squared deviations from a mean') rather than 's ratio of an association () against just one variable's own spread ( alone). This is exactly why can be , or , or — any real number at all, in whatever units the question happens to use — while is always trapped between and : nothing about 's formula bounds it, and everything about 's formula does.
Beyond the spec
Spec 4.3's own guidance says outright that derivations are 'not required' for this unit — a student can apply 'r is unaffected by (linear) coding,' the exact phrase a real mark scheme credits, without ever seeing why it's true. This is why, and it also surfaces the one genuine subtlety that credited phrase glosses over — a subtlety the reviewed archive gives no sign of the exam actually testing, but one worth understanding rather than trusting blindly.
Suppose and are coded as and , for constants — exactly the shape spec 4.2's own guidance names ('linear change of variable may be required'). Since for every data point (the shift cancels the moment a mean is subtracted, because too), squaring and summing gives — and by identical reasoning, . The cross term picks up both scale factors at once: .
Finishing the derivation — where the sign question actually comes from
Substituting all three into the PMCC formula for the coded data: .
Everything before the final fraction is just — the PMCC of the original, uncoded data. The final fraction, , is whenever and share a sign and whenever they don't — it can never be anything else, since it's a magnitude divided by the exact same signed number. That's the whole result: , with the sign flipping if and only if exactly one of the two coding DIVISORS ( or — not or , the two shifts, which never appear in the final result at all) is negative.
Both coding examples this facts bank documents divide by a positive constant — the Jan 2024 Q2 coding (verified directly against the real mark scheme: a Fahrenheit-to-Celsius-shaped transform, dividing by 9 and multiplying by 5, both positive) and the Section 2.2 variance-coding example (Jun 2022 Q3(e), dividing by 2 — coding exists to shrink unwieldy numbers, not to flip their sign). Neither is a full audit of every coding question across all 14 series reviewed, so this is suggestive rather than exhaustive, but in both real examples this pass has checked, the credited phrase 'r not affected by (linear) coding' carries no qualification because it doesn't need one — sign included. The edge case here — one negative divisor, flipping the sign — is real mathematics, not a trick, but it's flagged as background understanding rather than exam content: nothing in the reviewed archive shows it being tested, and manufacturing a negative-divisor coding question to drill against would mean inventing a trap this pass's evidence doesn't actually show exists on this paper.
Complete it yourself
Complete the chain — does r need recalculating after coding?
- 01
A shop records, for 10 days, the average outdoor temperature that day (, °C) and the number of hot chocolates it sells (). To keep the arithmetic manageable, the data is coded as and , and the PMCC of the coded data is correctly found to be .
- 02
The shop wants the PMCC of the actual, uncoded temperature and sales figures — without recomputing , and from scratch using the original numbers.
Named traps
- sqrt-omitted-in-pmcc-denominator
- Confirmed on the real question this comes from — Jun 2024 Q4(b), the actual PMCC step of the same question whose part (c) is a 'show that' for the regression line (verified directly against the official Pearson mark scheme, WST01_01_2406_MS, and examiner report, WST01_01_2406_ER, during this lesson's own review, not just against the facts bank's secondhand write-up of it): *'Part (b) was answered well with many students able to calculate a correct value of the product moment correlation coefficient. Common error included the omission of the square root in the denominator.'* The failure isn't an accuracy slip — it's dividing by directly instead of , which doesn't just lose precision, it produces a genuinely different (and, since the one step that keeps bounded between and has been skipped, often an out-of-range) number.
- pmcc-recomputed-unnecessarily-after-coding
- A real mark scheme credits a clean, quotable fact directly: *'r not affected by (linear) coding'* (Jan 2024 MS, Q2(d)). The trap is spending time — and sometimes marks — undoing a coding scheme that never needed undoing: recomputing , and from the original, uncoded and values after already finding from the coded data, on the mistaken assumption that needs 'converting back' the way the regression coefficient genuinely does. scales with the coding constants; doesn't, by construction (see the mechanism and derivation earlier in this lesson) — and a script that recalculates from scratch for the uncoded data isn't doing extra-safe working, it's demonstrating it hasn't understood the property the mark scheme is actually crediting.
Marked, line by line
A tutor records, for six students, the number of hours spent revising in the week before a test () and the student's score on that test out of 100 (). Summary statistics: , , . From an earlier part of the question, the regression line of on has already been found to be . You are further given . (a) Use the regression line to find , without recomputing it from the raw data. (2) (b) Hence find the product moment correlation coefficient , giving your answer to 3 significant figures. (3) (c) Give a brief interpretation of your value of , in the context of this data. (2) — VERIDIAN-original question, dataset and target values (Sxx = 30, the line ŷ = 40 + 3x, Syy = 500, and the resulting r = 0.735). Built in the same general question shape real WST01 papers use for this topic — a fitted regression line already in hand, used to shortcut straight to a needed sum rather than recomputing it from raw data — though the specific combination tested here (derive Sxy from the line, then compute r) is a VERIDIAN construction, not a reproduction of Jun 2024 Q4's actual part order: that real question finds r in part (b) before it 'shows that' the regression line's own gradient in part (c) — see this lesson's closing flag block for the correction this made to WST01-verified-facts.md's original mischaracterisation of that question.
7 marks available
(a) — 2 marks
- 01M1
Method mark for rearranging b = Sxy/Sxx to make Sxy the subject, then substituting the already-known gradient (b = 3, from the regression line found earlier) and Sxx (= 30) — using the fitted line to shortcut straight to Sxy, rather than recomputing it from raw x,y data.
- 02A1
cao — correct value, dependent on the method above.
(b) — 3 marks
- 101M1
Method mark for correct substitution into the PMCC formula (given in full in the booklet — not required from memory) using their own Sxy from part (a) and the given Syy.
- 102A1
(kept unrounded, or to several more figures than the final answer needs), so
Accuracy mark for carrying enough precision through the square root before dividing — applying the same discipline Pearson documents for 'show that' calculations, including this exact paper's own regression-line gradient step: give values to at least one more decimal place than the target needs, and round only at the very last step.
- 103A1
(3 s.f.)
cao — correct final answer, dependent on the accuracy mark above.
(c) — 2 marks
- 201B1
is fairly close to 1, showing a reasonably strong POSITIVE linear relationship between revision hours and test score — consistent with the positive gradient (b = 3) already found for the regression line.
Independent mark for connecting the SIGN and the relative STRENGTH of r to the two named variables in context — not a bare 'strong correlation' with nothing attached to what's actually being correlated.
- 202B1
This supports fitting a straight-line model to this data, though r alone says nothing about whether the same tight linear pattern would continue for far more revision hours than any student in this sample actually put in.
Independent mark for the LIMITATIONS half of spec 4.3's own wording ('its use, interpretation and limitations') — r describes how well a line fits the data it was computed from, not a guarantee that fit extends to x-values outside the study, the same caution taught for predictions made from the line itself.
In your own words
In one sentence: why does rounding a square root to the same number of significant figures the final answer needs — before dividing by it — put that final answer at genuine risk, when rounding it to one extra figure first would not?
Retrieval — with feedback on every choice
For a dataset, , , . What is the product moment correlation coefficient ?
A dataset is coded using and to simplify the arithmetic. The PMCC of the coded data, , is found to be . What is for the original, uncoded and data?
A question asks: 'Find r, giving your answer to 3 significant figures,' and your working reaches the step . What's the safest way to evaluate this?
A scatter diagram shows a clear U-shaped relationship between and : as increases, first decreases, then increases. The PMCC for this data works out as . Which is the most defensible conclusion?
- PMCC: r = Sxy / √(Sxx × Syy). Syy = Σy² − (Σy)²/n — same shape as Sxx, with y in place of x.
- r always lies between −1 and 1. Getting |r| > 1 out of a calculation means an arithmetic mistake, not an unusual dataset.
- Sign of r always matches sign of b — same numerator, Sxy.
- r unaffected by (linear) coding. b DOES scale with the coding constants — r doesn't. Never recompute r from original data after already finding it from coded data.
- Keep a square-root denominator unrounded (or to 1+ extra figure) until the final line — rounding it early can change the final significant figure, not just its precision.
Not affiliated with or endorsed by Pearson Edexcel. This lesson's own adversarial review caught and corrected a misattribution in WST01-verified-facts.md itself: that file's 'PMCC use and limitations (4.3)' write-up files the Jun 2024 accuracy quote ('too often students lost the final A mark... give answers to at least one more decimal place than the given value') as evidence about a 'show that r = …' question that derives Sxx/Sxy from an already-found regression line. Fetching and reading the real Pearson mark scheme (WST01_01_2406_MS) and examiner report (WST01_01_2406_ER) directly shows that's wrong: that quote is Q4(c), a 'show that [the regression line is] g = −42.3 + 0.722d' step computing the gradient b = Sdg/Sdd = 0.7218…→0.722 — the exact real numbers (Sxy = 12105.12, Sxx = 16769.78) the show-that-answer-discipline.ts sibling lesson quotes at greater length, and correctly attributes to the regression line's gradient, not to r. The real PMCC step in that same question is the earlier Q4(b) — not a 'show that' question at all (the mark scheme awards 'awrt 0.98', no printed target) — and its own documented common error, now correctly cited in this lesson's trap-taxonomy in place of the earlier false attribution, is a different one: 'the omission of the square root in the denominator.' The premature-rounding worked example below is kept, because the underlying accuracy discipline is real and Pearson-documented — for this same paper's regression-line gradient step, and repeatedly as general 'show that' guidance elsewhere in the archive — but it is now presented honestly as extending that documented principle to a step the reviewed archive does not show being tested this way, not as itself a directly observed r-specific exam trap. The Jan 2024 MS Q2(d) coding quote ('r not affected by (linear) coding') was independently re-verified against the real mark scheme during this same review and required no correction; that check also confirmed the coding transform it credits (a Fahrenheit-to-Celsius-shaped change of variable) uses positive constants throughout, consistent with — though not, on its own, a full-archive proof of — the coding-invariance discussion later in this lesson. Every number in this lesson's own worked example (Sxx = 30, the line ŷ = 40 + 3x, Sxy = 90, Syy = 500, r = 0.735, and the wrong-path r = 0.738 from rounding √15000 to 122 first) and the chain-drill's coding scenario (u = (x−10)/2, v = (y−50)/5, r_uv = −0.78) are VERIDIAN-original — independently checked with a Python script (90/√15000 = 0.7348469… → 0.735; 90/122 = 0.7377049… → 0.738) before being written into this file. The 'use, interpretation and limitations' teach block (linear-only measurement, correlation vs causation, the range caveat) is standard subject-matter content required by spec 4.3's own explicit wording, not drawn from a quoted examiner report — the reviewed archive contains no standalone example testing r's limitations directly, and this lesson says so in its own text rather than presenting general knowledge as documented exam evidence. The coding-invariance derivation (beyond-spec) and its sign-flip edge case go beyond what spec 4.3 requires ('derivations... will not be required') on purpose, and the sign-flip subtlety specifically is flagged as background understanding with no sign of being tested in the reviewed archive, not as exam content in its own right. The per-line mark allocations attached to the original worked example are modelled on the verified mark-scheme conventions in WST01-verified-facts.md §4 (M/A/B marks, cao), not transcribed from a real mark scheme, which for an original question does not exist.
For a dataset, , , . What is the product moment correlation coefficient ?
Correct. .
- B
This divides by directly, skipping the square root the formula actually requires. 's denominator is , not itself — a genuinely different, much larger number.
- C
This computes — the formula for the GRADIENT , not for . It's also a useful self-check in its own right: can never exceed 1 in size, so getting a value like 1.125 out of a correlation calculation is itself a sign something has gone wrong, before even checking which formula was used.
- D (3 s.f.)
This is — the correct pieces, combined the wrong way up. As with option C, a result bigger than 1 is a red flag worth checking against before moving on, since a genuine PMCC can never exceed 1 in magnitude.
Traps tested: Sqrt omitted in pmcc denominator · Pmcc confused with gradient · Pmcc fraction inverted
A dataset is coded using and to simplify the arithmetic. The PMCC of the coded data, , is found to be . What is for the original, uncoded and data?
- — the same value, because both coding divisors (5 and 2) are positive
Correct. PMCC is unaffected by a linear change of variable, and here neither coding constant negates a variable, so the sign carries through unchanged as well as the magnitude — the coded value already IS the answer for the original data.
- B
This treats r as if it rescales the way the gradient b would — but r isn't a rate with units to convert back, it's a bounded, unit-free number. 6.4 is also, on its own, an immediate red flag: r can never exceed 1 in magnitude, whatever calculation produced it.
- C
The same error as option B, applied in the opposite direction — dividing by the coding constants instead of multiplying. Either direction assumes r needs converting back through the coding scheme at all, which it doesn't.
- DIt cannot be found without recomputing Sxx, Syy and Sxy from the original, uncoded data
This is exactly the unnecessary extra work the coding was meant to avoid. The whole point of the credited property — r unaffected by linear coding — is that the coded calculation already answers the question for the original variables too.
Traps tested: Pmcc treated as scaling like gradient · Coding invariance not recognised
A question asks: 'Find r, giving your answer to 3 significant figures,' and your working reaches the step . What's the safest way to evaluate this?
- Keep √15000 unrounded (or carry at least one extra significant figure) through the division, and round only the final value of r
Correct — this is the exact discipline this lesson's worked example is built around, and it's the difference between reaching the correct r = 0.735 and the wrong r = 0.738.
- BRound √15000 to 3 s.f. (122) first, then divide
This is precisely the trap named in this lesson's own worked example: rounding the square root to the same precision the final answer needs, before dividing by it, changes the final 3 s.f. figure from 0.735 to 0.738 — a different answer, not just a less precise one.
- CRound √15000 to the nearest whole number for a quick estimate, then refine afterwards if time allows
A quick estimate can be a useful sanity check on the SIZE of an answer, but it isn't a substitute for the actual working a mark scheme credits — and 'refine afterwards if time allows' risks the rounded estimate simply becoming the final submitted answer.
- DIt makes no difference, since a calculator's square root function is exact
A calculator's own internal value is effectively exact — the risk is entirely in what gets written down as the intermediate working and then divided by on paper. A written-down rounded value carries the rounding error forward regardless of how precisely the calculator itself computed it.
Traps tested: Premature rounding in pmcc denominator · Calculator precision assumed to remove rounding risk
A scatter diagram shows a clear U-shaped relationship between and : as increases, first decreases, then increases. The PMCC for this data works out as . Which is the most defensible conclusion?
- close to 0 shows very little LINEAR association — but the scatter diagram makes clear there's a strong, genuinely real relationship that just isn't a straight line. PMCC only measures straight-line association, and this pattern is exactly what it's built to miss.
Correct. This is the limitation named in this lesson's teach block: r measures a specific SHAPE of relationship (linear), not the general presence or absence of one — a strong curved pattern can coexist with an r value close to 0 without any contradiction.
- B shows that and are essentially unrelated
The scatter diagram directly contradicts this — a clear, describable U-shaped pattern is not 'unrelated'. r measures one specific kind of relationship; a low value rules that kind out, not every kind.
- CThe data must contain an error, since a real underlying relationship should always produce a high value of r
This assumes every genuine relationship is linear, which is exactly the assumption r's own limitation exists to warn against. A clean, strong, entirely correct curved relationship can produce a low r with nothing wrong with the data at all.
- D indicates a weak negative correlation
0.04 is positive, not negative — a small but real misreading of the sign. More importantly, a value this close to 0 in either direction says almost nothing about linear association at all, which is the actual point this question is testing.
Traps tested: Small r mistaken for no relationship · High r assumed necessary for real relationship · Sign misread
Practice this for real
This site teaches the mechanism; the exam is sat on Pearson's own real questions. Go find and attempt these yourself — nothing here substitutes for actually sitting a timed paper.
Pearson's official past-papers portalSelect International Advanced Level → Mathematics → any series, then look for WST01.
Up next
Regression — gradient interpretation and extrapolation/reliability
Every one of the five examiner reports read for this unit flags the same failure, worded a different way each time. Computing b = S_{xy}/S_{xx} is not where the marks go missing — candidates who can barely finish a S_{xx} calculation still reliably reach the right line. What goes missing is the sentence after it: saying, in words, what one extra unit of x actually buys you in y, and knowing the difference between a prediction the data has already tested and one it is only guessing at.
50 min