Applying It — University Applications and Real Research
Where the source-tier and evidence-tier discipline actually pays off: an EPQ or extended essay, a supercurricular reading list, and any specific claim in a personal statement
7 min read
The edge is discrimination, not volume
Every earlier lesson in this course built one habit: before you let a claim into your own writing, ask two separate questions — how strong is the underlying evidence, and who is telling you, with what incentive. This lesson is about spending that habit somewhere concrete. Three places in a university application reward it directly, and none of them reward what most applicants actually do instead, which is pile up citations without discriminating between them. A reader — an EPQ examiner, an admissions tutor, an interviewer — who has seen thousands of files can tell the difference between four sources cited at the right tier and evaluated honestly, and twelve sources cited because a longer list looks more serious. The first reads as real thinking. The second reads as a bibliography.
Where an examiner is literally grading this
If you're doing an EPQ, this isn't a metaphor — it's the mark scheme. AQA's own Extended Project specification splits the qualification into four assessment objectives, and the second, "Use Resources," is worth 20% of your final mark on its own. Its own top mark band doesn't reward research volume: it asks for detailed research that includes evaluating a wide range of sources, with critical analysis and clear links to the underlying ideas. The bottom band explicitly penalizes the opposite failure — research that shows little or no evaluation of what was found, however much of it there is. A long reference list with no discrimination between a peer-reviewed source and a blog post repeating it secondhand does not climb this mark scheme; a shorter list where you can say, in your own words, why one source is stronger than another does. The IB extended essay runs on the same underlying logic under different criterion names, and so, less formally, does any personal-statement claim you'd be asked to defend at interview: the question underneath all three is identical to the one this course opened with, just asked by someone whose job is to notice whether you actually know the answer.
The supercurricular reading list: read for tier, not just topic
A reading list built by topic alone — five books and articles that are all "about" your subject — is the manufactured version of supercurricular work, in the same sense Building a Spike describes manufactured versus high-signal evidence more broadly: it's easy to produce, and a trained reader has seen it a thousand times. The higher-signal version tags each item by both axes as you read it, not just before an interview: is this the primary paper, a working researcher's own account, or a secondary retelling several steps removed from the original data — and separately, is the underlying claim resting on a large controlled study, a small uncontrolled one, or mostly a compelling mechanism nobody has actually tested end to end. An interviewer's favourite question in exactly this territory — "what's wrong with that argument?" — is unanswerable if you've only ever read the topic and never rated the sources on it. It's a five-minute answer if you have.
The personal-statement trap: the claim that's clean because it's wrong
The specific failure mode worth naming: the claim that made it into a thousand other personal statements did so because it's clean, memorable, and slightly wrong, not despite it. A finding stops being contested and starts being folklore exactly when it gets simple enough to survive being retold by someone who never read the original study — which means the version circulating in TED talks, popular books, and other applicants' drafts is systematically the version stripped of its actual caveats. Essay Frameworks covers the same discipline applied to a school's own named programs — verify the specific claim against a primary source before it goes in a final draft, because an informed reader who knows the real version notices instantly. The identical rule applies to a factual claim about your field, and the rest of this lesson is one real example, worked end to end, of exactly that check.
Worked example: "children who delay gratification grow up more successful"
This is a genuine claim that shows up constantly in personal statements and EPQs about psychology, education, self-discipline, and behavioural economics — usually attached to the "marshmallow test." Run it through both axes.
The version most applicants would cite. Popularised retellings — Malcolm Gladwell–style writing, self-help books, motivational talks — present it as a settled, general finding: delay ability at four predicts life success. Source tier: press or surface, several steps from the data. Evidence tier, as retold: unstated, because the retelling doesn't carry the original study's own caveats forward.
The primary literature it actually traces to. Walter Mischel and colleagues ran the original delay-of-gratification studies on children at a preschool on Stanford's campus from 1968–1974, most fully written up as Shoda, Mischel, & Peake (1990), Developmental Psychology, 26(6), 978–986. Source tier: primary, peer-reviewed literature — as good as source tier gets. But look at the evidence tier before treating that as settling anything: of the original pool, researchers could trace only 185 of 653 children for adolescent follow-up, and the specific, widely-quoted correlations — delay time against SAT scores of r = .57 (math) and r = .42 (verbal) — come from a subsample of just 35 to 48 children, all from one university's own community, with no adjustment for the family background those children shared. Same source tier as a landmark finding gets; a genuinely thin evidence tier underneath it.
The correction the field produced on its own. Watts, Duncan, & Quan (2018), Psychological Science (DOI: 10.1177/0956797618761661), reran the question using the NICHD Study of Early Child Care and Youth Development — a ten-site, socioeconomically diverse US cohort, not one campus preschool — with a focal subsample of 552 children whose mothers hadn't completed college, a sample the authors themselves note is roughly ten times the size of Shoda and colleagues' original one. Same source tier again: peer-reviewed, published in the same flagship journal. Completely different evidence tier: the raw relationship they found was already about half the size of the original correlations before any adjustment, and shrank by roughly two-thirds once family background, early cognitive ability, and home environment were controlled for. Most of what predictive power survived came from clearing a very low bar — being able to wait at least twenty seconds — not from the dramatic full-length wait the story usually turns on. A 2024 follow-up by much of the same team, Sperber, Vandell, Duncan, & Watts, Child Development, 95(6), 2015–2029, pushed the same children into adulthood and found the relationship still didn't reliably hold.
The lesson, stated plainly. Source tier alone would have told you to trust both the 1990 and the 2018 papers equally — they're both peer-reviewed primary literature. Only the evidence-tier question — how large and representative was the sample, and were the obvious confounds controlled — explains why one of them should change your mind more than the other. A student who writes "studies show delaying gratification predicts success" has cited the folklore version. A student who writes that the original, widely-cited finding rested on a few dozen children from one university community and didn't survive a fifteen-fold larger, controlled replication has cited the same underlying story at the tier it actually deserves — and has, not incidentally, produced a more interesting sentence. If effect sizes and what "controlling for" actually does are unfamiliar mechanics rather than just vocabulary, the site's own article on reading a study without a stats degree walks through exactly that.
Carrying the discipline forward
None of this is specific to psychology or to this one study. It's the same two questions, asked of whatever claim you're about to put your name behind: how strong is the evidence, and who benefits from you believing it, stated separately and never merged into one impression of "credible." This site runs on the identical rubric for its own research pages at /protocol — evidence quality and source incentive, tracked as two independent labels on every claim, for exactly the reason this course exists: a claim can be popular and wrong, obscure and solid, or peer-reviewed and still thin, and only checking both axes tells you which. That's the whole competitive advantage this course has been building toward. Not more sources. The right ones, at the tier they've actually earned.
Up next
Sources and provenance
Every real source behind this course's claims about source tiers, evidence tiers, and reading research — ranked by the same rubric it teaches, with the honest gaps named rather than hidden
8 min