How to use this course

Two questions, asked in order, on every claim you read from here on: how good is the evidence, and who is telling you

8 min read

Scope, stated once and held to throughout: this course teaches a filter, not a reading list. It covers how to separate the strength of a piece of evidence from the trustworthiness of whoever is presenting it, how a paper is actually structured and what its abstract will not tell you, and the specific, recurring tells that mark a source as decoration rather than research. It does not teach statistics as a subject — Reading a study without a stats degree already does that job on this site and this course leans on it rather than repeating it. And it makes no claim about any specific school's admissions weighting; what it teaches is the discipline that a strong EPQ, a strong personal-statement claim, and a strong supercurricular research project all reward, verified below rather than asserted.

Where this course's own filter comes from

This course did not invent the two-axis system it teaches. It is lifted directly from how research is actually labelled elsewhere on this site, in /protocol — the standing dossiers on training, longevity science and building a business, where every claim already carries two separate tags. That system was built first, under real stakes (readers making training, health and financial decisions off it), and this course exists to teach the reasoning behind it explicitly, so you can run it yourself on whatever you're reading — a paper for an EPQ, a claim in your personal statement, a source for a supercurricular project, or, later, anything else.

The confidence-tag system: two axes, not one

Here is the single most common failure in this genre of writing, named plainly by the source code this site's dossiers are built from: treating "how good is this evidence" and "who is telling me this" as one combined trust score. They are not the same question, and collapsing them is how bad information launders itself into looking credible.

Evidence asks how strong the underlying data is for the claim itself, independent of who is reporting it:

TagWhat it means
rctHuman randomised controlled trial, or a meta-analysis of them. The bar.
humanHuman data, but observational, small, or from a single lab.
mechanismMechanistically coherent; animal or in-vitro only — no human outcome data yet.
contestedLive, genuine disagreement between credible researchers.
refutedWas believed at one point, now substantially undermined by later work.
speculativeInteresting, unproven, and honest about being unproven.

Tier asks who is telling you this, and what they stand to gain from you believing it:

TagWhat it means
primaryPeer-reviewed literature itself, or a regulator's own filing.
researcherA working scientist writing inside their own field.
operatorA practitioner with a real, verifiable, numbers-backed track record.
companySelf-reported by a party with a commercial interest in the answer.
pressJournalism or an aggregator, without direct access to the primary material.
surfaceContent marketing, an SEO farm, a course funnel.

Notice these live as two separate fields in the actual code, not one — evidence?: Evidence and tier?: Tier are independent optional properties on the same entry, because a single combined "trust: 7/10" field would hide exactly the distinction that makes the system useful. A claim can carry the best possible tier and a weak evidence tag at the same time — a peer-reviewed journal reporting a finding that is real but animal-only — and that combination is not a contradiction to resolve. It is information. Losing it by averaging the two into one number is the failure mode this whole course exists to train you out of.

The same finding, asked both questions at once

Take one real example, from this site's own quarterly research dossier on aging science, The Frontier. A 2025 paper in Nature reported that bowhead whale cells overexpress a protein called CIRBP and repair broken DNA both more often and more accurately than human, mouse or dolphin comparator cells — with the causal chain closed by knockdown experiments, biochemical reconstitution, and a cross-species lifespan rescue in fruit flies. That claim earns mechanism on the evidence axis, because the organism-level result is not in humans, and primary on the tier axis, because it is the actual peer-reviewed paper, not someone's summary of it. High tier, modest evidence — both true, both worth knowing, and neither one substitutes for the other.

Now take a different claim from the same tracked field: a company called Rubedo Life Sciences announcing, in a press release, specific Phase 1 numbers for a topical drug candidate — a named percentage reduction in skin thickness in a named condition. That claim earns human on the evidence axis — real people, a real trial — but company on the tier axis, because the only account of it is the company's own announcement, with no peer-reviewed publication behind the numbers yet. Here the evidence tag looks stronger than the whale example and the tier tag is weaker. A one-axis "trust score" would struggle to even represent these two claims as different from each other. Two axes represent them correctly on the first try.

The thesis this course is built on

State it plainly, because it cuts against how most people are trained to read: a well-known, unglamorous claim with strong RCT evidence beats a novel one with a beautiful mechanism and no outcome data, every time. Popularity is not weakness. Skepticism belongs on evidence quality — not on how many people have already heard a claim.

This isn't a throwaway line. It is a testable prediction about what the actual strongest evidence in any well-researched domain looks like, and it holds up: in that same dossier, the interventions carrying the strongest evidence tag on the site — cardiorespiratory fitness, resistance training, sleep regularity, protein sufficiency, not smoking, sun protection — are also the ones every reader has already heard a hundred times, next to newer, more exciting claims (a repurposed diabetes drug, a gut-microbiome metabolite, a partial-cell-reprogramming therapy) that carry real but thinner evidence. The dossier's own verdict on this is worth quoting exactly, because it is the whole thesis in one line: "the strongest evidence in this entire dossier points at interventions that are free, widely known, and boring. That is not a disappointing finding — it is the finding."

The instinct this course trains you against is the one that reaches for the unfamiliar claim because it feels like more of a discovery to cite. A reader — a teacher, an EPQ supervisor, an admissions tutor — who has seen a thousand essays knows that instinct too, and a source list built to fight it rather than indulge it reads differently on the page.

Why this is directly useful for a university application

This is not a vague claim about "transferable skills." One piece of it is checkable against a real, public mark scheme, so it was checked: AQA's Extended Project Qualification specification (7993, version 1.4) allocates a full, separate Assessment Objective — AO2, "Use Resources," worth 20% of the total mark — to exactly this skill, defined in the spec's own words as research that "critically select[s], organise[s] and use[s] information" and "select[s] and use[s] a range of resources." That is a quarter of the entire qualification's mark scheme sitting directly on the discipline this course teaches: not finding a source, but weighing one correctly. AO4, "Review" — another 20% — asks you to evaluate your own outcomes against your original objectives, which is the same honesty this system asks of a claim, turned on your own work.

The same discipline shows up less formally everywhere else an application gets read. A personal-statement claim backed by a source you can place on both axes — "a primary-tier paper found X, with this specific limitation" — reads as a different order of claim from one backed by nothing, or by a blog post repeating a study's headline without its effect size. A supercurricular research project stops being a reading list and starts being evidence of how you think the moment it shows you weighing sources against each other rather than simply citing the first five results a search returns. None of that requires a stats background. It requires running the two questions above on every source before it goes in the file — exactly the habit Building a Spike, on this site's applications course, argues is the actual differentiator in a strong application: one deep, verifiable, evidence-backed claim beats ten shallow, unverifiable ones.

What's ahead

  1. Foundations — the two axes in full: every evidence tag and every source tier, what each can and cannot support, and the specific failure of collapsing "well-known" with "weak."
  2. Reading a paper — abstract, methods, effect size, confidence interval and limitations, in the order that actually tells you what a paper found, not the order the abstract presents them in.
  3. Spotting the four tells — the four recurring signatures of a surface-tier source dressed up as research, worked against real examples of each.
  4. Applying it — running both axes on a live claim end to end, from an EPQ topic through a personal-statement line to a supercurricular project source list.
  5. Reference — every source this course itself draws on, tiered the same way it asks you to tier your own.

This course teaches a method for weighing evidence, not a verdict on any specific claim you bring to it. Every external figure and citation above was checked against a primary source at the time of writing — the AQA specification directly, the CIRBP and Rubedo findings against this site's own dossier and its own underlying citations — and every specific number, fee, or mark weighting drifts over time in the way this course itself warns about; verify anything you intend to rely on against its current primary source before you do.

Research Literacy · progress saved in this browser · sign in to sync across devices

Up next

Evidence tiers

The six-value scale this site's own research pages run on in real code, and why the least flattering tag on it is the one that proves the system is honest

7 min