Evidence tiers
The six-value scale this site's own research pages run on in real code, and why the least flattering tag on it is the one that proves the system is honest
8 min read
Every claim with a citation attached is actually answering two separate questions, and almost nothing you read online keeps them apart: how good is the underlying evidence, and who is telling you, and what do they want from you. A supplement company can accurately summarise a real trial. A working scientist can be speculating well past their own data. Collapsing those two questions into one — treating a confident voice as proof of a well-tested claim — is the single most common failure in science writing aimed at a general reader, and it's exactly what this site's own /protocol research dossiers are built, in actual TypeScript, to prevent:
// content/protocol/types.ts
export type Evidence =
| "rct" // Human RCT / meta-analysis. The bar.
| "human" // Human data, but observational, small, or single-lab.
| "mechanism" // Mechanistically coherent; animal or in-vitro only.
| "contested" // Live disagreement between credible researchers.
| "refuted" // Was believed, now substantially undermined.
| "speculative"; // Interesting, unproven, and honest about it.
This lesson is about that scale — Evidence, how good the data is, independent of who's saying it. (A second axis, Tier, who's talking and what they stand to gain, is tracked separately in the same code and covered in its own place.) Every entry across a dozen-plus dossiers on this site carries one of these six tags, checked against the real research behind it. None of the six examples below are invented for this lesson — they're all live, on the site, right now.
RCT / meta-analysis — the bar
The strongest evidence this scale recognises: a randomised controlled trial, or better, a meta-analysis pooling several. Most of what circulates as health advice never clears this bar. The site's own recurring example is Kodama et al., "Cardiorespiratory Fitness as a Quantitative Predictor of All-Cause Mortality and Cardiovascular Events in Healthy Men and Women: A Meta-Analysis," JAMA, 2009 — pooled evidence that each 1-MET increase in fitness (about 3.5 ml/kg/min of VO₂max) associates with roughly a 13% reduction in all-cause mortality. It's used in both the Mechanism Ledger and the Measurement Ranges dossier, and in both places it's kept carefully separate from a second, distinct study — Mandsager et al., JAMA Network Open, 2018, N=122,007, which found a categorical fitness–mortality gap rather than a per-unit slope. Two real findings, from two different papers. The tier below shows what happens when a site merges them into one.
Human — real data, imperfect design
One rung down: genuine human data that wasn't produced by a randomised trial — an observational cohort, a small sample, or a result that's only come out of one lab. The site's Mechanism Ledger tags sauna bathing this way for a precise reason: the Finnish cohort data showing dose-dependent reductions in cardiovascular and all-cause mortality with sauna frequency is large and long-running, but nobody randomly assigned thousands of people to sauna or not — it's association, not experiment. What earns it human rather than speculative is a coherent mechanism sitting underneath the correlation: a haemodynamic load genuinely comparable to moderate exercise, expanded plasma volume, upregulated endothelial nitric oxide synthase. Good mechanism plus imperfect study design is exactly what this tier means — promising, not proven.
Mechanism — coherent, but never tested in a person
Below human is evidence that never touched a human being at all — cells in a dish, or an animal model, with a causal chain you could draw on a whiteboard. The site's Frontier dossier tags the bowhead whale CIRBP finding this way (Firsanov, Zacher, Gorbunova, Seluanov et al., Nature, October 2025): bowhead whale cells overexpress a DNA-repair protein that measurably improves the accuracy of double-strand-break repair, and the paper closes the causal chain about as tightly as this kind of biology allows — overexpression in human cells, knockdown in whale cells, and a cross-species lifespan-extension rescue in fruit flies. Every link in that chain is real and demonstrated. None of it happened in a person. mechanism tier isn't weak evidence — it's evidence of the wrong kind to license a claim about your own body, however elegant the story.
Contested — a live argument, not a settled one
Some claims aren't wrong or unproven so much as genuinely unresolved, with credentialed researchers on both sides. The Mechanism Ledger flags "neuromechanical matching" — the idea that muscle recruitment order shifts with joint angle rather than following a strict size-based hierarchy — exactly this way: strength researcher Chris Beardsley is among its more prominent proponents, the direct human evidence for it is thin, and a competing account explains most of the same observations without it. The page states plainly that its own training model leans on the side that hasn't won the argument, so a reader knows precisely where that assumption is load-bearing. That's the discipline contested exists to enforce: not picking a winner where the field hasn't picked one, and saying so in the text instead of hiding it behind confident prose.
Refuted — the tier that teaches the most
The other five tiers describe evidence quality at a single point in time. refuted is different — it describes a claim's history: something the field genuinely believed, that later evidence substantially undermined. That movement, from believed to refuted, is not a failure of science. It's what science actually looks like while it's working, as opposed to the popular image of it as a settled list of facts.
The site's clearest live example is taurine. A 2023 paper in Science reported that circulating taurine declines with age and that supplementing it extended healthspan in mice and monkeys — a genuinely elegant mechanism that drove enormous public uptake, some of it at doses far beyond anything tested. Later human cohort data, tracked in the site's own record of the reversal, found circulating taurine flat or rising with age across multiple populations — the opposite of the original premise — and the paper's own lead author has since stopped recommending supplementation. Nothing about the 2023 result was fabricated or dishonest. It was a real finding, in a real journal, that additional evidence overturned. That is what "refuted" means on this scale, and it's worth being precise about what it does not mean: it doesn't mean the original researchers were incompetent, or that peer review failed, or that nothing in the paper can be trusted. It means the field did exactly what it's supposed to do when new data arrives — updated, in public, on the record.
This site keeps a permanent, append-only page for exactly this movement: /protocol/discontinued. Its own stated reasoning is worth reading directly, because it's the actual argument for tracking refutation rather than quietly deleting the old claim: "A protocol that only shows its current state looks like one that has never been wrong, and no such protocol exists. So nothing here gets deleted — it gets moved." The page keeps three kinds of entry separate on purpose — claims this site itself got wrong (misattributed a mortality statistic to the wrong paper, inverted a named researcher's actual position), protocol elements it retired for a better one without either being false, and reversals the field itself produced, taurine among them. Keeping those three apart matters: a citation error is a different failure than a genuine scientific reversal, and merging them would make the record less useful, not more honest.
The page's most useful single lesson sits inside its own retraction record, not the field's: one entry describes fabricating a "2026 follow-up study" that didn't exist, then hedging about how reliable that phantom study was — and names the failure precisely: "a hedge is not a substitute for checking. Flagging uncertainty about a fabricated citation looks like discipline and is the opposite." That's worth carrying into your own writing directly. Sounding appropriately cautious is not the same thing as having verified the thing you're being cautious about.
Speculative — named honestly, not hidden
The lowest tier that still earns a mention at all — not because it's disqualified, but because pretending an idea doesn't exist is worse than naming it plainly and saying what's actually known. The site's Reference Protocol dossier tags AAV-delivered follistatin gene therapy this way: a real, mechanistically coherent idea (follistatin blocks myostatin, the same pathway responsible for the extra muscle mass seen in myostatin-null cattle), performed outside the ordinary regulatory system, with durability, off-target expression, and long-term immune consequences in a healthy adult genuinely unknown. The entry doesn't round that uncertainty off in either direction — not oversold as promising, not dismissed as nonsense. speculative is the tag for exactly that: interesting, unproven, and honest about both halves of that sentence.
Why this is directly useful past biology
None of this is specific to training or longevity research. The same six questions apply to a claim in an EPQ, a personal statement's "did you know" opener, or a source in an extended essay bibliography: is this an RCT or meta-analysis, or one small study? Human data, or a mouse? A settled finding, or one side of a live argument? Something once believed and since walked back? A university reader — or a supervisor marking a research project — can tell within a paragraph whether a claim has been checked against its actual evidential weight or borrowed on confidence alone. Citing a press release as if it were the paper, or a single mouse study as if it settled a question in humans, reads instantly as the second kind of writing. Sourcing to the right tier, and saying plainly when you're not sure which tier you're in, reads as the first — and it's a habit worth building well before you need it for an application.
Up next
Source tiers
Who is telling you this, and what do they want from you — a question separate from how good the evidence is
10 min