Red flags

The four checkable tells behind a claim that only sounds well-supported — and exactly what to look at to find out

9 min read

"Be skeptical" isn't a method. It's a mood, and moods don't scale to the moment you're actually holding a paper, a headline, or a citation in an essay and need to decide whether to trust it. What scales is a short list of specific, checkable things — not vibes about whether a study "feels" rigorous, but concrete places in a real paper (or a real absence of one) that predict, on their own, whether a claim survives scrutiny. Four of them cover most of what goes wrong: publication bias, conflict of interest, small or unreplicated findings treated as settled, and correlation dressed up as causation. Each one below gets a real, named case — not a hypothetical — and a specific thing to go and look at, that you can actually do with the source open in front of you.

Tell 1 — Publication bias: the file drawer problem

A study that finds nothing interesting is far less likely to get published than one that finds something striking, and — this is the part people miss — it's often less likely to even get submitted, because researchers correctly predict a null result is a harder sell to a journal. Robert Rosenthal named this precisely in "The File Drawer Problem and Tolerance for Null Results," Psychological Bulletin 86(3), 1979: 638–641 — the paper that gave the phenomenon its name, arguing that journals fill up with the small fraction of studies that cross a significance threshold while the much larger set that didn't sits in a literal filing cabinet, unseen. [Established] No fraud is required for this to distort what you read: if you only ever see the studies that got published, you're looking at a biased sample by construction, and a single striking result is more likely to be an overestimate — or a fluke — than the honest average of everything that was actually tried.

How to actually check this: don't weigh a single study as if it were the whole literature. Search for whether a meta-analysis or systematic review exists on the specific claim (Google Scholar or PubMed, the claim's keywords plus "meta-analysis"), and if one does, check whether it reports a funnel-plot asymmetry test or an Egger's test — the standard statistical checks for exactly this bias, and something any competent meta-analysis states explicitly. For a clinical claim specifically, search the trial itself on ClinicalTrials.gov or the WHO's International Clinical Trials Registry Platform: a registered trial that never produced a published result is the file-drawer problem sitting in the open record, and it's directly checkable rather than something you have to take on faith.

Tell 2 — Conflict of interest and industry funding

A funder doesn't need to falsify a single data point to shape a result. They can fund the comparator that flatters their product, commission a literature review and set its terms without disclosing it, or simply decline to fund the follow-up that would have muddied the story. Two real, extensively documented cases show this happening at the level of an entire field's public understanding, not one bad paper. Kearns, Schmidt, and Glantz, "Sugar Industry and Coronary Heart Disease Research: A Historical Analysis of Internal Industry Documents," JAMA Internal Medicine 176(11), 2016: 1680–1685, used internal Sugar Research Foundation records to show the industry funded and helped shape a 1965 literature review in the New England Journal of Medicine — undisclosed at the time — that singled out dietary fat as the cause of coronary heart disease while downplaying sugar's own role. [Established] Separately, reporting by Anahad O'Connor in The New York Times and by NPR in August 2015, later corroborated by internal emails, showed Coca-Cola had given roughly $1.5 million to launch, and close to $4 million total to the founders of, the Global Energy Balance Network — a group whose public message was that inactivity, not diet, drove obesity. The organization dissolved within months, on November 30, 2015, once the funding was exposed. [Established — reported and corroborated by internal documents]

How to actually check this: read the paper's own Conflict of Interest or Funding declaration before you read its conclusion. This is a standard, mandatory part of any journal that follows the International Committee of Medical Journal Editors' disclosure recommendations, which covers most reputable biomedical and psychology journals — the section is usually labeled "Conflict of Interest," "Funding," "Financial Disclosures," or "Declarations," sitting near the end of the paper, and every author is required to complete it. A declared conflict doesn't automatically invalidate a finding, but it should raise the bar for wanting independent, differently-funded replication before you treat the claim as settled. And if a paper simply has no such section at all, that's itself worth noticing — it's expected practice at any journal worth citing, not an optional courtesy.

Tell 3 — A small, unreplicated finding treated as settled

A single study, especially a small one, can produce a striking result purely from chance, or from the ordinary flexibility researchers have in choosing which variables to control for, which subgroup to report, or when to stop collecting data. Whether a finding survives an independent replication, run at proper statistical power, is the actual test of whether it's real — and psychology ran that test on itself directly. The Open Science Collaboration's "Estimating the Reproducibility of Psychological Science," Science 349(6251), 2015: aac4716, had independent teams attempt to directly replicate 100 studies originally published in three of the field's own top journals. 97% of the original studies had reported a statistically significant effect; only 36% of the replications did, and the replicated effect sizes were, on average, about half the size of the originals. [Established] Nutrition research has its own version of the same problem, for a related but distinct reason — it leans heavily on large observational cohorts rather than small experiments, which makes it vulnerable to confounding rather than pure sampling noise. John Ioannidis makes exactly this case in "The Challenge of Reforming Nutritional Epidemiologic Research," JAMA 320(10), 2018: 969–970, showing that taken at face value, the hazard ratios reported across nutrition cohort meta-analyses imply effects for individual foods on lifespan that are simply not biologically plausible — itself evidence that the underlying associations are inflated, confounded, or both. [Established] This site's own protocol pages track a live version of this same pattern when it happens to a specific supplement claim: /protocol/discontinued is the append-only record of exactly this, including a mechanistically elegant 2023 finding on taurine that later human cohort data walked back.

How to actually check this: look at the sample size in the Methods section before you weight a finding — an N in the dozens or low hundreds isn't enough to reliably detect a modest but genuine effect, and a single small significant result is a hypothesis worth testing further, not an answer. Then go looking for a replication specifically: search the claim on Google Scholar with "replication" appended, or check the Open Science Framework (osf.io) if the field has an active preregistration culture. Prefer a meta-analysis or systematic review that pools multiple independent studies over any single paper, and when you find one, check whether it flags its own heterogeneity between studies — a meta-analysis that admits the studies it pooled don't agree with each other is more trustworthy than one that doesn't mention it.

Tell 4 — Correlation mistaken for causation

Two things moving together doesn't tell you which one caused the other, or whether something else entirely is driving both. The clean, real illustration of this is the breakfast-and-grades claim. Smith, Blizzard, McNaughton, et al., "Skipping Breakfast Among 8–9 Year Old Children Is Associated with Teacher-Reported But Not Objectively Measured Academic Performance Two Years Later," BMC Nutrition 3, 2017: 86, is unusually clean precisely because it tested both kinds of measurement in the same children. Teachers rated breakfast-skippers as doing meaningfully worse in reading, maths, and overall academic performance — a genuine, statistically significant gap. But the same children's actual standardized test scores showed only a small difference, and once the researchers statistically adjusted for socioeconomic status and other confounders, most of even that small gap disappeared. [Established] Children who skip breakfast disproportionately come from households facing other disadvantages that plausibly affect academic outcomes on their own — the breakfast itself, in this analysis, wasn't doing the work the teacher ratings implied it was doing.

How to actually check this: read the Methods section, not just the abstract, and identify the study design before anything else. Was it a randomized controlled trial, where participants were actually assigned to a condition, or an observational, cohort, or cross-sectional study comparing groups that already existed? Only the former can support a causal claim on its own. If it's observational, check specifically which confounders the authors adjusted for — a study on an outcome that plausibly tracks household income or parental education, but that never mentions either, has done noticeably less work than one that does. And watch the verb: compare the actual phrase in the abstract ("associated with," "correlated with") against how the same finding gets described anywhere else you encounter it — a press release, a news headline, a course selling you something. "Associated with" quietly becoming "improves" or "causes" somewhere between the paper and the headline is one of the most common tells there is, and one of the easiest to catch once you're looking for it.

None of this requires assuming bad faith

That's the part worth holding onto. Publication bias, funding influence, unreplicated small findings, and correlation read as causation are all structural features of how research gets funded, run, and reported — not necessarily evidence that any one person involved did something dishonest. Which is exactly why each one has a specific, checkable trace rather than a tone you have to intuit: a funding declaration, a trial registry entry, a replication record, a stated study design. And they compound. A small, industry-funded study that adjusted for no confounders, found something exciting enough to get published on the first try, and has never been independently replicated can carry all four flags at once — and every one of those four things is something you can go and look at directly, in the time it takes to read past the abstract.

Research Literacy · progress saved in this browser · sign in to sync across devices

Up next

Applying It — University Applications and Real Research

Where the source-tier and evidence-tier discipline actually pays off: an EPQ or extended essay, a supercurricular reading list, and any specific claim in a personal statement

6 min