Reading a paper without a stats degree
Five sections, in a specific order, using one real paper that became a TED talk, a decade of confident advice, and its own lead author's public statement that she no longer believed it
9 min read
Reading a study without a stats degree gives you the three numbers that matter — effect size, confidence interval, sample size — and the one question that catches most of the rest: was this the outcome the study actually set out to measure. That's the toolkit. What it doesn't have room for is where those three numbers physically sit inside an actual paper, because a real PDF is not a news article. It has a fixed shape, the shape hides the numbers several sections past where most people stop reading, and learning the shape is a five-minute skill, not a statistics course.
This lesson teaches that shape end to end, against one real paper, chosen because it demonstrates every part of it: what the abstract oversells, what the Methods section actually reveals, what the real effect sizes looked like once you found them, why this particular paper had almost no confidence interval to read at all, and what its own limitations discussion did — and didn't — say. Then it shows what happened five years later, and what the paper's own first author said about it, in public, under her own name.
The abstract: written last, built to be quoted
An abstract runs roughly 150 to 250 words, and it's written after the study is finished, once the authors already know what they found. Its job is to get an editor, a reviewer, or a reader to keep going — not to weigh the evidence for you. That's not a flaw in abstracts; it's what they're for. But it means an abstract routinely states its central claim with no sample size, no confidence interval, and no effect size attached, because none of those numbers make for a better sentence.
The paper: Dana Carney, Amy Cuddy, and Andy Yap, "Power Posing: Brief Nonverbal Displays Affect Neuroendocrine Levels and Risk Tolerance," Psychological Science 21(10), 2010: 1363–1368. It's a strong test case precisely because of how far it traveled — the basis for Amy Cuddy's TEDGlobal 2012 talk, which was for years the second-most-watched talk in TED's own history, and for popular advice, repeated across a decade of coaching and self-help content, that two minutes in an expansive stance before a job interview or a negotiation would measurably shift your hormones and your nerve.
The abstract opens by noting that humans and animals signal power through open postures and powerlessness through closed ones, then asks the actual research question: can adopting the posture cause the state, rather than just express it. It answers yes — two minutes of an expansive "high-power" pose, it says, raised testosterone, lowered cortisol, and increased both felt power and risk tolerance, with the reverse in the closed "low-power" pose. It generalizes from there to a claim about embodiment broadly, and closes on the hook: that two simple poses are enough for someone to "instantly become more powerful."
Read it again and count the numbers. There are none — no N, no p, no effect size, no interval. Every sentence is stated as though it holds unconditionally, which is exactly the register that should send you looking for the Methods section instead of stopping here.
Methods: how many, and what they actually measured
This is where "how many" gets answered, along with how people were assigned to a condition and what was actually measured — not the plain-language claim in the abstract, but the specific instrument or task standing in for it.
Here: 42 participants, 26 women and 16 men, randomly assigned to hold either two "high-power" poses or two "low-power" poses for one minute each. Felt power was a self-report scale. Risk tolerance was a gambling task — keep $2 safely, or risk it for a 50/50 shot at $4. Testosterone and cortisol came from saliva samples, taken before and roughly seventeen minutes after the pose manipulation. Forty-two people is not a large study for a field-defining claim, and the earlier lesson's own warning applies directly: in an underpowered study, the significant results are disproportionately the ones that got lucky, and they run large. [Established] — that exact mechanism, measured at scale across 730 neuroscience studies pooled from 49 meta-analyses, is the finding behind Button, Ioannidis, Mokrysz, et al., "Power Failure: Why Small Sample Size Undermines the Reliability of Neuroscience," Nature Reviews Neuroscience 14(5), 2013: 365–376 — a median statistical power across that literature of just 21%.
Effect size: the number Cohen's benchmarks were built for
Three sections down, the paper's actual numbers: testosterone rose in high-power posers relative to low-power posers, F(1, 39) = 4.29, p < .05, r = .34. Cortisol fell, F(1, 38) = 7.45, p < .02, r = .43. 86% of high-power posers took the gambling risk against 60% of low-power posers, χ²(1, N = 42) = 3.86, p < .05. Self-reported feelings of power were higher, F(1, 41) = 9.53, p < .01, r = .44.
Those r values are Jacob Cohen's benchmarks again, in their correlation form rather than the d form the earlier lesson used for its sleep-aid example — Cohen's own 1988 textbook gives .10, .30, and .50 as small, medium, and large for a correlation coefficient: same source, same caveat about treating them as defaults rather than rules. [Established] By that yardstick, .34 to .44 reads as a solid, unremarkable "medium" effect on the page — which is exactly the range where a genuine effect and a lucky small-sample overestimate look identical from the outside.
The confidence interval that wasn't there
The earlier lesson's advice on a confidence interval is to look at the unimpressive end rather than the middle. This paper makes that advice hard to follow in a useful way: it doesn't report a confidence interval for any headline result at all. What its figures show instead is the standard error of the mean — and a standard-error bar is visually about half the width of the 95% interval for the same data, since a 95% interval runs roughly ±1.96 standard errors while the pictured bar runs ±1. A chart with tight-looking SE bars is not the same claim as a chart with a tight confidence interval, and a paper that shows you the first without the second hasn't shown you what you'd actually want to see before trusting the picture.
Limitations: the paragraph almost nobody reads
The Discussion section is written by the people with the most information about what could be wrong with their own study, and it's the paragraph almost everyone skips once the exciting part is behind them. Here, it isn't cautious: it praises the pose manipulation as simple and elegant, argues it's ready to be taken directly into real-world settings, and closes by calling the implications for everyday life substantial, without once raising sample size as a reason for care. The one caveat it does offer is narrow and easy to miss: it points to a second, previously unpublished study of 49 more participants as reassurance that the effect isn't specific to these exact poses — but that additional study covers only felt power and risk-taking, not testosterone or cortisol. The hormone claims, the ones that made the finding sound biological rather than merely psychological, were never the part that got the second look inside the paper itself.
What happened after
In 2015, a separate team — Eva Ranehill, Anna Dreber, Magnus Johannesson, Susanne Leiberg, Sunhae Sul, and Roberto Weber — ran a close replication in the same journal with 200 participants, nearly five times the original. [Established] They found no significant effect on testosterone, on cortisol, or on any of three behavioral risk measures. The one thing that held up was the softest measure in the original study: people who held the expansive poses still said they felt more powerful.
In 2016, Dana Carney — the original paper's first author — posted a public statement on her Berkeley faculty page saying she no longer believed power-pose effects were real, that her lab had stopped researching them, and that the original effects had been small to begin with. She didn't think the paper should be retracted; nothing in it was fabricated or falsified. She'd simply updated her own belief as the evidence accumulated since 2010 turned against it — a rarer thing to watch happen in public, under a named author's own byline, than it should be. [Established] It's one instance of a pattern common enough across the field that this course's own Red Flags module tracks it at scale, using a different, much larger study.
Why this is worth five minutes of your time
None of this required computing anything. A reader in 2010 who read past the abstract into the Methods and Results sections would have seen a sample of 42, a set of p-values landing just inside the conventional cutoff, and a Discussion section that never once raised its own sample size — every reason for measured skepticism, years before either the replication or Carney's own statement arrived.
That's the same discipline an EPQ's assessment objective for using resources, or a strong personal-statement citation, actually rewards — not finding a source, but showing what you checked it against. Citing "a 2010 study found power posing changes your hormones" is an abstract-level claim, and by 2016 it was one its own author had walked back. Citing the paper's real sample size, the range its effect sizes fell in, and the fact that a larger, later study found the hormonal claims didn't hold — that's the same source, read three sections further than the sentence built to be quoted.
A five-minute version, for any paper
- Skip the abstract's verdict; find the sample size in Methods.
- Find the effect size in Results, in real units if there are any — not the abstract's paraphrase of it.
- Check whether there's a confidence interval or only a standard error next to it, and look at the unimpressive end if there is one.
- Check whether the headline outcome is the one the study said upfront it was built to measure.
- Read the Discussion or Limitations paragraph the authors wrote themselves, and notice what it doesn't raise as a concern.
None of it requires a statistics degree. It requires going a few sections further than the sentence that was built to be quoted.
Up next
Red flags
The four checkable tells behind a claim that only sounds well-supported — and exactly what to look at to find out
7 min