Every week a study becomes a headline, the headline becomes a post, and the post becomes something you half-believe about your own body. The chain has three or four links and each one loses information. By the time it reaches you, “the receptor did something interesting in twelve mice” has become “scientists discover the switch behind weight loss.”
You do not need a statistics degree to break that chain. You need about ten minutes and a fixed order of questions. This is that order. Run it before you change a supplement, a training block, a diet or an argument you plan to have at dinner.
The seven gates
Work them in sequence. Most headlines fail one of the first three, which means you are usually done in under two minutes.
Gate 1 — Who was it done in? Mice, cells in a dish, or humans. This is the single highest-yield question and almost no headline answers it. Mouse metabolism has misled obesity research repeatedly; a compound that extends life in a worm has cleared a bar roughly no higher than “is a chemical.” A mouse or cell study is a hypothesis about humans, not a finding in them. It can be excellent science and still be irrelevant to your Tuesday. If it is not in humans, stop. File it under “interesting, watch the space,” change nothing.
Gate 2 — How many people? Find the n. Not the number enrolled in the parent trial — the number actually analysed for the result being reported. Under about 30 in a nutrition or performance study, treat any effect as provisional no matter how large it looks. Small samples do not produce wrong answers so much as unstable ones: run the same study again with a different 20 people and the result can move a long way.
Gate 3 — Was there a control group, and were people randomised to it? If everyone in the study got the treatment, then everything else that happened during those eight weeks — the season, the diet advice they were given at intake, the fact that they knew they were being watched — is baked into the result and cannot be separated from it. No control arm means the study can describe a change but cannot attribute it. Randomisation is what makes the comparison fair; without it you are usually looking at an observational association and should downgrade accordingly.
Gate 4 — Did they measure the thing you care about, or a stand-in for it? Studies often report surrogate endpoints: gene expression, a blood marker, an inflammation score, a change on an fMRI. These are cheap and fast, and they are hypotheses about long-run outcomes, not the outcomes themselves. A quieter inflammatory transcript is not a prevented heart attack. Ask: did anyone actually live longer, lift more, sleep better, get sick less? If the answer is “we measured a marker that is associated with that,” you have a mechanism study.
Gate 5 — How big is the effect, in absolute terms? “Cuts your risk by 50%” is a relative number and it is nearly meaningless on its own. If the risk was 2 in 1,000 and becomes 1 in 1,000, that is a 50% relative reduction and a 0.1 percentage-point absolute one. Both are true. Only one tells you whether to care. Always convert: what was the rate before, what is it after, what is the difference in real units? For performance studies, do the same with the raw number — a “significant improvement in sleep quality” might be eleven minutes.
Gate 6 — Who funded it, and who supplied the product? Funding does not invalidate a study. It does belong in the sentence. Look for two things in the paper’s disclosure section: the money, and the materials. An industry body supplying the juice, the supplement or the device is a different kind of exposure from a public research council writing a cheque. Note it, weight it, move on — do not use it as an excuse to dismiss inconvenient findings you would have accepted from a friendlier funder.
Gate 7 — Is this one study, or the twentieth? A single paper is a data point. The question that matters is whether it agrees with the existing body of evidence or contradicts it. A result that confirms twenty prior studies is boring and probably true. A result that overturns them is exciting and probably wrong — not certainly, but the prior is against it. Headlines invert this ranking completely, because the boring one does not travel.
The 90-second source trace
Almost every health headline you see is three steps removed from the paper. Walking back up the chain takes a minute and a half.
- Headline → written by someone optimising for clicks, often without reading the paper.
- Aggregator (ScienceDaily and similar) → usually a near-verbatim reprint of the press release. Useful because it names the journal, the institution and the first author.
- University press release → written by a communications office whose job is attention. This is where mouse studies lose the word “mice.”
- The abstract → free, one paragraph, and contains the species, the n and the actual effect size. This is the stop you are looking for.
Aggregator pages almost always print the DOI and journal reference at the bottom. Search the paper title, read the abstract, check the methods line for n and design. Ninety seconds. You now know more than everyone who shared the post.
The scorecard
Score one point per gate passed. Keep it somewhere you will actually use it.
- Humans (not mice, not cells) — 1
- n above ~100, or above ~30 for a tightly controlled crossover — 1
- Randomised, with a control or placebo arm — 1
- Real outcome measured, not a surrogate marker — 1
- Absolute effect is large enough to notice in your own life — 1
- No direct commercial interest in the result — 1
- Consistent with prior evidence, or is itself a meta-analysis — 1
6–7: worth acting on, if the change is cheap and reversible. 4–5: worth remembering, not worth restructuring anything for. 0–3: entertainment. Read it, enjoy it, change nothing.
Three worked examples
The orange juice study. A São Paulo trial found daily orange juice shifted the activity of 1,705 genes in volunteers’ immune cells. Run the gates: humans, yes (1). n of 20 analysed (0). No control arm — everyone drank the juice (0). Endpoint is gene expression, a surrogate (0). Effect size is real but unconvertible into anything you would feel (0). Juice supplied by a citrus growers’ research fund, though the funding was public (0). One study (0). Score: 1/7. Genuinely interesting molecular work; not a reason to add 40 grams of sugar to your morning.
The GIPR brain-circuit study. Cambridge researchers showed that activating a receptor in the brainstem and blocking it in the hypothalamus both reduce food intake, explaining why two opposite classes of obesity drug both work. Gates: mice (0) — and that is where you stop for personal purposes. Score: 0/7 as guidance. But note what the scorecard is for. As a signal about where the pharmaceutical industry is heading, this study is extremely informative. The filter tells you not to change your behaviour; it does not tell you the research is unimportant. Those are different verdicts and you should keep them separate.
The repetition-bias work. A TU Dresden team pooled 15 datasets and more than 700 participants to show that simply repeating a choice raises how highly people rate that option afterwards. Gates: humans (1). n above 700 (1). Controlled experimental designs (1). The outcome is the behaviour itself, not a proxy (1). Effect is modest but directly interpretable (1). No commercial interest (1). Pooled across 15 datasets, which is the whole point (1). Score: 7/7. Small finding, high confidence — which is exactly the trade you want. Act on it: put a real review step on your renewals rather than trusting that you still like the thing you keep re-choosing.
Failure modes
- Using the filter as a weapon. The gates are for findings you like as much as findings you do not. If you only audit the studies that contradict what you already do, you have built a bias engine, not a filter.
- Confusing “not proven” with “false.” A 2/7 study is unproven. It is not debunked. Hold it loosely, do not campaign against it.
- Demanding 7/7 before doing anything. If a change is free, safe and reversible — walking after dinner, going to bed twenty minutes earlier — the evidence bar can be low, because the cost of being wrong is nearly zero. Scale the bar to the cost, not to the excitement.
- Stopping at the press release. The abstract is one click further and answers the two questions the release omitted.
- Ignoring the money question because the answer is inconvenient. Also ignoring everything else because of it.
Scripts
When someone sends you the headline:
“Interesting — do you know if that one was in people or in mice? The write-up doesn’t say and it usually matters more than the finding.”
When you are tempted by a supplement built on one paper:
“What was the absolute change, and how many people did they measure it in? If I can’t answer both, I’m not buying it this month.”
When you have to brief someone else on a study:
“Here’s the finding, here’s the species, here’s the n, here’s who paid. My read is [act / watch / ignore].” Four facts and a verdict. Anything longer will not survive the retelling anyway.
The standing rules
- Never change a protocol on the strength of a single study, whatever it scores.
- Species first, n second, control arm third. Most headlines die in the first three questions.
- Convert every percentage into an absolute number before you react to it.
- Match the evidence bar to the cost of being wrong — not to how exciting the claim is.
- “Mechanism found” and “outcome improved” are different announcements. Almost nothing in your feed makes the second one.
The point of the filter is not scepticism for its own sake. It is that your attention and your habits are finite, and the studies worth spending them on are a small fraction of the ones that reach you. Ten minutes at the front end buys back the months you would otherwise spend on something that was only ever true in twelve mice.
