How to Read Health Evidence: The One Skill Behind Every Other Claim
Summary
Almost every health argument turns on one skill: knowing what kind of evidence a claim rests on and how much weight it can bear — randomised trials can show causation, observational studies mostly show correlation (and are riddled with confounding), mechanism stories prove nothing on their own, and "a study showed…" is nearly meaningless until you ask what kind, how big, funded by whom, and replicated? This is the literacy that lets you weight every other entry in this knowledge base.
Why Foundational
Tier 0.5 (Foundational) because this is not a contestable health claim but the method for evaluating health claims — the near-axiomatic basis on which every other entry's tier is assigned. The hierarchy (randomisation controls confounding; mechanism/anecdote are weak; replication matters) is settled philosophy-of-evidence, not an empirical finding that could be overturned.
Practical takeaway
• Translate every claim into its evidence tier before acting on it. "RCT-confirmed and replicated" earns a behaviour change; "observational association" earns interest, not action; "mechanism/anecdote" earns curiosity, not belief.
• Demand the absolute effect. "Cuts risk 50%" means little without the baseline — 2-in-1000 to 1-in-1000 is a 50% relative reduction and a 0.1% absolute one (see relative_vs_absolute_risk).
• Distrust the single study and the press release. Wait for replication; read what was measured, not the headline.
• Ask who benefits. Funding and conflicts predict findings (see cui_bono_industry_funding_bias) — in both directions (pharma and supplement/wellness industries).
• **For "does it work for me?", run an honest self-experiment rather than guessing (see n_of_1_self_experimentation).
• Hold uncertainty openly.** "We don't know yet" is a valid, common answer — Realised's tiers (Foundational → Experimental) exist precisely to mark how sure we are.
Evidence detail
Why This Entry Exists
You will be told, confidently and constantly, that X causes Y, that a supplement "boosts" Z, that a food is "linked to" disease. Most of it is built on weak evidence dressed as strong. The single highest-leverage thing a person can learn for their health is not a fact — it's how to tell strong evidence from weak, so they can update appropriately instead of being whipsawed by every headline and influencer.
This is the hub for that skill. It gives the evidence hierarchy and the core questions to ask; the spokes drill into the specific traps — how risk numbers mislead (relative_vs_absolute_risk), why the published literature overstates effects (publication_bias_and_evidence_distortion), how to test things on yourself honestly (n_of_1_self_experimentation), and how to follow the money (cui_bono_industry_funding_bias). It's Tier 0.5 (Foundational) because it underwrites the tiering of everything else.
What bad thinking this protects against:
• "A study showed…" → treating any study as settled, regardless of design, size, or funding.
• "It's linked to / associated with…" → reading correlation as causation.
• "It makes sense mechanistically, so it works" → trusting a plausible story over outcome data.
• Headline whiplash → updating hard on single studies that won't replicate.
EVIDENCE (the hierarchy itself)
1. The rough strength order (with caveats):
• Systematic reviews / meta-analyses of RCTs — highest, if the underlying trials are good and unbiased (garbage in, garbage out; see publication bias).
• Randomised controlled trials (RCTs) — can establish causation because randomisation balances known and unknown confounders. The gold standard for "does X cause Y."
• Cohort / observational studies — track real people over time; show associations, vulnerable to confounding (the famous "moderate drinkers live longer" was largely sicker people avoiding alcohol). Useful for hypotheses, long-term/rare outcomes, and where RCTs are impossible — but rarely settle causation alone.
• Case-control, cross-sectional — weaker; prone to recall and selection bias.
• Case reports, animal/in-vitro, mechanism, expert opinion, anecdote — hypothesis-generating at best; mechanism and anecdote prove nothing about human outcomes on their own.
2. Design beats size. A large observational study can be confidently wrong (confounded at scale); a smaller well-run RCT can be right. "n = 500,000" impresses but doesn't fix confounding.
3. Replication is the real test. Single studies — even RCTs — frequently fail to replicate. Ioannidis's "Why Most Published Research Findings Are False" (2005) and the broad replication crisis mean a finding isn't trustworthy until it's repeated by independent groups.
4. Surrogate ≠ outcome. A drug that improves a marker (cholesterol number, bone density, a scan) hasn't been shown to improve what you care about (living longer, fewer heart attacks) until measured directly — many marker-movers failed on hard outcomes.
5. The questions that do the work. For any claim, ask: What design? (RCT vs observational vs mechanism) · How big is the effect, in absolute terms? (see relative_vs_absolute_risk) · Replicated? · Who funded it, and were outcomes pre-registered? (see cui_bono_industry_funding_bias, publication_bias_and_evidence_distortion) · Real-world outcome or surrogate? · Does the headline match what the study actually measured?
MECHANISM (why the hierarchy is shaped this way)
Why randomisation is special. Confounding is the core problem of observational data: people who do X differ from people who don't in a thousand other ways (wealth, health-consciousness, underlying illness). Randomisation assigns X by chance, so on average the groups differ only in X — including on factors no one thought to measure. That's why an RCT can claim causation and a cohort usually can't.
Why mechanism misleads. A plausible biological story ("antioxidants neutralise free radicals, so they should prevent disease") can be entirely real at the cellular level and still fail in the body (high-dose antioxidant trials showed no benefit, beta-carotene harmed smokers). Biology is too complex for "it should work" to substitute for "it was shown to work."
Why the published record is skewed. Positive, novel, industry-favourable results are more likely to be run, completed, published, and promoted — so the literature you see overstates effects (the spoke publication_bias_and_evidence_distortion).
How this hub connects (the evidence-literacy cluster)
This hub anchors the evidence-literacy cluster — the skills for weighting any health claim:
• relative_vs_absolute_risk — how risk numbers are inflated; the single most decision-relevant stats skill.
• publication_bias_and_evidence_distortion — why the published literature overstates effects (file-drawer, industry trials, outcome-switching).
• n_of_1_self_experimentation — testing an intervention on yourself without fooling yourself.
• cui_bono_industry_funding_bias — following the money; conflicts of interest, both pharma and wellness.
• surrogate_endpoints_vs_outcomes — when a lab marker moves but the patient outcome does not follow.
• regression_to_the_mean — why before-after improvement fakes a treatment effect.
• healthy_user_bias — why the people who take the supplement were already going to live longer.
• Closely related: belief_effects_and_honest_framing (placebo/expectancy — why something can feel effective without being effective).
RISKS AND CONTRAINDICATIONS (how this skill gets misused)
• Evidence nihilism — "studies can show anything, so ignore them all." Wrong lesson: the hierarchy exists because not all evidence is equal; use it, don't abandon it.
• Weaponised skepticism — demanding RCT-level proof only for claims you dislike while accepting anecdote for ones you like. Apply the standard symmetrically.
• "No RCT, so it's false." Absence of a trial ≠ disproof; for many lifestyle questions RCTs are impractical, and we reason from the best available evidence while marking the uncertainty.
• Over-updating on mechanism — the most common consumer error: believing a tidy story over outcome data.
Cross-Pillar Connections
This is a cross-pillar foundation: it governs how to read claims in every pillar (a sleep supplement, a diet, a training method, a mental-health intervention). Any entry that says "Tier 1 / Strong" vs "Tier 3 / Emerging" is applying this hub's logic; the spokes are the specific tools.
Industry bias note
Evidence-literacy is itself a target: pharma funds and promotes favourable trials and emphasises relative risk; the supplement/wellness industry leans on mechanism stories, anecdote, and cherry-picked single studies because it rarely has RCTs. Both exploit low evidence-literacy. This hub and its spokes are deliberately symmetric — the same scrutiny applied to a drug claim and a "natural" claim — which is Realised's core editorial stance encoded as a skill.
Sources (5)
- Ioannidis JPA (2005), "Why Most Published Research Findings Are False," PLoS Medicine.↗
- Evidence-based-medicine hierarchies (Sackett; Oxford CEBM Levels of Evidence; GRADE working group) — study-design strength ordering and certainty rating.↗
- Confounding and the "healthy-user"/"sick-quitter" effect literature (e.g., alcohol-mortality reanalyses).↗
- Surrogate-endpoint failures (e.g., antiarrhythmics/CAST; niacin/HDL; antioxidant-vitamin trials) — marker improvement ≠ outcome improvement.↗
- Funding notation: anchored on independent evidence-methodology literature (Ioannidis, CEBM/GRADE) — none selling a product. This is the method, applied symmetrically to all claims regardless of who profits.*↗