Strong Cross-Pillar

Regression to the Mean: Why You Feel Better After You Hit Bottom (Even If the Treatment Did Nothing)

Summary

When you measure something at its extreme — your worst pain day, your highest blood pressure reading, your lowest mood — the next measurement will, on average, drift back toward your own baseline even if you did nothing, because the extreme was partly noise and noise doesn't repeat. People seek help at their worst, naturally drift back, and credit whatever they did in between. This statistical artefact, regression to the mean, fakes the appearance of benefit across self-experiments, clinics, and before-after studies. The honest read of an uncontrolled improvement isn't "it worked" or "it did n

Why Strong

Strong Evidence. Regression to the mean is a mathematically guaranteed consequence of imperfect test-retest correlation — a theorem, not an empirical hypothesis — demonstrated continuously since Galton 1886 and codified as the design rationale for control groups across epidemiology. This is settled, foundational methodology. It sits at Strong rather than Foundational only because the Foundational tier in this knowledge base denotes recovery-first lifestyle foundations, whereas this is an epistemics-and-method spoke under the evidence-literacy hub.

Practical takeaway

How to spot it. Ask one question of any before-after claim: was the starting point chosen because it was extreme? If you began the treatment, the study recruited the patient, or you started tracking because things were unusually bad (or a marker was unusually high), regression is in play and the improvement is partly guaranteed before any treatment acts.

What to do instead of trusting the first improvement:
• Treat "I felt better after X" as a hypothesis, not a verdict — including your own logs. Part of the gain is you returning toward your own baseline.
• Don't start the clock at your worst. Establish a stable baseline first (a week or two of typical days), then begin — so the comparison isn't anchored on an outlier.
• Cycle on and off. If an effect is real, it should reappear across multiple stable-baseline starts and fade when you stop. One improvement at one low point proves nothing; a reproducible on/off pattern is the evidence that defeats the regression objection for you specifically. (See n_of_1_self_experimentation for how to run this without fooling yourself.)
• For markers, expect the bounce. A single high blood-pressure or glucose reading will usually look "improved" on retest regardless of action. Judge marker changes against typical variation, not against the worst reading. (See biomarker_tracking.)
• When you can't compare, say so. The honest conclusion from an uncontrolled improvement is "this design can't tell us," not "it worked" and not "it's placebo."

What to do instead of trusting the testimonial. Before-after photos, "I tried it at rock bottom and recovered" reviews, and single-arm clinic results are the exact shape regression manufactures. They are interest, not evidence. Ask for a comparison group before changing behaviour.

Evidence detail

Why This Entry Exists

The single most convincing thing in health is also one of the least trustworthy: "I was at rock bottom, I tried X, and I got better." That story feels like proof. It is not. A large part of "getting better after your worst" is just being measured at your worst — and the worst, by definition, is unusually far from your typical state and tends not to repeat. You would have improved, on average, regardless of what you did.

This is the spoke that explains why a control group exists. It is the statistical engine underneath placebo groups, miracle-cure testimonials, and your own n-of-1 logs. The skill here is to hold two things at once: regression to the mean is real, ubiquitous, and mathematically guaranteed — and it does not prove your treatment was useless. It proves that an uncontrolled before-after curve is uninterpretable, not that the intervention is inert. The fix is comparison, not cynicism.

What bad advice this protects against, in all directions:
• "I took it at my worst and recovered, so it works" → crediting an intervention for drift you'd have had anyway.
• "The before-after numbers dropped, so the program is effective" → mistaking a selection artefact for a treatment effect.
• "It's all just regression to the mean, nothing works" → the symmetric error: using the artefact to dismiss every real change.
• "Better instruments will fix this" → thinking measurement precision removes regression. It doesn't; only design does.
• "That change is just regression" applied to a routine, unselected measurement → over-applying the rule where the extreme-selection structure isn't present.

EVIDENCE (the canonical cases)

1. Galton 1886 — the origin of the word "regression." Francis Galton's Regression towards Mediocrity in Hereditary Stature studied 1,078 families and found that sons of unusually tall or short mid-parents deviated from the population mean by only about two-thirds of the parental deviation. In his words, "the average regression of the offspring is a constant fraction of their respective mid-parental deviations." Extremes were followed, predictably, by less-extreme values. This is the canonical demonstration and the source of the term itself. (Strong Evidence — the foundational, universally-cited case. Note: Galton's original causal "reversion" framing was later corrected to a purely statistical one — a useful lesson that regression is a property of measurement, not a biological force.)

2. Barnett, van der Pols & Dobson 2005 — the mechanism is guaranteed, not empirical. This standard methodology reference states the condition plainly: whenever you select a follow-up sub-sample because the baseline was extreme, the follow-up regresses toward the mean. It lays out the design and analysis fixes and is the citation that anchors "why a control group exists." (Strong Evidence — the modern methods reference; states the selection-on-baseline condition and remedies explicitly.)

3. Blood pressure — the textbook clinical worked example (Davis 1976; corroborated for home BP by the BP GUIDE study, Wang et al. 2020). Programs that recruit people for high baseline blood pressure and then compare before/after systematically overstate benefit, because high readings selected at baseline decline on retest from regression alone — within-subject variation plus measurement error — independent of any treatment. The established correction is to subtract the model-predicted regression decline from the observed reduction to get the net effect. Critically, regression in home BP is of similar magnitude to office BP: it is not fixed by a better cuff. (Strong Evidence — BP is the standard clinical illustration; the correction method is established.)

4. Mood and depression — the "I felt better at my lowest" case, directly on-scope (Hengartner 2020). People enter depression trials at their nadir; many improve regardless of what is done. Regression to the mean plus spontaneous remission may account for most of the improvement seen in placebo groups. The both-ways point is essential: this does not prove antidepressants are useless. It shows that the placebo-group improvement is largely non-specific drift — which is exactly why drug efficacy is judged by the drug-minus-placebo difference, not by the raw before-after. (The regression / spontaneous-remission mechanism is Strong Evidence and uncontested; the stronger "placebo effect may be near-zero" reinterpretation is Emerging Evidence — one author's contested reassessment. Hengartner carries a known anti-antidepressant prior, so we borrow his mechanism, not his efficacy verdict. The placebo question itself belongs to belief_effects_and_honest_framing.)

5. Kahneman's flight instructors — regression fabricates causal beliefs (Kahneman 2011, from his 1970s work with Israeli Air Force instructors). Instructors noticed that cadets praised after an excellent maneuver did worse next time, while cadets screamed at after a terrible one improved — and concluded that punishment works and praise backfires. Both shifts are pure regression: an exceptional performance (partly luck) is followed by a more typical one regardless of feedback. Kahneman called recognising this "one of the most satisfying eureka experiences of my career." It shows regression doesn't just fake treatment effects — it manufactures confident causal beliefs about whatever we happened to do at the extreme. (Strong Evidence as the canonical cognitive illustration; the best bridge from "statistics" to "why your own n-of-1 conviction is untrustworthy at the extreme.")

6. The fix is comparison, not cynicism (Barnett et al. 2005; depression-trial consensus). A genuine effect and pure regression produce the identical single-arm before-after curve. The only way to separate them is (a) a randomised control group — efficacy equals the treated improvement minus the control improvement — or (b) at the n-of-1 level, a stable or multiple-baseline / on-off design so you aren't starting the clock at an extreme. The aim of research is to determine whether the treated group improves more than regression alone can explain. Better instruments do not remove regression, because it survives reduced measurement error through genuine within-person variation. Only design does. (Strong Evidence — randomisation-as-regression-control is settled methodology and the design rationale for the entire trial apparatus.)

MECHANISM (the statistical logic)

Any single measurement is part signal (your true underlying level) and part noise (today's random fluctuation, measurement error, a bad night, a tense cuff reading). When you pick out an extreme value, you are disproportionately picking moments where the noise pushed hard in one direction. The signal persists on retest; the noise, being random, does not repeat. So the next measurement sits closer to your true level — closer to the mean.

The size of the effect is fully determined by the test-retest correlation r. The predicted standardised follow-up equals r times the standardised baseline. So:
• When r = 1 (no noise, perfect reliability), there is no regression — the extreme repeats exactly.
• When r approaches 0 (all noise), regression is maximal — the follow-up collapses to the mean.
• Real measurements sit between, so any extreme baseline is on average followed by a value closer to the mean.

Two conditions make it bite, and both describe the structure of "I started treatment because I was at my worst": measurement error or genuine variability (so r < 1), and selection on an extreme baseline (so you're sampling the tail). This is why it is a theorem, not a hypothesis — it follows from imperfect correlation by arithmetic, with no biology required. And it is why it cannot be falsified the way an empirical claim can: it is a property of how measurement and selection interact.

Regression is not the same thing as spontaneous remission, natural history of an illness, or placebo and expectancy effects. In real before-after data these co-occur and add together. Regression is specifically the statistical-selection component: the part driven by having chosen an extreme starting point. (Placebo and expectancy are the province of belief_effects_and_honest_framing.)

RISKS AND CONTRAINDICATIONS (the over-correction failure mode)

The main hazard with this entry is wielding it as a universal solvent for any change you'd rather dismiss.
• Cynicism creep. "It's just regression to the mean" is not a proof that a treatment is inert. Regression indicts the uncontrolled inference, not the intervention. A real effect and pure regression look identical in a single arm — so "we can't tell from this" is the correct verdict, not "it does nothing." Don't let the dramatic examples tip into nihilism.
• Over-application. Regression only operates when a measurement was selected or observed because it was extreme. A pre-planned, fixed-schedule measurement on an unselected person or day is not subject to selection-driven regression — ordinary noise still applies, but the artefact that fakes treatment effects does not. Someone who answers every reported change with "that's just regression" is misusing the concept; the extreme-selection structure has to be present.
• Conflating distinct artefacts. Regression is the statistical-selection component only. Spontaneous remission, the natural course of an illness, and placebo/expectancy are separate phenomena that co-occur in before-after data. Naming the wrong one weakens the analysis. Defer the placebo/expectancy piece to belief_effects_and_honest_framing.
• Importing a citation's prior. The depression example draws on a source with a directional anti-antidepressant stance. Use it for the (uncontested) regression and spontaneous-remission mechanism; do not import its stronger, contested claim that the placebo effect is near-zero.

Controversy

Nature of the disagreement. There is essentially no dispute about whether regression to the mean is real — it is a mathematical consequence of imperfect correlation, demonstrated continuously since 1886. The live tension is over interpretation: how much of any given before-after improvement is regression versus a genuine effect, and how readily to invoke it.

Position A — regression manufactures the appearance of benefit. Whenever a moment or a person is selected because a measurement is extreme, the next measurement will on average be less extreme even if nothing is done. People seek help at their worst, drift back toward baseline, and credit whatever they did. This is precisely why uncontrolled before-after data — self-experiments, "I tried X and felt better," single-arm studies — cannot establish that a treatment works, and why the control group exists.

Position B — it does not mean nothing works. Regression means uncontrolled improvement is uninterpretable, not that the treatment is inert. A real effect and pure regression produce the identical single-arm curve; the only way to tell them apart is comparison — a randomised control or a stable multiple-baseline. The correct response is "this design can't tell us," not blanket cynicism. The fix is a comparison group, not abandoning the intervention.

Funding. The core methodology carries low direct bias and runs against sellers' interests: regression explains away exactly the before-after "results" that supplement, device, and wellness marketing depend on. The commercial incentive runs the other way — the uncontrolled-testimonial economy has every reason to ignore regression. One citation needs a flag: the depression source (Hengartner 2020) carries a directional anti-antidepressant prior and is used here only for the uncontested regression / spontaneous-remission mechanism. Galton, Davis, Barnett, and Kahneman carry no product stake.

Realised Position: Realised treats every uncontrolled "I felt better after X" — including the user's own n-of-1 logs — as a hypothesis, not a verdict. When a user starts a supplement or protocol at their worst and improves, the honest read is: part of that is you returning toward your own baseline, and we cannot credit the intervention from before-after alone. This is why Realised pushes toward comparison — on/off cycling, stable-baseline timing, the engine's deterministic before/after windows — rather than ratifying first-improvement enthusiasm. And it is why it never flips to "so it's all placebo, nothing works," which is the symmetric error. We distinguish "we can't tell from this design" from "it doesn't work," and we teach users to build the on/off evidence that would actually settle it for them.

Cross-Pillar Connections

Regression to the mean is genuinely cross-pillar: it operates on any measurement selected at its extreme — a worst sleep week, a flare of joint pain, a low-mood nadir, a high glucose or blood-pressure reading. It is a spoke of the evidence-literacy hub (rct_vs_observational_evidence), the statistical cousin of healthy_user_bias (another reason uncontrolled comparisons mislead), and a precondition for honest self-testing in n_of_1_self_experimentation. It shapes how to read biomarker_tracking (single extreme readings bounce back) and overlaps with — but is distinct from — placebo and expectancy in belief_effects_and_honest_framing. It pairs with surrogate_endpoints_vs_outcomes as a pair of "why the obvious before-after read is wrong" tools.

What would change our mind

Falsifiability: explicit upgrade/downgrade criteria from source

Nothing would overturn regression to the mean itself — it follows from imperfect correlation by arithmetic and is not falsifiable the way an empirical claim is. What would sharpen this entry:

1. A documented case where a claimed "regression artefact" was shown, via proper control, to be a genuine treatment effect that naive analysts had wrongly dismissed as "just regression." This would strengthen the anti-cynicism (Position B) side and is worth hunting for.
2. Quantitative decomposition studies that partition before-after improvement into regression versus spontaneous-remission versus specific-effect for a Realised-relevant domain (sleep, pain, mood) — letting us give magnitudes rather than "a large proportion."
3. A Realised user's n-of-1 on/off cycling that consistently reproduced an effect across multiple stable-baseline starts — which would correctly defeat the regression objection for that user. The entry's job is to teach users to build exactly that evidence rather than to either trust or dismiss the first improvement.

Industry bias note

Structural incentives the evidence base may reflect

Direct bias in the core methodology is low, and the incentive structure is revealing. Regression is uncomfortable for sellers: it explains away the before-after "results" — testimonials, single-arm program data, "I took it at my lowest and recovered" reviews — that supplement, device, and wellness marketing depend on. So the bias runs the other way: the entire uncontrolled-testimonial economy has a commercial interest in ignoring regression and letting natural drift read as product efficacy. Both pharma and wellness exploit low literacy here — pharma via favourable single-arm framing where convenient, wellness via testimonials at the extreme. The corrective is symmetric: demand a comparison group regardless of who's selling. The one source needing a bias flag is the depression citation (directional anti-antidepressant prior), used only for the uncontested mechanism.

Sources (6)

Open in the Library: search, filter, every entry →

We set no cookies and run no ad trackers. We count visits with Cloudflare's cookieless, privacy-first analytics. The only thing stored on your device is which example you last viewed.