N-of-1: How to Test Something on Yourself Without Fooling Yourself
Summary
Population studies tell you what works on average; they can't tell you whether a given thing works for you — that's an n-of-1 question, and you can answer it honestly with a few cheap rules (define one outcome, change one thing at a time, use washouts and ideally blinding, run several on/off cycles, and beware regression to the mean and placebo) instead of the usual "I tried it and felt better" self-deception.
Why Strong
Tier 1 (Strong). N-of-1 is an established design, and the confounders it controls (placebo, regression to the mean, natural variation) are well-documented. It's a Strong, concrete method under the Foundational evidence-literacy hub. (The quality of any given person's n-of-1 depends on how well they apply the controls.)
Practical takeaway
A simple, honest n-of-1 protocol:
1. Pick ONE outcome, measured the same way — e.g., sleep latency, a 1–10 energy score logged at a fixed time, a symptom count, or a biomarker (see biomarker_tracking). Pre-decide what change would count as "worth it."
2. Change ONE thing — hold everything else as steady as you can (sleep, training, other supplements).
3. Use on/off blocks with washout — e.g., 2 weeks on, 2 weeks off, repeated 2–3×; choose block length to fit how fast the effect should appear and reverse (see relative_vs_absolute_risk's cousin: judge by the expected response time).
4. Blind yourself if you can — identical-looking capsules (one active, one placebo) prepared/labelled by someone else, or at least don't peek at which block you're in until after logging. This is the step that separates real effects from placebo.
5. Track, don't remember — log daily; don't reconstruct from memory.
6. Judge across cycles — a real effect tracks the on/off pattern repeatedly; a one-off coincidence won't.
Then: keep what reliably tracks; drop what doesn't (most things won't beat placebo — that's a useful result, saving money and effort). Treat your result as personal, not a recommendation for others.
Evidence detail
Why This Entry Exists
A huge share of real health decisions are individual: does this supplement help my sleep? does cutting dairy clear my skin? does this nootropic actually do anything for me? Group averages can't settle these — individual response varies, and "average benefit" hides responders and non-responders. The honest tool is a structured self-experiment (n-of-1), which is a recognised study design, not just biohacker theatre. But naive self-testing ("I started X and felt great") is one of the easiest ways to fool yourself — placebo, expectancy, natural fluctuation, and starting-when-you're-worst all conspire to manufacture fake wins. This entry gives the method that separates a real personal effect from a story.
It's a spoke of rct_vs_observational_evidence: it applies the same logic (controls, blinding, replication) at the scale of one person.
What bad thinking this protects against:
• "I tried it and felt better, so it works (for me)." → the classic uncontrolled n-of-1 fooled by placebo + timing.
• "It worked for me, so it'll work for you." → over-generalising a personal result.
• Changing five things at once → never knowing which (if any) did anything.
Evidence
1. N-of-1 is a legitimate, formal design (Tier 1). N-of-1 trials — multiple, ideally randomised and blinded, crossover periods within a single person — are recognised in evidence-based medicine (some hierarchies rank a well-run n-of-1 highly for the individual). They're used clinically (e.g., to test whether a specific drug helps a specific patient) and are the rigorous version of self-experimentation.
2. The confounders that fake personal "wins" (Tier 1):
• Placebo/expectancy — believing something works produces real felt improvement (see belief_effects_and_honest_framing); without blinding you can't separate this from a true effect.
• Regression to the mean — you usually start an intervention when symptoms are worst; they'd often drift back toward baseline anyway, and you credit the intervention.
• Natural fluctuation / seasonality / co-interventions — sleep, mood, skin, energy vary for many reasons; a single before/after can't attribute the change.
• Recall/confirmation bias — you remember the good days and the data that fit your hope.
3. Why one cycle isn't enough (Tier 1). A single on→off can be coincidence. Repeated on/off (ABAB…) cycles, with the outcome tracked the same way each time, are what turn a hunch into a signal — if the effect reliably tracks the intervention across cycles, it's probably real.
4. What it can and can't tell you (Tier 1). It can establish whether you respond to something with a reasonably fast, reversible, measurable effect (sleep, energy, a symptom, a biomarker). It can't assess slow, cumulative, or one-directional outcomes (cancer risk, bone density over years), rare harms, or anything you can't measure — and it never generalises beyond you.
Mechanism
Why structure beats intuition. The brain is a story-making machine that defaults to "I did X, then felt Y, therefore X caused Y" — ignoring placebo, regression, and noise. Structure neutralises each: a defined outcome stops goalpost-moving; one variable at a time isolates cause; washout periods stop carry-over; blinding (where possible) removes expectancy; repeated cycles rule out coincidence; pre-committing to what counts as success stops post-hoc rationalisation.
Why it's the right tool for "does it work for me." Population RCTs estimate an average treatment effect; real individuals scatter around that average (true responders and non-responders exist for many interventions). For a fast, reversible, measurable outcome, a personal crossover directly answers your question in a way no group study can.
Risks And Contraindications
• Don't n-of-1 dangerous or irreversible things — stopping essential medication, extreme protocols, anything with serious downside; reserve it for low-risk, reversible interventions.
• Some outcomes can't be self-tested — long-term/cumulative risks, rare harms, hard endpoints; for those, defer to population evidence (the hub).
• Don't over-generalise — "worked for me" is an n-of-1, not evidence for anyone else (the inverse of dismissing population data because you didn't respond).
• Beware stacking — running several changes at once forfeits the whole point.
• Not a substitute for medical diagnosis/treatment of significant conditions.
Controversy
Little methodological controversy — n-of-1 is an accepted design. The friction is cultural: the biohacking scene runs lots of uncontrolled, unblinded self-experiments and over-trusts them, while skeptics dismiss self-experimentation wholesale. Realised's position threads it: structured, blinded-where-possible, repeated-cycle n-of-1 is genuinely informative for you; the naive "I tried it and felt great" version is mostly placebo + regression and shouldn't drive lasting decisions.
Cross-Pillar Connections
• Hub (rct_vs_observational_evidence): the same controls (randomisation, blinding, replication) at the scale of one person; population evidence sets your prior, n-of-1 tests your response.
• belief_effects_and_honest_framing: placebo/expectancy — the main thing blinding defends against.
• relative_vs_absolute_risk: judging whether an observed personal change is meaningful, and respecting expected response times.
• biomarker_tracking: choosing and measuring an objective outcome for the experiment.
Industry bias note
The supplement/wearable/biohacking market thrives on uncontrolled self-experimentation: testimonials ("I tried it and…") are cheap, persuasive marketing, and devices encourage tracking-as-engagement without controls. Teaching proper n-of-1 is partly a defence against that — most products won't beat a blinded placebo block, and finding that out protects the user's wallet. No independent party profits from teaching rigorous self-experimentation, which is why it's under-taught.
Sources (4)
- N-of-1 trial methodology (Guyatt et al.; evidence-based-medicine texts; some EBM hierarchies rank well-conducted n-of-1 highly for individual decisions).↗
- Regression to the mean and placebo/expectancy literature (and belief_effects_and_honest_framing) — the core confounders in uncontrolled self-testing.↗
- Crossover/washout design principles (carry-over effects, repeated periods) from clinical trial methodology.↗
- Funding notation: anchored on independent clinical-trial methodology — none selling a product. The method is, in part, a consumer defence against testimonial-driven marketing.*↗