Strong Diet

Diet Experimentation

Summary

You can learn whether a dietary change actually does something for your body — but only if you run the trial so that placebo, expectation, getting-better-anyway, and ordinary day-to-day noise can't masquerade as a result: change one thing at a time, give it a defined window, blind the comparison where you can, and never trust a single before-and-after.

Why Strong

Tier 1 for the METHOD because: the design principles (change one variable, defined response window, washout to kill carry-over, replicate/multiple cycles to beat day-to-day noise, blind the subjective comparison, pre-commit the verdict) are core, uncontested trial-design science, and they have been applied specifically to diet in published n-of-1 work (WE-MACNUTR) and in the blinded FODMAP-rechallenge literature. The reasons a careless trial fools you (regression to the mean, expectancy/nocebo, self-limiting course, single-variable confounding) are equally well-established.

Tier 2 for the PAYOFF (that running this will find a real, useful personal effect) because: the evidence that personal dietary responses are large and detectable is good (PREDICT, CGM cohorts) but bounded by high within-person variability, and the most common honest result is a null. We know the method is sound; we are appropriately uncertain about how often it pays off for a given user/question.

NOT Tier 1 overall because: the headline reason-to-bother depends on the personal-response-magnitude evidence, which is Tier 2 and partly entangled with industry-funded work; and we will not let the strong commercial personalisation claim borrow the method's Tier-1 credibility.

NOT Tier 3 because: this is not emerging or speculative — it is established methodology with direct dietary validation, not a hopeful hypothesis. The only genuinely uncertain part is yield, not validity.

The strong commercial "one test → lifelong prescription" personalisation claim is explicitly OUTSIDE the tiering — located in Controversy as over-claim, not endorsed at any tier.

Practical takeaway

Before you start — pick experiments worth running. The discipline below has a real time cost, so spend it where the answer changes what you do. Good n=1 candidates: a suspected food trigger (dairy, gluten, FODMAPs, histamine, caffeine, alcohol), a supplement you can't tell is doing anything, a timing question (carbs at night vs day, fasted vs fed training), a swap you're on the fence about. Bad candidates: anything with a known safety floor (don't "experiment" your calories below 1500/1200 or your sleep below 6h to see what happens), and anything you already comply with well — let the system track those.

The protocol (scale the rigour to the stakes):

1. Write the question and the verdict rule first. "Does cutting dairy reduce my bloating?" → "I'll score morning bloating 0–10 daily; 'yes' = the dairy-free average is ≥2 points lower across at least two on/off repeats." Decide this now, while you can't see the data.
2. Change exactly one thing. Hold sleep, training, alcohol, caffeine and (if relevant) cycle phase as steady as real life allows. Note them anyway — they're your confound log.
3. Set the window by what you're testing:
• Acute (glucose, GI symptoms, sleep that night, energy crash): 3–7 days per condition.
• Subjective steady-state (mood, satiety, energy, skin, regularity): 2–3 weeks per condition — long enough to clear the novelty bump.
• Body composition / weight: 4–8 weeks minimum, judged on the trend line, not any single weigh-in.
4. Wash out between conditions. Roughly 3–7 days of return-to-baseline before flipping, so the old condition's tail doesn't contaminate the new one.
5. Repeat the comparison. This is non-negotiable for anything subjective. On/off/on (and ideally on/off/on/off). One cycle is a hunch; three concordant cycles is a finding. If conditions 1 and 3 (both "on") agree with each other and differ from condition 2 ("off"), you have something. If they don't, you don't.
6. Blind it if the outcome is subjective and the input is blindable. Capsules → have someone fill identical containers, half active half placebo, and tell you the code only at the end. Single suspected food → identical-looking prepared versions. If you genuinely can't blind, label the result "unblinded — expectation not excluded" and trust it less.
7. Measure something you can compare, not just a vibe. A daily 0–10 score in a notes app beats memory, which reconstructs the past to fit the present belief. Objective where cheap (weight, a CGM if you already have one, resting HR, a symptom count).
8. Read it honestly. Look at the average difference across repeats and ask: is it bigger than my normal day-to-day swing? Did it reproduce? Did I keep the one-thing rule? A small, non-reproducing difference is a null — say so.

What "working" looks like (i.e. you ran it well): you can state the result in one sentence with a number and the number came from repeated, comparable measurements — "dairy-free mornings averaged ~3 points less bloating across three on/off cycles, sleep held steady" — or you can state an honest null — "no reproducible difference; stopping the restriction." Either is a win. The only failure is "I think it helped" with no comparison behind it.

What to track: the one variable you're testing; the outcome score (daily); your confound log (sleep, training, alcohol, illness, cycle phase); and the dates of each switch and washout.

A realistic expectation, set up front: per WE-MACNUTR, the single most likely result of a clean dietary n=1 is no clear personal effect (~46% of people, for the macro question they tested). That is not a failed experiment — finding out a change does nothing for you frees you from maintaining it. Recovery, not optimisation: the win is often subtraction.

Evidence detail

Why This Entry Exists

The most useful thing you can find out about your own diet is rarely in a textbook — it's whether this specific change moves your numbers or symptoms. Does dairy actually bloat you, or did you just notice it on a bad gut week? Does creatine do anything you can feel? Did cutting gluten fix your energy, or did you start sleeping better the same week? These are good, answerable questions. The problem is that almost everyone answers them badly, and the two ways of answering badly are mirror images.

Failure mode one: believing a thing works because you felt better after starting it. This is the engine of the entire supplement industry. You start a regimen when you feel worst (that's why you went looking), so you were always going to drift back toward your average regardless — that's regression to the mean. You expected it to help, so you noticed the good and discounted the bad — that's expectancy. The condition was self-limiting and was going to lift anyway. By the time the bottle's empty you're convinced, and you buy the next one. "It worked for me" is the most expensive sentence in wellness.

Failure mode two: dismissing a real effect because the test was too messy to see it. The mirror image. You "tried cutting sugar for three days and nothing happened," so you conclude it does nothing — but three days is inside the noise, you also changed your sleep that week, and you never measured anything you could compare. A real signal was there and your sloppy trial couldn't detect it. This failure is quieter but just as costly: it's how people abandon the one change that would actually have helped.

This entry exists to give the user the small set of disciplines that separate a personal experiment that tells you something from one that just confirms what you already believed. It is not about turning your kitchen into a lab. It's about knowing the four traps well enough that you stop falling into them — which, for most people, is the single biggest upgrade available to their relationship with food.

The boundary with the platform's own machinery (so the user isn't confused): Realised's engines already run a version of this logic in the background — they track compliance, compare against your baseline, and wait a defined response window before judging whether a change is working. That's the system doing structured observation for you. This entry is about the skill you run when you want to test something the system isn't tracking, or when you want to understand why a verdict is trustworthy. Same principles; different operator.

Evidence

1. Personal dietary responses are real and sometimes large — this is what makes n=1 worth doing at all. (Tier 2.)
• **PREDICT 1 (Berry, Valdes, Spector et al., Nature Medicine 2020; foundational data in Zeevi/Segal/Elinav, Cell 2015).** When many people eat the same standardised meal, their blood-glucose responses differ roughly two-fold, and the differences are reproducible enough that clinical + microbiome features predict them better than calories or carbs alone. The headline is genuine: there is real person-to-person variation in how bodies handle the same food. (Funding: PREDICT — Zoe Ltd [industry, commercialises the result] plus academic centres incl. Mass General, King's College London, Stanford, Harvard Chan [government/academic]. The industry conflict is direct and material — see Controversy.)
• **WE-MACNUTR (Tian et al., AJCN 2021) — the cleanest dietary n-of-1 design published. 28 adults, each their own control, three cycles of 6 days high-fat-low-carb vs 6 days low-fat-high-carb with a 6-day washout between. Result that matters for self-experimenters: 32% responded better (glucose) to high-carb, 21% to high-fat, and 46% were non-responders.** Individual response was real and the most common outcome was no clear personal difference. (Funding: National Natural Science Foundation of China + provincial/postdoctoral funds [government].)

**2. The single most important methodological finding: within-person, day-to-day noise to the SAME food is large — large enough to fake or hide an effect in a one-shot trial. (Tier 1 — this is why you must replicate.)**
• Continuous-glucose-monitoring studies in non-diabetic people show that the response to identical meals on different days within the same person varies substantially — driven by prior sleep, stress, illness, recent exercise, time of day, and what you ate before (CGMap, Cell Metabolism 2023; CGM reproducibility analyses, Scientific Reports 2023; Nature Medicine 2024 intrapersonal fasting-glucose variability). Day-to-day fasting glucose alone carries a within-person SD of roughly 7–8 mg/dL before you change anything.
• Implication, stated plainly: if you eat food A today and food B tomorrow and compare once, a large part of any difference you see is which day it was, not which food. A single A-vs-B comparison is usually measuring noise. This is the empirical reason the formal trials use multiple cycles — WE-MACNUTR's authors state outright that "more than 2 intervention pairs" are needed for a stable estimate of an individual effect.

3. Confounds beat self-experimenters constantly; the fix is built into trial design. (Tier 1.)
• **Regression to the mean + self-limiting course + expectancy together largely constitute the placebo response** in self-tracking (statistical analyses of placebo-vs-natural-history, e.g. McGill Office for Science & Society reviews; AJRCCM 2023 showing "placebo" on a walk test was essentially regression to the mean). People measure themselves when they feel worst, then improve toward baseline and credit the intervention.
• Nocebo / expectation runs the other way too — and it's measurable in food trials. In the blinded FODMAP rechallenge work (Singh et al., Gastroenterology 2024; van Lanen feasibility studies), symptoms recurred in ~85% of genuine FODMAP challenges but the glucose control also triggered symptoms in a meaningful fraction of patients — i.e. people react to what they believe they ate. This is why the gold-standard elimination protocols blind the rechallenge (identical-looking powders/brownies, FODMAP vs glucose, dose-escalated over a 3-day challenge + 3-day washout, ~9 weeks to test all groups).
• Single-variable validity. Self-tracking reviews are blunt: n=1 studies "suffer from severe shortcomings" precisely because environmental variables (sleep, stress, the menstrual cycle, training load, travel, illness) move the outcome at the same time as the thing you're testing. The menstrual-cycle literature in particular shows symptom burden and phase shifting sleep, recovery, mood and glucose independently of diet — so for cycling women, phase is a confound that must be tracked or matched, not ignored.

4. Aggregated n-of-1 trials are real evidence — the method scales. (Tier 1, context.) Pooling many properly-run n-of-1 trials via hierarchical/Bayesian models yields population-level effect estimates comparable to parallel RCTs, with better power per participant (systematic review/meta-analysis, J Clin Epidemiol 2016; arXiv Bayesian-adaptive-n-of-1 work 2019). The discipline isn't a poor man's science — done properly it is science, at the resolution of the individual.

Mechanism

There is no biological mechanism to explain here — diet experimentation is a method, and the "mechanism" is epistemic: why a careless personal trial fools you and a careful one doesn't.

Why the careless trial fools you. Your perceived outcome on any given day is a sum of (a) the real effect of the thing you changed, (b) where you happened to be on your own natural up-and-down that day, (c) every other thing that changed (sleep, stress, illness, cycle phase, what you ate yesterday), and (d) what you expected to feel. A naive before-and-after reads the whole sum and attributes it to (a). Because (b), (c) and (d) are often larger than (a), the attribution is usually wrong — and it's wrong in whichever direction your hope or fear pointed.

What each design discipline actually neutralises:
• Change one thing at a time isolates (a) from (c). If you cut dairy and start the gym the same Monday, the experiment is dead on arrival — no result it produces can be assigned.
• A defined response window handles the fact that different changes declare themselves on different timescales: a glucose or acute-GI effect shows in days; satiety and energy shifts in 1–2 weeks; weight and body-composition signal needs weeks-to-months because daily scale noise (water, glycogen, gut contents) swamps real fat change over short windows. Judging too early manufactures false negatives; judging endlessly manufactures false positives by waiting for a good day.
• A washout between conditions removes carry-over (c-from-the-past): the previous diet's effects don't always stop the instant you switch. The dietary trials use ~6 days; the FODMAP rechallenge uses ~3 days per food. The principle: leave enough gap that you're measuring the new condition, not the tail of the old one.
• Replication / multiple cycles is the one most people skip and the one that matters most, because it's the only thing that beats (b) — the day-to-day noise. One A-then-B is a coin flip. A-B-A-B-A-B, or simply repeating the comparison several times and looking at the average difference, is what lets a real signal rise above the noise floor. If the effect only shows up once and never again, it was the day, not the diet.
• Blinding where you can is the only thing that touches (d), expectation. You usually can't blind "I cut carbs." But you often can blind a supplement (capsule vs identical placebo a friend assigns), or a specific food trigger (have someone prepare two outwardly-identical versions). When blinding is impossible, the honest move is to down-weight a positive result, not pretend expectancy isn't operating.
• Pre-commit to the outcome and the threshold — decide before you start what you'll measure and what counts as "worked" (e.g. "average bloating score drops by ≥2 points across three repeats"). This is the defence against the most human confound of all: moving the goalposts to match the result you wanted.

The whole method is one idea: make the real effect the only thing that's allowed to vary, give it room to show up, and force it to show up more than once.

Risks And Contraindications

• Do not experiment past a safety floor. Self-experimentation is for swaps and additions, not for testing how low you can push calories (hard floor 1500 kcal male / 1200 female), sleep, or essential intake. "Let's see what happens" is not a licence to undereat. Restriction experiments in anyone with a history of disordered eating are contraindicated without clinical support — the structure of an "elimination trial" can become a socially-acceptable cover for restriction (see eating_disorder_body_image_diagnostic).
• Elimination diets carry a real cost; don't run them open-endedly. Cutting whole food groups (gluten, dairy, FODMAPs) narrows the diet, can reduce fibre/microbiome diversity, and — if never properly reintroduced — leaves people needlessly restricted on a belief the trial never actually tested. The discipline of reintroduction (the rechallenge) is the part that makes elimination safe and informative; skipping it is how people end up on a 6-food diet for a problem dairy never caused. For gut symptoms, run this with a dietitian where possible (see ibs_diagnostic_lifestyle).
• Rule out the conditions that need diagnosis, not experimentation. Suspected coeliac disease must be tested before removing gluten (removal invalidates the test). Persistent or alarm-feature symptoms (weight loss, blood, nocturnal symptoms, anaemia) are a doctor's office, not an n=1.
• The expectation trap is itself a mild risk. A strongly-held belief that a food is harmful can generate genuine symptoms on (even blinded) exposure — real distress, wrong cause. The nocebo finding in FODMAP trials is the evidence. Treat a "reaction" to a blinded control as information, not failure.
• Minimal physical risk for the method itself when applied to ordinary food swaps and reversible supplements within their normal dose range; the risks above are about what you experiment on, not the experimenting.

Controversy

The controversy: how much can a single person actually learn about their own optimal diet from self-experimentation — and how much is the "personalised nutrition" industry overselling it?

Position A — strong personalisation ("test yourself, get your number"). Commercial personalised-nutrition (the Zoe/PREDICT lineage, CGM-for-the-healthy products) argues individual responses are large, reproducible, and predictable from a one-time test + algorithm, and that you should eat to your numbers. Evidence in favour: the genuine ~2-fold between-person variation in glucose response to identical meals; the predictive models that beat carb-counting.

Position B — n=1 nihilism ("it's all noise and placebo"). The skeptical position holds that within-person day-to-day variability is so large, and confounds and placebo so pervasive, that individual self-experiments mostly produce noise dressed as insight — and that commercial personalisation is overfitting sold as precision (the "garbage in → garbage out" critique of glucose-prediction personalisation).

The funding/bias dimension (both directions, this is the crux):
• Toward over-claim: the personalised-nutrition industry profits when you believe a single test reveals a stable, actionable personal truth — it sells the test, the app subscription, and the supplements keyed to "your" results. The PREDICT data is real and academically solid; the product claim built on it (one test → durable prescription) outruns the reproducibility data.
• Toward dismissal: the same within-person-variability finding that limits the commercial claim is also convenient for incumbents who'd rather you accept generic guidance and not look too closely at what specific foods do to you.
• And the classic Realised signal: the supplement industry's entire growth model depends on users being bad at n=1 — on "it worked for me" surviving because nobody ran a clean comparison. Teaching clean self-experimentation is directly adversarial to that revenue, and accordingly under-taught.

Realised Position: Both extremes are wrong, and the honest answer sits between them — which is exactly why the method matters more than the verdict. Real personal differences exist (Position A is right that variation is real). The within-person noise floor is genuinely high and a single test is often measuring nothing (Position B is right about one-shot trials and about commercial overfitting). The resolution is not to pick a side but to raise the rigour until the question becomes answerable: change one thing, give it a window, wash out, replicate, blind where you can, and pre-commit the verdict. Run that way, n=1 is legitimate evidence at the resolution that matters most to you — your own body. Run carelessly, it's a placebo-detection machine pointed at your wallet. We claim the method confidently and hold the typical yield honestly: most clean dietary experiments end in a useful null, and that is a feature.

Cross-Pillar Connections

• Diet (individual_metabolism_variation_and_personalised_nutrition): that personal variation exists is the premise; this entry is the method for testing it on yourself without being fooled. (diet_adherence): adherence is about sustaining a change; experimentation is about deciding whether the change is worth sustaining — run the experiment first, then sustain what passed. (blood_sugar_regulation): the glycemic-response data is the cleanest worked example of both real personal variation and high within-person noise. (ibs_diagnostic_lifestyle, histamine_intolerance_and_mast_cell_activation): elimination-and-rechallenge is the canonical, highest-stakes self-experiment — and the one most often run badly (no reintroduction, no blinding).
• Mental (belief_effects_and_honest_framing): expectancy and nocebo are the confounds blinding exists to defeat; the same honest-framing discipline applies. (goal_setting_psychology): pre-committing the outcome and threshold before seeing data is goal-setting applied to self-knowledge.
• Physical / cross-cutting (caffeine_stimulant_guidance): caffeine is both a frequent self-experiment target and a confound to hold steady during other trials.

What would change our mind

Falsifiability: explicit upgrade/downgrade criteria from source

We would upgrade (toward "a single n=1 reliably finds your personal optimum"):
• Evidence that within-person day-to-day variability to identical foods is small enough that one well-controlled A/B cycle is reliably conclusive for common dietary questions (current CGM data says the opposite).
• Independent (non-industry) replication showing one-time personalisation tests produce stable predictions that hold months later and improve hard outcomes more than generic guidance.

We would downgrade (toward "even careful n=1 is mostly noise"):
• Evidence that confounds and expectancy are uncontrollable enough in real-world self-tracking that even multi-cycle, pre-registered personal trials don't beat chance for subjective outcomes.
• Demonstration that aggregated/properly-run n-of-1 designs fail to recover effects that parallel RCTs detect.

**Either way, the methodology tier (Tier 1) is robust independent of the magnitude question** — washout, single-variable change, replication and blinding are not in dispute; only how often they'll find a real personal effect is.

Industry bias note

Structural incentives the evidence base may reflect

Bias risk: HIGH, and unusually two-sided.
• The intervention (clean self-experimentation) is unpatentable, free, and revenue-threatening. It costs nothing, requires no product, and its main effect on a user is to make them stop buying things that don't work for them. The supplement and functional-food industries' growth depends materially on the opposite: on "it worked for me" testimonials that survive because no one ran a controlled comparison. This is the classic Realised signal — a high-value, low-revenue practice that is therefore under-taught. There is essentially no commercial actor funding consumer education in clean n=1, which is why most people have never been told the four traps.
• The adjacent product space (personalised nutrition / CGM-for-the-healthy) has a direct, material conflict. The PREDICT science is genuine and largely academic/government-co-funded; the commercial layer built on it (Zoe and similar) profits from the strong interpretation — that a one-time test yields a durable, actionable personal prescription — which outruns the within-person reproducibility data. We cite the science and discount the product claim accordingly.
• The honest middle is in no one's commercial interest, which is precisely why a truth platform should hold it: real personal variation exists and the noise floor is high and the method is what bridges them. Both the over-claimers and the dismissers have a reason to flatten that nuance.

Sources (15)

Open in the Library: search, filter, every entry →

We set no cookies and run no ad trackers. We count visits with Cloudflare's cookieless, privacy-first analytics. The only thing stored on your device is which example you last viewed.