Notebook chart and pill organiser used to track nootropic effectiveness over a six-week testing protoco

How to Track Nootropic Effectiveness

How to Track Nootropic Effectiveness — At a Glance

What this coversA six-week method for establishing whether a nootropic is doing anything measurable for you, using a multi-point baseline, objective testing and a defined threshold for what counts as real change.
Strongest available designA blinded multi-crossover self-experiment (on–off–on–off) using an objective cognitive battery. This is only practical for fast-acting compounds; slow-onset compounds require a weaker withdrawal design.
Key findingIn a 14-day sleep-restriction experiment, objective performance declined steadily while participants’ own sense of sleepiness levelled off and could not distinguish six hours in bed from four (Van Dongen et al., 2003, n=48).
Critical caveatA single-arm before-and-after test is not an N-of-1 trial and cannot separate a compound’s effect from practice effects, regression to the mean or seasonal change. It is a starting point, not proof.
Evidence standard usedPeer-reviewed human research on expectation effects, test–retest reliability and reliable-change statistics. Where a study reported a null result, that null result is stated here rather than omitted.
Biggest mistake to avoidTaking one baseline measurement. One measurement tells you nothing about your own week-to-week variability, and without that number you have no way of knowing whether a change is signal or noise.

Educational information only. This article describes a self-tracking methodology and is not medical advice. It does not diagnose, treat or assess any condition. Consult a qualified healthcare provider before starting any supplement regimen, particularly if you have a pre-existing condition or take prescription medication.

Nootropic effectiveness is the single hardest thing to measure in this field, and the overwhelming majority of people never measure it at all. They take something for a fortnight, notice a good Tuesday, and conclude it works. Or they notice nothing in particular, conclude it doesn’t, and move on to the next compound. Both conclusions are drawn from the same source: a vague retrospective impression of how the last few weeks felt.

That impression is not evidence. It is the least reliable instrument you own, and there is a substantial body of research explaining exactly why. Before you spend another few hundred pounds working through the nootropics and supplements shortlist, it is worth learning how to tell a genuine effect from a story you have told yourself.

In 18+ years of testing compounds on myself, the methodology has cost me far more time than the compounds ever did — and it has been worth every hour. The version below is the one I currently use. It is not the strongest possible design, and I will be explicit about where it falls short, because a method that oversells itself is no better than the impressions it replaces.

Why “I Feel Sharper” Is the Weakest Data You Have

There is a study that ought to be required reading for anyone who takes anything for cognition. Researchers gave healthy adults a nasal spray. One group was told it contained a performance-enhancing drug; another was told it contained something that would impair them. Both sprays were inert. Winkler and Hermann (2019) found that the group expecting enhancement rated their own cognitive change substantially more positively than the group expecting impairment, with a large effect size, and also reported feeling less tired.

Here is the part that matters most, and the part usually left out when this study is cited: expectation did not change their actual cognitive performance at all. The authors had hypothesised that it would. It didn’t. Objective task scores were unmoved in both directions while self-assessment swung widely.

That dissociation is the whole argument for structured testing. Belief reliably moves the subjective needle and leaves the objective one alone. If your only instrument is self-report, you are measuring the thing most sensitive to what you were hoping for when you swallowed the capsule.

The dissociation runs the other way too

Self-report doesn’t only produce false positives. It also misses real, substantial impairment. In a fourteen-day controlled sleep-restriction experiment, Van Dongen and colleagues (2003) restricted 48 healthy adults to four, six or eight hours in bed per night. Objective performance in the restricted groups degraded cumulatively across all tasks, day after day. Subjective sleepiness ratings, by contrast, rose sharply at first and then largely flattened — and crucially, they did not meaningfully distinguish the six-hour condition from the four-hour condition.

People getting four hours a night were measurably falling apart and felt roughly the same as people getting six. Your sense of your own cognitive state does not track your cognitive state. This is also why sleep architecture deserves attention before any compound does — and why sleep must be logged throughout any test you run.

The Four Things That Manufacture a Result

Even once you switch to objective testing, four mechanisms will happily hand you an improvement that has nothing to do with what you’re taking. A protocol is really just a set of countermeasures against these four.

1. Practice effects

You get better at cognitive tests by taking cognitive tests. Rijnen and colleagues (2018) tested 158 healthy Dutch adults on the CNS Vital Signs battery, then retested at three months and again at twelve. Scores rose significantly on cognitive flexibility, processing speed and reaction time at the three-month retest.

Two details are commonly misreported. First, the gain did not then reverse — between the second and third administrations there was no significant further change. The improvement plateaued at a new, higher level rather than washing out, which means you cannot simply discard your first retest and assume the effect has gone. Second, those intervals were three and twelve months. If you are retesting weekly or fortnightly, as most self-experimenters do, your practice effect is plausibly larger than what that study observed, not smaller. Nobody has characterised it well at those intervals, and I would treat any short-interval improvement in the first fortnight with real suspicion.

2. Regression to the mean

People start taking things when they feel at their worst. That timing alone guarantees an apparent improvement. As Barnett, van der Pols and Dobson (2005) describe it, unusually extreme measurements tend to be followed by measurements closer to the average, and this effect becomes more pronounced the noisier your measurement is and the more you have selected your starting point on the basis of a low baseline value. They conclude that regression to the mean should always be considered as a possible cause of an observed change.

Which is precisely the situation of someone who buys a nootropic during their worst month of the year. The countermeasure is to build your baseline from several measurements spread over time, not from the one bad week that prompted the purchase.

3. Confounders you didn’t log

Sleep, alcohol, training load, illness, workload and stress all move cognitive scores more than most nootropics plausibly do. If you are not recording them, any of them can masquerade as your compound. Sleep duration and timing, caffeine dose and timing, alcohol, and a stress proxy such as morning HRV are the minimum viable log. Change one variable at a time — the same discipline that governs sensible stacking.

4. Expectation

Covered above, and the hardest to eliminate, because you know what you took. Objective testing blunts it. Blinding removes it, and blinding is more achievable at home than most people assume — I describe how below. The Harvard Health overview of placebo effects is a useful primer on why the ritual of taking something is itself active.

What Actually Counts as a Real Change

This is the question almost every self-experimentation guide skips, and it has a formal answer. Clinical psychology has spent decades on exactly this problem — deciding whether an individual patient’s score has genuinely moved or merely wobbled within measurement error. The standard tool is the Reliable Change Index, introduced by Jacobson and Truax (1991). It is the difference between your two scores divided by the standard error of that difference; a value beyond roughly 1.96 indicates change unlikely to be measurement noise alone.

This is not an obscure academic point. It is the explicit recommendation of the Rijnen practice-effect study cited above, whose stated contribution was providing reliable-change formulae precisely because imperfect test–retest reliability and practice effects otherwise make individual change scores uninterpretable.

🧮 Worked example — turning your baseline spread into a verdict

Most home testing platforms don’t publish the reliability coefficients you’d need for a formal calculation, so use your own variability as the yardstick instead.

Take four baseline measurements before changing anything. Note your highest and lowest scores on the domain you care about. That spread is your personal noise floor. A post-intervention score only counts as a candidate signal if it sits clearly outside that spread — not at its edge — and stays outside it across repeat measurements.

Say your four baseline scores for processing speed span 96 to 112. During the intervention weeks, a score of 114 is not a result — it’s your Tuesday, sitting just inside your own noise. A score of 128 that repeats on retest is worth taking seriously.

Then comes the step that turns a number into a verdict — the withdrawal week. If that 128 drifts back down toward the 96–112 band once the compound is gone, the decline is evidence the compound was doing something. If it stays at 128 with nothing on board, you were most likely watching a practice effect climb, not a real effect. Same score, opposite conclusion — which is exactly why the withdrawal phase, not the intervention phase, carries most of the information.

This is a deliberately conservative rule and it will cause you to miss small genuine effects. I accept that trade willingly. In this field the cost of a false positive — years of money and attention spent on something inert — is far higher than the cost of missing a marginal benefit. NIH’s Know the Science resources on interpreting health research make a similar case for scepticism as the default posture.

Choose Your Design by How Fast the Compound Acts

A genuine N-of-1 trial is not a before-and-after comparison. As Zenner, Böttinger and Konigorski (2022) define it in their work on tools for user-led self-experiments, an N-of-1 trial is a multi-crossover design — you alternate repeatedly between conditions, which is what allows an effect to be separated from a trend over time.

That distinction determines everything about your protocol, and it depends on how quickly the compound acts and clears.

Track A — Fast-acting compounds: a true crossover

If effects appear within an hour and clear within a day — the L-theanine and caffeine stack being the obvious case — you can run the strong design. Alternate compound days and placebo days in blocks, test on each, and repeat the cycle four or more times.

You can blind yourself with about twenty minutes of setup: buy empty capsules, prepare an equal number of active and inert ones, have someone else load them into a fourteen-day pill organiser and write the sequence down without showing you. You break the code only after all testing is complete. This is the single highest-value step in the entire protocol and almost nobody bothers.

Track B — Slow-onset compounds: reversal, honestly labelled

Compounds like Bacopa monnieri need weeks to reach effect and weeks to wash out. Rapid crossover is impossible. The best available compromise is a reversal design: extended baseline, intervention period, then a withdrawal period long enough for the compound to clear, with testing throughout.

Be clear-eyed about what this buys you. It is stronger than a simple before-and-after, because a genuine effect should partially reverse on withdrawal and a practice effect should not. It is still substantially weaker than a blinded crossover: it is unblinded, it has one reversal rather than several, and a seasonal or workload trend running across the whole period can still fool it. I run this design routinely and I do not describe its outputs as proof of anything.

Measurement Methods Ranked by Strength

MethodStrengthWhat it controls for
Blinded multi-crossover, objective battery🟢 Strongest availableExpectation, practice, regression to the mean, day-to-day noise
Unblinded reversal (baseline → on → off)🟡 ModeratePractice effects and some noise; not expectation, not long trends
Multi-point baseline vs. post-test🟡 ModerateRegression to the mean and normal variability
Single baseline vs. single retest🔴 WeakAlmost nothing — practice effect fully uncontrolled
Retrospective impression (“it felt good”)🔴 Not evidenceNothing; maximally sensitive to expectation
Protocol

The NeuroEdge Reversal Protocol

Six weeks. Four phases. Written for slow-onset compounds; Track A substitutions noted where they differ.

Phase 1 · Weeks 1–2

Establish your noise floor

Four cognitive test sessions, same time of day, at least three days apart. Take nothing new. Log sleep duration, caffeine dose and timing, alcohol, and morning HRV daily. Record your highest and lowest score per domain — this spread is the threshold everything later is judged against.

Phase 2 · Weeks 3–5

Single-variable intervention

Introduce one compound at one fixed dose and time. Change nothing else — not training, not sleep schedule, not another supplement. Test weekly. Keep logging. For Track A, this phase becomes alternating blinded on/off blocks with testing in each block.

Phase 3 · Week 6

Withdrawal

Stop the compound. Keep everything else identical. Test twice. This is the phase most people skip and it carries most of the information: a real effect should decay, a practice effect should not. Slow-clearing compounds may need longer than a week here.

Phase 4 · Interpretation

Apply the threshold

A result requires three things together: intervention scores clearly outside your Phase 1 spread, a partial decline during withdrawal, and no confounder in your log that explains it. Two out of three is not a result. Discard the compound or repeat the trial.

Peter Benson, Cognitive Enhancement Researcher

Peter’s Testing Notes

I have run some version of this structure more than sixty times. The most useful thing I can tell you is how often it returns nothing, because that is the part nobody publishes.

My most recent complete run was Bacopa monnieri, 300mg of the 55% bacosides extract from Nootropics Depot, taken with breakfast at 06:40 daily, against a four-session Creyos baseline collected across the preceding fortnight. Nothing distinguishable from baseline through week four. Verbal memory moved outside my baseline spread at week six and held there at week seven. During the withdrawal week it dropped back, but not to baseline — which is exactly the ambiguous outcome this design produces most often, because Bacopa’s clearance is slower than a one-week washout allows. I extended withdrawal to three weeks and it returned to within the baseline range by the end. That is the closest to a positive result I have recorded from an unblinded trial, and I still hold it loosely.

Two failures worth reporting. I once ran a six-week trial that produced a beautiful improvement curve; the log showed my mean sleep had risen by 41 minutes a night across the same period, following a change in my evening schedule. The compound got the credit it hadn’t earned until I checked. And my first three years of testing used single baselines, which I now regard as having produced no usable data whatsoever — 26 months of Oura data has made it very clear how wide my ordinary week-to-week variation is.

I have never managed to blind myself properly for a slow-onset compound. I am not going to pretend otherwise. Everything in Track B carries that limitation.

Updated July 2026.

Key Takeaways

Expectation reliably shifts how you rate your own cognition while leaving objective performance unchanged. Self-report measures your hopes, not the compound.

Subjective sense also misses real impairment: sleep-restricted participants felt roughly the same at four and six hours while objectively degrading day after day.

Practice effects plateau rather than disappear, so you cannot discard an early retest and assume the inflation has cleared.

A change only counts if it sits clearly outside the spread of four baseline measurements — one baseline gives you no way to distinguish signal from an ordinary good day.

Blind yourself where the compound’s speed allows it. Where it doesn’t, run a withdrawal phase and call the result what it is — suggestive, not conclusive.

Tracking Nootropic Effectiveness — FAQ

How long does it take to know if a nootropic is working?

Six weeks minimum for a slow-onset compound, structured as two weeks of baseline measurement, three weeks of intervention and one week of withdrawal. Fast-acting compounds such as an L-theanine and caffeine combination can be assessed in two to three weeks using alternating on and off blocks. Anything shorter than a fortnight tells you almost nothing, because you will not have established your own week-to-week variability, and without that figure you cannot distinguish a real change from an ordinary good day.

Can I trust how I feel instead of taking cognitive tests?

No, and the research on this is unusually clear. In a 2019 trial using inert nasal sprays, participants told they were receiving a cognitive enhancer rated their own performance change far more positively than those told they were receiving an impairing substance — while their measured performance did not differ. Separately, in controlled sleep-restriction work, subjective sleepiness ratings failed to distinguish six hours in bed from four even as objective performance kept declining. Self-report is sensitive to expectation and insensitive to genuine change.

Won’t my scores improve just from taking the test repeatedly?

Yes. In a study of 158 healthy adults on the CNS Vital Signs battery, cognitive flexibility, processing speed and reaction time all rose significantly at a three-month retest with no intervention. Importantly, the gain then plateaued rather than reversing — between the second and third administrations there was no further significant change. So repeated testing inflates your scores to a new level and holds them there. The countermeasure is a multi-session baseline before you start, so that the practice inflation is already priced in.

Is this an N-of-1 trial?

Only the fast-acting version is. A genuine N-of-1 trial is a multi-crossover design in which you alternate repeatedly between conditions, which is what allows a compound’s effect to be separated from a trend over time. The six-week reversal protocol described here is a single baseline–intervention–withdrawal sequence, which is stronger than a simple before-and-after but weaker than a true crossover. Anyone describing an unblinded, single-reversal self-test as an N-of-1 trial is overstating what it can show.

How large does a change need to be before it counts?

Larger than your own measurement noise. Clinical research uses the Reliable Change Index, which divides the difference between two scores by the standard error of that difference and treats values beyond roughly 1.96 as unlikely to be error alone. At home, without published reliability coefficients, use your baseline range as a proxy: collect four baseline sessions, note the highest and lowest score in each domain, and only treat a post-intervention score as meaningful if it falls clearly outside that range and stays outside it on retest.

🧠

Get “7 Days to a Sharper Brain” — Free

Peter Benson’s personal daily protocol, rebuilt from 18 years of testing

Seven evidence-based interventions, in the exact order that makes each one more effective — from sleep foundation to your complete assembled daily stack.

Day 1 — Sleep foundation + Magnesium Glycinate
Day 2 — L-Theanine + Caffeine focus stack
Day 3 — Brain nutrition timing for stable energy
Day 4 — BDNF movement protocol
Day 5 — 90-60-30 sleep environment sequence
Day 6 — Stress resilience + cognitive load framework
Day 7 — Neuroplasticity, Lion’s Mane + your complete assembled daily stack

Join 2,000+ readers optimising their cognitive performance. Unsubscribe anytime.

Peter Benson, Cognitive Enhancement Researcher

Peter Benson

Cognitive Enhancement Researcher | 18+ Years Independent Research

Peter has run structured self-experiments on cognitive compounds since 2008, using objective cognitive testing alongside 26 months of continuous sleep and recovery data. He publishes negative results alongside positive ones.

Last reviewed: July 2026

Scientific References

  1. Winkler, A., & Hermann, C. (2019). Placebo- and nocebo-effects in cognitive neuroenhancement: when expectation shapes perception. Frontiers in Psychiatry, 10, 498. PMID 31354552
  2. Van Dongen, H. P. A., Maislin, G., Mullington, J. M., & Dinges, D. F. (2003). The cumulative cost of additional wakefulness: dose-response effects on neurobehavioral functions and sleep physiology from chronic sleep restriction and total sleep deprivation. Sleep, 26(2), 117–126. PMID 12683469
  3. Rijnen, S. J. M., van der Linden, S. D., Emons, W. H. M., Sitskoorn, M. M., & Gehring, K. (2018). Test-retest reliability and practice effects of a computerized neuropsychological battery: a solution-oriented approach. Psychological Assessment, 30(12), 1652–1662. PMID 29952595
  4. Jacobson, N. S., & Truax, P. (1991). Clinical significance: a statistical approach to defining meaningful change in psychotherapy research. Journal of Consulting and Clinical Psychology, 59(1), 12–19. Full text (PDF)
  5. Barnett, A. G., van der Pols, J. C., & Dobson, A. J. (2005). Regression to the mean: what it is and how to deal with it. International Journal of Epidemiology, 34(1), 215–220. PMID 15333621
  6. Zenner, A. M., Böttinger, E., & Konigorski, S. (2022). StudyMe: a new mobile app for user-centric N-of-1 trials. Trials, 23, 1045. PMC9793632
  7. National Center for Complementary and Integrative Health, National Institutes of Health. Know the Science: how to make sense of health research. nccih.nih.gov
  8. Harvard Health Publishing, Harvard Medical School. The power of the placebo effect. health.harvard.edu

Similar Posts