Skip to main content

Nonsensia.pl

A molecule can be active in a cell-free assay, produce a clean result in mice, and show nothing measurable in a randomised trial. None of those three outcomes contradicts the other two. They are answers to different questions asked under different conditions, and the recurring error in this literature is treating them as one accumulating body of proof. What counts as evidence for a cognition compound depends entirely on which question a given experiment was able to ask, and on how much room the design left for a result to appear by chance.

This page is the map for the cognition research category. It sets out the evidence hierarchy, the characteristic failure mode at each level, and the reason most published mechanisms remain hypotheses rather than established facts. It is about how research claims are built and how they break, not about use in humans; the scope of supply is set out in our terms of supply.

What counts as evidence for a cognition compound

Evidence is not a quantity that accumulates into certainty. It is a set of separate claims, each with a scope fixed by the experiment that produced it, and a claim cannot be promoted to a wider scope by being repeated more often or more confidently.

Six claims are routinely collapsed into one sentence. That a molecule binds an isolated target. That it changes something in cultured cells. That it alters a measured behaviour in rodents. That it produces a physiological change in humans. That it changes a cognitive test score in humans. And that an independent group has reproduced any of the above. These are ordered by how much they constrain what happens in an intact organism, and the distance between the first and the last is where most overstatement lives.

The word that hides the collapse is usually a verb. A compound that inhibits an enzyme in a purified preparation has not been shown to inhibit that enzyme in tissue. A compound that changes escape latency in a maze has not been shown to affect a human cognitive construct. The experiment fixes the verb’s scope, and any summary that widens it is making a claim the data does not carry.

The evidence hierarchy from cell-free assay to independent replication

Each rung answers a narrower question than the one the reader usually has in mind. Setting them side by side makes the substitutions visible.

Level of evidence What the experiment can establish What it cannot establish Characteristic failure mode
Cell-free or binding assay That the molecule interacts with an isolated target under chosen conditions That the interaction happens in living tissue at attainable concentrations Activity reported at concentrations no organism would reach
Cell culture That a response occurs in one cell type in a defined medium Relevance to an intact nervous system with barriers and metabolism Cell line drift; no absorption, distribution or clearance
Rodent in vivo study That a treated group differed from a control group on a specific task That the task maps onto any human cognitive construct Small groups, unreported randomisation, unblinded scoring
Small human trial Whether an effect was detected in one sample under one protocol A stable effect size, or generalisation beyond that sample Low power, many outcome measures, selective reporting
Independent replication That the finding survives a different team and different hands The mechanism behind the finding Rarely attempted, and null results are rarely published

Two properties of this ladder matter more than its order. Each rung can be climbed only by doing new work, never by arguing from the rung below. And the volume of published material shrinks sharply towards the bottom, so the loudest claims tend to rest on the weakest rungs.

Why in vitro activity is the weakest rung and the loudest one

A binding assay measures an interaction between purified components in a buffer chosen by the experimenter. That is a real measurement, and it constrains real chemistry. What it does not include is everything that stands between a molecule and a target in a living animal: absorption, distribution, metabolism, clearance, protein binding, and the question of whether the compound reaches the relevant tissue at all.

The difference between in vitro and in vivo evidence is therefore not a difference in quality but in what the system contains. An in vitro system is defined by what the experimenter put into it. An in vivo system contains everything the experimenter did not control, which is precisely where most compounds fail.

This is why a receptor binding assay does not predict a behavioural effect. Affinity for a target is a necessary condition for a target-mediated effect, not a sufficient one, and the conditional runs in only one direction. Reported potency also has to be read against concentrations an organism could plausibly reach; activity demonstrated far above that range describes chemistry rather than pharmacology.

What rodent studies can and cannot carry across to humans

The question of why animal studies fail to replicate in humans has been studied directly rather than merely asserted. A systematic review in the BMJ compared treatment effects in animal experiments with the results of clinical trials for interventions where the human evidence was unambiguous, and found the agreement uneven in both directions. Corticosteroids showed no benefit in clinical trials of head injury, yet the pooled animal data indicated benefit, with an odds ratio for adverse functional outcome of 0.58 and a confidence interval from 0.41 to 0.83. Tirilazad was associated with a worse outcome in patients with ischaemic stroke, while in animal models it reduced infarct volume by 29 percent and improved neurobehavioural scores by 48 percent.

Discordance of that size is not explained by species differences alone. A large part of it is a reporting and publishing problem, and that part has been quantified. An analysis of the animal literature on acute ischaemic stroke, drawing on sixteen systematic reviews covering 525 unique publications, found that only ten of those publications, about 2 percent, reported no significant effect on infarct volume, and only six, about 1.2 percent, failed to report at least one significant finding. A literature in which almost nothing is negative is not a literature reporting everything it found. Trim-and-fill analysis in the same study indicated that publication bias might account for roughly a third of the efficacy reported in systematic reviews, with reported efficacy falling from 31.3 percent to 23.8 percent after adjustment, and estimated a further 214 unreported experiments alongside the 1,359 identified.

The methodological response has been reporting standards rather than new statistics. A workshop convened by the US National Institute of Neurological Disorders and Stroke in June 2012 set a minimum: studies should report sample-size estimation, whether and how animals were randomised, whether investigators were blind to the treatment, and how data were handled. The ARRIVE guidelines, first issued in 2010 and updated as ARRIVE 2.0 in 2020, turn the same principles into a checklist, with a prioritised core known as the ARRIVE Essential 10. The 2020 update states plainly that adherence to the original guidelines had been inconsistent and that the anticipated improvement in reporting quality had not been achieved.

The practical reading of a rodent paper follows from this. An animal study that does not state how many animals were used, how they were allocated, and whether the person scoring the outcome knew the group assignment has not reported enough for its result to be weighed.

Why small human trials produce unstable numbers

Low statistical power is usually described as a reduced chance of detecting a true effect. The less familiar half of the problem is that low power also reduces the probability that a statistically significant result reflects a true effect at all. An analysis in Nature Reviews Neuroscience made that argument and reported that the average statistical power of studies in the neurosciences is very low, with overestimates of effect size and poor reproducibility as direct consequences.

The published figure for median power in that analysis is not reproduced here, because it appears in the body of a paywalled article rather than in any open record confirmed during writing. The qualitative finding, stated in the abstract, is that average power in the field is very low.

The structural version of the argument is older and broader. A 2005 essay in PLoS Medicine set out the conditions under which a published finding is less likely to be true: when studies in a field are smaller, when effect sizes are smaller, when more relationships are tested with less preselection, when there is greater flexibility in designs, definitions, outcomes and analytical modes, when financial or other interests are stronger, and when more teams chase statistical significance in the same area. Its simulations concluded that for most study designs and settings, a research claim is more likely to be false than true. Cognition research meets several of those conditions at once, because cognitive test batteries produce many outcome measures and few of them are fixed in advance.

How badly this can bite was measured in psychology. A large collaborative effort replicated 100 experimental and correlational studies from three journals using high-powered designs and, where available, original materials. Ninety-seven percent of the original studies had reported statistically significant results; 36 percent of the replications reached statistical significance. Replication effects were half the magnitude of the originals, 47 percent of original effect sizes fell within the 95 percent confidence interval of the replication effect, and 39 percent of effects were subjectively rated as having replicated.

Design adds its own instability. Crossover trials, common in this area because each participant serves as their own control, are vulnerable to order and practice effects when the outcome is a cognitive test that improves with exposure; without counterbalancing and an adequate washout, the sequence itself becomes a variable. A systematic review and meta-analysis of controlled trials on two prescription stimulant-type drugs frequently discussed in this context concluded that expectations regarding their effectiveness exceed their actual measured effects, and that only studies with sufficient extractable data could be included in the statistical analysis at all.

Why most proposed mechanisms are still hypotheses

A mechanism is a causal chain: the compound reaches a site, engages a target there, that engagement changes a process, and that process change produces the measured outcome. Published evidence usually covers the first or second link and infers the rest.

A proposed mechanism remains a hypothesis until the chain is closed experimentally. Closing it takes a specific kind of work: demonstrated target engagement in the living system rather than in a homogenate, a concentration-response relationship that behaves as the model predicts, loss of the effect when the proposed target is blocked or removed, and independent reproduction of the whole pattern. Reviews frequently describe a mechanism as established when only the first step has been shown, and the phrase is then copied forward until its origin disappears.

Plausibility is the trap. A story that fits known biology is easier to believe and no more likely to be true, and a mechanism that explains an effect which itself failed to replicate explains nothing. Where the evidence is thin, saying so is the accurate report: for a large share of compounds discussed in cognition research, the honest summary is that a target interaction has been observed in vitro, some rodent behavioural data exist, and the human literature is small, heterogeneous and largely unreplicated.

How to read a claim about a cognition research compound

Five questions separate a supported claim from a repeated one.

  • Which rung is this? A cell-free result and a randomised trial are different claims, regardless of how the summary is phrased.
  • Was the comparison randomised and blinded? If a paper does not say, assume it was not, because reporting standards exist for exactly that reason.
  • Was the outcome fixed in advance? A significant result on one measure out of many, chosen after the data arrived, is a hypothesis rather than a finding.
  • Has an independent group reproduced it? Replication by the originating laboratory answers a weaker question than replication by strangers.
  • Does the summary widen the verb? Compare the claim in the abstract against the claim the design could support; the gap between them is the overstatement.

Frequently asked questions

What is the difference between in vitro and in vivo evidence?

In vitro evidence comes from isolated components in a system the experimenter assembled, such as purified protein or cultured cells. In vivo evidence comes from an intact organism, where absorption, distribution, metabolism and clearance all apply. The difference is not rigour but content: an in vivo system contains the variables that most often stop an in vitro result from generalising.

Does a receptor binding assay predict a behavioural effect?

No. Binding affinity is a necessary condition for a target-mediated effect, not a sufficient one. A compound can bind an isolated target and never reach that target in a living animal, be cleared too quickly, or engage other targets with opposing consequences. Binding data constrains which mechanisms are possible; it does not establish that any behavioural effect occurs.

Why do animal studies fail to replicate in humans?

Partly species differences, and substantially reporting and publication practice. A BMJ systematic review comparing animal experiments with clinical trials found discordance in both directions, including interventions that were beneficial in animal models and harmful or ineffective in patients. Where animal studies omit randomisation, blinding and sample-size reporting, their effect sizes cannot be weighed reliably.

How does publication bias affect animal research?

It inflates apparent efficacy by removing negative results from the record. In the animal literature on acute ischaemic stroke, across 525 publications identified through sixteen systematic reviews, only about 2 percent reported no significant effect on infarct volume. Trim-and-fill analysis indicated that bias could account for roughly a third of the efficacy reported in reviews of that field.

Why is statistical power low in neuroscience studies?

Small sample sizes combined with modest true effect sizes. An analysis in Nature Reviews Neuroscience reported that average power across the neurosciences is very low, and explained that low power both reduces the chance of detecting a real effect and lowers the probability that a significant result is real. Surviving significant results are therefore systematically inflated in magnitude.

How many replications reached statistical significance in the large psychology replication study?

Thirty-six percent. In a collaborative replication of 100 experimental and correlational studies from three psychology journals, 97 percent of original studies had reported statistically significant results, while 36 percent of the replications did. Replication effect sizes were on average half the magnitude of the originals, and 39 percent of effects were subjectively rated as having replicated.

Why is a proposed mechanism still called a hypothesis?

Because a mechanism is a chain of causal links, and most published work demonstrates only the first one. Establishing it requires target engagement shown in the living system, a concentration-response relationship consistent with the model, loss of the effect when the target is blocked, and independent reproduction. Until those exist together, the mechanism is a proposal that fits the data.

References

  1. Ioannidis JPA, Why Most Published Research Findings Are False, PLoS Medicine, 2005
  2. Button KS et al., Power failure: why small sample size undermines the reliability of neuroscience, Nature Reviews Neuroscience, 2013
  3. Estimating the reproducibility of psychological science, Science, 2015
  4. Perel P et al., Comparison of treatment effects between animal experiments and clinical trials: systematic review, BMJ, 2006
  5. Sena ES et al., Publication Bias in Reports of Animal Stroke Studies Leads to Major Overstatement of Efficacy, PLoS Biology, 2010
  6. Landis SC et al., A call for transparent reporting to optimize the predictive value of preclinical research, Nature, 2012
  7. Percie du Sert N et al., The ARRIVE guidelines 2.0: Updated guidelines for reporting animal research, PLOS Biology, 2020
  8. Repantis D et al., Modafinil and methylphenidate for neuroenhancement in healthy individuals: A systematic review, Pharmacological Research, 2010

Research use only. Nonsensia Lab supplies analytical reference standards for laboratory and research applications. This article is published for scientific and educational purposes. It is not medical advice, it does not describe any use in humans, and nothing in it should be read as a recommendation to administer any substance to a person or animal.

Filed under: Nootropics Research

Nonsensia Lab supplies the compounds discussed in this guide as analytical reference standards for laboratory and research use.