Concept

Replicability Crisis

Also known as: Replication Crisis

methodology behavioural-insights evidence

Many well-known, widely-cited findings in social psychology and behavioral science — including some the show has covered as settled fact — fail to replicate when tested again, especially at larger scale or with more diverse samples. Michael Hallsworth argues practitioners should care about what actually works in the real world, not about defending a favorite theory, and should update or drop ideas that don't hold up under scrutiny.

The Stanford marshmallow experiment (1972) found that children who resisted eating one marshmallow immediately, in order to get two later, went on to have better life outcomes (test scores, educational attainment, lower BMI). A follow-up study with a far larger and more diverse sample found no such correlation once economic background was accounted for — children from poorer, food-scarce backgrounds were simply more likely to take the immediate marshmallow, regardless of later life outcomes (see Stanford Marshmallow Experiment and Failed Replication).

Honesty-priming via signature timing — the widely-cited finding that signing a declaration before (not after) filling out a form increases honesty — was later re-examined by its own original authors, who, drawing partly on the Behavioural Insights Team's own real-world field trials, found they could not reliably replicate the effect (see Honesty Priming (Signature Timing)).

The good news: an independent academic audit by UC Berkeley and the US Office of Evaluation Sciences, reviewing 165 of BIT's own randomized trials (349 interventions, reaching over 24 million people), found a robust, statistically well-powered average improvement of 8.1% — evidence that low-cost behavioral interventions, evaluated rigorously and transparently (including unpublished null results), can hold up under scrutiny even as individual high-profile findings elsewhere don't (see BIT-Berkeley 165-Trial Nudge Meta-Analysis).

Jason Collins adds two more data points. A large-scale audit by Brian Nosek and colleagues, attempting to replicate roughly 100 published psychology studies, found only about 36% held up — with priming research (including the "Florida effect" elderly-word-priming study) among the weakest-replicating subfields (see Nosek 100-Studies Replication Project).

A full episode devoted to priming's non-replication, with the Florida effect's own failed-replication history (Episode 245). The Florida effect (John Bargh's 1996 finding that priming with elderly-associated words makes people walk more slowly) was first challenged by a 2012 study by Doyen and colleagues, which replicated the design at larger scale and found no effect at all — and suggested the original result may have partly reflected unconscious bias from researchers who knew what result to expect while timing participants with a stopwatch (see Doyen Florida-Effect Replication Failure Study (2012), Bargh Chen Burrows Florida-Effect Walking-Speed Study (1996)). Daniel Kahneman, who had featured the study prominently in Thinking, Fast and Slow, responded in a 2012 open email conceding he feared priming research was "a train wreck looming." A follow-up analysis by the site Replicability Index found 11 of the 12 studies in Kahneman's own priming chapter were unreliable. Phill Agnew's own five-part attempt to replicate well-known priming findings (politeness words, an Apple-vs-IBM logo creativity test, guilt-word priming, salad/cheesecake semantic anchoring, and a reading-speed version of the Florida effect) failed to reproduce any of them as originally reported — see Priming for the full breakdown. But replication failures aren't always the final word: the famous "NBA hot hand doesn't exist" finding, which stood unchallenged for roughly 30 years despite players and fans insisting otherwise, turned out to rest on a subtle statistical error in the original 1985 analysis — corrected, the hot hand is real, and fairly strong (see NBA Hot Hand Reanalysis). Practical guidance for testing new claims: prioritize testing where the cost of being wrong and the cost of running the test are both low (a quick A/B test on a high-traffic page), and be more cautious where redesigning a whole process or form is expensive to test and reverse.

Richard Shotton offers a practical way to work with an uneven evidence base rather than being paralysed by it: hold findings on a continuum from replicated to one-off, and match your spend to where a finding sits. Extremeness aversion sits at the robust end, with studies across electronics, popcorn, beer and coffee all agreeing (see Tversky Camera Extremeness-Aversion Study). At the other end he places magnitude congruence — the Coulter brothers' claim that people conflate the physical font size a price is printed in with its magnitude, so an original price set large and a discounted price set small makes the saving feel bigger — which he believes is a single study, and therefore "still in the world of hypotheses."

His rule: if implementing an idea is costly, use only the well-replicated findings; if it's a costless change to your website, the less-proven ideas are worth testing. That, he says, is where behavioural science moves from a science to an art.

A far more severe case, beyond ordinary non-replication: outright data fabrication (Episode 141). Francesca Gino, a prominent Harvard behavioral scientist, was found by the research-audit blog Data Colada to have apparently fabricated or manipulated data in four separate published papers — not a failure to replicate, but evidence the original results may never have been real (see Data Colada Gino Investigation). Pete Judo names several other areas he considers weakened by unreliable research quite apart from the Gino case: priming generally ("the most egregious" term to avoid using unqualified), the claim that ego depletion can be "replenished" with a sugary drink (largely disproven under stricter testing conditions — see Ego Depletion (Decision Fatigue)), and Amy Cuddy's power-posing research (also disproven, by Cuddy's own later account — see Amy Cuddy). He frames the Gino case as consistent with a pattern of "bad apples" rather than a reason to doubt the field as a whole, pointing to researchers like Katie Milkman actively working to raise the field's methodological standards in response.

A full episode devoted to this pattern (Episode 195). Diederik Stapel fabricated dozens of studies over nearly a decade at Tilburg University, including a flagship Science paper claiming littered environments make people more racially biased — a study he never actually ran (see Stapel Utrecht Train Station Fabrication Case). Karen Ruggiero fabricated an entire discrimination-attribution study published in a top journal (see Karen Ruggiero Fabricated Discrimination-Attribution Study). And — importantly, upgrading a caveat already on this wiki — the "sign at the top of the form" honesty-priming study (see Honesty Priming (Signature Timing)) is not merely a failed replication: a 2021 statistical audit found the underlying dataset itself shows clear signs of fabrication (an implausibly uniform mileage distribution, and duplicate rows disguised with added random noise), the same Ariely-co-authored study this podcast originally cited as real in Episode 7. Amy Cuddy's power-posing research is revisited with its actual original figures, underscoring how a 42-person study produced claims that reached 71 million viewers before failing to replicate (see Amy Cuddy Power-Pose Study). These cases are framed together as evidence for Truth Bias (Truth-Default Theory): confident claims from a credentialed author, published in a trusted venue, and unsurprising enough not to invite scrutiny, can go unquestioned for years.

A full retrospective on the Gino case once it had fully played out (Episode 221). The investigation actually began with PhD student Zoe Zanni in 2018, whose supervisors discouraged her from pursuing it before she brought it to Data Colada herself. Harvard's own 11-month internal investigation (2021) found 28% of one dataset's survey responses manually edited, fabricated "dummy" responses added, and tampered cells still visibly highlighted in grey — confirmed in full by a 1,300-page Harvard report released in 2024. A separate investigation found apparent plagiarism across several of Gino's published books (see Gino Plagiarism Allegations). Gino sued Harvard and the three Data Colada researchers for $25 million; a federal judge dismissed the defamation claims against Data Colada in October 2024 as protected speech, while her breach-of-contract claim against Harvard proceeded (see Gino v. Harvard and Data Colada Lawsuit). Agnew closes on a self-correction framing via Yuval Noah Harari's Nexus: like markets and elections, science earns trust through demonstrated ability to catch and correct its own failures — a scandal properly exposed and resolved is evidence the system works, not proof it can't be trusted.

More fraud- and precision-driven cases, plus another famous non-replicating priming finding (Episode 202). Dan Simons adds physicist Jan Hendrik Schön's fabricated superconductivity results, uncovered only after over a hundred labs failed to replicate them (see Elizabeth Bik Fraud-Detection Cases); the "positivity ratio" claim, whose precise-looking 2.9013 figure was traced to a nonsensical fluid-dynamics model (see Positivity Ratio Fabricated-Precision Case); a 2003 first-person-shooter/cognition study (17 participants, ~3,500 citations) that has never replicated (see First-Person Shooter Cognition Replication Case); and his own high-powered failed replication of the widely-cited "warm coffee makes people seem warmer" finding, originally published in Science in 2008 (see Williams Bargh Warm-Coffee Study — which also notes a possible conflation with a similar claim attributed to Diederik Stapel's fabricated work in Episode 195).

Discussed in

  • Episode 44 — 44- Do Nudges Really Work
  • Episode 75 — 75- Beware of Behaviour Science BS
  • Episode 138 — 138-listen-to-exactly-17-minutes-and-42-seconds-of-this-episode
  • Episode 141 — 141-emergency-pod-harvard-fake-data-scandal
  • Episode 195 — 195-the-lying-psychologist-who-fooled-the-world
  • Episode 202 — 202-the-trade-secrets-con-men-don-t-reveal
  • Episode 221 — 221-francesca-gino-scandal-what-really-happened (full Gino-case retrospective: origin, Harvard's report, plagiarism, lawsuit, self-correction framing)
  • Episode 245 — 245-i-debunked-psychology-s-greatest-myth (Doyen's Florida-effect replication failure; Kahneman's open letter; Phill Agnew's own five failed priming replications)

Related