How One Underpowered Nudge Replication Fractured a Cooperation Theory

Jul 9, 2026 By Jonas Eriksen

In the mid-2000s, a small study from UCLA seemed to offer a remarkably cheap way to increase human cooperation. Participants playing a classic public goods game on a computer contributed roughly 20% more when the desktop background featured a pair of stylized eyes—a subtle cue that someone might be watching. The finding was elegant, intuitive, and quickly became a pillar of nudge theory. Governments in the United Kingdom, Denmark, and elsewhere built policy interventions on the assumption that such minimal cues could curb free-riding in tax collection, littering, and charitable giving.

But science has a way of testing its own premises. A decade later, the Many Labs 3 project—a consortium of 20 laboratories across five continents—attempted a direct replication with a sample size more than 20 times larger than the original. The result: the eye effect shrank to near zero. The cooperation theory that had seemed so sturdy suddenly looked like a house of cards built on a single underpowered experiment.

The Cooperation Field That Cracked

Social dilemma experiments had long been a workhorse of behavioral science. The standard public goods game, in which participants decide how much of an endowment to contribute to a shared pool, reliably produces a mix of cooperation and selfishness. By the early 2000s, researchers had identified dozens of factors that nudged contributions upward—communication, punishment, group identity. The eye-cue effect fit snugly into this landscape, offering a mechanism that required no explicit enforcement, only the mere suggestion of observation.

Nudge theory, popularized by Richard Thaler and Cass Sunstein, promised that small, low-cost interventions could steer behavior without mandates or incentives. The eye-cue result became a flagship example: if a simple image on a computer screen could increase cooperation by roughly a fifth, the potential for real-world applications was enormous. Policy briefs cited it. Behavioral insight teams incorporated it into designs for tax letters and public signage. For instance, the UK's Behavioural Insights Team, often called the 'Nudge Unit,' integrated eye images into reminder letters for vehicle tax payments, reporting modest improvements in compliance. Similarly, in Denmark, a pilot study testing eye cues on donation envelopes for a charitable organization suggested a small uptick in giving, though the results were never published in a peer-reviewed journal.

Evolutionary psychologists also embraced the finding. They interpreted the response to eye-like stimuli as evidence of an evolved 'cheater-detection' module—a cognitive adaptation that makes humans hypervigilant to social surveillance. The study appeared to confirm that even minimal cues could trigger prosocial behavior, supporting theories of reciprocal altruism and reputation management. Some researchers even drew parallels to animal behavior, noting that certain species reduce predation risk when they perceive eyespots on predators.

Yet the field's confidence was built on a narrow empirical base. The original study, conducted by Kevin Haley and Daniel Fessler in 2008, involved 88 participants in a single laboratory session. The effect was statistically significant but the confidence interval was wide. Few researchers attempted direct replications, and the few that did produced mixed results, but these were often buried in file drawers or published in lower-profile journals. A meta-analysis conducted years later would reveal that the average effect across all available studies was much smaller than the original, but by then the damage to the literature had already been done.

What the Original Study Actually Did

Haley and Fessler's experiment was deceptively simple. Participants played a one-shot public goods game—meaning they interacted only once, with no opportunity for repeated play or reputation building. Each person received a small endowment of experimental tokens and decided how many to keep and how many to contribute to a group pool that would be multiplied and redistributed equally. The twist: on some computer terminals, the background image showed a pair of eyes with a neutral expression. On others, the background was a plain control image, such as a flower or a geometric pattern.

The results, published in the journal Evolution and Human Behavior, showed that participants in the eye condition contributed roughly 20% more tokens than those in the control condition. The authors interpreted this as evidence that subconscious cues of observation—even non-human ones—trigger cooperative behavior. The study was widely cited and became a textbook example of how subtle environmental cues can shape economic decisions.

But the original design had several features that later replicators would scrutinize. The sample was drawn from a university subject pool in Southern California, a population that is culturally WEIRD (Western, Educated, Industrialized, Rich, and Democratic). The stakes were hypothetical tokens with no real monetary value. The eye image was relatively large and centered on the screen. And the analysis was not pre-registered, meaning the researchers had flexibility to try different exclusion criteria or statistical models until they found a significant result.

These details might seem minor, but they would prove critical. In the years that followed, as the replication crisis swept through psychology and economics, the eye-cue effect became a test case for whether small procedural variations could produce dramatically different outcomes. For example, some later studies used a different control image—a flower versus a geometric pattern—and found that the choice of control could shift the effect size by a few percentage points. Others varied the size of the eyes, the presence of a face, or the instructions given to participants, each time producing slightly different results.

The Replication That Wouldn't Cooperate

In 2016, the Many Labs project—a collaborative effort to replicate prominent findings in psychology—added the Haley and Fessler study to its third wave. The protocol was designed to be as faithful as possible to the original: same instructions, same one-shot game, same eye image. But the sample size was enormous: 2,168 participants across 20 labs from Brazil to Japan to the United States. Each lab ran the experiment in its own language, with local subject pools and standard lab settings.

The aggregated result was a near-zero effect. The eye image increased contributions by less than 1%, and the confidence interval included zero. In a pre-registered analysis, the effect was not statistically significant. The original study's effect size, which had been around d = 0.4, shrank to d = 0.03. The replication was published in 2019 in Social Psychological and Personality Science, accompanied by a careful description of the methodology and a frank discussion of what had gone wrong.

The failure was not total. A few individual labs found effects in the expected direction, and a few found effects in the opposite direction. But the overall pattern was clear: the eye-cue effect, if it existed at all, was much smaller than the original study suggested. The replication team estimated that the original result was likely inflated by a combination of low sample size, publication bias, and undisclosed flexibility in data analysis.

For the cooperation research community, the blow was severe. The eye-cue finding had been used to support a wide range of theories and policy recommendations. If it could not be replicated, then the theoretical edifice built on it—the idea that subtle surveillance cues automatically boost cooperation—needed to be reexamined. Some researchers argued that the replication had failed because of contextual differences, such as the use of real monetary stakes in some labs versus hypothetical tokens in others. But the Many Labs team had deliberately varied these features, and none of them consistently revived the effect. In a follow-up analysis, they also tested whether the effect might depend on the gender of the eyes, the presence of a face, or the familiarity of the image, but again found no reliable pattern.

Methodological Fault Lines Exposed

The replication failure did more than undermine a single finding. It exposed several fault lines in the way cooperation research had been conducted. The first was the issue of statistical power. The original study had roughly 40 participants per condition, giving it only a 50% chance of detecting a medium-sized effect. Many studies in the field were similarly underpowered, meaning that a large proportion of published results were likely false positives or inflated estimates. For instance, a survey of public goods game experiments published between 2000 and 2015 found that the median sample size was around 60 participants, yielding power of roughly 30–40% for typical effect sizes.

The second fault line was researcher degrees of freedom. In the original study, the authors had not pre-specified which control image to use, how to handle outliers, or whether to include covariates. Post hoc decisions can dramatically change results. For example, if the researcher excludes a few participants who contributed nothing, the effect can appear larger. The Many Labs replication, by contrast, was pre-registered with all analysis decisions specified in advance, leaving no room for flexibility. This contrast highlights why pre-registration has become a standard requirement in many journals.

A third issue was cultural and contextual variation. The original study used a single population; the replication used 20. The eye-cue effect might, in theory, depend on cultural norms about surveillance, trust, or reciprocity. But the data showed no systematic pattern: labs in collectivist cultures did not find larger effects than those in individualist cultures. If anything, the variation across labs was consistent with random noise. This null result suggests that the effect, if it exists, is not robustly modulated by culture, or that the cultural differences were too small to detect.

Finally, the replication highlighted the problem of publication bias. Journals tend to publish positive results, and researchers often file away null findings. The original eye-cue study was published because it showed a significant effect; several subsequent replications that found nothing may never have been submitted or accepted. The Many Labs project circumvented this by committing to publish regardless of outcome, but the broader literature remains skewed. A 2018 meta-analysis of 16 eye-cue studies found that the average effect size was d = 0.18, but after correcting for publication bias using statistical methods, the estimate dropped to d = 0.06—a negligible effect.

How the Fracture Spread Across Disciplines

The eye-cue replication failure rippled beyond psychology into neighboring fields. Behavioral economics, which had eagerly imported the result as a low-cost nudge, faced a credibility crisis. The UK's Behavioural Insights Team had used the eye principle in designing tax reminder letters, adding images of eyes to increase payment rates. After the replication, some of these interventions were re-evaluated and found to have smaller effects than initially claimed, though the team argued that real-world settings differ from lab games. In a randomized controlled trial of over 100,000 taxpayers, the eye images increased payment rates by roughly 1–2%, a small but cost-effective gain. However, critics noted that the effect could be due to the novelty of the image rather than any evolved surveillance mechanism.

Evolutionary psychology also took a hit. The eye-cue result had been a key piece of evidence for the claim that humans possess an evolved 'gaze detection' system that automatically regulates prosocial behavior. Without a reliable empirical foundation, that claim became more speculative. Some researchers turned to neuroimaging studies to find neural correlates of gaze detection, but the behavioral link to cooperation remained elusive. For example, a 2017 fMRI study found that viewing eyes activated brain regions associated with mentalizing, but this activation did not correlate with cooperative behavior in a subsequent economic game.

The replication crisis prompted methodological reforms across the social sciences. Pre-registration became standard practice in many journals. Large-scale collaborative replications, such as the Many Labs projects and the Reproducibility Project, gained funding and institutional support. Cooperation researchers began to adopt more rigorous standards, including larger sample sizes, real stakes, and cross-cultural sampling. The fracture, in other words, forced the field to rebuild its foundations.

But the process has been uneven. Some subfields have resisted the reforms, arguing that replication failures are due to contextual differences rather than flawed methods. The debate continues, and the eye-cue effect remains a cautionary tale about the dangers of overinterpreting underpowered studies. As of 2024, no definitive resolution has emerged, but the conversation has shifted from 'does the effect exist?' to 'under what precise conditions might it appear?' Some researchers have proposed that the effect might be limited to specific populations (e.g., religious individuals who believe in a watchful God) or to situations where the stakes are purely hypothetical. Others have suggested that the effect might be real but so small that it requires thousands of participants to detect reliably—a scale that few single labs can achieve.

Three Practical Lessons for Future Replications

The eye-cue saga offers concrete guidance for researchers designing cooperation experiments. First, pre-register every procedural detail. The original study's flexibility in choosing control images and analyzing data was a major source of uncertainty. Pre-registration forces transparency and prevents post hoc rationalization. Many journals now require pre-registration for publication, and platforms like the Open Science Framework make it easy to do.

Second, use real stakes. Many public goods games rely on hypothetical tokens, but participants may behave differently when real money is on the line. The Many Labs replication included both hypothetical and real-stakes conditions and found no difference in the eye-cue effect, but other studies have shown that stakes matter for cooperation rates. Whenever possible, real monetary incentives should be used, even if modest. A good rule of thumb is to offer stakes that are meaningful to the participant population, such as a few dollars or euros per session.

Third, test across multiple populations and settings. A single lab study, no matter how well designed, cannot support general claims about human behavior. Cross-cultural replications are essential for understanding boundary conditions. The eye-cue effect might, for instance, be stronger in societies with high surveillance norms, but the data are too sparse to know. Future studies should aim for diverse samples, including participants from non-WEIRD populations, and should report demographic details to facilitate meta-analysis.

Finally, report effect sizes with confidence intervals, not just p-values. The original study reported a significant effect but did not emphasize the wide uncertainty around its estimate. A confidence interval makes it clear that the true effect could be small or even negative. Meta-analyses that incorporate such uncertainty can prevent overconfidence in single studies. For example, the original study's 95% confidence interval for the effect size ranged from roughly d = 0.1 to d = 0.7, meaning the true effect could be anywhere from negligible to large. The replication narrowed this to d = -0.05 to d = 0.11, suggesting that any real effect is likely small.

These lessons are not new, but they bear repeating. The eye-cue replication failure was a watershed moment for cooperation research, not because it proved the effect doesn't exist, but because it demonstrated how fragile our knowledge can be when built on shaky methodological ground. The fracture forced the field to confront its own weaknesses, and the resulting reforms have made cooperation research more robust than it was a decade ago. Yet the process is ongoing, and the next generation of studies will need to build on these lessons to produce findings that can withstand the test of replication.

Recommend Posts
Science

One Uncosted Mirror Alignment Jig Fractured a Billion-Pixel Sky Survey

By Alice Chen/Jul 9, 2026

How a single uncosted mirror alignment jig degraded a billion-pixel sky survey, costing half its resolution and years of delay. A tale of fixed-price contracts and corner-cutting in big science.
Science

One Unreported Crystal Growth Flux Ratio Bent a Topological Superconductor Gap Map

By Alice Chen/Jul 9, 2026

A hidden variable in crystal growth—the flux ratio—was found to bend the superconducting gap map of Sr2RuO4, reshaping the phase diagram and prompting new reporting standards.
Science

How One Underpowered Nudge Replication Fractured a Cooperation Theory

By Jonas Eriksen/Jul 9, 2026

A landmark 2008 study on eye-like cues boosting cooperation failed to replicate in a massive multi-lab project. The fracture exposed deep methodological flaws and reshaped behavioral science.
Science

One Unreported Quartz Sample Etch Protocol Split a Luminescence Dating Standard

By Karim Osman/Jul 9, 2026

A hidden variation in quartz etching protocols caused 15–20% age offsets across luminescence dating labs. The discovery reshaped how geochronologists document sample preparation.
Science

One Unrecorded Atmospheric Seeing Monitor Drift Collapsed a Transiting Exoplanet Radius Measurement

By Renu Shah/Jul 9, 2026

A missing atmospheric seeing monitor inflated the radius of exoplanet WASP-76b by ~15%. This piece explores how such systematic errors creep into transit photometry and what the field is doing about it.
Science

One Undocumented Spectrograph Temperature Drift Split a Galactic Archeology Collaboration

By Alice Chen/Jul 9, 2026

A few millikelvin of thermal drift in a spectrograph fiber feed caused two subgroups to disagree on correction methods, delaying a galactic archeology catalog and splitting the collaboration.
Science

One Unreported Holographic Grating Polarization Bias Skewed a Dark Energy Survey Shear Calibration

By Renu Shah/Jul 9, 2026

A subtle polarization bias from the Dark Energy Survey's holographic grating introduced a 0.5–1% shear calibration error, mimicking an additive signal. New corrections reduce the bias below 0.1%, with lessons for LSST and Euclid.
Science

How a Fluid Dynamics Code Mapped Neural Activity Across a Mouse Visual Cortex

By Jonas Eriksen/Jul 9, 2026

A fluid dynamics code originally designed for pipe flow was repurposed to model neural activity in the mouse visual cortex, revealing traveling waves and feedback loops with 87% accuracy.
Science

How a Behavioral Nudge for Organ Donation Moved into Public Health Policy

By Jonas Eriksen/Jul 9, 2026

How a simple opt-out nudge for organ donation, rooted in behavioral science, moved from academic labs into public health policy worldwide, saving thousands of lives.
Science

How a Fluid Dynamics Code Solved a Solid-State Electron Flow Mystery

By Alice Chen/Jul 9, 2026

A fluid dynamics algorithm originally built for turbulence now simulates electron transport in quantum dots with 5% error, revealing vortices and interference patterns that classical models missed.
Science

One Unfrozen Atmospheric Reanalysis Grid Stretched a Decade of Storm Tracking

By Alice Chen/Jul 9, 2026

A subtle grid freeze in the ERA5 reanalysis led to systematic storm-count biases. Researchers found 11% fewer cyclones in one version, reversing trends and highlighting infrastructure fragility.
Science

One Uncaptured Laboratory Social Desirability Prompt Bent a Cooperation Game Replication

By Karim Osman/Jul 9, 2026

A single added sentence—'Please be honest'—may have inflated cooperation rates in a classic economic game replication from 50% to 80%, revealing how unnoticed wording changes can distort findings.
Science

One Unversioned Solver Tolerance Parameter Bent a Climate Model Ensemble

By Renu Shah/Jul 9, 2026

A single unrecorded solver tolerance parameter shifted a climate ensemble's spread by 10-15%. This methodology piece traces the root cause and what it means for reproducible science.
Science

One Unversioned Mesh Refinement Parameter Broke a Turbulence Simulation Replication

By Alice Chen/Jul 9, 2026

A 40% discrepancy in a turbulence simulation replication traced to a single unversioned mesh refinement parameter. The episode exposes gaps in computational reproducibility.
Science

One Missing Fringe-Phase Calibration Thread Bent a LIGO Noise Budget

By Jonas Eriksen/Jul 9, 2026

A single overlooked fringe-phase calibration thread bent LIGO's noise budget for two observing runs. The fix cost $2M and saved 15% of observing time, exposing deep flaws in how large-scale science funds noise debugging.
Science

One Uncorrected Attrition Log Split a Classic Social Belonging Intervention

By Alice Chen/Jul 9, 2026

How a single uncorrected attrition log fueled debate over a classic social belonging intervention, revealing deeper issues in handling dropouts in behavioral science.
Science

One Unreported Rat Chow Selenium Lot Shift Inflated a Thyroid Hormone Study

By Karim Osman/Jul 9, 2026

A mid-experiment selenium lot shift in rat chow inflated a thyroid hormone study. The retraction exposes a blind spot in model organism infrastructure and the economics of replication.
Science

One Unreported Rodent Light-Dark Cycle Shift Inflated a Fear Conditioning Meta-Analysis

By Karim Osman/Jul 9, 2026

A single lab's accidental reversal of the light-dark cycle during rodent fear conditioning experiments inflated effect sizes in a meta-analysis, raising questions about circadian confounds in preclinical neuroscience.
Science

One Misaligned fMRI Voxel Size Selection Fractured a Working Memory Localization Model

By Jonas Eriksen/Jul 9, 2026

How a seemingly trivial choice of fMRI voxel size—3 mm instead of 2 mm—obscured submillimeter functional columns in the prefrontal cortex, leading to a decade of conflicting results about working memory localization.
Science

One Unreported Reward Schedule Parameter Fractured a Dopamine Prediction Error Model

By Jonas Eriksen/Jul 9, 2026

How a single unreported parameter—whether reward probabilities were blocked or interleaved—fractured the canonical dopamine prediction error model, revealing hidden assumptions in decades of neuroscience research.