How One Underpowered Nudge Replication Fractured a Cooperation Theory
In the mid-2000s, a small study from UCLA seemed to offer a remarkably cheap way to increase human cooperation. Participants playing a classic public goods game on a computer contributed roughly 20% more when the desktop background featured a pair of stylized eyes—a subtle cue that someone might be watching. The finding was elegant, intuitive, and quickly became a pillar of nudge theory. Governments in the United Kingdom, Denmark, and elsewhere built policy interventions on the assumption that such minimal cues could curb free-riding in tax collection, littering, and charitable giving.
But science has a way of testing its own premises. A decade later, the Many Labs 3 project—a consortium of 20 laboratories across five continents—attempted a direct replication with a sample size more than 20 times larger than the original. The result: the eye effect shrank to near zero. The cooperation theory that had seemed so sturdy suddenly looked like a house of cards built on a single underpowered experiment.
The Cooperation Field That Cracked
Social dilemma experiments had long been a workhorse of behavioral science. The standard public goods game, in which participants decide how much of an endowment to contribute to a shared pool, reliably produces a mix of cooperation and selfishness. By the early 2000s, researchers had identified dozens of factors that nudged contributions upward—communication, punishment, group identity. The eye-cue effect fit snugly into this landscape, offering a mechanism that required no explicit enforcement, only the mere suggestion of observation.
Nudge theory, popularized by Richard Thaler and Cass Sunstein, promised that small, low-cost interventions could steer behavior without mandates or incentives. The eye-cue result became a flagship example: if a simple image on a computer screen could increase cooperation by roughly a fifth, the potential for real-world applications was enormous. Policy briefs cited it. Behavioral insight teams incorporated it into designs for tax letters and public signage. For instance, the UK's Behavioural Insights Team, often called the 'Nudge Unit,' integrated eye images into reminder letters for vehicle tax payments, reporting modest improvements in compliance. Similarly, in Denmark, a pilot study testing eye cues on donation envelopes for a charitable organization suggested a small uptick in giving, though the results were never published in a peer-reviewed journal.
Evolutionary psychologists also embraced the finding. They interpreted the response to eye-like stimuli as evidence of an evolved 'cheater-detection' module—a cognitive adaptation that makes humans hypervigilant to social surveillance. The study appeared to confirm that even minimal cues could trigger prosocial behavior, supporting theories of reciprocal altruism and reputation management. Some researchers even drew parallels to animal behavior, noting that certain species reduce predation risk when they perceive eyespots on predators.
Yet the field's confidence was built on a narrow empirical base. The original study, conducted by Kevin Haley and Daniel Fessler in 2008, involved 88 participants in a single laboratory session. The effect was statistically significant but the confidence interval was wide. Few researchers attempted direct replications, and the few that did produced mixed results, but these were often buried in file drawers or published in lower-profile journals. A meta-analysis conducted years later would reveal that the average effect across all available studies was much smaller than the original, but by then the damage to the literature had already been done.
What the Original Study Actually Did
Haley and Fessler's experiment was deceptively simple. Participants played a one-shot public goods game—meaning they interacted only once, with no opportunity for repeated play or reputation building. Each person received a small endowment of experimental tokens and decided how many to keep and how many to contribute to a group pool that would be multiplied and redistributed equally. The twist: on some computer terminals, the background image showed a pair of eyes with a neutral expression. On others, the background was a plain control image, such as a flower or a geometric pattern.
The results, published in the journal Evolution and Human Behavior, showed that participants in the eye condition contributed roughly 20% more tokens than those in the control condition. The authors interpreted this as evidence that subconscious cues of observation—even non-human ones—trigger cooperative behavior. The study was widely cited and became a textbook example of how subtle environmental cues can shape economic decisions.
But the original design had several features that later replicators would scrutinize. The sample was drawn from a university subject pool in Southern California, a population that is culturally WEIRD (Western, Educated, Industrialized, Rich, and Democratic). The stakes were hypothetical tokens with no real monetary value. The eye image was relatively large and centered on the screen. And the analysis was not pre-registered, meaning the researchers had flexibility to try different exclusion criteria or statistical models until they found a significant result.
These details might seem minor, but they would prove critical. In the years that followed, as the replication crisis swept through psychology and economics, the eye-cue effect became a test case for whether small procedural variations could produce dramatically different outcomes. For example, some later studies used a different control image—a flower versus a geometric pattern—and found that the choice of control could shift the effect size by a few percentage points. Others varied the size of the eyes, the presence of a face, or the instructions given to participants, each time producing slightly different results.
The Replication That Wouldn't Cooperate
In 2016, the Many Labs project—a collaborative effort to replicate prominent findings in psychology—added the Haley and Fessler study to its third wave. The protocol was designed to be as faithful as possible to the original: same instructions, same one-shot game, same eye image. But the sample size was enormous: 2,168 participants across 20 labs from Brazil to Japan to the United States. Each lab ran the experiment in its own language, with local subject pools and standard lab settings.
The aggregated result was a near-zero effect. The eye image increased contributions by less than 1%, and the confidence interval included zero. In a pre-registered analysis, the effect was not statistically significant. The original study's effect size, which had been around d = 0.4, shrank to d = 0.03. The replication was published in 2019 in Social Psychological and Personality Science, accompanied by a careful description of the methodology and a frank discussion of what had gone wrong.
The failure was not total. A few individual labs found effects in the expected direction, and a few found effects in the opposite direction. But the overall pattern was clear: the eye-cue effect, if it existed at all, was much smaller than the original study suggested. The replication team estimated that the original result was likely inflated by a combination of low sample size, publication bias, and undisclosed flexibility in data analysis.
For the cooperation research community, the blow was severe. The eye-cue finding had been used to support a wide range of theories and policy recommendations. If it could not be replicated, then the theoretical edifice built on it—the idea that subtle surveillance cues automatically boost cooperation—needed to be reexamined. Some researchers argued that the replication had failed because of contextual differences, such as the use of real monetary stakes in some labs versus hypothetical tokens in others. But the Many Labs team had deliberately varied these features, and none of them consistently revived the effect. In a follow-up analysis, they also tested whether the effect might depend on the gender of the eyes, the presence of a face, or the familiarity of the image, but again found no reliable pattern.
Methodological Fault Lines Exposed
The replication failure did more than undermine a single finding. It exposed several fault lines in the way cooperation research had been conducted. The first was the issue of statistical power. The original study had roughly 40 participants per condition, giving it only a 50% chance of detecting a medium-sized effect. Many studies in the field were similarly underpowered, meaning that a large proportion of published results were likely false positives or inflated estimates. For instance, a survey of public goods game experiments published between 2000 and 2015 found that the median sample size was around 60 participants, yielding power of roughly 30–40% for typical effect sizes.
The second fault line was researcher degrees of freedom. In the original study, the authors had not pre-specified which control image to use, how to handle outliers, or whether to include covariates. Post hoc decisions can dramatically change results. For example, if the researcher excludes a few participants who contributed nothing, the effect can appear larger. The Many Labs replication, by contrast, was pre-registered with all analysis decisions specified in advance, leaving no room for flexibility. This contrast highlights why pre-registration has become a standard requirement in many journals.
A third issue was cultural and contextual variation. The original study used a single population; the replication used 20. The eye-cue effect might, in theory, depend on cultural norms about surveillance, trust, or reciprocity. But the data showed no systematic pattern: labs in collectivist cultures did not find larger effects than those in individualist cultures. If anything, the variation across labs was consistent with random noise. This null result suggests that the effect, if it exists, is not robustly modulated by culture, or that the cultural differences were too small to detect.
Finally, the replication highlighted the problem of publication bias. Journals tend to publish positive results, and researchers often file away null findings. The original eye-cue study was published because it showed a significant effect; several subsequent replications that found nothing may never have been submitted or accepted. The Many Labs project circumvented this by committing to publish regardless of outcome, but the broader literature remains skewed. A 2018 meta-analysis of 16 eye-cue studies found that the average effect size was d = 0.18, but after correcting for publication bias using statistical methods, the estimate dropped to d = 0.06—a negligible effect.
How the Fracture Spread Across Disciplines
The eye-cue replication failure rippled beyond psychology into neighboring fields. Behavioral economics, which had eagerly imported the result as a low-cost nudge, faced a credibility crisis. The UK's Behavioural Insights Team had used the eye principle in designing tax reminder letters, adding images of eyes to increase payment rates. After the replication, some of these interventions were re-evaluated and found to have smaller effects than initially claimed, though the team argued that real-world settings differ from lab games. In a randomized controlled trial of over 100,000 taxpayers, the eye images increased payment rates by roughly 1–2%, a small but cost-effective gain. However, critics noted that the effect could be due to the novelty of the image rather than any evolved surveillance mechanism.
Evolutionary psychology also took a hit. The eye-cue result had been a key piece of evidence for the claim that humans possess an evolved 'gaze detection' system that automatically regulates prosocial behavior. Without a reliable empirical foundation, that claim became more speculative. Some researchers turned to neuroimaging studies to find neural correlates of gaze detection, but the behavioral link to cooperation remained elusive. For example, a 2017 fMRI study found that viewing eyes activated brain regions associated with mentalizing, but this activation did not correlate with cooperative behavior in a subsequent economic game.
The replication crisis prompted methodological reforms across the social sciences. Pre-registration became standard practice in many journals. Large-scale collaborative replications, such as the Many Labs projects and the Reproducibility Project, gained funding and institutional support. Cooperation researchers began to adopt more rigorous standards, including larger sample sizes, real stakes, and cross-cultural sampling. The fracture, in other words, forced the field to rebuild its foundations.
But the process has been uneven. Some subfields have resisted the reforms, arguing that replication failures are due to contextual differences rather than flawed methods. The debate continues, and the eye-cue effect remains a cautionary tale about the dangers of overinterpreting underpowered studies. As of 2024, no definitive resolution has emerged, but the conversation has shifted from 'does the effect exist?' to 'under what precise conditions might it appear?' Some researchers have proposed that the effect might be limited to specific populations (e.g., religious individuals who believe in a watchful God) or to situations where the stakes are purely hypothetical. Others have suggested that the effect might be real but so small that it requires thousands of participants to detect reliably—a scale that few single labs can achieve.
Three Practical Lessons for Future Replications
The eye-cue saga offers concrete guidance for researchers designing cooperation experiments. First, pre-register every procedural detail. The original study's flexibility in choosing control images and analyzing data was a major source of uncertainty. Pre-registration forces transparency and prevents post hoc rationalization. Many journals now require pre-registration for publication, and platforms like the Open Science Framework make it easy to do.
Second, use real stakes. Many public goods games rely on hypothetical tokens, but participants may behave differently when real money is on the line. The Many Labs replication included both hypothetical and real-stakes conditions and found no difference in the eye-cue effect, but other studies have shown that stakes matter for cooperation rates. Whenever possible, real monetary incentives should be used, even if modest. A good rule of thumb is to offer stakes that are meaningful to the participant population, such as a few dollars or euros per session.
Third, test across multiple populations and settings. A single lab study, no matter how well designed, cannot support general claims about human behavior. Cross-cultural replications are essential for understanding boundary conditions. The eye-cue effect might, for instance, be stronger in societies with high surveillance norms, but the data are too sparse to know. Future studies should aim for diverse samples, including participants from non-WEIRD populations, and should report demographic details to facilitate meta-analysis.
Finally, report effect sizes with confidence intervals, not just p-values. The original study reported a significant effect but did not emphasize the wide uncertainty around its estimate. A confidence interval makes it clear that the true effect could be small or even negative. Meta-analyses that incorporate such uncertainty can prevent overconfidence in single studies. For example, the original study's 95% confidence interval for the effect size ranged from roughly d = 0.1 to d = 0.7, meaning the true effect could be anywhere from negligible to large. The replication narrowed this to d = -0.05 to d = 0.11, suggesting that any real effect is likely small.
These lessons are not new, but they bear repeating. The eye-cue replication failure was a watershed moment for cooperation research, not because it proved the effect doesn't exist, but because it demonstrated how fragile our knowledge can be when built on shaky methodological ground. The fracture forced the field to confront its own weaknesses, and the resulting reforms have made cooperation research more robust than it was a decade ago. Yet the process is ongoing, and the next generation of studies will need to build on these lessons to produce findings that can withstand the test of replication.