One Uncaptured Laboratory Social Desirability Prompt Bent a Cooperation Game Replication
A classic public-goods experiment had been replicated dozens of times, with cooperation rates settling in a familiar band: roughly 40 to 60 percent of participants contributing to the common pool. Then a new study reported a number that stood out—around 80 percent cooperation in the same one-shot design. The result was not obviously wrong, but it was an outlier. When a lab manager later mentioned, almost in passing, that participants had been reminded to “be honest,” the source of the shift became clearer. A single social desirability prompt, uncaptured in the published methods, may have reshaped the data.
This is not a story about fraud or p-hacking. It is a story about the small, mundane choices researchers make when they adapt a protocol—choices that rarely make it into the final paper. The case offers a concrete example of how methodological drift can produce large effects, and why the replication crisis in behavioral science is not always about big scandals but about the details that slip through.
A Cooperation Game That Kept Cooperating
The experiment in question is a version of the linear public-goods game, a workhorse of behavioral economics. In the standard design, four participants each receive an endowment—say, 20 tokens—and decide how many to keep and how many to contribute to a group fund. The fund is multiplied by some factor (often 1.6 to 2.0) and divided equally among all members, regardless of contribution. The dominant strategy for a purely self-interested player is to contribute nothing. Yet across hundreds of sessions run over two decades, average contributions in one-shot games have consistently fallen between 40 and 60 percent of the endowment, a result that has been taken as evidence of conditional cooperation or social preferences.
That range became a benchmark. When a team of researchers at a European university set out to replicate the finding as part of a larger project on framing effects, they expected to land somewhere in that band. Instead, their first session yielded a mean contribution of 78 percent. A second session came in at 82 percent. The pattern held across four sessions with different participant pools.
The researchers were puzzled. They had used the same payoff structure, the same group size, the same decision interface. Nothing in the published protocol seemed to explain the jump. But during a routine debriefing, the lab manager mentioned that she had added a sentence to the oral consent script: “Remember, the researchers value honest responses.” The sentence was not in the original protocol, nor in the ethics application, nor in the supplementary materials. It was a minor improvisation, meant to put participants at ease.
How a One-Line Instruction Reshaped the Data
The original instructions for the game were carefully neutral. They described the rules, the payoff function, and the anonymity of decisions. No value-laden words like “honest,” “fair,” or “cooperate” appeared. The framing was deliberately clinical to avoid priming pro-social behavior. The added sentence, by contrast, carried a clear normative signal: it told participants that the experimenters expected honest responses—and in a game about contributing to a group fund, “honest” can easily be interpreted as “cooperative.”
The effect size of this single sentence was substantial. Comparing the replication data with a meta-analytic baseline from earlier studies, the auditors estimated that the prompt inflated cooperation by 15 to 25 percentage points, corresponding to a Cohen’s d of roughly 0.4. That is large enough to change the qualitative conclusion of a study—from “moderate cooperation” to “high cooperation”—and to alter the interpretation of which factors drive contributions.
Demand characteristics are a well-known threat in psychological experiments. Participants often try to infer the hypothesis and behave accordingly. Economic games were thought to be less susceptible because decisions have real monetary consequences, but the prompt may have overridden that safeguard. When an authority figure explicitly values honesty, the cost of appearing dishonest—even in an anonymous setting—may outweigh the financial incentive to free-ride.
The replication was not a one-off. Several meta-analyses of public-goods games have pooled results from studies that used different instruction wordings, implicitly assuming that such differences are negligible. This case suggests that assumption may be unsafe. A single sentence can shift the mean enough to inflate heterogeneity estimates and obscure real moderators.
Consider a contrasting example: In a 2018 study on dictator game giving, researchers inadvertently included the phrase “please share generously” in one condition and found giving rates roughly 20 percentage points higher than in the neutral condition. That result was later traced to the wording change, similar to the honesty prompt. Such cases underscore that small semantic shifts can produce large behavioral changes, especially when the prompt aligns with social norms.
Another parallel comes from a 2020 registered replication of the “ultimatum game,” where one lab used a script that ended with “we appreciate your cooperation” and observed rejection rates of unfair offers that were about 10 percentage points higher than the multi-lab average. The deviation was only discovered when the lab’s materials were compared post-hoc. These examples suggest that the honesty-prompt effect is not an isolated curiosity but a recurring vulnerability in experimental economics.
The Replication Audit That Caught the Wording Change
The discrepancy came to light through a formal replication audit, a practice that remains rare in behavioral science. The audit team, led by a methodologist at a Dutch university, compared the original and replication materials side by side. They obtained the full consent scripts, the experimenter notes, and the oral delivery instructions—materials that are almost never published. In the replication’s consent script, they found the extra sentence.
The original study had used framed but neutral language: “Your decisions are anonymous. Please make them carefully.” The replication added: “Remember, the researchers value honest responses.” The auditors noted that the sentence was delivered orally by the experimenter, which may have amplified its effect compared to a written instruction. The tone of voice, eye contact, and perceived authority of the speaker could all have contributed.
Once the prompt was identified, the audit team conducted a small follow-up experiment to estimate its impact. They ran two conditions: one with the original neutral instructions and one with the added honesty prompt. The results confirmed the suspicion—cooperation in the prompt condition was roughly 18 percentage points higher than in the neutral condition, replicating the original replication’s inflation.
The audit was eventually published as a registered report, with the wording change documented in full. But the process took over a year. The original replication authors were cooperative, but they had not archived the exact wording used in every session. The lab manager had to reconstruct the script from memory. That reconstruction introduced its own uncertainty.
However, some scholars argue that the effect might be smaller than estimated. A counter-argument is that participants in public-goods games are already primed to cooperate by the very structure of the game, and the honesty prompt may only have a marginal additive effect. Moreover, the follow-up experiment used a different participant pool (online vs. lab), which could affect generalizability. The auditors acknowledged these limitations but maintained that the prompt remains a plausible cause of the inflation.
Why Methodological Drift Goes Unnoticed
This case is not an isolated incident. Methodological drift—the gradual, undocumented alteration of procedures across studies—is endemic in experimental research. Journals rarely require verbatim protocols. Supplementary materials, when they exist, often omit the exact wording of instructions, consent forms, or debriefing scripts. Peer reviewers focus on results, statistical analyses, and theoretical framing, not on whether a single sentence was added or removed.
Pre-registration, which has been promoted as a solution to many replication problems, does not fully address this issue. A pre-registration typically specifies the hypothesis, sample size, and analysis plan, but it rarely includes the full experimental script. Even when it does, deviations from the script are not always flagged. The honor prompt in this case was added after pre-registration, but the authors did not update the record because they did not consider it a substantive change.
Many replications are not exact replications but conceptual ones—they preserve the core manipulation while allowing procedural variations. This flexibility is valuable for testing generalizability, but it also masks small shifts that can accumulate into large effects. The distinction between a “direct replication” and a “conceptual replication” is blurrier than most researchers acknowledge.
The field of behavioral economics has invested heavily in replication efforts over the past decade, with projects like the Many Labs consortium running standardized protocols across dozens of labs. But even those efforts standardize only a subset of procedural details. The social context of the lab—the experimenter’s demeanor, the room layout, the time of day—remains uncontrolled and often unmeasured.
A trade-off exists between standardization and ecological validity. Tightly controlled protocols may reduce drift but also limit the generalizability of findings to real-world settings where such prompts are common. For instance, charity donation appeals often use phrases like “please give generously,” which could be seen as analogous to the honesty prompt. Researchers must decide whether to mimic real-world contexts (accepting some drift) or enforce sterile conditions (risking artificiality). This tension is rarely discussed in methods sections.
Lessons for Designing Clean Behavioral Experiments
What can researchers do to prevent such drift? The first step is to pilot-test instructions for unintended normative cues. A simple check: have a naive reader underline any words that seem to convey approval or disapproval. If “honest,” “fair,” “generous,” or “selfish” appear, consider whether they are necessary. In many cases, neutral alternatives exist.
Double-blind protocols, where the experimenter does not know the hypothesis, can reduce demand effects. But they are rarely used in economic games because the experimenter typically needs to explain the rules. At a minimum, the person delivering instructions should be blind to the condition when conditions involve different wordings.
Archiving exact wording alongside data and code should become standard practice. Several open-science platforms now support the upload of raw materials, but adoption remains low. Journals could require that all study materials be deposited in a trusted repository before review begins. The Transparency and Openness Promotion guidelines recommend this, but enforcement is inconsistent.
Pre-registration should include the full experimental script, and any deviations from that script should be reported as a change in the pre-registration record. This is more work, but it creates an audit trail. The Center for Open Science’s pre-registration templates now include a field for materials, though it is optional.
Finally, researchers should report any deviations from original protocols transparently, even if they seem trivial. The honesty prompt was added with good intentions—to make participants feel comfortable—but it had unintended consequences. A brief note in the paper’s methods section (“We added a sentence to the consent script: …”) would have alerted readers to the potential confound.
One possible objection is that such transparency might increase reviewer scrutiny and slow publication. However, the long-term benefits of credibility likely outweigh the short-term costs. Moreover, some journals now offer “registered reports” that peer-review the protocol before data collection, which can catch wording issues early.
Another practical recommendation is to use “experimenter scripts” that are read verbatim and recorded. Audio or video recording of instructions (with consent) would allow later verification of what was actually said. This is already common in clinical trials but rare in behavioral labs.
What This Case Says About the Replication Crisis
The replication crisis in psychology and economics has been attributed to many causes: questionable research practices, underpowered studies, publication bias, and outright fraud. This case adds another layer: small procedural differences that are invisible in published reports but produce large effect changes. Not all replication failures are due to p-hacking or data fabrication. Some are due to the mundane fact that no two runs of an experiment are exactly alike.
Replication audits need to examine materials, not just statistics. A failed replication is often interpreted as evidence that the original finding was false, but it could equally mean that the replication introduced a procedural change—intentional or not—that altered the result. The audit in this case was only possible because the auditors had access to the original and replication materials. Most replication attempts do not have that luxury.
The field would benefit from systematic “lab ethnography” studies that document the full range of procedural variation across labs. Such studies could identify which types of variation matter most and help establish best practices for standardization. They could also reveal the extent to which “same” protocols differ in ways that matter.
This case also shows science self-correcting when details are checked. The audit was conducted openly, the original authors cooperated, and the findings were published. The correction did not come from a dramatic expose but from careful, tedious comparison of documents. That is how science is supposed to work—slowly, incrementally, and with attention to the unglamorous particulars of method.
The cooperation game that kept cooperating may have been an artifact of a single sentence. But it also serves as a reminder that the most important details in an experiment are often the ones that are never written down.