One Uncorrected Attrition Log Split a Classic Social Belonging Intervention

Jul 9, 2026 By Alice Chen

In the early 2010s, a brief writing exercise designed to bolster a sense of social belonging among college students seemed to produce a remarkable result: African American students who completed the exercise earned higher grades than their peers who did not. The study, led by Gregory Walton at Stanford University, became a touchstone for a generation of psychological interventions aimed at closing achievement gaps. But as replication attempts accumulated over the following decade, the effect proved elusive. Then, in 2022, a meta-analysis in Nature Human Behaviour turned attention to a detail that had been hiding in plain sight: an attrition log. One contributing study had lost more control participants than treatment participants, and how that missing data was handled—or not handled—became a flashpoint for a broader debate about rigor in behavioral science.

A Classic Intervention, a Replication Crisis, and an Attrition Log

The social-belonging intervention, pioneered by Walton and his colleague Geoffrey Cohen, was deceptively simple. Incoming college students read survey results suggesting that most students worry about belonging during their first year, and that these worries fade with time. The students then wrote an essay describing how their own experiences mirrored this pattern, and recorded a video message for future students. The goal was to normalize adversity and reframe it as temporary and shared. The entire exercise typically took about one hour.

The original 2011 study, published in Science, reported that African American students who completed the exercise earned a grade point average roughly 0.25 points higher over the next three years compared to a control group. For a one-hour activity, that effect size was striking. The sample was modest—92 students at a selective university—but the result resonated widely. Schools across the United States adopted versions of the intervention.

Yet when other labs tried to replicate the finding, results were mixed. Some found smaller effects; others found none. A large-scale replication project, Many Labs 3, published in 2019, tested the intervention across 20 sites and more than 4,000 participants. The overall effect shrank to about 0.07 GPA points, and heterogeneity across sites was high. The intervention seemed to work in some contexts but not others, and the reasons were unclear.

Enter the attrition log. In 2022, a team led by Paul T. von Hippel at the University of Texas at Austin published a meta-analysis in Nature Human Behaviour that re-examined the original studies. They noticed that one of the contributing experiments—a replication conducted at a different university—had a striking pattern: 28% of control participants dropped out of the study, compared to only 12% in the treatment group. Differential attrition can bias results if the reasons for dropout are related to the outcome. If struggling students are more likely to drop out, and they are unevenly distributed across conditions, the apparent treatment effect may be inflated. The authors of the original work pushed back. In a correspondence published in PNAS, they argued that the attrition coding was flawed—some participants classified as dropouts had simply not completed follow-up surveys but still had GPA data available. The dispute highlighted a deeper issue: attrition logs are often incomplete, and the rules for handling missing data are rarely pre-registered or standardized.

The Original Effect Size and Its Contested Ground

The 0.25 GPA boost reported by Walton and Cohen in 2011 was large enough to attract attention—and skepticism. For a brief writing exercise to produce an effect comparable to reducing class size or increasing instructional time seemed implausible to some. The study's small sample size (92 students) meant that the confidence interval around the effect was wide, and the result was fragile: removing a few participants could change the conclusion.

Replication attempts often failed to match that magnitude. A 2014 study by Walton and colleagues at a different university found a smaller effect, and a 2016 replication at an elite private university found no significant difference. By 2019, the Many Labs 3 consortium had pooled data from over 4,000 students and found an average effect of roughly 0.07 GPA points—a third of the original estimate. The heterogeneity across sites was substantial, with some sites showing positive effects and others showing null or even negative trends.

Why the variation? One possibility is that the intervention works best in contexts where belonging concerns are acute—for instance, at highly selective institutions where minority students feel particularly isolated. Another is that the delivery method matters: some replications used online modules instead of in-person sessions, or different prompts. But a third possibility, raised by the attrition analysis, is that differential dropout artificially inflated the original estimates.

The debate over effect size is not merely academic. School administrators and policymakers have invested resources in belonging interventions based on the promise of large effects. If the true effect is smaller and more context-dependent, the cost-benefit calculus shifts. Understanding what drives the variation is essential for deciding where—and whether—to implement these programs.

Where the Attrition Log Entered the Argument

The 2022 meta-analysis by von Hippel and colleagues did not set out to target the belonging intervention specifically. The team was interested in the broader problem of attrition bias in randomized experiments. They compiled a dataset of 35 studies from the social-belonging literature and found that differential attrition was common: in several studies, dropout rates differed by more than 10 percentage points between conditions. When they applied statistical corrections for missing data, the overall effect estimate shrank.

One study stood out. A replication conducted at a large public university had a control-group dropout rate of 28% versus 12% in the treatment group. The authors of that study had not reported attrition by condition in their original paper; the discrepancy only emerged when von Hippel's team requested the raw data. The log showed that many control participants had stopped responding to surveys, but their GPA data—the primary outcome—was still available from university records. The question was whether to include those participants in the analysis.

The original authors argued that they should be included, because GPA data were available regardless of survey completion. Von Hippel's team countered that the missing survey data could indicate disengagement, which might correlate with academic performance. Without knowing why participants stopped responding, the safest approach was to treat the missing data as potentially informative and conduct sensitivity analyses.

The correspondence in PNAS in 2023 laid out the disagreement in detail. Walton and Cohen pointed out that the attrition classification used by von Hippel was not the same as the one used in the original studies, and that the effect remained significant under several alternative coding schemes. Von Hippel responded that the key issue was not the coding but the lack of transparency: without a pre-registered attrition plan, researchers could choose the coding that best supported their hypothesis.

How Researchers Disagree on Handling Dropouts

The attrition log controversy is a microcosm of a larger methodological debate. In randomized experiments, the gold standard is intent-to-treat (ITT) analysis, which includes all participants regardless of whether they completed the intervention. But ITT can be misleading if dropout is differential. Per-protocol analysis, which includes only those who completed the study, can also be biased if completers differ from dropouts.

Missing data can be classified into three types: missing completely at random (MCAR), missing at random (MAR), and missing not at random (MNAR). MCAR is rare in practice. MAR is the assumption behind most multiple imputation methods. MNAR is the hardest to handle, because the missingness itself is related to the outcome—for example, if struggling students are more likely to drop out. In the belonging intervention, if control participants who were struggling academically were more likely to stop responding, their absence could inflate the control group's average GPA, making the treatment effect appear larger.

Sensitivity analyses can test how robust results are to different assumptions about missing data. For instance, researchers can impute missing outcomes under various scenarios—assuming dropouts have worse outcomes, or better outcomes—and see if the conclusion changes. But such analyses are rarely reported in original papers. A 2020 survey of psychology experiments found that fewer than 10% reported any sensitivity analysis for attrition.

Pre-registration of attrition rules is also uncommon. Without a pre-specified plan, researchers may be tempted to choose the analysis that yields the most favorable result. The belonging intervention debate has prompted calls for journals to require attrition flow diagrams, similar to those used in clinical trials, and for authors to pre-register their handling of missing data.

What Subsequent Large-Scale Replications Found

The Many Labs 3 project, published in 2019 in Social Psychological and Personality Science, was one of the largest replication efforts in psychology. It included the social-belonging intervention as one of its target studies, with 20 participating labs and over 4,000 participants. The overall effect on GPA was small—about 0.07 points—and not statistically significant after correcting for multiple comparisons. But the heterogeneity was striking: some sites found effects as large as 0.3 GPA points, while others found negative effects.

Why such variation? The Many Labs 3 team examined several moderators, including campus climate, demographic composition, and delivery format. None consistently explained the differences. One possibility is that the intervention's effectiveness depends on the specific concerns of the student population, which may vary across institutions and over time. Another is that the implementation fidelity varied: some sites may have delivered the exercise with more enthusiasm or in a more supportive setting.

A subsequent meta-analysis by the original authors, published in 2020 in Educational Researcher, included 14 studies and found an average effect of about 0.11 GPA points for underrepresented minority students. That estimate is smaller than the original 0.25 but still meaningful. However, the meta-analysis included several studies from the original lab, which may have introduced bias. Independent replications tended to show smaller effects.

The open science collaboration that produced Many Labs 3 also highlighted the importance of sharing raw data. Without access to the original data, the attrition log would never have been discovered. The project's data are publicly available, allowing other researchers to conduct their own analyses and test alternative assumptions.

Methodological Lessons for Belonging Interventions

The belonging intervention saga offers several lessons for behavioral science. First, effect sizes from small studies should be interpreted cautiously. The original 2011 study had a sample of 92 students, which is sufficient to detect large effects but not small or moderate ones. The confidence interval around the 0.25 estimate was wide, and the result was sensitive to the inclusion or exclusion of a few participants.

Second, context matters. The intervention may work well at elite, predominantly white institutions where minority students feel particularly isolated, but less so at diverse or less selective schools. Researchers should measure and report contextual factors—such as campus climate, demographic composition, and baseline belonging—to help understand when the intervention is likely to be effective.

Third, attrition logs should be public and coded blind to condition. Researchers should report flow diagrams showing how many participants were randomized, how many completed each follow-up, and why they dropped out. Ideally, attrition coding should be done by someone unaware of the condition assignment, to avoid unconscious bias.

Fourth, pre-registration of analysis plans should include rules for handling missing data. Researchers should specify whether they will use ITT, per-protocol, or some other approach, and what sensitivity analyses they will conduct. This prevents post-hoc decisions that could favor the desired result.

Finally, replication efforts must share raw data. The attrition log controversy would not have emerged without open data. Journals and funders should mandate data sharing as a condition of publication, with appropriate privacy protections.

Practical Takeaways for Researchers and Reviewers

For researchers designing randomized experiments, the first step is to minimize attrition. This means keeping follow-up surveys short, offering incentives, and maintaining contact with participants. But even with best efforts, some dropout is inevitable. For instance, in a typical longitudinal study in education, attrition rates of 10–20% per wave are common, and differential attrition of 5–10 percentage points can bias estimates if not addressed. The key is to document it thoroughly.

Report a flow diagram with reasons for dropout, as recommended by the CONSORT statement for clinical trials. Include the number of participants randomized, the number who completed each assessment, and the number with missing outcome data. If possible, compare dropouts and completers on baseline characteristics to assess whether attrition is selective.

Conduct multiple imputation to handle missing data under the MAR assumption, and report sensitivity analyses under MNAR assumptions. For example, researchers can assume that dropouts have outcomes that are 0.2 standard deviations worse than completers, and see if the conclusion changes. If the result is robust, confidence increases.

Editors and reviewers should require attrition disclosure as a condition of publication. Many journals already do, but enforcement is uneven. A simple checklist at submission could help: Did the authors report attrition by condition? Did they conduct sensitivity analyses? Did they pre-register their attrition plan?

Replication efforts must share raw data, as the Many Labs 3 project did. This allows other researchers to re-analyze the data with different assumptions and check for errors. The attrition log that sparked the debate over the belonging intervention was not a scandal—it was a diagnostic tool. It revealed a vulnerability in the evidence base that could be addressed through better methods.

Looking Ahead: Toward Better Attrition Reporting

The debate over the social-belonging intervention has not ended, but it has catalyzed important changes. Several journals now require attrition flow diagrams as part of their submission guidelines. Pre-registration templates often include a section for handling missing data. And funding agencies are increasingly mandating data sharing for large-scale projects.

Yet challenges remain. Many researchers still do not report attrition by condition, and sensitivity analyses are rare. The incentives for novelty and positive results can discourage thorough reporting. To shift norms, journals could adopt a badge system for transparency, similar to the Open Science Framework badges for data sharing and pre-registration. Reviewers could be trained to check for attrition disclosure during the review process.

Another promising development is the use of automated tools to detect attrition patterns. For instance, the “attrition-check” R package allows researchers to quickly identify differential attrition in their own data and generate sensitivity analyses. Such tools lower the barrier to good practice.

Ultimately, the story of the social-belonging intervention is not one of fraud or incompetence. It is a story of how science progresses through disagreement and self-correction. The attrition log, once uncorrected, is now a case study in the importance of transparency. As one researcher put it, “The log is not the enemy; the enemy is the assumption that missing data doesn’t matter.” Moving forward, the field must ensure that every study includes a clear, pre-registered plan for handling dropouts—so that the next controversy is averted before it begins.

Recommend Posts
Science

One Uncosted Mirror Alignment Jig Fractured a Billion-Pixel Sky Survey

By Alice Chen/Jul 9, 2026

How a single uncosted mirror alignment jig degraded a billion-pixel sky survey, costing half its resolution and years of delay. A tale of fixed-price contracts and corner-cutting in big science.
Science

One Unreported Crystal Growth Flux Ratio Bent a Topological Superconductor Gap Map

By Alice Chen/Jul 9, 2026

A hidden variable in crystal growth—the flux ratio—was found to bend the superconducting gap map of Sr2RuO4, reshaping the phase diagram and prompting new reporting standards.
Science

How One Underpowered Nudge Replication Fractured a Cooperation Theory

By Jonas Eriksen/Jul 9, 2026

A landmark 2008 study on eye-like cues boosting cooperation failed to replicate in a massive multi-lab project. The fracture exposed deep methodological flaws and reshaped behavioral science.
Science

One Unreported Quartz Sample Etch Protocol Split a Luminescence Dating Standard

By Karim Osman/Jul 9, 2026

A hidden variation in quartz etching protocols caused 15–20% age offsets across luminescence dating labs. The discovery reshaped how geochronologists document sample preparation.
Science

One Unrecorded Atmospheric Seeing Monitor Drift Collapsed a Transiting Exoplanet Radius Measurement

By Renu Shah/Jul 9, 2026

A missing atmospheric seeing monitor inflated the radius of exoplanet WASP-76b by ~15%. This piece explores how such systematic errors creep into transit photometry and what the field is doing about it.
Science

One Undocumented Spectrograph Temperature Drift Split a Galactic Archeology Collaboration

By Alice Chen/Jul 9, 2026

A few millikelvin of thermal drift in a spectrograph fiber feed caused two subgroups to disagree on correction methods, delaying a galactic archeology catalog and splitting the collaboration.
Science

One Unreported Holographic Grating Polarization Bias Skewed a Dark Energy Survey Shear Calibration

By Renu Shah/Jul 9, 2026

A subtle polarization bias from the Dark Energy Survey's holographic grating introduced a 0.5–1% shear calibration error, mimicking an additive signal. New corrections reduce the bias below 0.1%, with lessons for LSST and Euclid.
Science

How a Fluid Dynamics Code Mapped Neural Activity Across a Mouse Visual Cortex

By Jonas Eriksen/Jul 9, 2026

A fluid dynamics code originally designed for pipe flow was repurposed to model neural activity in the mouse visual cortex, revealing traveling waves and feedback loops with 87% accuracy.
Science

How a Behavioral Nudge for Organ Donation Moved into Public Health Policy

By Jonas Eriksen/Jul 9, 2026

How a simple opt-out nudge for organ donation, rooted in behavioral science, moved from academic labs into public health policy worldwide, saving thousands of lives.
Science

How a Fluid Dynamics Code Solved a Solid-State Electron Flow Mystery

By Alice Chen/Jul 9, 2026

A fluid dynamics algorithm originally built for turbulence now simulates electron transport in quantum dots with 5% error, revealing vortices and interference patterns that classical models missed.
Science

One Unfrozen Atmospheric Reanalysis Grid Stretched a Decade of Storm Tracking

By Alice Chen/Jul 9, 2026

A subtle grid freeze in the ERA5 reanalysis led to systematic storm-count biases. Researchers found 11% fewer cyclones in one version, reversing trends and highlighting infrastructure fragility.
Science

One Uncaptured Laboratory Social Desirability Prompt Bent a Cooperation Game Replication

By Karim Osman/Jul 9, 2026

A single added sentence—'Please be honest'—may have inflated cooperation rates in a classic economic game replication from 50% to 80%, revealing how unnoticed wording changes can distort findings.
Science

One Unversioned Solver Tolerance Parameter Bent a Climate Model Ensemble

By Renu Shah/Jul 9, 2026

A single unrecorded solver tolerance parameter shifted a climate ensemble's spread by 10-15%. This methodology piece traces the root cause and what it means for reproducible science.
Science

One Unversioned Mesh Refinement Parameter Broke a Turbulence Simulation Replication

By Alice Chen/Jul 9, 2026

A 40% discrepancy in a turbulence simulation replication traced to a single unversioned mesh refinement parameter. The episode exposes gaps in computational reproducibility.
Science

One Missing Fringe-Phase Calibration Thread Bent a LIGO Noise Budget

By Jonas Eriksen/Jul 9, 2026

A single overlooked fringe-phase calibration thread bent LIGO's noise budget for two observing runs. The fix cost $2M and saved 15% of observing time, exposing deep flaws in how large-scale science funds noise debugging.
Science

One Uncorrected Attrition Log Split a Classic Social Belonging Intervention

By Alice Chen/Jul 9, 2026

How a single uncorrected attrition log fueled debate over a classic social belonging intervention, revealing deeper issues in handling dropouts in behavioral science.
Science

One Unreported Rat Chow Selenium Lot Shift Inflated a Thyroid Hormone Study

By Karim Osman/Jul 9, 2026

A mid-experiment selenium lot shift in rat chow inflated a thyroid hormone study. The retraction exposes a blind spot in model organism infrastructure and the economics of replication.
Science

One Unreported Rodent Light-Dark Cycle Shift Inflated a Fear Conditioning Meta-Analysis

By Karim Osman/Jul 9, 2026

A single lab's accidental reversal of the light-dark cycle during rodent fear conditioning experiments inflated effect sizes in a meta-analysis, raising questions about circadian confounds in preclinical neuroscience.
Science

One Misaligned fMRI Voxel Size Selection Fractured a Working Memory Localization Model

By Jonas Eriksen/Jul 9, 2026

How a seemingly trivial choice of fMRI voxel size—3 mm instead of 2 mm—obscured submillimeter functional columns in the prefrontal cortex, leading to a decade of conflicting results about working memory localization.
Science

One Unreported Reward Schedule Parameter Fractured a Dopamine Prediction Error Model

By Jonas Eriksen/Jul 9, 2026

How a single unreported parameter—whether reward probabilities were blocked or interleaved—fractured the canonical dopamine prediction error model, revealing hidden assumptions in decades of neuroscience research.