One Misaligned fMRI Voxel Size Selection Fractured a Working Memory Localization Model
In the early 2010s, cognitive neuroscience faced a reproducibility puzzle. Researchers studying working memory—the brain's ability to hold and manipulate information over seconds—kept finding that the dorsolateral prefrontal cortex (DLPFC) lit up in some studies but not others. Meta-analyses showed a scattered pattern, and the field started to wonder whether the DLPFC was truly central to working memory or merely an artifact of task design. The debate grew heated, with preprints and rebuttals flying between labs.
What eventually emerged was not a grand theoretical failure but a technical one: the size of the fMRI voxel. Standard acquisitions used 3-mm isotropic voxels, a choice inherited from earlier scanners and software defaults. But the functional columns in the prefrontal cortex—the basic computational units—are roughly 2 mm wide. The mismatch meant that many studies were averaging signals across adjacent columns with opposite tuning, effectively cancelling out the very activity they sought to measure. One misaligned parameter had fractured an entire localization model.
A Voxel-Size Decision That Broke a Cognitive Map
Functional magnetic resonance imaging (fMRI) measures blood-oxygen-level-dependent (BOLD) signals, which reflect neural activity indirectly. The spatial resolution of the technique is determined by the voxel—the three-dimensional volume element from which the signal is sampled. Choosing a voxel size involves a classic trade-off: smaller voxels give finer spatial detail but reduce signal-to-noise ratio (SNR) and require longer scan times; larger voxels boost SNR at the cost of blurring across neural boundaries.
In the early 2000s, most fMRI studies of higher cognition settled on 3-mm isotropic voxels as a pragmatic default. For example, a 2004 study by Owen and colleagues on working memory load used 3.5-mm voxels, reflecting the prevailing practice at the time. The rationale was straightforward: whole-brain coverage with decent temporal resolution was paramount, and the cognitive processes in question—attention, memory, decision-making—were thought to involve relatively large, distributed networks. The idea that submillimeter-scale organization might matter seemed, at the time, a concern for visual or motor cortex studies, not prefrontal ones.
But working memory is not a monolithic function. It depends on a cascade of subprocesses—encoding, maintenance, manipulation, retrieval—each of which may engage different subregions of the prefrontal cortex (PFC). The PFC itself is anatomically heterogeneous, with distinct cytoarchitectonic areas (Brodmann areas 9, 46, 45, and 44) that differ in columnar width and laminar structure. For example, area 46 in the DLPFC has columnar widths around 2 mm, while area 44 in the ventrolateral PFC is closer to 2.5 mm. A 3-mm voxel cannot resolve these differences. The result was that many activation maps from the 2000s misattributed activity. When a 3-mm voxel straddled two functional columns with opposite selectivity—say, one tuned to spatial location and another to object identity—the BOLD signals canceled out, creating a false negative. Conversely, when a voxel happened to align with a single column, the signal might be exaggerated. The cognitive map being drawn was systematically warped.
How a 3-mm Voxel Obscured a 2-mm Circuit
The seminal work of Joaquin Fuster in the 1970s had established that neurons in the PFC show sustained firing during the delay period of delayed-response tasks—a hallmark of working memory. These electrophysiological recordings, with millimeter precision, revealed that individual neurons in the DLPFC maintain information about a stimulus even after it is removed. But when fMRI attempted to localize these same processes, the spatial blur introduced by 3-mm voxels made it difficult to distinguish between adjacent functional zones.
Consider a classic delayed-response paradigm: a subject sees a cue (e.g., a location on a screen), waits several seconds, and then makes a response based on that cue. During the delay, DLPFC neurons fire selectively to the remembered location. In a 3-mm voxel, however, the signal from neurons tuned to the left location mixes with that from neurons tuned to the right location, if both fall within the same voxel. The net BOLD response may be indistinguishable from baseline, even though many neurons are active.
This partial-volume effect is well known in neuroimaging but was often dismissed as negligible for higher-order regions. Evidence from ultra-high-field (7T) studies later showed that functional columns in the PFC are as narrow as 2 mm, and that 3-mm voxels introduce a cross-column cancellation that reduces sensitivity by as much as 40–60% for certain contrasts. The problem was not just missing activation; it was creating spurious patterns where the noise from mixed signals happened to align with a task condition.
In 2015, a reanalysis of a widely cited DLPFC study using the original data but resampled to 2-mm voxels found that the previously reported activation peak shifted by nearly 1 cm—a distance that could mean the difference between area 46 and area 9. The original conclusion that the DLPFC encoded spatial working memory was not wrong, but the precise coordinates were unreliable. The localization model built on those coordinates became a house of cards.
The 2012–2017 Preprint Wars Over Dorsolateral PFC
Between 2012 and 2017, the literature on working memory localization became increasingly contentious. A landmark 2013 meta-analysis by Duncan and Owen argued that the DLPFC was not specifically involved in working memory but rather in a general "multiple-demand" network activated by any cognitively demanding task. Their analysis aggregated hundreds of fMRI studies and found that the DLPFC appeared in tasks ranging from arithmetic to inhibitory control. This challenged the prevailing view that working memory had a dedicated neural substrate.
Proponents of the DLPFC-specific model, notably researchers at Yale and the University of California, countered that the meta-analysis was biased by the inclusion of studies with coarse resolution. They pointed out that many of the included experiments used 3.5-mm or even 4-mm voxels, which would smear the signal across regions. A reanalysis restricted to studies with voxel sizes ≤ 2.5 mm showed a much tighter cluster in the mid-DLPFC, consistent with the classic Fuster model.
The Sternberg task—a working memory paradigm where subjects maintain a variable-length list of items—became a battleground. Several labs attempted to replicate a 2010 finding that DLPFC activity scaled with memory load. Replication attempts using 3-mm voxels failed to find the effect, while those using 2-mm or finer resolution succeeded. A 2016 paper by Curtis and D'Esposito reviewed these discrepancies and concluded that voxel size was the most likely moderator, though they stopped short of calling it the sole cause.
Meanwhile, the N-back task—a widely used working memory paradigm—showed similar sensitivity. A 2017 meta-analysis of N-back studies found that effect sizes in the DLPFC were significantly larger in studies that reported voxel sizes below 3 mm. The authors estimated that roughly 30% of the variance in activation strength could be explained by resolution alone. The field began to realize that the "preprint wars" were not about theory but about a hidden methodological moderator.
From Software Defaults to Systematic Bias
The choice of voxel size is not made in isolation. It interacts with other preprocessing steps, particularly spatial smoothing. Most fMRI analysis pipelines apply a Gaussian smoothing kernel—typically 6–8 mm full width at half maximum (FWHM)—to increase SNR and reduce inter-subject anatomical variability. But smoothing further blurs the signal, exacerbating the partial-volume problem. A 3-mm voxel smoothed with a 6-mm kernel effectively creates an effective resolution of around 7–8 mm, far too coarse to resolve 2-mm columns.
Software defaults in packages like SPM and AFNI have historically encouraged 3-mm voxels and large smoothing kernels. SPM's default normalization template uses 3-mm voxels, and its recommended smoothing kernel is 8 mm. These defaults were set in an era when computational constraints and scanner gradients limited resolution. They persist in many labs today, even as hardware has improved. The result is a systematic bias toward false negatives in small structures and false positives at cluster edges.
A landmark 2016 paper by Eklund and colleagues revealed that cluster-based inference—the standard method for correcting multiple comparisons in fMRI—produced false-positive rates as high as 60% under common software defaults. While Eklund focused on autocorrelation models, subsequent work showed that voxel size was a contributing factor: larger voxels reduce the effective number of independent tests, inflating cluster sizes and leading to spurious activations. The combination of 3-mm voxels, 8-mm smoothing, and cluster inference created a perfect storm for irreproducibility.
Some labs began to argue that the entire working memory localization literature needed to be re-evaluated with tighter methodological standards. A 2018 commentary in Nature Reviews Neuroscience called for pre-registration of voxel size and smoothing parameters, noting that "the choice of spatial resolution is not a mere technical detail but a theoretical commitment about the scale of neural organization." The field was slowly waking up to the fact that defaults are not neutral.
Reanalysis: Smaller Voxels, Cleaner Signal
The Human Connectome Project (HCP), launched in 2010, adopted a 2-mm isotropic acquisition for its resting-state and task-fMRI data. This decision was driven by the need to track white-matter tracts and cortical myelin maps, but it inadvertently provided a testbed for the voxel-size debate. Researchers reanalyzed working memory data from the HCP—which used a 2-back task—and found that the DLPFC activation was more focal and more consistent across subjects than in earlier 3-mm studies. The peak coordinates clustered tightly in area 46, with minimal spread into adjacent regions.
Multi-echo EPI sequences, which acquire multiple echoes at different TEs, further improved the signal by separating BOLD from non-BOLD noise. When combined with 2-mm voxels, these sequences achieved SNR comparable to 3-mm single-echo acquisitions while preserving spatial detail. A 2019 study using multi-echo EPI at 3T found that the DLPFC activation during a delayed-recognition task was significantly more robust than in standard single-echo 3-mm data, with a 30% increase in effect size at the group level.
Within-subject test-retest reliability also improved. In a 2020 study, participants performed the same working memory task on two separate days. The intraclass correlation coefficient (ICC) for DLPFC activation was 0.72 with 2-mm voxels, compared to 0.45 with 3-mm voxels. This suggested that the earlier failures to replicate were not due to unstable neural responses but to measurement noise introduced by coarse sampling. The smaller voxels effectively "cleaned up" the signal by reducing cross-column contamination.
However, the transition to higher resolution is not without costs. Smaller voxels reduce temporal resolution because more slices are needed for whole-brain coverage, and they increase the severity of motion artifacts. In the HCP, this was mitigated by advanced motion correction and longer scan times. For many labs, the trade-off remains prohibitive, especially for clinical populations where scanning time is limited. The challenge is to find the optimal resolution for a given question, not simply to minimize voxel size.
As of late 2024, several large-scale initiatives—including the UK Biobank and the Adolescent Brain Cognitive Development (ABCD) study—have adopted 2-mm or finer resolutions for their fMRI protocols. The shift reflects a growing consensus that the 3-mm default is inadequate for studying fine-scale functional organization. But the change is gradual, and legacy data sets continue to be analyzed with old parameters, perpetuating the biases.
Lessons for the Next Generation of Localization Studies
The first lesson is that spatial resolution must be matched to the expected scale of the neural circuitry under study. For working memory, where functional columns in the DLPFC are around 2 mm, a voxel size of 2 mm or smaller is necessary. For larger-scale networks like the default mode network, 3 mm may be sufficient. The key is to justify the choice based on prior anatomical knowledge, not on convenience or tradition.
Second, pre-registration of resolution parameters is essential. Many labs now pre-register their analysis plans, but the acquisition parameters are often left unspecified. A pre-registration statement should include voxel size, smoothing kernel, and the rationale for these choices. This allows reviewers and readers to assess whether the resolution is adequate for the claims being made. Without such transparency, the field risks repeating the same mistake.
Third, the use of ultra-high-field (7T) MRI can provide additional benefits. At 7T, the increased SNR allows for submillimeter voxel sizes (e.g., 0.8 mm isotropic) that can resolve cortical columns directly. Several 7T studies of working memory have confirmed the columnar organization of the DLPFC and shown that activation patterns are more specific than at 3T. However, 7T is not widely available, and its adoption is limited by cost and technical challenges.
Finally, effect sizes should be reported per voxel class—that is, separately for voxels in the core of a region versus those at the boundary. This can reveal whether the activation is driven by a few well-sampled voxels or by a diffuse cloud. A 2021 study introduced a "resolution-adjusted effect size" metric that normalizes activation strength by the effective number of independent voxels in a region. Such metrics can help compare studies with different voxel sizes and provide a more honest assessment of the evidence.
These lessons are not limited to working memory. Similar issues have been documented in studies of the visual cortex, where 3-mm voxels obscure orientation columns, and in the hippocampus, where subfield boundaries are blurred. The broader implication is that neuroimaging is not a "one-size-fits-all" technology. The choice of voxel size is a theoretical commitment that should be made explicit and testable.
What One Parameter Taught Us About Consensus
The story of the misaligned voxel size is a case study in how scientific consensus can be shaped by hidden methodological choices. For nearly a decade, the working memory field was divided between those who believed the DLPFC was central and those who saw it as part of a general-purpose network. The division was not resolved by more data or better theory but by a technical parameter that had been treated as trivial.
The process of self-correction was slow and painful. It required reanalyses, meta-analyses, and a willingness to admit that earlier results were contingent on arbitrary defaults. The field now demands greater transparency in reporting acquisition parameters, and many journals require authors to justify their voxel size and smoothing choices. This is a step forward, but it remains an uphill battle against established practices and resource constraints.
One can draw a parallel to other methodological crises in science. In psychology, the replication crisis led to pre-registration and open data. In neuroimaging, the voxel-size crisis has led to a greater appreciation of spatial resolution as a core experimental variable. But unlike psychological priming studies, which were sometimes simply false, the working memory localization model was not wrong—it was misplaced. The truth was there, but at the wrong scale.
Yet significant challenges remain. Despite growing awareness, many fMRI studies still use 3-mm voxels as a default, and the cost of higher-resolution acquisitions limits their adoption in resource-constrained settings. Moreover, the field lacks a standardized framework for correcting voxel-size differences across studies, making meta-analyses difficult. Without such tools, the risk of repeating the same biases persists. The voxel-size crisis is not fully resolved; it is an ongoing process of methodological refinement. The takeaway is that no choice in experimental design is neutral. Every parameter, from the repetition time to the smoothing kernel, embeds assumptions about the nature of the signal. The scientific community is learning to treat these assumptions as hypotheses to be tested, not as defaults to be inherited. The voxel-size story is a reminder that progress often comes not from grand discoveries but from paying attention to the mundane details that shape what we can and cannot see.