A new study has produced a sobering finding for anyone who thinks brain-scan evidence speaks for itself: the same person, scanned on two different MRI machines within a week, can produce results measuring their brain’s internal wiring that look substantially different depending on which scanner did the scanning. The study, Cross-vendor reliability of functional and structural brain connectivity in a travelling cohort, was published in the journal Scientific Reports in 2026 by researchers at Ruhr University Bochum, led by Lionel Butry and Dr. Lara Schlaffke.
The researchers scanned ten healthy adults on both a Siemens Prisma 3T scanner and a Philips Achieva 3.0T scanner, roughly five days apart on average, using standard clinical protocols on each machine. They then compared two kinds of brain “connectome” maps that increasingly show up in neuroscience research and, at times, in litigation: functional connectivity, which measures how synchronized activity is between different brain regions during rest, and structural connectivity, which maps the physical white-matter wiring connecting those regions.
The results were not reassuring, particularly for functional connectivity. At the level of an individual person, agreement between the two scanners for functional connectivity was rated “poor” using the standard reliability scale researchers use for this kind of measurement, with an average reliability score of just 0.22 on a 0-to-1 scale where anything under 0.5 counts as poor. Structural connectivity fared better but still landed only in the “fair” range, around 0.43. Statistical modeling showed the choice of scanner alone accounted for roughly 12% of the variation researchers measured in functional connectivity and 8% in structural connectivity — more than double the scanner-driven variability an earlier, larger eight-site study had found using two different scanner brands. The researchers also tried applying a widely used statistical correction technique called neuroComBat, originally developed for genetic data and later adapted for brain imaging, to see if it could fix the problem. It worked well for comparing groups of people against each other, essentially erasing scanner-driven differences down to about 1%, but did almost nothing to fix the reliability problem for any individual person’s own scan — the exact level at which a scan would typically be used to evaluate one specific patient or claimant.
The Bochum team frames this as a caution for the wider field, particularly for the growing number of studies that pool brain scans from multiple research sites or hospitals, or that switch scanner equipment partway through a long-running study. But the practical stakes go beyond academic research. Functional and structural connectivity measures, including diffusion tensor imaging (a technique closely related to the structural connectivity method used in this study), have been offered as evidence in personal injury and disability litigation, particularly in cases involving mild traumatic brain injury, where plaintiffs sometimes point to subtle white-matter or connectivity abnormalities as objective proof of injury even when a standard CT or structural MRI looks normal.
That use has already generated real controversy in the legal literature. A widely cited analysis in the journal World Journal of Clinical Cases, Diffusion tensor imaging in the courtroom: Distinction between scientific specificity and legally admissible evidence, describes an ongoing conflict between how sensitive these imaging techniques are in a research setting and the legal standard courts are supposed to apply before admitting scientific evidence, warning that attorneys and juries without technical training can be poorly positioned to evaluate the strength of DTI-based claims. Other legal scholars have raised similar concerns specifically about functional MRI, cautioning that courts have historically been far more willing to admit conventional structural imaging (proof that a physical injury exists) than functional imaging purporting to show how a person’s brain is working or even whether they are being truthful, precisely because of open questions about reliability and what the underlying signals actually mean.
The Bochum study’s findings also echo a much larger reckoning that has been underway in neuroscience since 2022, when a team led by researchers at Washington University in St. Louis published a widely discussed paper in Nature, Reproducible brain-wide association studies require thousands of individuals. That study, using data from nearly 50,000 participants, found that most published brain-behavior studies had been conducted with far too few subjects to produce reliable results, and that typical sample sizes of a few dozen people were prone to turning up statistically significant but ultimately spurious associations. That paper reshaped funding and design expectations across the field and is now itself part of an ongoing debate, since more recent Nature-published research has argued that study design choices, not just raw sample size, can meaningfully improve reproducibility even in smaller studies.
For the Bochum researchers, the practical message is narrower but pointed: a single brain scan from a single scanner, evaluated for a single individual, carries real uncertainty that group-level statistical fixes cannot erase. The study’s authors recommend that any research or clinical program facing a scanner change collect “bridging” data on both machines before and after the switch, and that any given patient’s edge-level connectivity results, especially involving subcortical or limbic brain regions where the study found the weakest reliability, be interpreted with real caution rather than as a precise, stable measurement.
