On September 18, the National Institute of Standards and Technology and the National Institute of Biomedical Imaging and Bioengineering opened a public inquiry into the measurements behind medical imaging, devices, diagnostics, and therapy. A symposium follows on September 24, and public comments remain open through November 30. The agencies want to know which measurements and standards should be improved as scanners, software, and artificial intelligence change. The inquiry will inform a future roadmap. No imaging standard or clinical guideline changed when the notice appeared.
The timing matters for anyone comparing two scans. A report may say that a tumor shrank, a nodule grew, or a brain structure changed. That difference may be medically important. It may also depend partly on whether the two images were made and analyzed in the same way.
The number has a technical history
Metrology is the science of measurement. A reliable measurement should remain meaningful when another trained person repeats it with another calibrated instrument. In medical imaging, that goal becomes difficult because the result can depend on the scanner, acquisition settings, contrast method, reconstruction software, and analysis tool. Quantitative imaging tries to turn pictures into reproducible measurements. Instead of only describing a finding as mild, small, or subtle, a report may give a tumor diameter, brain volume, or tissue value. A number looks ready for comparison. The comparison works only if the measurement process remained compatible.
A NIST-led study published in 2021 brought together researchers from 11 institutions. They tested 27 MRI scanners from three vendors at nine clinical sites, using a calibrated object filled with materials that mimic tissue so the machines could be checked against the same reference. They measured one property called T1 relaxation time. The measurements showed significant bias and variation across scanners, without a consistent pattern by vendor. The researchers concluded that a diagnostic threshold developed on one MRI system could not simply be transferred to another system.
Katy Keenan, the NIST study leader, explained the distinction: "The pixel values themselves cannot be compared to pixel values in other datasets. It is not easy to compare T1-weighted data across datasets. In the case of carefully acquired T1, comparison of the pixel values is possible, because they have a quantitative meaning."
Her last sentence carries the practical point. Quantitative values can support comparison when teams acquire and check them carefully. A number alone does not prove that the underlying method stayed stable.
The study examined one MRI property using a tissue substitute rather than patients. It provides evidence of a measurement problem, but no general error rate for MRI diagnoses, CT, PET, or ultrasound.
Software can help and can add another variable
A 2025 study in Radiology: Artificial Intelligence examined brain-volume data from 20,864 participants, 11 research groups, 43 MRI scanners, and three continents. Researchers tested a machine-learning method designed to reduce scanner-related differences while preserving the biological differences associated with Alzheimer's disease.
The method performed better than the two comparison approaches, including when it encountered scanners outside its development data. That result suggests software can make measurements from different machines more comparable.
It also shows how hard the separation remains. Neuroradiologist Sven Haller described the clinical dilemma in an editorial accompanying the study: "When there are subtle changes such as subtle atrophy at early stages of neurodegenerative diseases like dementia, the question invariably arises of whether such changes are real or biased due to scanner differences."
A harmonization tool must remove the machine's influence without erasing the disease signal. Haller said there was no simple solution. Study co-leader Damiano Archetti said current methods cannot completely separate scanner-related variation from disease-related variation. The research concerned brain measurements in the Alzheimer's spectrum. It did not validate every harmonization tool or disease use.
Updates add another layer. FDA guidance issued in August 2025 says a manufacturer seeking approval for planned changes to an AI-enabled medical device should describe the modifications, explain how they will be developed and validated, and assess their impact. FDA can review that plan as part of a marketing submission so covered updates do not each require a separate submission. The guidance applies to manufacturers and regulators. It does not give a patient a direct right to an internal software history. It does show that an AI tool's update history can matter to safety and effectiveness. For a measurement used across months or years, the software version belongs in the evidence chain alongside the scanner and protocol.
What the evidence cannot tell us
These sources cannot determine whether any individual's serial scans are comparable. They do not show that a machine difference explains a specific change in a report. They also do not show that a radiologist missed a technical difference or that a patient should question a diagnosis without qualified help. The NIST study covered one quantitative MRI measurement. The later harmonization study covered brain volumetry in a defined research setting. FDA's guidance concerns planned changes to regulated AI-enabled device software. The new NIST inquiry is gathering needs for a roadmap that does not exist yet.
The evidence supports a narrower conclusion: technical changes can affect imaging measurements, and experts are still building better ways to preserve comparability across those changes.
Ask what changed between the scans
If an important measurement moved between two imaging reports, ask the radiologist or ordering clinician whether the scanner, acquisition protocol, contrast method, measurement software, or AI analysis changed between the studies. Then ask whether any difference affects the comparison.
That question leaves interpretation with the clinical team while making the measurement chain visible. A change in the number may reflect a change in the body. Before acting on the difference, it is reasonable to ask whether the measuring system changed too.
This is general orientation about medical measurements, not medical advice. Do not delay care or change treatment based on this article.
Public sources
- NIST/NIBIB medical metrology RFI, Federal Register, September 18, 2026
- Official Federal Register PDF, 91 FR 59111
- NIST summary of the multi-site quantitative MRI study
- Primary PLOS One study
- RSNA specialist report on MRI harmonization
- Primary Radiology: Artificial Intelligence study
- FDA guidance on change-control plans for AI-enabled devices
- Review: Reproducibility and quality assurance in MRI
