Abstract
INTRODUCTION
Centiloid (CL) scaling standardizes amyloid positron emission tomography (PET) quantification across tracers and platforms; however, variability across software implementations may affect diagnostic classification. This study evaluated inter‐software variability and diagnostic performance across five platforms using identical 1 8F‐florbetapir datasets.
METHODS
Retrospectively, 192 patients undergoing 1 8F‐florbetapir PET/computed tomography (CT) and magnetic resonance imaging (MRI) were analyzed. CL values were generated using four US Food and Drug Administration (FDA) ‐cleared platforms and an in‐house Centiloid standard pipeline. Agreement was assessed using intraclass correlation coefficient, with bias and limits of agreement evaluated by linear modeling and Bland–Altman analysis. Diagnostic performance was assessed using receiver operating characteristic (ROC) analysis and classification against visual interpretation.
RESULTS
Agreement was excellent (intraclass correlation coefficient [ICC] 0.969; 95%CI 0.961–0.975). Some platforms produced systematically higher CL values versus others slightly lower. LoA reached ± 30 CL. Diagnostic accuracy was high (area under the curve [AUC]: 0.949–0.975), with sensitivity 0.923–0.968 and specificity 0.667–0.806.
DISCUSSION
Despite excellent agreement, systematic differences persist and may affect classification near thresholds, supporting consistent use of a single processing pipeline.
Keywords: 18F‐florbetapir, Alzheimer's disease, Centiloid quantification, diagnostic performance, inter‐software variability
Highlights
Clinically relevant differences in Centiloid (CL) values may persist across US Food and Drug Administration (FDA) ‐approved commercial software platforms, which commonly use postitron emission tomography (PET) ‐based rather than MRI‐based spatial normalization.
CL quantification was compared across four FDA‐approved commercial platforms and one in‐house CL Standard reference, evaluating both PET‐based and MRI‐based approaches.
Despite excellent overall agreement in 18F‐florbetapir PET quantification, large inter‐platform CL differences in individual cases can occur.
Consistent use of the same software platform is recommended for longitudinal CL assessment to minimize non‐biological variability and avoid assumptions of cross‐platform interchangeability
1. BACKGROUND
Amyloid positron emission tomography (PET) plays a central role in the in vivo detection and quantification of cerebral β‐amyloid deposition in Alzheimer's disease (AD). 1 , 2 , 3 , 4 , 5 Quantitative assessment of amyloid burden is increasingly important for diagnosis, longitudinal monitoring, and determining eligibility for disease‐modifying therapies.
Early quantitative amyloid PET predominantly relied on standardized uptake value ratios (SUVRs), which are sensitive to tracer type, reference region, image preprocessing, and software implementation, limiting comparability across studies and centers. 6 , 7 , 8 , 9 , 10 The Centiloid (CL) Project was therefore introduced to standardize amyloid PET quantification using a common scale anchored to 0 CL for young cognitively normal controls and 100 CL for typical AD levels, originally defined using 11C‐Pittsburgh Compound B (PiB) PET. 8 Tracer‐specific calibration equations subsequently extended the CL framework to US Food and Drug Administration (FDA) ‐approved 18F tracers, including 18F‐florbetapir. 10 , 11 , 12 , 13 , 14
Despite this standardization, clinically relevant inter‐software differences may persist because of variations in spatial normalization, brain templates, region of interest (ROI) definition, and implementation details, 6 potentially influencing threshold‐based clinical interpretation and decision‐making.
Only limited studies have directly compared CL values generated by FDA‐cleared software platforms using 1 8F‐florbetapir PET datasets. 15 Most commercial implementations rely on PET‐based templates for spatial normalization, whereas magnetic resonance imaging (MRI) ‐based approaches, although potentially more anatomically accurate, are less widely available and less frequently evaluated. Previous studies have demonstrated excellent overall agreement across commercial and research software, while highlighting persistent platform‐dependent variability and the need for tracer‐specific validation. 10 , 11 , 15
This study presents a comprehensive head‐to‐head comparison of four FDA‐cleared commercial software platforms and an in‐house implementation of the CL standard reference pipeline using the same cohort of 1 8F‐florbetapir PET and MRI studies. We aimed to quantify inter‐software variability, assess agreement and systematic bias, and provide practical guidance for interpreting CL values across commonly used clinical platforms.
2. METHODS
2.1. Subjects and imaging data
This retrospective study received approval from the Mayo Clinic Institutional Review Board (IRB). Imaging data from clinical patients evaluated for AD were included, consisting of 1 8F‐florbetapir brain PET/CT and corresponding high‐resolution diagnostic MRI examinations. PET/CT imaging was performed using a 10‐minute single‐bed acquisition at 51 ± 3 minutes following administration of 372.6 ± 20.4 mCi of 1 8F‐florbetapir. MRI was acquired using a sagittal T1‐weighted MPRAGE sequence with a voxel size of 1 × 1 × 1.2 mm on 3T systems. All DICOM images were anonymized and de‐identified in compliance with institutional policies and applicable regulations governing human subjects research. Studies demonstrating significant head motion artifacts were excluded.
2.2. CL quantification
RESEARCH IN CONTEXT
Systematic review: The Centiloid framework standardizes amyloid positron emission tomography (PET) quantification across tracers and centers. Previous studies have validated tracer‐specific conversions and demonstrated overall consistency, largely using PET‐based normalization, whereas magnetic resonance imaging (MRI) ‐based approaches are less widely available in commercial software and have been evaluated less extensively. Most studies have assessed single pipelines or a limited number of platforms, with relatively few direct comparisons among widely used clinical software packages.
Interpretation: This study provides a head‐to‐head comparison of five platforms using identical datasets and shows that, despite excellent overall agreement, systematic differences persist across both PET‐ and MRI‐based approaches. These differences may lead to discordant classification near clinical thresholds, with implications for diagnosis, treatment eligibility, and trial enrollment in AD.
Future directions: Consistent use of a single processing pipeline is important for reliable clinical interpretation and longitudinal assessment. Future studies should further investigate harmonization strategies and multicenter validation to improve robustness and clinical utility.
A typical workflow of CL analysis with key processing steps is shown in Figure 1. CL quantification begins with a standardized brain template containing predefined ROIs, which may be PET‐ or MRI‐based depending on the software platform (Step 1). Patient PET images are co‐registered to the corresponding MRI when available (Step 2), and the individual PET (and MRI, when applicable) images are spatially normalized to the template space (Step 3). Standard cortical target and reference ROIs are then applied to calculate a global SUVR (Step 4), which is converted to a CL value using the tracer‐specific calibration equation (Step 5). The standard target region comprises six cortical regions (frontal, parietal, temporal, anterior cingulate, posterior cingulate, and precuneus), with the whole cerebellum used as the reference region.
FIGURE 1.

A schematic overview of the Centiloid analysis workflow and key processing steps.
2.3. Software platforms for CL quantification
CL quantification was performed using five image processing pipelines comprising four FDA‐cleared commercial software platforms and one in‐house implementation of the CL Standard reference pipeline. All 1 8F‐florbetapir PET and MRI datasets were processed independently according to the manufacturer‐recommended or in‐house validated procedures.
The evaluated software platforms (all using the whole cerebellum as the reference region) include:
MIMneuro (MIM Software, GE Healthcare): PET‐only workflow using a standardized PET template.
syngo.PET (syngo.PET Cortical Analysis, Siemens Healthineers): PET‐only workflow using a standardized PET template.
Combinostics PET‐only (Combinostics cPET, SyntheticMR): PET‐only workflow using a standardized PET template.
Combinostics PET/MRI (Combinostics cPET, SyntheticMR): Hybrid PET/MRI workflow using the subject's MRI, spatially normalized to a standardized MRI template.
Centiloid Standard (Mayo Clinic, Rochester): In‐house implementation of the published CL reference pipeline, 8 , 14 using a hybrid PET/MRI workflow with the subject's MRI, spatially normalized to a standardized MRI template.
The CL Standard pipeline replicated the published CL methodology using the publicly available GAAIN PiB datasets for Level‐1 validation before applying the published 18F‐florbetapir calibration equation. 12 The implementation used SPM8 for PET‐MRI co‐registration, MRI spatial normalization, and the standard CL volumes of interest (VOIs).
2.4. Statistical analysis
Patient demographics were summarized as mean ± standard deviation for continuous variables and total with percentage for categorical variables. Overall agreement of CL values across software platforms was assessed using the intraclass correlation coefficient (ICC3) via the icc function in the irr package in R, treating the five software platforms as fixed effects.
Systematic bias relative to the reference standard was quantified as the mean differences using a linear mixed‐effects regression model (lmer, lme4 package), with subject as a random effect to account for within‐subject correlation. Proportional bias was assessed using mixed‐effects regression of pairwise differences (Software−Standard) against the centered pairwise mean.
Pairwise agreement with the reference standard was further evaluated using Bland–Altman analysis and 95% limits of agreement (LoA), together with ICC3.
Mean differences (Software−Standard) from the linear mixed‐effects models were reported as the primary estimate of systematic bias. Within‐subject variability was quantified using the median absolute deviation (MAD; median(|x − median(x)|)) of the subject‐level residuals. The MAD was multiplied by 1.4826 to estimate the robust residual standard deviation.
Diagnostic performance relative to the binary gold standard was assessed using receiver operating characteristic (ROC) analysis and summarized by the area under the curve (AUC). Sensitivity and specificity were calculated using a CL threshold of ≥ 25. As an exploratory analysis, the optimal diagnostic threshold for each software was determined by maximizing the Youden index from ROC analysis, and differences relative to the CL Standard were compared descriptively with the corresponding Bland–Altman mean differences to assess whether inter‐software differences were consistent with a constant calibration offset.
p‐Values < 0.05 were considered statistically significant. All analyses were performed using R version 4.4.1 (R Foundation for Statistical Computing, Vienna, Austria).
2.5. Reference framework
Two reference frameworks were used in this study.
Reference standard: Quantitatively, the in‐house CL Standard pipeline served as the reference standard for overall CL comparison.
Gold standard: Clinically, the interpretation of amyloid PET, obtained from the electronic medical record, served as the diagnostic gold standard. PET studies had previously been reviewed by a consensus group of experienced neuroradiologists and nuclear medicine radiologists (including B.J.B. and D.R.J.) with subspecialty expertise in AD imaging as part of a local AD Treatment Center conference to verify clinical interpretation.
3. RESULTS
A total of 192 patients (72.0 ± 7.3 years; 56.2% female) were included in the final analysis after quality control and exclusion of two cases that failed the in‐house implemented CL Standard pipeline. Mean ± SD CL values obtained from MIMneuro, syngo.PET, Combinostics PET‐only, Combinostics PET/MRI, and the CL Standard were 69.55 ± 45.66 (range: −26.4 to 177.4), 76.16 ± 48.11 (range: −41.0 to 187.1), 75.95 ± 48.39 (range: −26.0 to 186.8), 76.70 ± 48.40 (range: −26.8 to 203.1), and 70.96 ± 46.88 (range: −34.7 to 189.2), respectively.
3.1. Overall agreement across software platforms
Overall agreement across the five software platforms was high, with an ICC of 0.969 (95% confidence interval [CI]: 0.961–0.975), indicating excellent consistency across methods.
A linear mixed‐effects model with software platform as a fixed effect and subject as a random effect was fitted using the CL Standard as the reference standard.
3.2. Pairwise agreement with reference standard
Table 1 summarizes pairwise agreement between each commercial software platform and the CL Standard, including ICC, mean difference, absolute difference, median absolute deviation, and the difference in Youden‐optimal threshold.
TABLE 1.
Pairwise agreement between commercial software platforms and the Centiloid Standard, summarized by ICC, MD, AD, MAD, and Δ in Youden‐optimal threshold.
| Centiloid Standard | |||||
|---|---|---|---|---|---|
| Commercial software |
ICC (95% CI) |
MD (95% LoA) (CL) |
AD, mean ± SD (range) (CL) |
MAD / robust SD (robust LoA) (CL) |
Δ Youden‐optimal threshold (CL) |
| MIMneuro |
0.958 (0.944–0.969) |
−1.410 (−27.848–25.028) |
9.599 ± 9.556 (0.149–63.208) |
0.909 / 1.348 (−2.518–2.764) |
0.31 |
| syngo.PET |
0.972 (0.963–0.979) |
5.195 (−16.730–27.120) |
8.993 ± 8.424 (0.098–46.074) |
2.538 / 3.763 (−7.675–7.077) |
12.65 |
| Combinostics PET‐only |
0.975 (0.967–0.981) |
4.983 (−15.977–25.943) |
8.881 ± 7.747 (0.145–44.674) |
2.316 / 3.434 (−6.798–6.664) |
4.25 |
| Combinostics PET/MRI |
0.985 (0.980–0.988) |
5.733 (−10.662–22.128) |
7.902 ± 6.343 (0.147–45.474) |
0.701 / 1.040 (−1.993–2.084) |
−0.95 |
Abbreviations: Δ, difference; AD, absolute difference; CI, confidence interval; CL, Centriloid; ICC, intraclass correlation coefficient; LOA, limits of agreement; MAD, mean absolute deviation; MD, mean difference; PET, positron emission tomography; SD, standard deviation.
Pairwise ICCs for MIMneuro, syngo.PET, Combinostics PET‐only, and Combinostics PET/MRI were 0.958 (95% CI: 0.944–0.969), 0.972 (95% CI: 0.963–0.979), 0.975 (95% CI: 0.967–0.981), and 0.985 (95% CI: 0.980–0.988), respectively. syngo.PET, Combinostics PET‐only, and Combinostics PET/MRI produced systematically higher CL values by approximately 5–6 CL, whereas MIMneuro produced CL values that were, on average, approximately 1 CL lower. The mean absolute differences (mean ± SD) were 9.60 ± 9.56 CL, 8.99 ± 8.42 CL, 8.88 ± 7.75 CL, and 7.90 ± 6.34 CL, respectively, with corresponding ranges of 0.15–63.21 CL, 0.10–46.07 CL, 0.15–44.67 CL, and 0.15–45.47 CL.
Bland–Altman plots are shown in Figure 2. The analysis demonstrated mean differences (bias) of −1.41 CL (95% LoA: −27.85 to 25.03 CL), 5.19 CL (95% LoA: −16.73 to 27.12 CL), 4.98 CL (95% LoA: −15.98 to 25.94 CL), and 5.73 CL (95% LoA: −10.66 to 22.13 CL) for MIMneuro, syngo.PET, Combinostics PET‐only, and Combinostics PET/MRI, respectively, with 95% of the differences lying between the LoA ranges. No evidence of proportional bias was observed (slope estimates near zero with 95% confidence intervals excluding clinically meaningful values).
FIGURE 2.

Bland–Altman plots showing pairwise agreement in Centiloid (CL) values between each commercial software platform and the Centiloid Standard. The solid blue line represents the mean bias, and the dashed red lines represent the 95% limits of agreement.
The robust residual standard deviation estimated from the median absolute deviation was small (approximately 1–3 CL), corresponding to robust residual LoAs of approximately ± 2 to ± 7 CL.
Compared with the CL Standard, the differences in the Youden‐optimal thresholds were 0.31 CL for MIMneuro, 12.65 CL for syngo.PET, 4.25 CL for Combinostics PET‐only, and −0.95 CL for Combinostics PET/MRI.
3.3. Diagnostic performance relative to visual interpretation
Among the 192 patients, clinical visual interpretation (the gold standard) classified 156 as amyloid‐positive and 36 as amyloid‐negative. Using a CL threshold of ≥ 25, MIMneuro, syngo.PET, Combinostics PET‐only, Combinostics PET/MRI, and the CL Standard classified 155, 160, 157, 160, and 151 patients as amyloid‐positive, respectively.
Across the cohort, discordant amyloid classification relative to visual interpretation occurred in approximately 1–5 patients (0.5%–2.6%), depending on the software platform, primarily in cases with CL values near diagnostic thresholds.
All platforms demonstrated high sensitivity (0.923–0.968), with Combinostics PET/MRI showing the highest sensitivity (0.968; 95% CI, 0.927–0.990). Specificity was lower and more variable (0.667–0.806), with the CL Standard showing the highest specificity (0.806; 95% CI, 0.640–0.918). Detailed values for all platforms are shown in Figure 3.
FIGURE 3.

Receiver operating characteristic (ROC) curves demonstrating the diagnostic performance of MIMneuro (A), syngo.PET (B), Combinostics PET‐only (C), Combinostics PET/MRI (D), and the Centiloid Standard (E), using clinical binary visual interpretation as the gold standard.
Figure 3 shows the ROC curve comparison. All software platforms indicated excellent diagnostic performance, with AUC values of 0.958, 0.949, 0.949, 0.975 and 0.964 for MIMneuro, syngo.PET, Combinostics PET‐only, Combinostics PET/MRI, and the CL Standard, respectively. The highest AUC was observed for Combinostics PET/MRI (95% CI: 0.956–0.994).
Table 2 summarizes individual cases with large inter‐platform discrepancies, defined as an absolute difference of approximately 30 CL or greater from the reference standard for at least one software platform. Among these cases, all but one were visually amyloid‐positive, indicating that large absolute inter‐platform discrepancies were observed predominantly in scans with high amyloid burden.
TABLE 2.
Individual cases of which at least one software platform showing an absolute bias of approximately 30 CL or more from the reference standard.
| Platform A | Platform B | Platform C | Platform D | Platform E | Clinical visual (P/N) | |||||
|---|---|---|---|---|---|---|---|---|---|---|
| Index | CL | Absolute bias (CL) | CL | Absolute bias (CL) | CL | Absolute bias (CL) | CL | Absolute bias (CL) | Reference standard | Gold standard |
| 1 | 78.0 | 29 | 65.3 | 17 | 63.7 | 15 | 62.1 | 13 | 48.7 | P |
| 2 | 104.9 | 8 | 132.2 | 36 | 138.5 | 42 | 109.3 | 13 | 96.5 | P |
| 3 | 67.5 | 25 | 85.9 | 43 | 75.7 | 33 | 28.5 | 14 | 42.4 | N |
| 4 | 132.7 | 30 | 144.2 | 19 | 167.4 | 5 | 171.4 | 9 | 162.8 | P |
| 5 | 137.9 | 34 | 128.8 | 25 | 129.4 | 25 | 122.8 | 19 | 104.1 | P |
| 6 | 86.0 | 33 | 99.6 | 20 | 133.2 | 14 | 127.6 | 8 | 119.4 | P |
| 7 | 177.4 | 63 | 121.9 | 8 | 129.2 | 15 | 121.2 | 7 | 114.2 | P |
| 8 | 54.0 | 25 | 46.5 | 18 | 58.1 | 29 | 55.1 | 26 | 28.6 | P |
| 9 | 49.2 | 31 | 89.3 | 9 | 77.3 | 3 | 88.0 | 8 | 80.3 | P |
| 10 | 97.0 | 14 | 127.1 | 44 | 110.2 | 27 | 90.8 | 8 | 82.9 | P |
| 11 | 158.0 | 35 | 137.4 | 14 | 129.9 | 7 | 128.1 | 5 | 122.9 | P |
| 12 | 38.4 | 41 | 85.9 | 7 | 93.3 | 14 | 92.7 | 14 | 79.1 | P |
| 13 | 106.0 | 50 | 109.9 | 46 | 111.3 | 45 | 110.5 | 45 | 156.0 | P |
| Mean | 99.0 | 105.7 | 109.0 | 100.6 | 95.2 | |||||
| Std | 43.0 | 29.8 | 33.1 | 37.3 | 40.9 | |||||
Abbreviations: CL, Centiloid; P/N, positive/negative.
4. DISCUSSION
In this study, we performed a comprehensive head‐to‐head evaluation of CL quantification across multiple FDA‐cleared commercial software platforms using the same 1 8F‐florbetapir PET/CT and MRI datasets. The analysis incorporated an in‐house implementation of the CL Standard reference pipeline as the quantitative reference and clinical visual interpretation as the gold standard for diagnostic performance analysis. Our results demonstrate that, while overall agreement across platforms is excellent, substantial platform‐dependent CL value differences may occur in individual cases, with important implications for threshold‐based clinical interpretation. These findings highlight that, despite nominal standardization under the CL framework, quantitative outputs are not fully interchangeable across software platforms in clinical practice. This has direct implications for patient‐level decision‐making, particularly in the context of emerging disease‐modifying therapies that rely on specific CL thresholds for treatment eligibility.
Commercial software platforms implement the CL framework using different proprietary methodologies. Although all platforms calibrate their measurements according to the published CL methodology, 8 , 16 they differ in template selection, spatial normalization, ROI definition, and image processing strategies. Consequently, inter‐platform variability is likely introduced primarily during SUVR calculation and ROI segmentation rather than tracer‐specific CL scaling. Because the implementation details of most commercial software remain proprietary, the relative contribution of each processing step cannot be determined. These findings suggest that future harmonization efforts should focus not only on tracer calibration but also on image‐processing pipelines.
The CL Standard pipeline provides a useful reference for interpreting inter‐software variability. Relative to this reference standard, syngo.PET, Combinostics PET‐only, and Combinostics PET/MRI produced modestly higher CL values (approximately 5–6 CL), whereas MIMneuro showed slightly lower CL values (approximately −1 CL). Mixed‐effects modeling demonstrated that between‐subject variability substantially exceeded residual measurement variability, suggesting that the observed differences are more likely related to systematic implementation differences between software platforms rather than random measurement noise.
Although the CL Standard was intentionally used as the quantitative reference because it represents the published CL methodology and serves as the accepted reference framework for CL scaling rather than as a clinical software platform, direct comparison among FDA‐cleared software platforms is also clinically relevant because these represent the tools available in routine clinical practice. Additional pairwise Bland–Altman analyses among the four FDA‐cleared software platforms demonstrated excellent overall agreement. MIMneuro produced CL values approximately 6–7 CL lower than the other three platforms on average, whereas syngo.PET, Combinostics PET‐only, and Combinostics PET/MRI exhibited minimal systematic bias relative to one another. Nevertheless, the 95% limits of agreement remained approximately ± 20–36 CL, indicating that even FDA‐cleared software platforms are not fully interchangeable when quantitative thresholds are applied.
Notably, the largest inter‐platform differences were predominantly observed in highly amyloid‐positive scans (Table 2), where CL values generally remained well above the diagnostic threshold and therefore had little impact on binary classification. While these large discrepancies may have limited impact at the time of diagnosis, they could become more relevant in longitudinal applications, such as monitoring changes in amyloid burden during anti‐amyloid therapy, where consistent quantitative measurements over time are essential. In contrast, clinically meaningful discordance was more likely to occur in borderline cases with CL values near the diagnostic threshold. For example, using a CL threshold of 25, the visually amyloid‐positive case shown in Figure 4A would be classified as negative by MIMneuro, Combinostics PET‐only, and the CL Standard, but positive by syngo.PET and Combinostics PET/MRI.
FIGURE 4.

Representative cases illustrating clinically relevant inter‐platform variability in Centiloid (CL) values. (A) 1 8F‐florbetapir positron emission tomography (PET) and magnetic resonance imaging (MRI) (Magnetization Prepared Rapid Gradient Echo) of a 56‐year‐old man with suspected Alzheimer's disease. Visual assessment was positive, with abnormal radiotracer uptake most prominent in the right greater than left posterior temporal and occipital lobes. The corresponding CL values were as follows: MIMneuro, 21.14; syngo.PET, 32.7; Combinostics PET‐only, 17.6; Combinostics PET/MRI, 35.7; Centiloid Standard, 12.0. (B) 1 8F‐florbetapir PET and MRI (Magnetization Prepared Rapid Gradient Echo) of a 65‐year‐old man with suspected Alzheimer's disease. The patient has a prominent cisterna magna (not shown), which is incorporated into the whole‐cerebellum reference ROI by the PET‐only workflows (MIMneuro, syngo.PET, and Combinostics PET‐only). This inclusion lowers the reference‐region activity and consequently increases the calculated CL values. Visual assessment was negative, indicating sparse to no amyloid neuritic plaques. CL values across the Platforms varied as follows: MIMneuro, 67.5; Syngo.PET, 85.9; Combinostics PET‐only, 75.7; Combinostics PET/MRI, 28.5; Centiloid Standard, 42.4.
A CL threshold range of 24–30 has been proposed by the Alzheimer's Association Research Roundtable as a practical cutoff for initiating anti‐amyloid therapy in patients with mild cognitive impairment or dementia due to AD. 17 The case shown in Figure 4A remained discordant across software platforms throughout this recommended threshold range. Similarly, in the AHEAD 3‐45 Study of lecanemab for AD prevention, intermediate amyloid is defined as 20–40 CL and elevated amyloid as > 40 CL. 18 Although the PET scan in Figure 4B was visually interpreted as amyloid‐negative, MIMneuro, syngo.PET, Combinostics PET‐only, and the CL Standard classified it as elevated amyloid (> 40 CL), whereas Combinostics PET/MRI classified it as intermediate amyloid (20–40 CL), resulting in different trial eligibility classifications. Together, these examples illustrate that even modest inter‐platform differences may lead to clinically meaningful discrepancies when fixed quantitative thresholds are applied for diagnosis, treatment eligibility, or clinical trial enrollment. This is particularly relevant because quantitative amyloid PET burden is increasingly integrated with other AD biomarkers, including CSF Aβ, in clinical decision‐making. 19
Despite these systematic offsets, overall agreement remained excellent (ICC = 0.96), indicating that all platforms preserved the relative ranking of subjects. However, as demonstrated by the mixed‐effects modeling and Bland–Altman analyses, high ICC does not imply numerical equivalence between methods. This distinction is especially relevant in clinical practice, where absolute CL values rather than relative ranking often guide diagnostic classification and therapeutic decision‐making.
Our findings should be interpreted in the context of a recent study evaluating CL quantification from 1 8F‐flutemetamol PET using seven commercial and research pipelines. 20 The 1 8F‐flutemetamol study reported excellent overall reproducibility across software implementations and concluded that the choice of quantification software should not impact patient management decisions in clinical practice. While our results similarly demonstrate excellent overall agreement, we specifically focused on individual‐level discrepancies and threshold‐based classification. We observed substantial inter‐software differences in selected cases that resulted in discordant classifications near clinically relevant decision thresholds. These findings suggest that strong population‐level agreement does not necessarily imply complete interchangeability of CL values at the individual patient level, particularly when fixed quantitative thresholds are used for clinical interpretation, treatment eligibility, or longitudinal assessment.
Comparison of the PET‐only (MIMneuro, syngo.PET, Combinostics PET‐only) and hybrid PET/MRI (Combinostics PET/MRI, CL Standard) pipelines provides insight into potential sources of inter‐platform variability. MRI‐based pipelines may offer more anatomically accurate segmentation using high‐resolution structural imaging, whereas PET‐only pipelines rely on PET‐derived templates. Despite these methodological differences, both approaches demonstrated excellent agreement and diagnostic performance, with AUCs exceeding 0.94 across all platforms. While MRI‐based normalization may provide theoretical anatomical advantages, the observed CL differences suggest that normalization strategy alone does not fully account for inter‐platform variability. Instead, the combined effects of template selection, ROI definition, and processing strategy likely play a larger role.
Threshold‐based analysis using CL ≥ 25 demonstrated consistently high sensitivity but greater variability in specificity, indicating a strong ability to detect amyloid‐positive cases, while systematic CL offsets near the decision boundary influence specificity. Consequently, modest differences in absolute CL values may result in discordant classification in borderline cases (Figure 4), particularly in multicenter studies where software‐related variability may introduce non‐biological differences not captured by ICC.
As an exploratory analysis, we compared differences in the Youden‐optimal thresholds with the corresponding Bland–Altman mean differences to evaluate whether systematic measurement offsets could explain inter‐software differences in optimal diagnostic thresholds. For Combinostics PET‐only and MIMneuro, the differences in the Youden‐optimal thresholds closely approximated the corresponding Bland–Altman mean differences, suggesting that the observed inter‐method differences were largely explained by systematic calibration offsets. In contrast, syngo.PET and Combinostics PET/MRI demonstrated larger discrepancies between the differences in the Youden‐optimal thresholds and the corresponding Bland–Altman mean differences, suggesting that factors beyond a simple calibration offset may contribute to method‐specific optimal thresholds. Because this analysis was exploratory, these findings should be interpreted cautiously.
Several limitations should be acknowledged. This was a retrospective single‐center study, which may limit generalizability. Visual interpretation served as the gold standard rather than histopathological confirmation. Although visual readings are widely used in clinical practice, autopsy correlation would provide additional confirmation of the biological validity of amyloid PET quantification. Additionally, only 1 8F‐florbetapir PET was evaluated, and the findings may not directly generalize to other amyloid tracers. Furthermore, several factors beyond differences in software implementation may contribute to inter‐platform variability; however, these were not specifically evaluated in the present study. For example, structural brain abnormalities commonly encountered in patients undergoing amyloid PET, including cerebral atrophy, chronic infarcts, communicating hydrocephalus, postoperative changes, or intracranial masses, may affect image registration and ROI segmentation, thereby contributing to quantitative variability. Accordingly, quantitative findings should always be interpreted in conjunction with careful visual assessment. Nevertheless, the use of identical datasets from multiple FDA‐cleared software platforms, the inclusion of an in‐house developed CL Standard reference pipeline, the direct comparison of PET‐based and MRI‐based normalization approaches, and the integration of both agreement and diagnostic performance analyses represent key strengths of this study.
Future studies should investigate the specific methodological factors contributing to inter‐platform variability, including differences in image registration, ROI definition, reference‐region delineation, and quantitative processing pipelines. Validation in larger multicenter cohorts and across additional amyloid PET tracers will further clarify the generalizability of these findings.
5. CONCLUSION
Despite substantial standardization achieved by the CL framework, residual variability across software implementations remains an underrecognized source of measurement uncertainty. Although overall agreement and diagnostic performance were excellent across FDA‐cleared platforms, systematic inter‐software differences persisted and may become clinically meaningful near quantitative decision thresholds. Accordingly, CL values are not fully interchangeable, underscoring the importance of using a consistent processing pipeline for reliable clinical interpretation and longitudinal assessment in both clinical practice and clinical trials.
AUTHOR CONTRIBUTIONS
J. Zhang, D. Johnson, and B. Burkett conceived and designed the study. C. Schwarz provided in‐house software implementation pipeline. J. Zhang, D. Johnson, B. Burkett and C. Schwarz were responsible for the methodology and investigation. N. Dizdar and J. Zhang contributed to data de‐identification, processing and analysis. M. Johnson contributed to statistical analysis. D. Johnson, B. Burkett and C. Bilgin performed clinical image reviews. B. Kemp was responsible for clinical PET/CT acquisition and QA/QC management. J. Durski, V. Lowe and D. Johnson provided supervision and scientific guidance. J. Zhang drafted the original manuscript, and all authors contributed to manuscript revision and approved the final version.
CONFLICT OF INTEREST STATEMENT
Brian Burkett did consulting work for Clario. Christopher G. Schwarz receives research funding from the NIH, outside this work. Derek Johnson does consultant and/or advisory board work: GE Healthcare, Telix, Novartis, Collectar, and Eli Lilly. All other authors declare no conflict of interest. Author disclosures are available in the Supporting Information.
ETHICS STATEMENT
This study was approved by the Mayo Clinic institutional review board, and all procedures were performed in accordance with relevant guidelines and regulations.
CONSENT STATEMENT
The requirement for informed consent was waived due to the retrospective nature of the study, in accordance with the Mayo Clinic institutional review board approval.
Supporting information
Supporting Information
ACKNOWLEDGMENTS
The authors have nothing to report.
REFERENCES
- 1. Chapleau M, Iaccarino L, Soleimani‐Meigooni D, Rabinovici GD. The role of amyloid PET in imaging neurodegenerative disorders: a review. J Nucl Med. 2022;63:13S‐19S. doi:10.2967/jnumed.121.263195 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2. Ruan D, Sun L. Amyloid‐β PET in Alzheimer's disease: a systematic review and Bayesian meta‐analysis. Brain Behav. 2023;13:e2850. doi:10.1002/brb3.2850 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3. Rabinovici GD, Gatsonis C, Apgar C, et al. Association of amyloid positron emission tomography with subsequent change in clinical management among Medicare beneficiaries with mild cognitive impairment or dementia. JAMA. 2019;321:1286‐1294. doi:10.1001/jama.2019.2000 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4. Bollack A, Pemberton HG, Collij LE, et al. Longitudinal amyloid and tau PET imaging in Alzheimer's disease: a systematic review of methodologies and factors affecting quantification. Alzheimers Dement. 2023;19:5232‐5252. doi:10.1002/alz.13158 [DOI] [PubMed] [Google Scholar]
- 5. Rabinovici GD, Knopman DS, Arbizu J, et al. Updated appropriate use criteria for amyloid and tau PET: a report from the Alzheimer's Association and Society for Nuclear Medicine and Molecular Imaging Workgroup. J Nucl Med. 2025;66(2):S5‐S31. doi:10.2967/jnumed.124.268756 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6. Pemberton HG, Collij LE, Heeman F, et al. Quantification of amyloid PET for future clinical use: a state‐of‐the‐art review. Eur J Nucl Med Mol Imaging. 2022;49:3508‐3528. doi:10.1007/s00259‐022‐05784‐y [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7. Collij LE, Bollack A, La Joie R, et al. Centiloid recommendations for clinical context‐of‐use from the AMYPAD consortium. Alzheimers Dement. 2024;20:9037‐9048. doi:10.1002/alz.14336 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8. Klunk WE, Koeppe RA, Price JC, et al. The Centiloid Project: standardizing quantitative amyloid plaque estimation by PET. Alzheimers Dement. 2015;11:1‐15. doi:10.1016/j.jalz.2014.07.003 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9. Meyer PF, McSweeney M, Gonneaud J, Villeneuve S. PET amyloid imaging across the Alzheimer's disease spectrum: from disease mechanisms to prevention. Prog Mol Biol Transl Sci. 2019;165:63‐106. doi:10.1016/bs.pmbts.2019.05.001 [DOI] [PubMed] [Google Scholar]
- 10. Battle MR, Pillay LC, Lowe VJ, et al. Centiloid scaling for quantification of brain amyloid with [1 8F]flutemetamol using multiple processing methods. EJNMMI Res. 2018;8:107. doi:10.1186/s13550‐018‐0456‐7 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11. Rowe CC, Doré V, Jones G, et al. 1 8F‐Florbetaben PET beta‐amyloid binding expressed in Centiloids. Eur J Nucl Med Mol Imaging. 2017;44:2053‐2059. doi:10.1007/s00259‐017‐3749‐6 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12. Rowe CC, Jones G, Doré V, et al. Standardized expression of 1 8F‐NAV4694 and 1 1C‐PiB β‐amyloid PET results with the Centiloid scale. J Nucl Med. 2016;57:1233‐1237. doi:10.2967/jnumed.115.171595 [DOI] [PubMed] [Google Scholar]
- 13. Coath W, Modat M, Cardoso MJ, et al. Operationalizing the Centiloid scale for [1 8F]florbetapir PET studies on PET/MRI. Alzheimers Dement (Amst). 2023;15:e12434. doi:10.1002/dad2.12434 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14. Navitsky M, Joshi AD, Kennedy I, et al. Standardization of amyloid quantitation with florbetapir standardized uptake value ratios to the Centiloid scale. Alzheimers Dement. 2018;14:1565‐1571. doi:10.1016/j.jalz.2018.06.1353 [DOI] [PubMed] [Google Scholar]
- 15. DiFilippo FP, Rao SM. Evaluation of two clinical Centiloid analysis products for 1 8F‐florbetapir PET in cognitively unimpaired elders. Alzheimers Dement (Amst). 2025;17(3):e70163. doi:10.1002/dad2.70163 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16. Schwarz CG, Tosakulwong N, Senjem ML, et al. Considerations for performing Level‐2 Centiloid transformations for amyloid PET SUVR values. Sci Rep. 2018;8:7421. doi:10.1038/s41598‐018‐25459‐9 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17. Farrar G, Weber CJ, Rabinovici GD, et al. Expert opinion on Centiloid thresholds suitable for initiating anti‐amyloid therapy. J Prev Alzheimers Dis. 2025;12:100008. doi:10.1016/j.tjpad.2024.100008 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18. Rafii MS, Sperling RA, Donohue MC, et al. The AHEAD 3‐45 Study: design of a prevention trial for Alzheimer's disease. Alzheimers Dement. 2023;19:1227‐1238. doi:10.1002/alz.12748 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19. La Joie R, Visani AV, Baker SL, et al. Multisite study of the relationships between amyloid PET and CSF biomarkers. Alzheimers Dement. 2019;15(3):311‐322. doi:10.1016/j.jalz.2018.09.010 [Google Scholar]
- 20. Bollack A, Schwarz AJ, Bourgeat P, et al. Comparability of Centiloid values from [1 8F]flutemetamol scans using seven commercial and research software. Neuroimage Rep. 2026;6(2):100343. doi:10.1016/j.ynirp.2026.100343 [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Supporting Information
