Abstract
Objective
Despite the universal calibration of commercial IGF1 immunoassays to WHO IS 02/254, substantial inter-assay variability persists, leading to inconsistent patient classification. Harmonization towards a higher-order analytical anchor may reduce such variability.
Methods
Four matrix-matched, multi-level serum reference materials (RMs) were prepared from donor serum and value-assigned using an LC-MS/MS method calibrated to WHO IS 02/254. Commutability was assessed according to IFCC recommendations across four immunoassays (Cobas, iSYS, Immulite, Liaison). Deming regression-based recalibration equations derived from commutable RMs were applied to patient samples and healthy donor samples. The primary quantitative endpoint was reduction in standard error of estimate (SEE) relative to the LC-MS/MS method. Age- and sex-specific LC-MS/MS-anchored reference intervals were constructed as a downstream application.
Results
Prior to recalibration, immunoassays showed positive bias relative to LC-MS/MS of up to 60%. All four RMs were commutable for Liaison and iSYS, whereas the lowest concentration RM was classified as non-commutable for Cobas and Immulite. Recalibration towards the LC-MS/MS anchor resulted in marked alignment towards the identity line and reduced pooled SEE from 7.82 to 4.89 nmol/L (−37.4%) in patient samples and from 7.34 to 2.09 nmol/L (−71.5%) in healthy samples. Although harmonization effects were assay-dependent at the individual platform level, overall cross-platform dispersion was substantially attenuated.
Conclusions
Matrix-matched, commutable serum RMs value-assigned by a higher-order LC-MS/MS procedure enable substantial reduction of inter-assay bias and variability among IGF1 immunoassays. Harmonization towards a higher-order analytical anchor is achievable in routine practice and provides a robust foundation for consistent cross-platform interpretation.
Keywords: IGF1, harmonization, LC-MS/MS, reference intervals
Introduction
Insulin-like growth factor 1 (IGF1) is central to the diagnosis and management of growth hormone disorders across the lifespan (1). Despite the universal calibration of commercial immunoassays to the WHO International Standard (IS 02/254) (2), substantial inter-method variability persists. Differences in antibody specificity, analyte dissociation procedures, matrix effects and susceptibility to interference contribute to systematic disagreement between platforms. Consequently, identical patient samples may yield divergent IGF1 concentrations and standard deviation scores (SDSs). In routine clinical practice, this discordance is further amplified by differences in method-specific reference intervals, or by the use of reference intervals established with a different method, potentially leading to inconsistent patient classification and therapeutic decisions across institutions (3, 4, 5, 6, 7, 8, 9).
Calibration traceability to a common international standard does not in itself guarantee equivalence of patient test results. Within a metrological traceability framework, routine immunoassays represent lower-tier measurement procedures that ideally align to higher-order methods through commutable reference materials. A higher-order LC-MS/MS method, characterized by superior molecular specificity, can function as a reference anchor within this hierarchy (10). Indeed, several IGF1 LC-MS/MS methods have been published in recent years, showing excellent analytical performance with respect to selectivity and imprecision, thereby supporting their suitability as candidate reference measurement procedures (11, 12, 13, 14, 15, 16, 17, 18). When matrix-matched, commutable serum reference materials (RMs) are value-assigned using such a higher-order procedure, calibration transfer can occur in a manner that preserves patient sample commutability and improves inter-assay alignment (19).
Although a single harmonization sample has previously been explored for growth hormone (GH1) (20) and found commutable for IGF1, this sample proved less effective in reducing the between-laboratory CV for IGF1 compared to GH within national harmonization efforts in the Netherlands. Furthermore, no systematic demonstration has shown that multi-level, commutable serum RMs spanning the physiological range can substantially reduce inter-assay bias and variability across commercial IGF1 platforms. As such, it remains unclear whether analytical harmonization provides a sufficiently stable framework to support unified interpretation using LC-MS/MS-anchored reference intervals.
The primary objective of this study was therefore to determine whether multi-level, matrix-matched commutable serum RMs, value-assigned by a higher-order LC-MS/MS procedure, can harmonize routine IGF1 immunoassays towards an LC-MS/MS reference anchor. Specifically, we aimed to demonstrate RM commutability across platforms according to the criteria and quantify the reduction in inter-assay bias and variability after recalibration. The construction of LC-MS/MS-anchored age- and sex-specific reference intervals was included as a downstream application of the harmonized analytical framework.
Materials and methods
Reference material preparation
Four multi-level serum RMs were prepared from human donor serum obtained from blood bank collections. Donors were selected to provide naturally varying IGF1 concentrations spanning the physiological range. No recombinant IGF1 spiking or artificial modification (e.g., protein depletion or charcoal stripping) was performed. To obtain sufficient volume while preserving matrix integrity, a maximum of two donations were pooled per concentration level. Serum aliquots (400 μL) were stored at −80°C until analysis. IGF1 concentrations of the RMs were value-assigned using a validated LC-MS/MS method calibrated against WHO International Standard 02/254 (11). Value assignment was performed within a single analytical run (10-fold replicates) to minimize the impact of between-run variability on the RM value assignment. The resulting concentration levels ranged approximately from 7 to 32 nmol/L (CV: 3.17–3.46%).
Patient and healthy cohorts
For the harmonization and commutability analyses, 42 anonymized patient serum samples (adults and children) were obtained as leftover material from routine IGF1 testing and anonymized. The exclusion criteria included documented growth disorders, significant hepatic or renal dysfunction, malignancy and use of medications known to affect IGF1 concentrations. Samples were aliquoted and stored at −20°C prior to measurement. Healthy volunteer samples (n = 40) were derived from the Lifelines population-based Biobank (21).
Immunoassays and LC-MS/MS analysis
Four commercial IGF1 immunoassays currently in use in the Netherlands were evaluated: Cobas (Roche Diagnostics, the Netherlands), iSYS (IDS, Germany), Immulite 2000 (Siemens Healthineers, Germany) and Liaison (DiaSorin S.p.A., Italy). All assays were calibrated according to manufacturers’ instructions and traceable to WHO IS 02/254. Each of the laboratories performed full calibration and internal quality control prior to sample measurement. The intact IGF1 LC-MS/MS method was previously validated and calibrated against WHO IS 02/254 (11). The method functions as a higher-order reference anchor within a traceability hierarchy but is not positioned as a formally recognized reference measurement procedure.
Commutability assessment and recalibration
Commutability was assessed according to IFCC Working Group recommendations (22) using the difference-in-bias approach. We used the Excel sheets provided in the data supplements of the online version of this article for calculations. In short, the expanded uncertainty of the difference in bias between each RM and the clinical sample bias was calculated according to IFCC recommendations. This value reflects the combined influence of analytical imprecision and variability observed among clinical samples and defines the confidence limits used for commutability classification. For recalibration, RMs classified as commutable were included in calibration modelling. RMs classified as non-commutable were excluded. Measurements followed a fixed protocol and were conducted in a single run, with the RMs placed between the patient samples and healthy volunteer samples. Each sample was measured in duplicate. For each immunoassay, Deming regression (GraphPad Prism, USA) was performed between immunoassay results and LC-MS/MS values. Two experiments were conducted: i) RMs measured alongside patient samples, reflecting real-world clinical conditions, and ii) RMs measured alongside healthy volunteer samples (six months later in different laboratories to assess robustness and to isolate more structural analytical differences). Deming regression was used to derive calibration equations between each immunoassay and the LC-MS/MS anchor based exclusively on commutable RMs. For Liaison and iSYS, RM1–4 was included. For Cobas and Immulite, RM1 was classified as non-commutable and therefore excluded. These equations were subsequently applied to patient and healthy donor immunoassay results to generate recalibrated IGF1 values.
Standard error of estimate (SEE)
The primary quantitative indicator of harmonization success was reduction in the standard error of estimate (SEE) relative to LC-MS/MS. SEE was calculated as the square root of the residual mean square from Deming regression using vertical residuals (defined as the observed immunoassay value (y) minus the value predicted (ŷ) from Deming regression against LC-MS/MS) and n-2 degrees of freedom,
SEE was computed before and after recalibration for each cohort and method separately.
Bootstrap analysis
To assess statistical robustness of SEE reduction, paired bootstrap resampling (10.000 iterations) was performed at the sample level. In each iteration, sample pairs were resampled with replacement and SEE values were recalculated for pre- and post-recalibration models. The distribution of ΔSEE (SEE_pre – SEE_post) was used to derive 95% confidence intervals (CIs) and corresponding P values.
Outlier definition
Outliers were predefined as observations with studentized residuals exceeding ±3 standard deviations from Deming regression. Cook’s distance was used as a sensitivity measure. Outliers were retained in the comparative analyses but removed when calculating the SEE.
Historical reference dataset and recalculation procedure
Previously, reference intervals for IGF1 were established using the Liaison IGF1 immunoassay and included 698 males and 901 females (28% aged 0–18 years and 72% aged > 18 years). These samples consisted of anonymized leftover patient materials used for screening purposes (classified as non-WMO research and therefore exempt from formal Medical Ethical Committee approval) as well as samples from healthy individuals participating as controls in clinical studies (with written informed consent obtained according to institutional regulations). Samples were collected in two clinical laboratories, and each analytical run included a previously developed harmonization serum sample (introduced in 2004). Because this historical harmonization sample was subsequently measured using the LC-MS/MS method applied in the present study, original Liaison-based IGF1 results could be recalculated to the LC-MS/MS reference anchor using established regression relationships. This recalculation enabled integration of the historical dataset within the LC-MS/MS-anchored analytical framework.
Reference interval modelling
For further construction of LC-MS/MS-anchored reference intervals, serum samples were obtained from the Lifelines Biobank (the Netherlands), a large prospective population-based cohort study comprising 167,729 participants across three generations (21). A total of 1.520 samples (760 males and 760 females; 58% aged 8–18 years and 42% >18 years) were selected from apparently healthy volunteers aged 8–94 years. At the Lifelines site, samples were aliquoted and stored at −80°C until use. Continuous frozen storage was maintained throughout handling to preserve sample integrity. Age- and sex-specific reference intervals were constructed using the LMS method (23, 24). L, M, and S parameters were estimated for each year of age stratified by sex, enabling calculation of IGF1 standard deviation scores (SDSs). Statistical comparison between recalculated historical data and newly generated Lifelines data was made using the Mann–Whitney U test for non-normally distributed differences. The Lifelines combined IGF1 normative dataset is publicly available through the Zenodo repository (25).
Results
Inter-method bias relative to LC-MS/MS
For both the patient and healthy cohort, IGF1 concentrations measured by LC-MS/MS were consistently lower than those obtained with immunoassays. Depending on platform and concentration range, mean positive bias of immunoassays relative to LC-MS/MS ranged from negligible to approximately 60%. Regression slopes and intercepts for Cobas, iSYS and Liaison demonstrated substantial deviation from the identity line prior to recalibration (Fig. 1), confirming persistent inter-method disagreement despite common calibration to WHO IS 02/254. Detailed Deming regression parameters are provided in Table 1.
Figure 1.

Baseline comparison of immunoassay IGF1 concentrations versus LC-MS/MS reference anchor. Orange symbols represent patient samples; blue symbols represent healthy donor samples. The dashed line indicates the line of identity (y = x), the solid lines indicate Deming regression for both sample types. (A) Cobas, (B) Immulite, (C) iSYS and (D) Liaison.
Table 1.
Deming regression parameters of each immunoassay relative to the LC-MS/MS reference anchor before recalibration, presented separately for healthy and patient cohorts. Slopes and intercepts (with 95% confidence intervals and standard errors (SEs)) quantify baseline proportional and constant bias.
| Cohort | Method | Slope | Range of slope | SE slope | Intercept (nmol/L) | Range of intercept |
|---|---|---|---|---|---|---|
| Healthy | Cobas | 1.43 | 1.34–1.53 | 0.047 | −3.49 | −5.27 to −1.71 |
| Immulite | 0.84 | 0.81–0.88 | 0.017 | 1.76 | 0.78–2.74 | |
| iSYS | 1.48 | 1.34–1.61 | 0.066 | 0.82 | −1.89 to 3.53 | |
| Liaison | 1.13 | 1.06–1.20 | 0.035 | 2.56 | 1.01–4.1 | |
| Patient | Cobas | 1.38 | 1.26–1.50 | 0.060 | −4.69 | −8.38 to −0.99 |
| Immulite | 0.99 | 0.83–1.16 | 0.081 | −3.78 | −0.89 to 3.66 | |
| iSYS | 1.26 | 1.04–1.47 | 0.108 | −3.78 | −9.7 to 2.21 | |
| Liaison | 1.43 | 1.19–1.66 | 0.118 | 0.90 | −4.17 to 5.96 |
Commutability of matrix-matched reference materials
Four serum-based RMs spanning approximately 7–32 nmol/L were evaluated according to IFCC Working Group recommendations. Commutability acceptance criteria were defined according to the IFCC difference-in-bias approach (22). Commutability was assessed by comparing the difference in bias of each RM relative to the mean bias of clinical samples at corresponding concentrations, including calculation of expanded uncertainty. Regression plots (Fig. 2, panels A, C, E, and G) showed consistent proportional relationships between immunoassays and LC-MS/MS RMs values for Liaison, Cobas and iSYS, whereas Immulite exhibited a distinct slope pattern, particularly in the lower concentration range. For the Liaison and iSYS platforms, all four RMs fulfilled predefined commutability criteria in the difference-in-bias analysis (Fig. 2, panels F and H). For the Cobas platform, three RMs met the commutability criterion, whereas the lowest concentration RM (RM1) demonstrated non-commutability, with its difference-in-bias interval exceeding the predefined commutability limits (Fig. 2, panel B). Similarly, for the Immulite platform, RM1 was non-commutable, while RM2–4 fulfilled the commutability criterion (Fig. 2, panel D).
Figure 2.
Assessment of commutability of matrix-matched serum reference materials (RM1–RM4) relative to the LC-MS/MS reference anchor. (A, C, E and G) Deming regression of immunoassay results versus LC-MS/MS for Cobas, Immulite, iSYS and Liaison, respectively. The dashed line represents the line of identity (y = x); the solid line represents the Deming regression fit. (B, D, F, H) IFCC difference-in-bias plots (ln scale), showing the bias of clinical samples (black dots) and RMs (coloured circles) as ln(immunoassay) – ln(LC-MS/MS) plotted against the mean concentration of both methods. The solid horizontal line indicates the mean clinical sample bias; dashed horizontal lines represent predefined commutability criteria. Vertical bars represent the expanded uncertainty of the difference in bias for each RM. RMs were classified as commutable, indeterminate or non-commutable according to IFCC difference-in-bias criteria.
These findings confirm that multi-level, matrix-matched donor serum materials can function as commutable calibrators across most platforms, while also highlighting assay-specific matrix sensitivities.
Recalibration towards the LC-MS/MS reference anchor
Using Deming regression, calibration equations were derived exclusively from commutable RMs (RM1–4 for Liaison and iSYS; RM2–4 for Cobas and Immulite) and applied to patient and healthy donor samples. Post-recalibration regression lines demonstrated markedly improved alignment towards the LC-MS/MS identity line for most immunoassay methods (Fig. 3).
Figure 3.

Immunoassay IGF1 concentrations after recalibration towards the LC-MS/MS reference anchor using Deming regression-derived calibration equations. Orange symbols represent patient samples; blue symbols represent healthy donor samples. The solid lines indicate the Deming regression fit after recalibration, and the dashed line represents the line of identity (y = x). (A) Cobas, (B) Immulite, (C) iSYS and (D) Liaison. SEE values before and after recalibration are reported in Table 2.
The primary quantitative endpoint was reduction in the standard error of estimate (SEE), reflecting inter-assay variability at the patient sample level (Table 2). Within-cohort, recalibration significantly reduced SEE, indicating improved agreement with the LC-MS/MS, for Cobas, Liaison, and iSYS in both cohorts (bootstrap P < 0.001; positive ΔSEE; CIs above 0). However, Immulite worsened in both cohorts (negative ΔSEE; bootstrap P = 1.0). Despite this heterogeneity, the pooled (all assays combined) analysis shows a large, statistically robust improvement after recalibration in both cohorts (ΔSEE = 5.25 (71.5% reduction) for the healthy cohort and 2.92 (37.4% reduction) for the patient cohort; both P < 0.001).
Table 2.
Standard error of the estimate (SEE) from Deming regression (λ = 1) comparing immunoassays with LC-MS/MS before (pre) and after (post) recalibration. Stratified by cohort and assay and for pooled assays within each cohort. ΔSEE is defined as SEE_pre – SEE_post; positive values indicate improved agreement after recalibration. Ninety-five percent confidence intervals (CIs) for ΔSEE were obtained by percentile bootstrap (10.000 resamples). P values are one-sided. SEE was calculated using vertical residuals with n – 2 degrees of freedom.
| Cohort | Method | n | Pre-SEE | Post-SEE | % Reduction | ΔSEE | Lower 95% CI | Upper 95% CI | P value |
|---|---|---|---|---|---|---|---|---|---|
| Healthy | Pooled | 156 | 7.34 | 2.09 | 71.5 | 5.25 | 4.37 | 5.97 | <0.001 |
| Patient | Pooled | 159 | 7.82 | 4.89 | 37.4 | 2.92 | 1.58 | 4.18 | <0.001 |
| All | Pooled | 315 | 7.64 | 3.76 | 50.7 | 3.87 | 2.96 | 4.73 | <0.001 |
| Healthy | Cobas | 39 | 2.49 | 1.69 | 31.9 | 0.79 | 0.56 | 0.96 | <0.001 |
| Patient | Cobas | 40 | 5.45 | 3.70 | 32.2 | 1.75 | 0.94 | 2.60 | <0.001 |
| Healthy | Immulite | 39 | 1.67 | 1.80 | −8.1 | −0.14 | −0.16 | −0.10 | 1.0 |
| Patient | Immulite | 41 | 5.43 | 6.03 | −11.1 | −0.60 | −0.98 | −0.23 | 1.0 |
| Healthy | iSYS | 38 | 2.61 | 1.90 | 27.2 | 0.71 | 0.49 | 0.84 | <0.001 |
| Patient | iSYS | 39 | 5.93 | 4.71 | 20.6 | 1.22 | 0.79 | 1.57 | <0.001 |
| Healthy | Liaison | 40 | 2.52 | 2.09 | 17.0 | 0.43 | 0.34 | 0.49 | <0.001 |
| Patient | Liaison | 39 | 6.61 | 4.12 | 37.8 | 2.50 | 1.46 | 3.36 | <0.001 |
Overall, these results demonstrate substantial reduction of analytical dispersion following calibration transfer via commutable RMs, indicating successful harmonization towards the higher-order LC-MS/MS anchor. The higher dispersion observed in patient samples likely reflects biological variability rather than analytical instability, underscoring the importance of evaluating harmonization in real-world clinical specimens.
LC-MS/MS-anchored reference intervals as downstream application
Following demonstration of analytical harmonization, age- and sex-specific reference intervals were constructed using LC-MS/MS data from the Lifelines cohort (n = 1,520) supplemented with recalculated historical data. LMS modelling generated centile curves and corresponding SDS calculations (Fig. 4). A comparison between recalculated historical data and newly generated Lifelines data showed no statistically significant difference in males (P = 0.061). In females, a modest difference (P = 0.01) was observed within the 40–60 year age range. This divergence likely reflects cohort-specific characteristics rather than analytical instability. Importantly, reference interval construction is interpreted as a logical extension of demonstrated analytical alignment. Once immunoassays are harmonized towards a common LC-MS/MS reference anchor with substantially reduced inter-assay variability, application of shared interpretative frameworks becomes clinically feasible and defensible.
Figure 4.

LC-MS/MS-anchored age- and sex-specific reference intervals derived using LMS modelling. (A) females; (B) males. The solid line represents the median (0 SDS), and dashed lines represent −1.97 and +1.97 standard deviation scores (SDSs), corresponding to the 2.5th and 97.5th percentiles.
Discussion
This study demonstrates that multi-level, matrix-matched serum reference materials (RMs) value-assigned by a higher-order LC-MS/MS procedure substantially reduce inter-assay bias and variability among most IGF1 immunoassays. Despite universal calibration to WHO IS 02/254, baseline inter-method differences of up to 60% were observed. Large inter-method discrepancies for IGF1 immunoassays have been reported previously, including differences in age-related reference datasets or SDS (9, 26, 27, 28), yet without implementation of harmonization interventions. In the present study, recalibration using commutable RMs reduced pooled standard error of estimate (SEE) by 37.4% in patient samples and 71.5% in healthy donor samples. These reductions reflect a clinically meaningful decrease in analytical dispersion at the patient sample level. Positioning LC-MS/MS as a higher-order analytical anchor within a traceability hierarchy provides a metrologically coherent framework for interpreting these findings. Calibration traceability to WHO IS 02/254 alone did not prevent substantial between-assay structural differences, as reflected by divergent slopes and intercepts at baseline. In contrast, value assignment of matrix-matched, commutable serum RMs enabled recalibration that markedly reduced pooled dispersion, particularly in the healthy cohort where structural assay differences dominate over biological variability.
In contrast, Cobas and Immulite demonstrated non-commutability for the lowest concentration RM. For Immulite, recalibration resulted in worsening SEE in both cohorts (negative ΔSEE), indicating that structural assay characteristics were not fully corrected by linear recalibration using the available commutable RMs. These findings suggest platform-specific susceptibility to matrix effects or dissociation inefficiency at lower concentrations and indicate that recalibration cannot universally compensate for intrinsic analytical design. In addition to assay architecture, calibration geometry likely influenced post-recalibration alignment. Recalibration was performed using Deming regression derived from a limited number of multi-level RMs. As with any regression-based recalibration, the distribution of calibration concentrations determines statistical leverage and therefore influences slope and intercept adjustment. Exclusion of the lowest concentration RM reduced low-end leverage, which may partly explain platform-specific heterogeneity in alignment. Thus, harmonization performance depends not only on commutability but also on the concentration structure of the calibration materials.
Analytical harmonization has direct consequences for clinical decision-making in disorders of the growth hormone-IGF1 axis. In acromegaly, IGF1 concentrations and corresponding SDS values are central to diagnosis, assessment of biochemical control and therapeutic monitoring. Inter-assay bias of up to 60%, as observed prior to recalibration, may shift patients across decision thresholds (e.g., from mildly elevated to normal range or vice versa), thereby influencing treatment initiation or adjustment (4, 5). The reduction of analytical dispersion through harmonization decreases the likelihood that assay choice alone determines classification. Similarly, in growth hormone deficiency, treatment titration frequently targets normalization of age-adjusted IGF1 SDS (29). When SDS calculations are derived from method-specific reference intervals that are not analytically aligned, dose adjustments may reflect analytical artefacts rather than true biological change. Harmonization towards a higher-order anchor reduces assay-dependent variability and supports more consistent longitudinal monitoring, particularly when patients are followed across institutions using different platforms. Although this study did not directly evaluate clinical outcomes, the demonstrated reduction in analytical variability provides a mechanistic basis for improved consistency in diagnostic classification and therapeutic monitoring.
The presence of extreme discrepancies in a subset of patient samples may reflect sample-specific interferences, including abnormal IGF-binding protein concentrations, partial dissociation inefficiency or heterophilic antibody effects (3, 30). Such phenomena represent inherent limitations of immunoassay methodologies rather than failure of the harmonization framework. Indeed, harmonization towards a higher-order anchor may facilitate identification of such anomalous samples by reducing baseline inter-assay variability.
Reference interval construction was intentionally positioned as a downstream application of demonstrated analytical alignment. Once inter-assay variability attributable purely to methodological differences is substantially reduced, the use of LC-MS/MS-anchored age- and sex-specific reference intervals becomes analytically defensible. Importantly, the present study does not claim that universal intervals are automatically transferable across all settings; rather, it demonstrates that harmonization towards a higher-order anchor provides the necessary analytical infrastructure for such implementation.
Several limitations warrant consideration. First, the LC-MS/MS method functions as a higher-order reference anchor but is not established as an internationally recognized reference measurement procedure. Second, the concentration range of the RMs, obtained from healthy persons is limited, while patients show a much wider range of IGF1 concentrations. Third, the patient cohort used for recalibration (n = 42) is modest in size, although sufficient to demonstrate directional harmonization effects. Fourth, the reference dataset partly incorporated recalculated historical samples that included screening-derived patient material whilst a reference interval derived exclusively from a strictly defined healthy population might yield slightly different centile estimates. Finally, external validation across additional platforms and international settings remains necessary to confirm generalizability.
Declaration of interest
The authors declare that there is no conflict of interest that could be perceived as prejudicing the impartiality of the work reported.
Funding
This work was supported by funding the Lifelines samples from Pfizer (Research Grant No. WP1702911). The Lifelines initiative has been made possible by subsidy from the Dutch Ministry of Health, Welfare and Sport, the Dutch Ministry of Economic Affairs, the University Medical Center Groningen (UMCG), University of Groningen and the Provinces in the North of the Netherlands (Drenthe, Friesland, Groningen).
Ethics statement
The Lifelines study was approved by the ethics committee of the University Medical Center Groningen. Informed consent was obtained from all participants included in the study. The use of anonymized leftover patient samples was classified as non-WMO research by the Medical Ethical Committee of the University Medical Center Utrecht, the Netherlands.
Acknowledgments
We are grateful to Dr Bart Ballieux, Dr Judith Bons, Prof. Dr Annemieke Heijboer, Dr Gideon Lansbergen, Prof. Dr Yolanda de Rijke and Dr Joost van de Ven for the measurements of the samples.
References
- 1.Blum WF, Alherbish A, Alsagheir A, et al. The growth hormone–insulin-like growth factor-I axis in the diagnosis and treatment of growth disorders. Endocr Connections 2018. 7 R212–R222. ( 10.1530/EC-18-0099) [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Burns C, Rigsby P, Moore M, et al. The first international standard for insulin-like growth factor-1 (IGF-1) for immunoassay: preparation and calibration in an international collaborative study. Growth Hormone IGF Res 2009. 19 457–462. ( 10.1016/j.ghir.2009.02.002) [DOI] [PubMed] [Google Scholar]
- 3.Huang R, Shi J, Wei R, et al. Challenges of insulin-like growth factor-1 testing. Crit Rev Clin Lab Sci 2024. 61 388–403. ( 10.1080/10408363.2024.2306804) [DOI] [PubMed] [Google Scholar]
- 4.Postma MR, van Beek AP, van der Klauw MM, et al. IGF-1 as screening tool for acromegaly and adult-onset growth hormone deficiency in the Netherlands. Clin Endocrinol 2024. 100 260–268. ( 10.1111/cen.15000) [DOI] [PubMed] [Google Scholar]
- 5.Clemmons DR & Bidlingmaier M. IGF-I assay methods and biologic variability: evaluation of acromegaly treatment response. Eur J Endocrinol 2024. 191 R1–R8. ( 10.1093/ejendo/lvae065) [DOI] [PubMed] [Google Scholar]
- 6.Clemmons DR & Bidlingmaier M. Interpreting growth hormone and IGF-I results using modern assays and reference ranges for the monitoring of treatment effectiveness in acromegaly. Front Endocrinol 2023. 14 1–12. ( 10.3389/fendo.2023.1266339) [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Varewijck AJ, van der Lely AJ, Neggers SJCMM, et al. Disagreement in normative IGF-I levels may lead to different clinical interpretations and GH dose adjustments in GH deficiency. Clin Endocrinol 2018. 88 409–414. ( 10.1111/cen.13491) [DOI] [PubMed] [Google Scholar]
- 8.Glińska M, Walczak M, Wikiera B, et al. Difficulties in interpreting IGF-1 levels in short stature children born small for gestational age (SGA) treated with recombinant human growth hormone (rhGH) based on data from six clinical centers in Poland. J Clin Med 2023. 12 4392. ( 10.3390/jcm12134392) [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Pokrajac A, Wark G, Ellis AR, et al. Variation in GH and IGF-I assays limits the applicability of international consensus criteria to local practice. Clin Endocrinol 2007. 67 65–70. ( 10.1111/j.1365-2265.2007.02836.x) [DOI] [PubMed] [Google Scholar]
- 10.Miida T, Hirayama S, Fukushima Y, et al. Harmonization of lipoprotein(a) immunoassays using A serum panel value assigned with the IFCC-endorsed mass spectrometry-based reference measurement procedure as A first step towards apolipoprotein standardization. J Atherosclerosis Thromb 2025. 32 580–595. ( 10.5551/jat.65238) [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Pratt MS, van Faassen M, Remmelts N, et al. An antibody-free LC-MS/MS method for the quantification of intact insulin-like growth factors 1 and 2 in human plasma. Anal Bioanal Chem 2021. 413 2035–2044. ( 10.1007/s00216-021-03185-y) [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Zeng HL, Lu J, Li H, et al. A validated liquid chromatography-tandem mass spectrometry assay for simultaneous quantitation of intact IGF-1 and IGF-2 in human serum. J Chromatogr B 2025. 1257 124572. ( 10.1016/j.jchromb.2025.124572) [DOI] [PubMed] [Google Scholar]
- 13.Tanna NN, Lame ME & Wrona M. Development of an UPLC/MS-MS method for quantification of intact IGF-I from human serum. Bioanalysis 2020. 12 53–65. ( 10.4155/bio-2019-0234) [DOI] [PubMed] [Google Scholar]
- 14.Thevis M, Bredehöft M, Kohler M, et al. Mass spectrometry-based analysis of IGF-1 and hGH. In Handbook of Experimental Pharmacology, pp 201–207. Springer International, 2010. ( 10.1007/978-3-540-79088-4_9) [DOI] [PubMed] [Google Scholar]
- 15.Coppieters G, Judák P, Van Haecke N, et al. A high-throughput assay for the quantification of intact insulin-like growth factor I in human serum using online SPE-LC-HRMS. Clin Chim Acta 2020. 510 391–399. ( 10.1016/j.cca.2020.07.054) [DOI] [PubMed] [Google Scholar]
- 16.Bystrom CE, Sheng S & Clarke NJ. Narrow mass extraction of time-of-flight data for quantitative analysis of proteins: determination of insulin-like growth factor-1. Anal Chem 2011. 83 9005–9010. ( 10.1021/ac201800g) [DOI] [PubMed] [Google Scholar]
- 17.Hines J, Milosevic D, Ketha H, et al. Detection of IGF-1 protein variants by use of LC-MS with high-resolution accurate mass in routine clinical analysis. Clin Chem 2015. 61 990–991. ( 10.1373/CLINCHEM.2014.234799) [DOI] [PubMed] [Google Scholar]
- 18.Bronsema KJ, Klont F, Schalk FB, et al. A quantitative LC-MS/MS method for insulin-like growth factor 1 in human plasma. Clin Chem Lab Med 2018. 56 1905–1912. ( 10.1515/cclm-2017-1042) [DOI] [PubMed] [Google Scholar]
- 19.Cox HD, Lopes F, Woldemariam GA, et al. Interlaboratory agreement of insulin-like growth factor 1 concentrations measured by mass spectrometry. Clin Chem 2014. 60 541–548. ( 10.1373/clinchem.2013.208538) [DOI] [PubMed] [Google Scholar]
- 20.Ross HA, Lentjes EWGM, Menheere PMM, et al. Harmonization of growth hormone measurement results: the empirical approach. Clin Chim Acta 2014. 432 72–76. ( 10.1016/j.cca.2014.01.008) [DOI] [PubMed] [Google Scholar]
- 21.Sijtsma A, Rienks J, van der Harst P, et al. Cohort profile update: lifelines, a three-generation cohort study and Biobank. Int J Epidemiol 2022. 51 E295–E302. ( 10.1093/ije/dyab257) [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Nilsson G, Budd JR, Greenberg N, et al. IFCC working group recommendations for assessing commutability part 2: using the difference in bias between a reference material and clinical samples. Clin Chem 2018. 64 455–464. ( 10.1373/clinchem.2017.277541) [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Cole TJ & Green PJ. Smoothing reference centile curves: the lms method and penalized likelihood. Stat Med 1992. 11 1305–1319. ( 10.1002/sim.4780111005) [DOI] [PubMed] [Google Scholar]
- 24.Cole TJ. The LMS method for constructing normalized growth standards. Eur J Clin Nutr 1990. 44 45–60. [PubMed] [Google Scholar]
- 25.Lentjes E & Vos MJ. IGF-1 LC-MS/MS Normative Datasets. Zenodo, 2025. ( 10.5281/zenodo.15799741) [DOI] [Google Scholar]
- 26.Lee JKY, Cradic K, Singh RJ, et al. Discordance of insulin-like growth factor-1 results and interpretation on four different platforms. Clin Chim Acta 2023. 539 130–133. ( 10.1016/j.cca.2022.11.034) [DOI] [PubMed] [Google Scholar]
- 27.Chanson P, Arnoux A, Mavromati M, et al. Reference values for IGF-I serum concentrations: comparison of six immunoassays. J Clin Endocrinol Metab 2016. 101 3450–3458. ( 10.1210/jc.2016-1257) [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Mavromati M, Kuhn E, Agostini H, et al. Classification of patients with GH disorders may vary according to the IGF-I assay. J Clin Endocrinol Metab 2017. 102 2844–2852. ( 10.1210/jc.2017-00202) [DOI] [PubMed] [Google Scholar]
- 29.Molitch ME, Clemmons DR, Malozowski S, et al. Evaluation and treatment of adult growth hormone deficiency: An Endocrine Society Clinical Practice Guideline. J Clin Endocrinol Metab 2011. 96 1587–1609. ( 10.1210/jc.2011-0179) [DOI] [PubMed] [Google Scholar]
- 30.Wauthier L, Plebani M & Favresse J. Interferences in immunoassays: review and practical algorithm. Clin Chem Lab Med 2022. 60 808–820. ( 10.1515/cclm-2021-1288) [DOI] [PubMed] [Google Scholar]

This work is licensed under a 