ABSTRACT
Background
Studies show that foetal and birthweight‐for‐gestational age centiles are poor predictors of serious neonatal morbidity and neonatal mortality (SNMM) in univariable models.
Objective
We assessed the predictive performance of multivariable SNMM models based on maternal/pregnancy characteristics, with and without birthweight centiles.
Methods
The study was based on all live births in the United States, 2019–2021, with data obtained from the period live birth‐infant death files of the National Center for Health Statistics. SNMM was defined as any one or more of the following: 5‐minute Apgar score < 4, seizures, assisted ventilation for> 30 or neonatal death. SNMM was modelled by log‐linear regression on maternal/pregnancy characteristics as predictors, with and without birthweight centiles. Models were developed for live births at 24–42 weeks' and 39 weeks' gestation to all women and those with hypertensive disorders or pre‐existing diabetes. Model performance was assessed using area under the curve (AUC).
Results
The study population included 10,487,243 live births and 221,728 SNMM cases (2.1 per 100 live births). The models with all live births at 24–42 weeks' gestation had AUCs of 0.83 (95% confidence interval [CI] 0.82, 0.83) based on maternal/pregnancy characteristics and 0.83 (95% CI 0.83, 0.84) based on maternal/pregnancy characteristics and birthweight centiles. However, AUCs of models based on all live births at 39 weeks' gestation were 0.66 (95% CI 0.64, 0.68) with maternal/pregnancy characteristics and 0.69 (95% CI 0.68, 0.71) with maternal/pregnancy characteristics and birthweight centiles. AUCs of the models with live births at 39 weeks' gestation to women with pre‐existing diabetes were 0.69 (95% CI 0.66, 0.72) based on maternal/pregnancy characteristics, and 0.77 (95% CI 0.74, 0.79) with the addition of birthweight centiles.
Conclusions
Birthweight centiles improve multivariable SNMM predictive performance in specific subpopulations, although evaluation of decision thresholds is required to determine the clinical importance of improvement in predictive ability.
Keywords: area under the curve, birthweight, birthweight‐for‐gestational age, neonatal mortality, prediction, serious neonatal morbidity
1. Background
Birthweight‐for‐gestational age centiles (birthweight centiles) quantify the weight of newborn infants in relation to pregnancy duration and reflect average foetal growth in utero relative to a standard or reference population. Such centiles are routinely used to categorise newborns into small‐for‐gestational‐age (SGA), appropriate‐for‐gestational‐age (AGA) and large‐for‐gestational‐age (LGA) groups, typically based on the 10th and 90th centiles of a standard or reference. These three growth‐status categories are used in clinical practice and perinatal research, with SGA, in particular, viewed as an important risk factor for adverse perinatal outcomes [1]. Despite the widespread use of these growth categories for assessing health status and prognosis in foetuses and newborns, it is understood that SGA categorisation (including SGA categorisation that is differentiated into preterm and term subtypes) [2] fails to distinguish constitutionally small foetuses/infants from those with pathological growth restriction [3, 4, 5]. Also, the majority of adverse perinatal outcomes occur in the AGA group due to the 8‐fold larger size of the AGA group relative to the SGA group [6]. Another problem with such SGA/AGA/LGA categorisation is the lack of consensus on the reference/standard population used for comparison [7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21].
Several recent studies have assessed the performance of continuous foetal and birthweight centiles, and of SGA/LGA categories, for predicting severe neonatal morbidity and neonatal mortality (SNMM). These studies report that the indices perform poorly, irrespective of the reference/standard used for estimating the continuous centiles and the cut‐off used to identify SGA/LGA categories [7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21]. One study that used three different estimated foetal weight‐for‐gestational age and three different birthweight‐for‐gestational age references and standards for identifying SGA observed ‘marked variation in classification and… similarly poor performance’ [20]. Our previous study, which corroborated those findings, showed that identifying newborn infants as SGA does not substantially enhance clinical or population‐level prediction of adverse neonatal outcomes [22]. Not surprisingly, these and other negative assessments have led some experts to call for abandoning the SGA concept altogether [23].
The above‐mentioned critical perspective on foetal growth indices is based on univariable models, which employed SGA (or birthweight centiles) as the sole predictor of adverse perinatal outcomes. We hypothesised that SGA/LGA categorisation and birthweight centiles could provide independent information for predicting adverse perinatal outcomes when used in multivariable models. We, therefore, carried out a study to assess the predictive performance of a multivariable model with maternal and pregnancy characteristics, with and without birthweight centiles as predictors of SNMM.
2. Methods
2.1. Data Source, Inclusion and Exclusion Criteria
The study was based on all live births in the United States from 2019 to 2021, with data obtained from the period‐linked birth‐infant death files of the National Center for Health Statistics, Centers for Disease Control and Prevention [24]. The study was restricted to live births with an obstetric estimate of gestational age between 24 and 42 weeks' gestation, as birthweight centiles for births over 42 weeks or under 24 weeks have not been published by the Intergrowth 21st Project or other standards. Live births with a birthweight < 400 g or > 6000 g, and those with implausible combinations of birthweight and gestational age were excluded [25]. Other exclusion criteria included infants with a congenital anomaly and infants who died following an accident or homicide. Live births with missing or unknown values for the variables of interest were also excluded from the study.
2.2. Populations of Interest
The study was carried out in two populations, namely, all live births at 24–42 weeks' gestation and all live births at 39 weeks' gestation. The latter restriction of the study population to 39 weeks' gestation was motivated by two concerns. First, this restriction was intended to eliminate gestational age as a predictor from the multivariable model, as gestational duration has been shown to be a strong predictor of SNMM [22]. Secondly, we focused on a single gestational week in the third trimester, as there is increasing clinical interest in third‐trimester ultrasonographic foetal health assessment at specific gestational weeks [26, 27, 28]. Three populations were considered for model development and included (i) all live births; (ii) live births to women with hypertensive disorders of pregnancy (including chronic hypertension, gestational hypertension and eclampsia); and (iii) live births to women with pre‐pregnancy diabetes. All available live births (in the particular population) from 2019 to 2021 were used to develop models for the infants of hypertensive and diabetic mothers, while 3% of all live births in 2019 to 2021 were randomly selected to develop the models for the ‘all live births’ population.
2.3. Outcome Variable
The outcome variable of interest was composite SNMM and included infants with any one of the following: 5‐min Apgar score less than 4, need for assisted ventilation for greater than 30 min, neonatal seizures or neonatal death (death within the first 28 days after birth). Our previous studies have also used a similar composite outcome to identify serious neonatal morbidity and neonatal mortality [6, 7, 22].
2.4. Statistical Analysis
We fit univariable models to quantify the discriminative power of each predictor—birthweight centiles and maternal/pregnancy characteristics [29]. Next, we fit multivariable models with maternal and pregnancy characteristics. Finally, we modelled maternal/pregnancy characteristics along with birthweight centiles and interaction terms [29]. This sequential approach followed best practices in prognostic research by identifying the baseline predictive ability of individual variables, evaluating their combined performance and ensuring that any added complexity yielded a meaningful improvement in risk stratification [30].
The multivariable prediction model was developed using log‐linear regression, with SNMM as the dependent variable and maternal/pregnancy characteristics and birthweight centiles as independent variables. Independent predictors were chosen based on the literature, clinical relevance, availability in the data source and predictor‐outcome associations observed in univariable log‐linear models. These included maternal age, parity, race/ethnicity, cigarette smoking during pregnancy, body mass index, chronic hypertension, pre‐pregnancy diabetes, plurality, use of assisted reproductive technology, infections during pregnancy, route of delivery and gestational age (among others). The list of the final 28 predictors used in the model is shown in Table 1; predictor values and reference values for parameterisation are shown in Table S1. Gestational age and birthweight centiles were represented as categorical variables (completed single week of gestation and single‐unit centile groups) in the final model. The predictor variables and all their first‐order interactions were used to create a parsimonious prediction equation based on standard backward selection methods for dealing with a large number of potential predictors [31]. The final model was evaluated and contrasted against a model without birthweight centiles added as a predictor.
TABLE 1.
Results of log‐linear regression showing the univariate relationships between predictors and serious neonatal morbidity and neonatal mortality (SNMM), live births at 24–42 weeks' gestation, United States, 2019–2021.
| Predictor | No SNMM (n = 10,265,515) % | SNMM (n = 221,728) % | Rate ratio (95% CI) | Area under the curve (95% CI) |
|---|---|---|---|---|
| Maternal age (years) | ||||
| < 20 | 4.4 | 4.7 | 1.09 (1.07, 1.12) | 0.51 (0.51, 0.52) |
| 20–24 | 18.4 | 18.2 | 1.00 (Reference) | |
| 25–29 | 28.5 | 27.4 | 0.97 (0.96, 0.98) | |
| 30–34 | 29.8 | 28.8 | 0.97 (0.96, 0.99) | |
| 35–39 | 15.6 | 16.4 | 1.06 (1.05, 1.08) | |
| ≥ 40 | 3.5 | 4.5 | 1.28 (1.25, 1.31) | |
| Race/ethnicity | ||||
| Non‐Hispanic White | 51.8 | 54.0 | 1.00 (Reference) | 0.56 (0.56, 0.56) |
| Non‐Hispanic Black | 14.3 | 19.7 | 1.31 (1.30, 1.33) | |
| Non‐Hispanic AIAN | 0.7 | 1.0 | 1.28 (1.23, 1.34) | |
| Non‐Hispanic Asian | 6.2 | 4.2 | 0.64 (0.63, 0.66) | |
| Non‐Hispanic NHOPI | 0.2 | 0.3 | 1.20 (1.11, 1.29) | |
| Non‐Hispanic other | 2.3 | 2.7 | 1.12 (1.09, 1.15) | |
| Hispanic | 24.4 | 18.1 | 0.72 (0.71, 0.73) | |
| Parity | ||||
| 0 | 38.6 | 41.3 | 1.00 (Reference) | 0.54 (0.53, 0.54) |
| 1 | 31.8 | 26.9 | 0.79 (0.79, 0.80) | |
| 2 | 16.9 | 16.2 | 0.89 (0.88, 0.91) | |
| ≥ 3 | 12.6 | 15.6 | 1.15 (1.13, 1.16) | |
| Smoking | 5.2 | 9.0 | 1.77 (1.75, 1.80) | 0.52 (0.52, 0.52) |
| Pre‐pregnancy BMI (kg/m2) | ||||
| < 18.5 | 2.8 | 2.7 | 1.15 (1.12, 1.18) | 0.55 (0.55, 0.56) |
| 18.5–24.9 | 40.2 | 33.4 | 1.00 (Reference) | |
| 25.0–29.9 | 27.2 | 25.5 | 1.12 (1.11, 1.14) | |
| 30.0–34.9 | 16.0 | 17.9 | 1.34 (1.32, 1.35) | |
| 35.0–39.9 | 8.0 | 10.6 | 1.58 (1.56, 1.60) | |
| ≥ 40.0 | 5.8 | 9.8 | 2.01 (1.98, 2.04) | |
| Previous preterm birth | 3.5 | 10.3 | 3.01 (2.97, 3.06) | 0.53 (0.53, 0.54) |
| Previous caesareans | ||||
| 0 | 84.7 | 78.5 | 1.00 (Reference) | 0.53 (0.53, 0.53) |
| 1 | 10.5 | 13.5 | 1.38 (1.36, 1.39) | |
| ≥ 2 | 4.8 | 8.0 | 1.77 (1.74, 1.79) | |
| Pre‐pregnancy diabetes | 1.0 | 3.6 | 3.63 (3.55, 3.71) | 0.51 (0.51, 0.51) |
| Gestational diabetes | 7.6 | 10.5 | 1.40 (1.39, 1.42) | 0.51 (0.51, 0.52) |
| Chronic hypertension | 2.4 | 7.2 | 3.07 (3.03, 3.12) | 0.52 (0.52, 0.53) |
| Pregnancy hypertension | 8.5 | 19.3 | 2.51 (2.48, 2.54) | 0.55 (0.55, 0.56) |
| Infertility treatment | 1.9 | 4.7 | 2.45 (2.40, 2.50) | 0.51 (0.51, 0.51) |
| Assisted rep. technology | 1.3 | 3.2 | 2.43 (2.38, 2.49) | 0.51 (0.51, 0.51) |
| Infections during pregnancy | ||||
| Gonorrhoea | 0.3 | 0.6 | 1.78 (1.69, 1.88) | 0.50 (0.50, 0.50) |
| Syphilis | 0.2 | 0.4 | 2.07 (1.93, 2.22) | 0.50 (0.50, 0.50) |
| Chlamydia | 1.8 | 0.6 | 1.37 (1.33, 1.41) | 0.50 (0.50, 0.50) |
| Hepatitis C | 0.5 | 1.2 | 2.63 (2.53, 2.73) | 0.50 (0.50, 0.51) |
| Labour induction | 31.6 | 21.0 | 0.58 (0.58, 0.59) | 0.55 (0.55, 0.55) |
| Labour augmentation | 22.1 | 14.6 | 0.61 (0.60, 0.61) | 0.54 (0.54, 0.54) |
| Antenatal steroids | 3.2 | 30.7 | 11.4 (11.3, 11.5) | 0.64 (0.64, 0.64) |
| Antibiotics | 25.2 | 41.9 | 2.11 (2.09, 2.13) | 0.59 (0.58, 0.59) |
| Chorioamnionitis | 1.5 | 3.6 | 2.36 (2.31, 2.42) | 0.51 (0.51, 0.51) |
| Plurality | ||||
| Singleton | 97.1 | 85.3 | 1.00 (Reference) | 0.559 (0.56, 0.56) |
| Twin | 2.8 | 13.7 | 5.07 (5.00, 5.13) | |
| Triplet or higher | 0.1 | 1.0 | 15.2 (14.6, 15.9) | |
| Fetal Presentation | ||||
| Cephalic | 95.5 | 83.1 | 1.00 (Reference) | 0.56 (0.56, 0.56) |
| Breech | 3.7 | 15.0 | 4.33 (4.28, 4.38) | |
| Other | 0.8 | 1.9 | 2.67 (2.59, 2.75) | |
| Delivery type | ||||
| Spontaneous | 65.8 | 35.3 | 1.00 (Reference) | 0.66 (0.66, 0.66) |
| Forceps | 0.5 | 0.6 | 2.16 (2.04, 2.28) | |
| Vacuum | 2.6 | 2.0 | 1.47 (1.43, 1.52) | |
| Caesarean | 31.1 | 62.1 | 3.60 (3.57, 3.63) | |
| Gestational age: 24–42 weeks | 39 a | 34 a | 0.72 (0.72, 0.72) | 0.79 (0.79, 0.79) |
| Infant sex | ||||
| Female | 49.0 | 43.6 | 1.00 (Reference) | 0.53 (0.53, 0.53) |
| Male | 51.0 | 56.4 | 1.24 (1.23, 1.25) | |
| Birthweight centile: 0–100th | 58 a | 54 a | 0.99 (0.99, 0.99) | 0.60 (0.60, 0.61) |
Note: Reference groups not shown for binary variables. Rate ratios for continuous variables represent the relative changes in the rate of SNMM per unit increase in the independent variable (1 week increase for gestational age and 1 centile increase for birthweight centiles). Parity was defined in relation to pregnancy and calculated as live birth order minus one. Thus parity ‘0’ refers to the women who had not delivered a live birth prior to this pregnancy; parity ‘1’ refers to the women who had delivered 1 live birth prior to this pregnancy, etc.
Abbreviations: AIAN, American Indian or Alaska Native; CI, confidence interval; NHOPI, Native Hawaiian or Other Pacific Islander.
Mean gestational age and mean birthweight centile.
2.5. Evaluation of Predictive Performance and Model Validation
The predictive performance of the model was evaluated based on its ability to correctly distinguish between positive and negative SNMM cases (specifically, to the likelihood that the model assigned a higher predicted probability to a randomly chosen positive case than to a randomly chosen negative case). This was quantified using the area under the receiver operating characteristic (ROC) curve (AUC), which is identical to using the concordance statistic (c‐index) for models with binary outcomes [32, 33, 34]. The performance of the model with regard to calibration (i.e., the agreement between the observed outcomes and the predictions made by the model, or alternatively, the accuracy of the model's predicted probabilities compared with the observed probabilities) was visualised using calibration plots [35]. The Brier score was used as a measure of the overall accuracy of the model, as it simultaneously assesses discrimination and calibration. It takes a value between zero and one (the square of the largest possible difference between a predicted probability and the actual [0/1] outcome) [36, 37]. The performance of the models with and without birthweight centiles was then contrasted, primarily with respect to discrimination based on the AUC, to determine if the addition of birthweight centiles improved predictive performance [38].
Model validation was carried out using similarly processed data on live births in the United States from 2017 to 2018. Analysis was carried out using SAS 9.4 software (SAS Institute, Cary, North Carolina, USA).
2.6. Missing Data
Of the 11,054,903 live births, 470,398 (4.3%) were found to have missing or unknown values; multiple imputation was deemed unnecessary as the proportion of missing values was less than 5%, and these records were excluded from the analysis. 1.8%, 0.9%, 0.4%, 0.4%, 0.2%, 0.2%, 0.1%, 0.1% and 0.1% of all live births had missing or unknown values for maternal body mass index, maternal race, APGAR score, smoking during pregnancy, parity, gonorrhoea, assisted ventilation > 6 h, fertility enhancing drugs and foetal presentation, respectively. Each of the other predictors was missing information for less than 0.05% of all live births.
2.7. Ethics Approval
Ethics approval for the study was not sought, as the data used were from an anonymised, publicly available data source.
3. Results
The study population included 10,487,243 live births after exclusions based on gestational age, birthweight, presence of congenital anomalies and specific causes of death (Figure 1). There were 221,728 SNMM cases (2.1%): 39,914 had a 5‐min Apgar score < 4, 175,426 received assisted ventilation for greater than 30 min, 3255 had neonatal seizures and 20,562 died in the neonatal period.
FIGURE 1.

The exclusions applied to all live births from 2019 to 2021 in the United States to create the study population to develop prognostic prediction equations, obtained from the period‐linked birth‐infant death files of the National Center for Health Statistics.
Table 1 shows the distribution of maternal and pregnancy characteristics in the study population by SNMM status and the univariable relationships between these predictor variables and SNMM. The strongest univariable relationships (as assessed by rate ratios) were seen for previous preterm birth, pre‐pregnancy diabetes, chronic hypertension, antenatal corticosteroid therapy, twin or triplet live birth, breech presentation, caesarean delivery and gestational age. However, the ability of these predictors to discriminate between the presence and absence of SNMM was modest, except for antenatal corticosteroid therapy, method/route of delivery and gestational age (Table 1). Birthweight centiles were weakly associated with SNMM and showed poor ability to discriminate between the presence and absence of SNMM.
The multivariable models that were developed for each population are shown in Tables S3 and S4. The AUC for the multivariable model, including all live births at 24–42 weeks' gestation and based on all maternal/pregnancy predictors, was 0.83 (95% CI 0.82, 0.83), and this increased minimally when birthweight centiles were added to the model (Figure 2). Calibration, as assessed by the visual inspection of calibration plots, and accuracy, as assessed with Brier Scores, showed similar performance for the two models (Table 2; Figure S1 shows the calibration curves). Restricting the population of live births to those born at 39 weeks' gestation reduced the AUC of the models: the AUC for the univariable model with birth weight centiles alone was 0.62 (95% CI 0.60, 0.64), while that of the model with maternal/pregnancy characteristics was 0.66 (95% CI 0.64, 0.68). However, the AUC of this model was markedly improved to 0.69 (95% CI 0.68, 0.71) following the addition of birthweight centiles (Figure 2). The calibration curves showed good calibration for both the latter models, and both models had low Brier scores, indicating similar accuracy (Table 2; Figure S1 shows the calibration curves, which end abruptly as there are no observed probabilities greater than 0.3).
FIGURE 2.

The area under the receiver operating characteristics curve (AUC) for models that include maternal and pregnancy characteristics, both with and without birthweight centiles, as well as for the model featuring birthweight centiles alone for all live births at 24–42 weeks' (A) and at 39 weeks' gestation (B), United States, 2019–2021.
TABLE 2.
Area under the receiver operating characteristics curve (AUC), with 95% confidence intervals (95% CI), and Brier scores for each model.
| Population (Gestational age) | Predictors | AUC (95% CI) | Brier score |
|---|---|---|---|
| All live births | |||
| 24–42 weeks | Centiles only | 0.60 (0.59, 0.61) | 0.02 |
| 24–42 weeks | Maternal/pregnancy characteristics | 0.83 (0.82, 0.83) | 0.02 |
| 24–42 weeks | Mat/pregnancy characteristics + centiles | 0.83 (0.83, 0.84) | 0.02 |
| 39 weeks | Centiles only | 0.62 (0.60, 0.64) | 0.01 |
| 39 weeks | Maternal/pregnancy characteristics | 0.66 (0.64, 0.68) | 0.01 |
| 39 weeks | Mat/pregnancy characteristics + centiles | 0.69 (0.68, 0.71) | 0.01 |
| Live births to women with hypertensive disorders | |||
| 24–42 weeks | Centiles only | 0.62 (0.61, 0.62) | 0.05 |
| 24–42 weeks | Maternal/pregnancy characteristics | 0.85 (0.84, 0.85) | 0.04 |
| 24–42 weeks | Mat/pregnancy characteristics + centiles | 0.85 (0.85, 0.86) | 0.04 |
| 39 weeks | Centiles only | 0.58 (0.57, 0.59) | 0.01 |
| 39 weeks | Maternal/pregnancy characteristics | 0.68 (0.67, 0.69) | 0.01 |
| 39 weeks | Mat/pregnancy characteristics + centiles | 0.69 (0.69, 0.70) | 0.01 |
| Live births to women with pre‐pregnancy diabetes | |||
| 24–42 weeks | Centiles only | 0.59 (0.58, 0.60) | 0.09 |
| 24–42 weeks | Maternal/pregnancy characteristics | 0.83 (0.83, 0.84) | 0.06 |
| 24–42 weeks | Mat/pregnancy characteristics + centiles | 0.84 (0.83, 0.84) | 0.06 |
| 39 weeks | Centiles only | 0.71 (0.68, 0.74) | 0.01 |
| 39 weeks | Maternal/pregnancy characteristics | 0.69 (0.66, 0.72) | 0.01 |
| 39 weeks | Mat/pregnancy characteristics + centiles | 0.77 (0.74, 0.79) | 0.01 |
Among live births to women with hypertensive disorders at 24–42 weeks' gestation, the model AUC was 0.62 (95% CI 0.61, 0.62) when birthweight centiles were the sole predictors, the AUC was 0.85 (95% CI 0.84, 0.85) when maternal/pregnancy characteristics were included in the model, and the AUC was marginally higher when birthweight centiles were added to the model with maternal/pregnancy characteristics (Figure 3). Both the latter models showed good calibration as assessed by the smoothness of the calibration curves and accuracy as assessed by Brier scores (Table 2; Figure S2 shows the calibration curves). The model for live births at 39 weeks' gestation had a low AUC based on birthweight centiles alone, a higher AUC based on maternal/pregnancy characteristics and a marginally higher AUC when birthweight centiles were added to the latter model. Calibration and accuracy based on Brier scores were comparable across these models without and with birthweight centiles (Table 2; Figure S2 shows the calibration curves, which end abruptly as there are no observed probabilities greater than 0.4).
FIGURE 3.

The area under the receiver operating characteristics curve (AUC) for models that include maternal and pregnancy characteristics, both with and without birthweight centiles, as well as for the model featuring birthweight centiles alone for live births to women with hypertensive disorders of pregnancy at 24–42 weeks' (A) and at 39 weeks' gestation (B), United States, 2019–2021.
The models for live births to women with pre‐pregnancy diabetes at 24–42 weeks' gestation were similar to those for live births to women with hypertension: the AUC of the model based on birthweight centiles alone was low, the AUC based on maternal/pregnancy characteristics was higher, and this was slightly improved following the addition of birthweight centiles (Table 2; Figure S3 shows the calibration curves). The inspection of the calibration plots for models with maternal/pregnancy characteristics (with and without birthweight centiles) indicated a problem with the fit due to their lack of smoothness, while the Brier scores for the models with and without birthweight centiles showed similar prediction accuracy. The model for live births at 39 weeks' gestation to women with pre‐pregnancy diabetes had an AUC of 0.71 (95% CI 0.68, 0.74) when based on birthweight centiles alone, an AUC of 0.69 (95% CI 0.66, 0.72) when based on maternal/pregnancy characteristics (Figure 4) and a substantially higher AUC of 0.77 (95% CI 0.74, 0.79) following the addition of birthweight centiles to the model (Table 2). Figure S3 shows the calibration curves, which end abruptly as there are no observed probabilities greater than 0.5. Calibration and accuracy of the two models were similar.
FIGURE 4.

The area under the receiver operating characteristics curve (AUC) for models that include maternal and pregnancy characteristics, both with and without birthweight centiles, as well as for the model featuring birthweight centiles alone for live births to women with pre‐pregnancy diabetes mellitus at 24–42 weeks' (A) and at 39 weeks' gestation (B), United States, 2019–2021.
Model validation showed minimal model optimism (i.e., minimal overestimation of a model's performance due to overfitting) for most models. For instance, the validation model for all live births at 24–42 weeks' gestation (including maternal/pregnancy characteristics and birthweight centiles) had an AUC of 0.83 (95% CI 0.83, 0.83) compared with the original AUC of 0.83 (95% CI 0.83, 0.84). However, validation showed a small degree of optimism for the model with all live births at 39 weeks' gestation. Table S2 shows the results of validation for the different models.
4. Comment
4.1. Principal Findings
Our results show that maternal and pregnancy characteristics yielded a multivariable prediction model for SNMM with high discriminative ability for live births at 24–42 weeks' gestation. Although this model's discrimination was far superior to that of birthweight centiles in a univariable model, the model's AUC was not meaningfully improved by the additional inclusion of birthweight centiles. On the other hand, models restricted to live births at 39 weeks' gestation and based on maternal and pregnancy characteristics showed a relatively modest discriminative ability. However, the addition of birth weight centiles to this model with maternal and pregnancy characteristics considerably improved the model's ability to discriminate. Similarly, the prediction ability of models based on live births at 39 weeks' gestation to women with pre‐existing diabetes increased when birthweight centiles were included in the model (compared with a model with only maternal and pregnancy characteristics). Model validation carried out on other comparable datasets generally showed minimal optimism (except for the pre‐pregnancy diabetes subgroup, among whom the higher degree of optimism likely indicates overfitting).
4.2. Strengths of the Study
The strengths of our study include its population‐based provenance, large size and the relatively detailed information on maternal and pregnancy characteristics. The large sample size enabled stratified analysis focused on high‐risk populations at a specific gestational week.
4.3. Limitations of the Data
Limitations of our study include the potential for errors in the data source, given the routine data collection methods. However, validation studies show that the information in the database is relatively accurate, especially for important variables, such as birthweight [39, 40]. However, studies also show that the accuracy of information varies considerably for specific predictors, ranging from highly accurate (e.g., sensitivity of information on parity 93%–98%; specificity 98%–99%) to less accurate (e.g., sensitivity of information on any hypertension 39%–76%; specificity 99%) [39, 40]. Such errors in the data are likely to have reduced the accuracy of our models, although they are less likely to alter the predictive power of birth weight centiles. We did not include some variables with a high proportion of missing values as predictors (e.g., weight gain in pregnancy and timing of prenatal care). Certain clinically important variables and their interaction terms may have been removed by the backward selection process used for model development, or liberal use of interaction terms and backward selection could have led to model overfitting [31, 41]. However, external validation carried out using a different data set showed minimal optimism. On the other hand, it is possible that alternative methodologies, such as machine learning techniques, may have identified more complex non‐linear relationships between predictors and SNMM that could have improved the predictive ability of the models. Such techniques may be especially useful in other settings, for instance, in the antenatal period, where models predicting serious neonatal morbidity and perinatal death could be based on detailed ultrasound imaging information. Finally, we did not attempt to optimise sensitivity, specificity, predictive values and other classification measures, as our purpose was to assess the utility of birthweight centiles in multivariable models (rather than to develop models intended for clinical use) [30].
4.4. Interpretation
Our study confirmed the strong discriminant ability of gestational age as a predictor of SNMM [22]. In our study, we built two models: the first model, which was based on live births at 24–42 weeks and included gestational age as a predictor, showed excellent discrimination, while the second model was restricted to live births at 39 weeks' gestation and showed a lesser ability to discriminate. However, the latter model is particularly relevant from a clinical standpoint, as foetal ultrasonography at specific gestational weeks in the third trimester is increasingly a focus of clinical interest [28, 29, 30]. Although we did not hypothesise a priori that birthweight centiles would contribute differentially to models based on live births at 24–42 weeks vs. 39 weeks' gestation, the improvement in AUC among all live births at 39 weeks, and especially among live births to diabetic women at 39 weeks' gestation, likely implies that foetal growth indices may have greater predictive utility in specific subpopulations.
Traditionally, gestational duration and foetal growth have been regarded as the primary determinants of neonatal health status [42]. This understanding regarding the pathways to adverse perinatal outcomes has led to a focus on preterm birth and foetal growth restriction from the standpoint of preventive intervention, and also for identifying women at high risk of SNMM. Nevertheless, operationalisation of the foetal growth restriction concept and its diagnosis presents significant challenges. Identification of SGA foetuses in utero is affected by measurement error, and as mentioned, categorisation of foetuses and newborns into SGA, AGA and LGA categories cannot distinguish the pathologically growth‐restricted fetus/infant from the constitutionally small fetus/infant [2, 3, 4]. Whereas foetal growth restriction is pathologic by definition and constitutes an ‘illness’, its diagnostic proxy, SGA, can only be considered a risk factor for illness and not an illness per se.
The prediction of SNMM based on maternal and pregnancy characteristics (as carried out in our study) has little clinical utility in the newborn period, as prognostication setting in this context can be substantially improved through the use of easily available information on various physiologic indices and early signs of morbidity (e.g., changes in breathing, feeding, hypothermia and skin colour). Nevertheless, our study is important because it demonstrates, in principle, the potential utility of maternal/pregnancy characteristics and weight centiles for the antenatal assessment of foetal health status and the antenatal prediction of SNMM, at least in some subpopulations. Much of the reduction in adverse perinatal outcomes in modern obstetrics has been achieved through the carefully timed early delivery of foetuses diagnosed as compromised in utero. Improving the prediction of in utero foetal compromise is critical and requires the development and validation of multivariable prediction models that could potentially include foetal growth measures such as foetal weight for age centiles. However, a few significant differences between antenatal prediction of SNMM and prediction in the newborn period need to be noted. Although antenatal prediction of SNMM cannot benefit from information on gestational age at birth (unless delivery is iatrogenic), incorporation of antenatal predictors, such as cervical length and Bishop's score, could serve to improve antenatal prediction. Additionally, model prediction at any gestational week in the third trimester (e.g., 39 weeks) would have to be based on all pregnancies delivering at that gestational week and beyond.
5. Conclusions
Weight‐for‐gestational age centiles have the potential to improve the prediction of SNMM when used in multivariable prediction equations along with other predictors such as maternal and pregnancy characteristics. The incorporation of birthweight centiles in such predictive models appears to be especially useful for predicting SNMM in specific high‐risk pregnant subpopulations, such as those with pre‐pregnancy diabetes. However, further assessments of their clinical impact and cost–benefit need to be conducted before weight for gestational age centiles can be incorporated into predictive models and medical decision‐making.
Author Contributions
K.S.J. proposed the study, and S.J. carried out the preliminary analysis. All authors reviewed the preliminary results and provided critical feedback. S.J. carried out the final analysis, S.J. and K.S.J. drafted the manuscript and all authors reviewed the manuscript for intellectual content. The final version of the manuscript was approved by all authors.
Conflicts of Interest
The authors declare no conflicts of interest.
Supporting information
Appendix S1: ppe70065‐sup‐0001‐AppendixS1.docx.
Acknowledgements
K.S.J.'s work was supported by an Investigator award from the BC Children's Hospital Research Institute.
John S., Joseph K. S., Fahey J., Liu S., Lisonkova S., and Kramer M. S., “Do Birthweight‐For‐Gestational Age Centiles Predict Serious Neonatal Morbidity and Neonatal Mortality?,” Paediatric and Perinatal Epidemiology 40, no. 3 (2026): 400–410, 10.1111/ppe.70065.
Funding: This study was funded by the Canadian Institutes of Health Research (MOP‐67125 and PJT153439). The funder played no role in the design and conduct of the research and was not involved in the writing of the paper.
Editor's note: A commentary based on this article appears on pages 411–413.
Data Availability Statement
The data used in this study are publicly available at https://www.cdc.gov/nchs/data_access/vitalstatsonline.htm.
References
- 1. Kristensen S., Salihu H. M., Keith L. G., Kirby R. S., Fowler K. B., and Pass M. A., “SGA Subtypes and Mortality Risk Among Singleton Births,” Early Human Development 83 (2007): 99–105. [DOI] [PubMed] [Google Scholar]
- 2. Lawn J. E., Ohuma E. O., Bradley E., et al., “Small Babies, Big Risks: Global Estimates of Prevalence and Mortality for Vulnerable Newborns to Accelerate Change and Improve Counting,” Lancet 401 (2023): 1707–1719. [DOI] [PubMed] [Google Scholar]
- 3. Mandy G. T., Fetal Growth Restriction (FGR) and Small for Gestational Age (SGA) Newborns (Wolters Kluwer, 2024), https://www.uptodate.com/contents/fetal‐growth‐restriction‐fgr‐and‐small‐for‐gestational‐age‐sga‐newborns#H17. [Google Scholar]
- 4. Morris R. K., Johnstone E., Lees C., Morton V., Smith G., and Royal College of Obstetricians and Gynaecologists , “Investigation and Care of a Small‐For‐Gestational‐Age Fetus and a Growth Restricted Fetus (Green‐Top Guideline No. 31),” BJOG: An International Journal of Obstetrics and Gynaecology 131 (2024): e31–e80. [DOI] [PubMed] [Google Scholar]
- 5. American College of Obstetricians and Gynecologists' Committee on Practice Bulletins‐Obstetrics and the Society for Maternal‐Fetal Medicine , “ACOG Practice Bulletin No. 204: Fetal Growth Restriction,” Obstetrics and Gynecology 133 (2019): e97–e109. [DOI] [PubMed] [Google Scholar]
- 6. John S., Joseph K. S., Fahey J., Liu S., and Kramer M. S., “Counterpoint: Are Abnormal Fetal Growth Indices Valid Predictors of Neonatal Morbidity and Mortality?,” Paediatric and Perinatal Epidemiology 38 (2024): 18–21. [DOI] [PubMed] [Google Scholar]
- 7. Liu S., Metcalfe A., Leon J. A., et al., “Evaluation of the INTERGROWTH‐21st Project Newborn Standard for Use in Canada,” PLoS One 12 (2017): e0172910. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8. Francis A., Hugh O., and Gardosi J., “Customized vs INTERGROWTH‐21(St) Standards for the Assessment of Birthweight and Stillbirth Risk at Term,” American Journal of Obstetrics and Gynecology 218 (2018): S692–S699. [DOI] [PubMed] [Google Scholar]
- 9. Hua X., Shen M., Reddy U. M., et al., “Comparison of the INTERGROWTH‐21st, National Institute of Child Health and Human Development, and WHO Fetal Growth Standards,” International Journal of Gynaecology and Obstetrics 143 (2018): 156–163. [DOI] [PubMed] [Google Scholar]
- 10. Boghossian N. S., Geraci M., Edwards E. M., and Horbar J. D., “Neonatal and Fetal Growth Charts to Identify Preterm Infants <30 Weeks Gestation at Risk of Adverse Outcomes,” American Journal of Obstetrics and Gynecology 219 (2018): 195.e1–195.e14. [DOI] [PubMed] [Google Scholar]
- 11. Lavin T., Nedkoff L., Preen D., Theron G., and Pattinson R. C., “INTERGROWTH‐21st v. Local South African Growth Standards (Theron‐Thompson) for Identification of Small‐For‐Gestational‐Age Fetuses in Stillbirths: A Closer Look at Variation Across Pregnancy,” South African Medical Journal 109 (2019): 519–525. [DOI] [PubMed] [Google Scholar]
- 12. Pritchard N. L., Hiscock R. J., Lockie E., et al., “Identification of the Optimal Growth Charts for Use in a Preterm Population: An Australian State‐Wide Retrospective Cohort Study,” PLoS Medicine 16 (2019): e1002923. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13. Saviron‐Cornudella R., Esteban L. M., Tajada‐Duaso M., et al., “Detection of Adverse Perinatal Outcomes at Term Delivery Using Ultrasound Estimated Percentile Weight at 35 Weeks of Gestation: Comparison of Five Fetal Growth Standards,” Fetal Diagnosis and Therapy 47 (2020): 104–114. [DOI] [PubMed] [Google Scholar]
- 14. Kabiri D., Romero R., Gudicha D. W., et al., “Prediction of Adverse Perinatal Outcome by Fetal Biometry: Comparison of Customized and Population‐Based Standards,” Ultrasound in Obstetrics and Gynecology 55 (2020): 177–188. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15. Pritchard N., Lindquist A., Siqueira I. D. A., Walker S. P., and Permezel M., “INTERGROWTH‐21st Compared With GROW Customized Centiles in the Detection of Adverse Perinatal Outcomes at Term,” Journal of Maternal‐Fetal and Neonatal Medicine 33 (2020): 961–966. [DOI] [PubMed] [Google Scholar]
- 16. Hiersch L., Lipworth H., Kingdom J., Barrett J., and Melamed N., “Identification of the Optimal Growth Chart and Threshold for the Prediction of Antepartum Stillbirth,” Archives of Gynecology and Obstetrics 303 (2021): 381–390. [DOI] [PubMed] [Google Scholar]
- 17. Hocquette A., Durox M., Wood R., et al., “International Versus National Growth Charts for Identifying Small and Large‐For‐Gestational Age Newborns: A Population‐Based Study in 15 European Countries,” Lancet Regional Health Europe 8 (2021): 100167. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18. Monier I., Ego A., Benachi A., et al., “Comparison of the Performance of Estimated Fetal Weight Charts for the Detection of Small‐ and Large‐For‐Gestational Age Newborns With Adverse Outcomes: A French Population‐Based Study,” BJOG: An International Journal of Obstetrics and Gynaecology 129 (2022): 938–948. [DOI] [PubMed] [Google Scholar]
- 19. Fay E., Hugh O., Francis A., et al., “Customized GROW vs INTERGROWTH‐21(St) Birthweight Standards to Identify Small for Gestational Age Associated Perinatal Outcomes at Term,” American Journal of Obstetrics and Gynecology MFM 4 (2022): 100545. [DOI] [PubMed] [Google Scholar]
- 20. Choi S. K. Y., Gordon A., Hilder L., et al., “Performance of Six Birth‐Weight and Estimated‐Fetal‐Weight Standards for Predicting Adverse Perinatal Outcome: A 10‐Year Nationwide Population‐Based Study,” Ultrasound in Obstetrics and Gynecology 58 (2021): 264–277. [DOI] [PubMed] [Google Scholar]
- 21. Liauw J., Mayer C., Albert A., Fernandez A., and Hutcheon J. A., “Which Chart and Which Cut‐Point: Deciding on the INTERGROWTH, World Health Organization, or Hadlock Fetal Growth Chart,” BMC Pregnancy and Childbirth 22 (2022): 25. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22. John S., Joseph K. S., Fahey J., Liu S., and Kramer M. S., “The Clinical Performance and Population Health Impact of Birthweight‐For‐Gestational Age Indices at Term Gestation,” Paediatric and Perinatal Epidemiology 38 (2024): 1–11. [DOI] [PubMed] [Google Scholar]
- 23. Wilcox A. J., Snowden J. M., Ferguson K., Hutcheon J., and Basso O., “On the Study of Fetal Growth Restriction: Time to Abandon SGA,” European Journal of Epidemiology 39 (2024): 233–239. [DOI] [PubMed] [Google Scholar]
- 24. CDC. National Center for Health Statistics , “Vital Statistics Online Data Portal 2024,” cited January 5, 2024, https://www.cdc.gov/nchs/data_access/vitalstatsonline.htm.
- 25. Alexander G. R., Himes J. H., Kaufman R. B., Mor J., and Kogan M., “A United States National Reference for Fetal Growth,” Obstetrics and Gynecology 87 (1996): 163–168. [DOI] [PubMed] [Google Scholar]
- 26. Roma E., Arnau A., Berdala R., Bergos C., Montesinos J., and Figueras F., “Ultrasound Screening for Fetal Growth Restriction at 36 vs 32 Weeks' Gestation: A Randomized Trial (ROUTE),” Ultrasound in Obstetrics and Gynecology 46 (2015): 391–397. [DOI] [PubMed] [Google Scholar]
- 27. Al‐Hafez L., Chauhan S. P., Riegel M., Balogun O. A., Hammad I. A., and Berghella V., “Routine Third‐Trimester Ultrasound in Low‐Risk Pregnancies and Perinatal Death: A Systematic Review and Meta‐Analysis,” American Journal of Obstetrics and Gynecology MFM 2 (2020): 100242. [DOI] [PubMed] [Google Scholar]
- 28. Mustafa H. J., Javinani A., Muralidharan V., and Khalil A., “Diagnostic Performance of 32 vs 36 Weeks Ultrasound in Predicting Late‐Onset Fetal Growth Restriction and Small‐For‐Gestational‐Age Neonates: A Systematic Review and Meta‐Analysis,” American Journal of Obstetrics and Gynecology MFM 6 (2024): 101246. [DOI] [PubMed] [Google Scholar]
- 29. Collins G. S., Reitsma J. B., Altman D. G., and Moons K. G., “Transparent Reporting of a Multivariable Prediction Model for Individual Prognosis or Diagnosis (TRIPOD): The TRIPOD Statement,” BMJ 350 (2015): g7594. [DOI] [PubMed] [Google Scholar]
- 30. Moons K. G., Altman D. G., Reitsma J. B., et al., “Transparent Reporting of a Multivariable Prediction Model for Individual Prognosis or Diagnosis (TRIPOD): Explanation and Elaboration,” Annals of Internal Medicine 162, no. 1 (2015): W1–W73. [DOI] [PubMed] [Google Scholar]
- 31. Royston P., Moons K. G., Altman D. G., and Vergouwe Y., “Prognosis and Prognostic Research: Developing a Prognostic Model,” BMJ 338 (2009): b604. [DOI] [PubMed] [Google Scholar]
- 32. Hanley J. A. and McNeil B. J., “The Meaning and Use of the Area Under a Receiver Operating Characteristic (ROC) Curve,” Radiology 143 (1982): 29–36. [DOI] [PubMed] [Google Scholar]
- 33. Austin P. C. and Steyerberg E. W., “Interpreting the Concordance Statistic of a Logistic Regression Model: Relation to the Variance and Odds Ratio of a Continuous Explanatory Variable,” BMC Medical Research Methodology 12 (2012): 82. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34. F. E. Harrell, Jr. , Lee K. L., and Mark D. B., “Multivariable Prognostic Models: Issues in Developing Models, Evaluating Assumptions and Adequacy, and Measuring and Reducing Errors,” Statistics in Medicine 15 (1996): 361–387. [DOI] [PubMed] [Google Scholar]
- 35. Steyerberg E. W., Vickers A. J., Cook N. R., et al., “Assessing the Performance of Prediction Models: A Framework for Traditional and Novel Measures,” Epidemiology 21 (2010): 128–138. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36. Rufibach K., “Use of Brier Score to Assess Binary Predictions,” Journal of Clinical Epidemiology 63 (2010): 938–939. author reply. [DOI] [PubMed] [Google Scholar]
- 37. Gneiting T. and Raftery A. E., “Strictly Proper Scoring Rules, Prediction, and Estimation,” Journal of the American Statistical Association 102 (2007): 359–378. [Google Scholar]
- 38. DeLong E. R., DeLong D. M., and Clarke‐Pearson D. L., “Comparing the Areas Under Two or More Correlated Receiver Operating Characteristic Curves: A Nonparametric Approach,” Biometrics 44 (1988): 837–845. [PubMed] [Google Scholar]
- 39. Martin J. A., Wilson E. C., Osterman M. J., Saadi E. W., Sutton S. R., and Hamilton B. E., “Assessing the Quality of Medical and Health Data From the 2003 Birth Certificate Revision: Results From Two States,” National Vital Statistics Reports 62 (2013): 1–19. [PubMed] [Google Scholar]
- 40. Dietz P., Bombard J., Mulready‐Ward C., et al., “Validation of Selected Items on the 2003 U.S. Standard Certificate of Live Birth: New York City and Vermont,” Public Health Reports 130, no. 1 (2015): 60–70. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41. Bursac Z., Gauss C. H., Williams D. K., and Hosmer D. W., “Purposeful Selection of Variables in Logistic Regression,” Source Code for Biology and Medicine 3 (2008): 17. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42. Helps A., Leitao S., Gutman A., Greene R., and O'Donoghue K., “National Perinatal Mortality Audits and Resultant Initiatives in Four Countries,” European Journal of Obstetrics, Gynecology, and Reproductive Biology 267 (2021): 111–119. [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Appendix S1: ppe70065‐sup‐0001‐AppendixS1.docx.
Data Availability Statement
The data used in this study are publicly available at https://www.cdc.gov/nchs/data_access/vitalstatsonline.htm.
