Abstract
Accurate estimation and forecasts for neonatal mortality rates (NMRs) in low- and middle-income countries is an urgent problem. Much of child mortality is preventable, and understanding temporal trends is of great interest when evaluating past performance and planning future policy or programming. In countries without robust vital registration, we rely on modeled estimates based on survey data to understand trends. A toolkit of compelling temporal models exists, but these methods have not been comprehensively evaluated for their application for the estimation of the NMR in low- and middle-income countries using household survey data. Using Demographic and Health Surveys (DHS) and Multiple Indicator Cluster Surveys (MICS) data from 41 countries in sub-Saharan Africa, we estimate and forecast the national-level NMR for 1970–2030 separately with random walk, auto-regressive, penalized spline, natural spline, and logit-linear latent temporal models. We examine the statistical behavior of these temporal models with both an out-of-sample analysis using the DHS and MICS data and a simulation study. We find that the second-order random walk and the penalized spline have the least bias, and short-term forecasts from the penalized spline tend to have narrower intervals with better out-of-sample performance. From the analysis of the NMR in sub-Saharan Africa, we estimate that 6 or fewer of the 41 countries included are on track to achieve the Sustainable Development Goals target of 12 neonatal deaths per 1000 live births by 2030.
Keywords and phrases: random walk, penalized spline, autoregressive, temporal latent effects, Bayesian, neonatal mortality rate
1. Introduction.
The neonatal mortality rate (NMR) – the number of deaths in the first 28 days of life per 1000 live births – has been consistently declining in most countries, with large gains made in the last century. However, the NMR is not declining globally at the same pace as the under-five mortality rate (U5MR). There is some evidence that the pace of decline of the NMR is slowing in some countries, and large inequities remain between and within countries. A large majority of these deaths are preventable; the countries with the lowest NMR, such as Japan and Singapore (https://childmortality.org), have fewer than 1 death per 1000 live births and many of the leading causes of neonatal death globally are infectious diseases or neonatal disorders with well established interventions. And yet, the gap between the countries with the lowest NMR and the countries with the highest NMR was nearly 40 deaths per 1000 live births in 2022 (United Nations Children’s Fund (UNICEF), 2025).
The Sustainable Development Goals (SDGs) set by the United Nations in 2015 target 12 or fewer neonatal deaths per 1000 live births, in every country, by 2030. Hence, monitoring progress towards this goal is of particular interest. According to a study by Hug et al. (2019), two-thirds of countries projected to have NMR above the SDG target in 2030 are in sub-Saharan Africa, and therefore, our study will focus on this region. Unfortunately, many of the same countries with higher NMR do not have high quality vital registration systems, and so we must rely on statistical models based largely on survey data to estimate past and present NMR.
Household surveys, such as the Demographic and Health Survey (DHS) series and the Multiple Indicator Cluster Survey (MICS) series, are a rich source of data on many population health indicators. Full birth histories collected by these surveys include child-level data which can be used to estimate child mortality rates. However, survey data do not typically capture every year we aim to estimate, and weighted estimates often have high variance. We also need to make a single summary for each year, and some years may be captured by more than one survey. Researchers typically adopt the following strategy, which is cataloged in more detail by Susmann, Alexander and Alkema (2021). First, weighted estimates are computed, and their likelihood is defined to form the data model. Then, a latent temporal model is applied to smooth estimates in observed years, fill in missing years, and allow forecasts beyond observed years. In this way, we exploit temporal structure to make the best possible use of the collected data. However, the choice of latent temporal model can have large implications for years without data, and there is no consensus regarding which model should be used.
Estimates of national-level NMR have been produced by organizations such as the United Nations Inter-agency Group for Child Mortality Estimation (UN-IGME) (United Nations Children’s Fund (UNICEF), 2025; Alexander and Alkema, 2018) and the Global Burden of Disease Study (GBD) (Schumacher et al., 2024; Paulson et al., 2021). UN-IGME uses a Bayesian B-spline model (Alkema and New, 2014; Eilers and Marx, 1996) to capture the latent temporal trend, whereas the GBD uses a collection of ad-hoc regression steps in combination with Gaussian process regression. UN-IGME has also produced subnational estimates of child mortality (Mercer et al., 2015; Wakefield et al., 2019), and this work has utilized Gaussian Markov Random fields (GMRFs) for the latent temporal model, including autoregressive and random walk models (Rue and Held, 2005). To our knowledge, neither UN-IGME nor the GBD conducted a rigorous comparison of latent temporal models to guide their model selection. While recognizing that every country may be different, we aim to comprehensively evaluate latent temporal models frequently used for estimating the NMR across all countries in sub-Saharan Africa, to provide evidence to support model selection in estimation endeavors such as those conducted by UN-IGME and the GBD.
In addition to interpolation, GMRF and spline models are capable of making predictions of short-term forecasts, such as forecasting to the SDG target year of 2030. Surveys are not typically collected every year, so short-term forecasts may also be used to make estimates for the present when the last survey is more than one year old. We evaluate estimates and short-term forecasts on the basis of bias (how accurately a model can predict new observations); coverage of prediction intervals, which should be close to nominal coverage; and log conditional predictive ordinance (logCPO) (Gneiting and Raftery, 2007). We chose logCPO because it is a proper scoring rule which is appropriate for evaluating predictive distributions.
Longer-term forecasting frequently involves more complex models, which include more structure or auxiliary information. For example the Lee-Carter model is popular for forecasting age-specific mortality rates (Lee and Carter, 1992). The literature on modeling epidemics is rich with mechanistic models which incorporate complex relationships between many parameters. Another approach is to present scenario-based forecasts, such as in Hug et al. (2019) and Foreman et al. (2018). Intrinsic GMRFs are GMRFs with singular precision matrices, such as random walk models. Intrinsic models are non-stationary and so will produce credible intervals for longer-term forecasts which are too large to be useful. Hence, our focus is on interpolation and short-term forecasting.
While this paper focuses on the NMR, understanding past and present trends of relevant indicators facilitates informed decision making generally speaking, and our findings could be extended to other applications. For population health practitioners and policy makers, such indicators might include HIV prevalence or vaccination coverage.
In Section 2 we start by defining the problem statement, the survey data available for estimating the NMR, and the data aggregation into weighted estimates. In Section 3, we describe our framework for Bayesian temporal models for the NMR, and define the candidate latent temporal models we consider. In Section 4, we study the temporal models using simulated survey data. Finally, in Section 5, we present results from an analysis of national-level NMR, for 1970–2030, using DHS and MICS surveys from 41 countries in sub-Saharan Africa. The results based on real data include both summaries of progress towards the SDG target for NMR and out-of-sample validation for the temporal models.
2. Goals and data sources.
Neonatal mortality rate is the probability of death in the first 28 days of life. For presentation, the NMR is reported in units of deaths per 1000 live births. In this paper we aim to produce national-level estimates of the NMR for countries indexed by , and for consecutive years or time periods indexed by , from a collection of country-specific surveys which we index by . Let denote the true NMR in year and country and denote the corresponding logit-transformed NMR.
Full birth histories from household surveys are the primary source of data on child mortality in low- and middle-income countries. Interviewers survey women of reproductive age and collect detailed records of each of the woman’s live births, including birth year and date and age at death for any children who died. In some countries, summary birth histories are also available. Summary birth histories are faster and cheaper to collect than full birth histories because they do not have information about each child, only the total number of children ever born to each mother and the number of those who have died. However, it is challenging to make meaningful inference about child mortality rates from these data and previous work has shown that they do not contribute substantially to estimates when full birth histories are also available (Wilson and Wakefield, 2021). For this reason, we do not include summary birth histories in this study.
All countries in sub-Saharan Africa with available full birth histories from 1970–2023 from either DHS or MICS, as of August 2025, are included in this study (41 in total). The countries are: Angola, Benin, Burkina Faso, Burundi, Cameroon, Central African Republic, Chad, Comoros, Congo, Côte d’Ivoire, Democratic Republic of the Congo, Eswatini, Ethiopia, Gabon, Gambia, Ghana, Guinea, Guinea-Bissau, Kenya, Lesotho, Liberia, Madagascar, Malawi, Mali, Mauritania, Mozambique, Namibia, Niger, Nigeria, Rwanda, Sao Tome and Principe, Senegal, Sierra Leone, Somalia, South Africa, Sudan, United Republic of Tanzania, Togo, Uganda, Zambia, and Zimbabwe. No survey from sub-Saharan Africa was excluded. To see which surveys were included, see the country-specific results in the Supplementary Materials (Paulson et al., 2025).
The structure of child mortality survey data differs from other binary indicators available from household surveys. The full birth histories recorded in each single survey include information about all children born during a retrospective period. This enables the estimation of temporal trends in the NMR even when only one survey is available. From each survey, we extract information about births from a retrospective period of 20 years prior to the survey. While there are records of births greater than 20 years before the survey, these are excluded because they come from mothers who were youngest at the time of birth and so will produce biased estimates, since young maternal age at birth is associated with high NMR (Wu et al., 2021; Finlay, Özaltin and Canning, 2011; Noori et al., 2022).
In addition to bias coming from younger mothers, full birth histories may be subject to other limitations. For example, survivorship bias may occur if relatively high mortality rates for mothers is associated with high NMR, and this effect may be stronger for estimation years farther away from the survey year. A special case of this is HIV/AIDs, where mothers who died of AIDS will not be available to be surveyed and their children are more likely to have died. Walker, Hill and Zhao (2012) proposed a correction for HIV/AIDs for under-five mortality estimation but it is not common practice to apply an HIV/AIDs correction for NMR because infected children are unlikely to die from HIV/AIDs during the neonatal period (Alexander and Alkema, 2018; Mahy, 2003). Misreporting and a lack of a definitional consensus for stillbirths and live births may also contribute to bias in estimates of the NMR. Full birth history data may also be subject to age-heaping, whereby age at death is often rounded to even intervals, but this is less likely to impact the NMR than it is to impact other age groups, such as infant mortality rate (Croft et al., 2023; Prieto, Verhulst and Guillot, 2021).
In an ideal world, we would have sufficient observations for each year of interest so that we have a weighted estimate for every year, and each weighted estimate has relatively low variance allowing conclusions to be drawn directly for many countries. However, weighted estimates from household surveys frequently have large variances due to both sampling variance and between-survey variation due to fieldwork issues. We want to identify the signal through the noise to make inference about some underlying time trend in the mean. Additionally, we want to make estimates for years with no data; either interpolation for years in-between observation years or short-term forecasts for years beyond the last observation. In these cases, we can borrow strength across time by selecting and fitting a latent temporal model, treating the weighted estimates as the observations. Figure 1 illustrates this estimation problem, including the weighted estimates with their sampling error and the predictions of the NMR, for four example countries; Malawi, Sierra Leone, Ethiopia, and Nigeria.
Fig 1.

Illustration of estimation of the NMR in Malawi, Sierra Leone, Ethiopia, and Nigeria, for 1970–2030. Points with bars are weighted estimates from DHS or MICS surveys with 95% intervals based on estimates of the sampling variance. A line and band are included to indicate the posterior median and 90% credible interval as estimated by UN-IGME. Separately, the posterior median, posterior 20% credible interval, and posterior 90% credible interval as estimated using a P-spline temporal model are indicated with an overlapping line, dark band, and lighter band, respectively.
In the remainder of this paper, let all parameters be country-specific, and drop the subscript for country for legibility. Given full birth histories from an arbitrary household survey from year , and an arbitrary year , let be the set of all live births occurring in year to mothers interviewed in survey . This set is considered random under design based inference. For each child indexed by , let be a binary indicator which is 1 if child died during the neonatal period and 0 if they did not. One limitation of these full birth history data is that they record only the birth and interview month, so a small number of newborns who are alive at the time of interview but less than 28 days in age will be counted as having a whole person-month of exposure. Also note that our choice to assign observations to birth year means that, for example, a neonatal death in January occurring to a child born in late December would not be assigned to the same year as the death.
Based on these data, the weighted (Hájek) estimate of the NMR in year is:
where is the design weight for individual (Hájek, 1971). The design weight is equal to the mother’s women’s individual sample weight, per convention (Croft et al., 2023). The jackknife is used to get an estimate of the variance of (Pedersen and Liu, 2012). Using the delta method, we estimate the variance of and call this . For this study, the weighted estimates and their variance are computed using the survey package in R (Lumley, 2023, 2010). Under repeated surveys, we make the assumption that the logit-transformed weighted estimate for year from survey is approximately normal (Breidt and Opsomer, 2017), according to:
| (1) |
where is the unknown survey-specific mean in year . We use the logit to improve the asymptotic sampling distribution of the estimates.
3. Bayesian temporal models for NMR.
3.1. Model description.
We start with the likelihood defined in Equation (1) and assume the survey-specific mean is the sum of the true logit-transformed NMR in year , , some survey-specific bias, , as in Mercer et al. (2015) and Wakefield et al. (2019), and some unstructured temporal effect :
for and . For this study, we treat as a fixed effect given by . The variance was selected so that the standard deviation of the survey effect would correspond to approximately 5% of the median weighted estimate of the NMR. In other words, for a true NMR of 40, a survey effect one standard deviation from the mean would result in a survey-specific mean of 40±2 deaths per 1000 live births. A sum-to-zero constraint is applied for across . The unstructured temporal effect is an independent and identically distributed (iid) effect given by . We include to allow time-specific shocks from the smooth temporal trend, differences between prevalence and risk, and possible excess unstructured survey error.
Each candidate latent temporal model for the true mortality risk has the form:
for , where is an intercept, is a linear slope on time, and is a structured temporal effect for year selected from the list described in the remainder of this section. In the following sub-sections we describe five candidate models for included in this paper in more detail. A simple linear model where is also included in this paper as a baseline for comparison. Details on prior selection for all models are given in Section 3.8. The posterior distribution for is used to generate predictions of the NMR for every year . We exclude the survey error and the iid effect from predictions of the NMR because we aim to estimate the trend in mortality risk in the absence of shocks or survey error.
For this study, the computation of weighted estimates and the modeling are carried out on a yearly scale. In contrast, UN-IGME processes weighted estimates for five-year periods, and treats these observations as if they come from the midpoint of the interval for the purposes of incorporation into a yearly model (United Nations Inter-agency Group for Child Mortality Estimation (UN IGME), 2024). This is suboptimal since the analysis model needs to acknowledge them as averages over the period to produce proper variance estimates (Godwin and Wakefield, 2021).
3.2. Autoregressive process of order 1.
We begin by describing the autoregressive process and the random walk – both of which are Gaussian Markov random fields (Rue and Held, 2005). This is the class of models utilized by UN-IGME for subnational estimation of child mortality rates (Wakefield et al., 2019). GMRFs have sparse precision matrices, allowing for fast computation with integrated nested Laplace approximation (INLA) (Rue, Martino and Chopin, 2009).
The autoregressive order 1 (AR1) model is appealing because it is stationary, so the variance is stable even in forecasts. For a stationary AR1 process, we have the initial state for :
with and . For the AR1 process can be written as:
A larger is associated with stronger temporal dependence, and a smaller implies that more of the data variability is explained by the AR1 process. For the AR1 model, our latent temporal model can be seen as penalizing deviations from the linear trend described by . Posterior medians from forecasts beyond the last observation will converge to this linear base model. In most cases this is preferable to an AR1 model with fixed (i.e. with the linear effect in time omitted), which would produce forecasts which are flat in time. However, the fact that the base model is the average linear trend in the data is a limitation of the AR1 model when the slope in the most recent years differs from .
3.3. Random walk of order 1.
The random walk of order 1 (RW1) model is a first-order intrinsic GMRF defined by iid Gaussian first-order differences, given by:
for , where . The RW1 model is a simpler model than the AR1, but it lacks the stationarity of the AR1. The base model for forecasts using the RW1 process has a constant mean equal to the posterior mean for where is the last year with an observation. Specifically, forecasts of the RW1 process have the conditional distribution:
Hence, as with the AR1 process, the temporal trends in the forecasts of based on the RW1 model will be driven by the average linear trend described by . A feature of intrinsic models is that they are improper and are invariant to changes in level, and as such, we apply the constraint that the sum for identifiability of a separate intercept and the RW1 model.
3.4. Random walk of order 2.
The random walk of order 2 (RW2) model is a second-order intrinsic GMRF defined by independent and normally distributed second-order differences, given by:
for , where . The base model for forecasts using the RW2 process is centered on the linear trend from the last two years with observed data. The conditional distribution is:
for . Therefore, the most recent trends in the period with observed data, as opposed to the average linear trend, will dominate the trend in the forecasts based on the RW2 model. The RW2 prior is also improper, and again we apply a sum-to-zero constraint so that we are able to identify the intercept, . For the RW2 model we also include the constraint so that we are able to identify the slope, .
3.5. Penalized spline.
The penalized spline (P-spline) model is a popular spline model for estimating smooth non-linear mean functions. It overcomes the major challenge of using splines – selecting the number and placement of knots – by using a large number of equally-spaced knots and placing a penalty on second-order differences for the spline coefficients to avoid over-fitting. Formally, the P-spline model is based on equally-spaced B-splines, with a RW2 prior on the spline coefficients (Lang and Brezger, 2004; Ventrucci and Rue, 2016):
for and , where , is a basis matrix with third-degree B-spline functions over the domain of , and is a vector of spline coefficients. The amount of shrinkage towards a linear trend is controlled by the precision . The P-spline model is especially relevant because of its connections to the spline model used by UN-IGME for the B3 model (Alkema and New, 2014). In their specification, the B3 model chooses one knot every 2.5 years. For estimates from 1980 to 2030, for example, this would be 20 knots. We also choose to place one knot every 2.5 years.
3.6. Natural cubic spline.
The natural cubic spline (N-spline) is also composed of B-spline basis functions, and can be written as:
with the constraint that is a linear function beyond the last knot. This linear constraint may be desirable, because without it, cubic splines may produce extreme projections which are not realistic. We choose one knot 15 years before the most recent observation, and one knot every 15 years within the period containing data. We found this frequency to achieve adequate flexibility for observation years and an appropriate level of sensitivity of forecasts to recent trends. However, we emphasize that results using this method vary greatly based on decisions about knot frequency and placement and analysts should proceed with caution when using the N-spline. This method is implemented with INLA (Gómez-Rubio, 2020).
3.7. Hyperparameters.
For each temporal model, we used penalized complexity (PC) priors, which have interpretable hyperparameters that are comparable across models, as described by Simpson et al. (2017). If is the parameter the prior describes, a PC prior has two hyperparameters, and , and gives us the prior probability where is an interpretable function of the parameter. For example, the PC prior for a precision parameter, , gives . We select universally to interpret as the prior median.
In all models, the intercept is given a normal prior with mean zero and precision zero, and the linear trend is given a normal prior with mean zero and precision 0.001. We used a PC prior with for the precision of the iid temporal effect so that the median standard deviation of corresponds to approximately 2 deaths per 1000 live births when the NMR is 40.
To make our candidate latent temporal models comparable with respect to the marginal variance implied by the priors, we simulated from the prior distributions for the structured term, , and selected hyperparameters to tune those priors for consistency, as follows. For each prior, we sampled replicates of for each . Then, we computed the geometric mean of the yearly empirical variances for . Call this measure , and define it as:
where the superscript is used to denote the sample number and is the sample mean. Finally, we selected the hyperparameter for each prior so that is approximately 0.018. The choice of is based on an analysis of residuals from country-specific linear models for the relationship between logit-transformed weighted estimates and time, which is described in Supplement B (Paulson et al., 2025).
The AR1 model is stationary, and hence the number of estimation years, , does not impact the marginal variance. For the random-walk models, we use scaling in accordance with Sørbye and Rue (2014) so that the marginal variance implied by the prior will not be modified by . However, for the P-spline model, the exercise described here needs to be repeated for each considered, since the number of years impacts the marginal variance implied by the P-spline model. We present results for and which align with the simulation and real data analyses in this paper, respectively.
The PC priors in the candidate latent temporal models and the corresponding hyperparameters selected are summarized in Table 1. For the AR1 model, a PC prior is used for the marginal variance with so that . A PC prior is also applied to the correlation, , and this PC prior has the interpretation . The prior on does not affect the marginal variance. It was selected based on an initial study of posterior medians of in initial fits to the DHS data using default priors in R-INLA. For the RW1 and RW2 models, a PC prior is used for with . For the P-spline model, we choose a PC prior for with equal to either 0.012 for or 0.009 for . For the N-spline model, the precision is fixed at 11 for the simulation study and is dynamically selected for each country based on the country-specific knot placement to achieve of approximately 0.018.
Table 1.
Geometric mean of yearly variance by model, parameter, and prior.
| Model | Parameter | Prior | |
|---|---|---|---|
| AR1 | marginal precision | 0.0182 | |
| correlation | |||
| RW1 | precision | 0.0177 | |
| RW2 | precision | 0.0174 | |
| P-spline ( years) | precision | 0.0169 | |
| P-spline ( years) | precision | 0.0181 |
3.8. Computation.
For this paper we use integrated nested Laplace approximations; an efficient and popular choice for models with Gaussian latent fields (Rue, Martino and Chopin, 2009; Rue et al., 2017). INLA is a deterministic method based on optimization and numerical integration, and hence is more efficient than stochastic alternatives such as Markov chain Monte Carlo (MCMC) based computation. Another option for computation is empirical Bayes with Template Model Builder, which can be a good choice for complex models that do not have the Gaussian latent field structure required by INLA (Kristensen et al., 2016).
4. Simulations.
One challenge of model validation using real data is that we can never compare our model to the true NMR, only to observations of the NMR. Therefore, we start by simulating true NMR and observations of that truth under an assumed data generating model. Our analysis of bias and coverage using these simulated data complements our findings based on DHS and MICS data in sub-Saharan Africa which are presented in Section 5.
4.1. Simulation methods.
Let be a smooth parametric function over time , for the the domain [1, 41], corresponding to country . Four functions were selected to represent four hypothetical countries – call them countries up-down, down-flat, down-down, and down-up (Figure 2a). We aimed to capture a range of observed temporal trends for the NMR. For example, countries down-flat and down-up have decreasing NMR between 1990 and 2010 with a leveling off or slight increase in NMR in more recent years, whereas country down-down has a more consistent decline in NMR and country up-down has an atypical trend which is a period of increasing NMR followed by a period of decreasing NMR. Congo and Madagascar are two examples where there is a period of increasing NMR followed by decreasing NMR. Liberia is an example where the NMR had a period of decline but appears to be stagnant in the last 10–15 years. Zimbabwe is one example where there appears to be a period of decline followed by increasing NMR. The majority of countries have patterns similar to the down-down country, so that the NMR has been consistently declining, but the pace of decline is variable. The expression for for each country is given in Supplement C (Paulson et al., 2025).
Fig 2.

(a) Time series of the NMR used for simulations for countries up-down, down-flat, down-down, and down-up, 1990–2030, (b) bias of the posterior median, as an estimator of the NMR, and (c) coverage of 80 percent credible intervals for from the autoregressive order 1 (AR1), random walk order 1 (RW1), random walk order 2 (RW2), penalized B-spline (P-spline), natural spline (N-spline), and linear latent temporal models. Simulations included 1000 replicates of observations for a 2020 survey with the NMR forecasted to 2030.
For an arbitrary country, , suppose the true logit-NMR in year , , is the logit-transformed polynomial . Then, conditional on , define the following data-generating model for a weighted estimate of the NMR for year , :
| (2) |
where is unstructured Gaussian variation which may represent stochastic deviation in prevalence from the underlying trend in mortality risk or small shocks, and is an iid term representing sampling error with known variance. Here, for simplicity, we are assuming the survey is unbiased so that the weighted estimate is centered at the truth – we will introduce the survey effect again in the next section. Note that in this simplified data-generating model, all weighted estimates have the same sampling variance.
Using this data-generating model, we first sampled one replicate of the sequence for each country, taking to represent the year 1990 and to represent the year 2030. Then, we created 1000 replicates of the observation for each country by sampling 1000 times and using Equation 2, for the years , so that the last 10 years are unobserved. We chose and so that there is more sampling variance than unstructured temporal variation. The sampling variation of 0.2 is the median sampling variation of the logit-transformed direct estimates from the DHS and MICS surveys used in this paper. For each of the replicates, we fit the AR1, RW1, RW2, P-spline, N-spline, and linear models as described in Section 3, omitting the fixed effect for survey-specific bias. Summaries of bias (average error over replicates) of the posterior median, , and coverage of 80% credible intervals by year were used to assess the performance of these models for estimating the true NMR in each country and year, .
4.2. Simulation results.
Our results suggest the RW2 and P-spline models are preferable to the AR1 and RW1 models for interpolation and estimation in observation years, because they have less bias but similar coverage in these years across the four country-types we have included. Bias of the posterior median as an estimator for and coverage of 80% credible intervals are presented in Figure 2. In the period with observations (between 1990 and 2020), the RW2 and P-spline models have the least bias. The AR1 model has a larger degree of negative bias for countries up-down and down-down during the 2000–2010 period, and the RW1 model has the same behavior but to a lesser extent. The N-spline and linear models both have more bias and poorer coverage than the RW2 and P-spline models in our simulations. Each of the temporal models has some bias at the front tail, even where data are available. The forecasted period between 2020 and 2030 is where we see both the largest amount of bias and the largest differences across country-types and temporal models.
None of the temporal models accurately captures the fall in the NMR in country up-down: all models over-estimate the NMR between 2020 and 2030, with overestimates of over 40 deaths per 1000 live births in 2030 according to the AR1, RW1, and linear models and over 30 deaths per 1000 live births in 2030 according to the RW2, P-spline, and N-spline models. The sampling variance is sufficiently large that there is not enough signal to identify the upward trend before 2010 or the downward trend after 2010. Example fits are included in Supplement E (Paulson et al., 2025).
For country down-flat, the RW2 and P-spline models provide largely unbiased forecasts. The AR1 and RW1 models, however, underestimate the NMR because the forecasts return to the average linear trend in the observed years, which is declining, and the linear model has the same behavior. The trend for country down-down is a consistent decline in the NMR, with some slowing in the pace of decline after 2010. All models are overly sensitive to the signs of slowing and hence over-estimate the NMR in the forecasts, but the models differ with respect to the magnitude of the bias. For the AR1, RW1, and linear models, this result is likely impacted by the relatively slow rate of change before 2000, so that the average slope is less steep than recent years, even if the pace of decline has begun to slow down by 2020. The RW2, P-spline, and N-spline models have the least bias for country down-down in 2030 (just below 10 deaths per 1000 live births), followed by the AR1, RW1, and linear models (just above 10 deaths per 1000 live births). Country down-up actually has a slight increase in the NMR between 2010 and 2030. The RW2 and P-spline models provide relatively little bias in these forecasts, the AR1, RW1, and linear models have substantial downward bias (over 20 deaths per 1000 live births in 2030), and the N-spline is in between these two groups.
Except in cases where there is substantial bias, such as for the AR1 model for country down-down between 2000 and 2010, the coverage of 80% credible intervals for predicting in observation years is approximately 90% for the AR1, RW1, RW2, and P-spline models. Hence, under the prior distributions selected for this study, all four of these candidate temporal models result in over-coverage of the truth. For forecasts, the AR1 and RW1 have increasingly poor coverage over time. For estimation years farther from the last observation, the bias is larger and hence the coverage is smaller, in some cases reaching as low as 0% by 2030. The RW2 and P-spline models have coverage which increases to nearly 100% by 2030 for the down-flat and down-up countries where it produces unbiased forecasts, with the degree of over-coverage being a somewhat greater issue for the RW2 model than the P-spline model. For the up-down and down-down countries where the RW2 and P-spline models over-estimate the NMR in forecasted years, coverage ranges between 30% and 70% in 2030. Again, coverage is lower for the P-spline model and for country-year-model combinations with larger bias. These results suggest that the P-spline model may be preferable to the RW2 model for forecasts if the forecasts are unbiased because they will have smaller more informative credible intervals. However, if the forecasts are biased, the larger RW2 intervals will more frequently cover the truth. Furthermore, we might expect biased forecasts to not be a rare occurrence, given down-down is a common temporal pattern which in our simulations resulted in slight bias but substantial under-coverage for forecasts using the RW2 and P-spline models. Both the N-spline and the linear models achieve nominal coverage (80%) in most country-years where they accurately estimate the NMR but almost always have extremely low coverage otherwise in our simulations. Additionally, the linear model has low coverage around 60% in the down-up country even for the years with no bias.
Supplemental Figure S1 shows the width of 80% credible intervals across replicates, by country and year (Paulson et al., 2025). The interval widths for observation years are largest for the AR1 and RW1 and smallest for the N-spline and linear models, with values ranging between 5 and 35 deaths per 1000 live births. In the forecasts, the interval widths for the RW2 and P-spline are the largest. For all but the down-down country, the RW2 and P-spline surpass interval widths of 50 deaths per 1000 live births before the 10th forecasted year (2030).
5. Analysis of country-level NMR in sub-Saharan Africa.
We consider all available full birth histories from DHS and MICS for 41 countries in sub-Saharan Africa as listed in Section 2. For each survey, we computed 20 weighted estimates of NMR – one for each of the 20 years leading up to and including the survey year. Using these weighted estimates, and the models described in Section 3, we produced estimates of the national NMR for each country for 1970–2030.
To compare the AR1, RW1, RW2, P-spline, N-spline, and linear models we conducted a validation exercise with two parts. First, we used leave-one-out cross validation to evaluate the ability of the models to interpolate and reconstruct missing years. Second, we used -ahead predictions, holding out the last five year of data, to evaluate forecasts. Recent literature discussing leave-group-out and leave-one-out cross validation for dependent data such as time series supports our decision to use a two-part validation in this way (Bürkner, Gabry and Vehtari, 2020; Liu and Rue, 2022; Adin et al., 2024). The following analyses are conducted separately by country. All parameters are country-specific but we omit a country subscript for parsimony of notation.
5.1. Interior predictions.
Separately for each in , where corresponds to 1970 and corresponds to 2030, we fit all candidate models while excluding weighted estimates from all available surveys for the NMR in year . Then, we computed validation metrics against the collection of holdouts using the posterior predictive distribution:
where is the set of weighted estimates from survey for all years excluding , is the joint likelihood for the weighted estimates as given by Equation (1), and is the posterior distribution for given the data from all years excluding . We took 1000 samples from the posterior predictive distribution and computed the 10th and 90th quantiles to get an 80% credible interval. Coverage is the proportion of the observed weighted estimates which fall inside these intervals. We also computed the logCPO (Gneiting and Raftery, 2007):
5.2. Forecasting a 5-year period.
The models which behave the best at interpolation may not be the same models which provide desirable short-term forecasts, and vice versa. So, we evaluated forecasting separately by holding out data from the most recent 5-year period, fitting the same models as described above, and performing out-of-sample validation. A forecast period of 5 years was selected so that there would be sufficient data remaining to estimate the survey effect after the hold outs are removed.
To get the truth to compare the results to, we did the following. For the most recent survey, , collected in the year , we directly computed the weighted estimates for the 5-year period , which we will denote , as:
where the notation is borrowed from the description of weighted estimates in Section 2. As with the single-year weighted estimates, these were computed using the survey package so that we also get estimates of the sampling variance.
To sample from the posterior predictive distribution for a 5-year period, based on our model which defines time-effects for single years, we first sampled from the posterior distribution for single-years then used the following aggregation:
where the five years included in the forecasted period are denoted by for , and is an index for sample number. Finally, we obtained samples from the posterior predictive distribution by adding samples from a distribution where is equal to the sampling variance of . Note that this is an approximation which assumes the number of live births is not changing rapidly year-over-year during any 5-year period — the true aggregation would be weighted by the proportion of live births in the period occurring in a given year. Based on these period estimates, we computed logCPO as previously described.
5.3. Out-of-sample validation results.
For the purpose of interpolation and estimation of the NMR in observation years, our results for interior 1-year out-of-sample coverage do not provide much evidence to support the universal use of one temporal model over another. There is substantial variability in coverage across countries relative to the variability across temporal models (Figure 3a). Median coverage of 80% predictive intervals for interior 1-year out-of-sample models is below nominal coverage for all models, but is highest for the AR1. For most countries, coverage is between 70% and 85%, with some exceptions between 60% and 70% (Angola, Sierra Leone) or above 85% (Sudan). Angola and Sierra Leone both have overlapping surveys with substantial differences, suggesting the survey effect may be too small, leading to relatively poor coverage.
Fig 3.

(a) Out-of-sample coverage of 80% intervals from the posterior predictive distribution, and (b) five-year period logCPO, for autoregressive order-1 (AR1), random walk order 1 (RW1), random walk order 2 (RW2), penalized spline (P-spline), natural spline (N-spline), and linear temporal models, for DHS and MICS surveys from 39 countries. The overlayed boxplots include median and interquartile range (IQR), with whiskers showing the value that is both farthest from the median and also less than 1.5 times the IQR from the lower or upper quartile.
Our results support the use of the P-spline model over the alternatives included in this study for short-term forecasts. For the five-year period forecasts, the P-spline model has the best logCPO, and the AR1, RW1, RW2, and N-spline have similar ranges of logCPO (Figure 3b). The linear model has a similar logCPO for some countries but in countries where the trend is obviously non-linear, such as Rwanda and Zimbabwe, it fares much worse than the alternatives included in this study (Figure S5). Single-year logCPO is included in Supplement A (Figure S2) (Paulson et al., 2025). For the AR1 model, the second-worst logCPO is from Niger, where there is slightly decreasing NMR between 1970 and 1990 followed by more rapidly declining NMR, according to surveys conducted between 1992 and 2012. Democratic Republic of the Congo has the worst logCPO for all models except the P-spline. The latest survey (2023 DHS) indicates a sharp turn in the temporal trend such that the NMR is increasing in the last decade, and these models struggle to predict the extent of that increase in the last five observed years.
5.4. Summary of progress towards SDG target.
Posterior median temporal trends in the NMR for a selection of countries with relatively large differences across temporal models are presented in Figure 4, and detailed country-specific results are included in Supplement F (Paulson et al., 2025). Country-specific figures include UN-IGME 2024 estimates of the NMR for reference, downloaded from https://childmortality.org/ (United Nations Children’s Fund (UNICEF), 2025). Of the 41 countries included, only Sao Tome and Principe and Comoros have posterior medians below the SDG target in 2030 according to the AR1 model. The list of countries with posterior median below the SDG target in 2030 for the other models is only Sao Tome and Principe for the RW1 model; Niger and Sao Tome and Principe for both the RW2 and the P-spline model; Angola, Burundi, Ghana, Niger, Rwanda, and Sao Tome and Principe for the N-spline model; and Angola, Comoros, and Sao Tome and Principe for the linear model. A total of 10 countries are estimated to have increasing NMR between 2020 and 2030 according to at least one of the temporal models. Across all countries, the mean percentage change in the NMR for a 10-year period was the largest in magnitude for the period after the Millennium Development Goals (MDGs) were established, between 2000 and 2010, when the average percentage change in the NMR was between −16% and −19% according to all models excluding the linear model. The linear model is excluded because its logit-linear nature with respect to time makes it an inaccurate estimator of how rates of change have varied over time. The average percentage change for 2010–2020 was estimated to be between −12% and −15% depending on the temporal model – values much closer to the pre-MDG era (1990–2000) estimates of −12% to −16%. The posterior probability of achieving the SDG target by 2030 is below 10% for the majority of countries; 35 of the 41 countries according to the AR1 model, 33 countries for the RW1 model, 30 countries for the RW2 model, 31 countries for both the P-spline and the N-spline models, and 37 countries for the linear model (Figure 5). To draw comparisons across countries with respect to their progress towards the SDG target for the NMR, we also examined the relative ranking of the NMR between these 41 countries in sub-Saharan Africa from the lowest to the highest NMR (as in Msemburi et al. (2023)), and how those rankings have changed over time (Figure 6). According to the P-spline model, there is evidence that rank is improving in the SDG era (2015-present) in Angola, Benin, Burundi, Cameroon, Ethiopia, Gambia, Ghana, and Liberia, suggesting progress in these countries since 2015 has been greater than in other countries. On the other hand, there is evidence that rank is increasing in the SDG era in Democratic Republic of the Congo, Eswatini, Gabon, Madagascar, Mozambique, Namibia, Uganda, and Zambia. Analogous plots to Figure 6 for the AR1, RW1, RW2, N-spline, and linear temporal models are included in Supplement D (Paulson et al., 2025). This rankings analysis is based on posterior samples for the rank, and not a simple ranking of posterior medians, to account for the uncertainty in the country-specific estimates.
Fig 4.

Posterior median and 20% credible interval for the NMR from the autoregressive order-1 (AR1), random walk order 1 (RW1), random walk order 2 (RW2), penalized B-spline (P-spline), natural spline (N-spline), and linear temporal models, 1970–2030, for a selection of countries with relatively large differences across temporal models.
Fig 5.

Posterior probability of the NMR falling below the SDG target of 12 deaths per 1000 live births, by 2030, according to the autoregressive order-1 (AR1), random walk order 1 (RW1), random walk order 2 (RW2), penalized B-spline (P-spline), natural spline (N-spline), and linear temporal models.
Fig 6.

Posterior probability of receiving a given ranking of NMR within the 41 countries included, by country and year, according to the P-spline model. Rankings are aggregated so that results are plotted for rank 1–4, 5–9, 10–14,…, 35–39, and 40–41. A lower rank corresponds to a lower relative NMR.
The five countries with the biggest relative difference in the posterior median for 2030 when comparing the highest to the lowest of the six candidate temporal models are Mozambique, Democratic Republic of the Congo, Niger, Burundi, Madagascar, and Uganda. In each case, the apparent rate of change from the most recent 5–10 years of observed data differs from the average rate of change over the full set of observations. Therefore, the AR1 and RW1 models, which estimate a trend closer to the average trend in the projections, produce different forecasts than the RW2 and P-spline models, which project the rate of change in more recent years to continue. The believability of either scenario should be assessed in a country-specific context given expert knowledge and judgment.
5.5. Comparison to UN-IGME.
The UN-IGME model for the NMR differs from ours in several key ways, including (a) the two sets of estimates have different data inclusions and exclusions in some cases, (b) UN-IGME estimates the NMR indirectly by modeling the ratio to the U5MR, (c) UN-IGME uses a hierarchical model which incorporates global information whereas our model is fit independently for each country, and (d) the UN-IGME model makes projections by taking a weighted average of global and country-specific trends (Alkema and New, 2014; Alexander and Alkema, 2018; United Nations Children’s Fund (UNICEF), 2025).
Data inclusions and exclusions are an important part of mortality estimation, and in some cases can have large impacts on the results. Two examples of countries where our results differ from UN-IGME’s for this reason are Niger and South Africa. Niger is an important example because based on the DHS data alone we project Niger to achieve the SDG target with the P-spline, RW2, and N-spline models, but the UN-IGME results – which incorporate two sources which are not either DHS or MICS from 2010 and 2021 – estimate the NMR to be more stagnant at a level well over the SDG target.
The UN-IGME estimates suggest a mean percent change of −13%, −18%, and −15% for the 1990–2000, 2000–2010, and 2010–2020 periods, respectively, across the same countries we have included in this study. Across the temporal models, we generally estimated slightly faster progress in 2000–2010 and slightly slower progress in 2010–2020. In other words, the UN-IGME estimates are more optimistic than ours about a continued pace of decline in the 2010s. One possible explanation for this is that the UN-IGME model is based on the relationship between the NMR and the U5MR, and greater progress has been made since 2010 in reducing the U5MR than in reducing the NMR (Paulson et al., 2021). Nigeria is an example of a country where our estimates of the post-2010 NMR are above the UN-IGME estimates of the NMR in the same period for all six models. The posterior median for the projections is increasing in our results when we use the RW2, P-spline, or N-spline model and decreasing in our results based on the RW1, AR1 model, or linear model, and also decreasing in the UN-IGME results. It is possible our increasing projections from the RW2, P-spline, and N-spline models are too extreme, but the data do not appear to support a decreasing trend in the NMR either. The 2024 Nigeria DHS, for which data are not yet available but a key indicators report has been published (Federal Ministry of Health and Social Welfare of Nigeria (FMoHSW), National Population Commission (NPC) [Nigeria], and ICF, 2024), indicates stagnant NMR despite decreasing U5MR. In cases like this, we suggest analysts take expert knowledge into account when choosing the latent temporal model, and we also advise against over-interpretation of the posterior median, with a focus instead of the prediction intervals. Zimbabwe is another example where the data support declining U5MR after 2008 but not declining NMR. UN-IGME projects the NMR in Zimbabwe to be declining after 2008 whereas we estimate the NMR to be increasing using the RW2, P-spline, N-spline, or linear model and stable using the AR1 or RW1 model. The fact that UN-IGME projections are a weighted average of global and country-specific trends may also result in forecasts of the NMR which are declining even when the country-specific data do not provide much evidence of decline.
Country-specific models such as ours can be better than global hierarchical models such as the UN-IGME model with respect to both computation time and reproducibility. However, there are some limits to this approach when the data are sparse. For example, there are a few countries such as Lesotho where we make estimates for years before the first included observation, and these backwards projections have a posterior median which increases over time. This finding runs counter to our strong expectation that the NMR is decreasing over time in most countries. Incorporating global information in some way would be one approach to encouraging a model to take a non-increasing trend unless there is strong evidence in the data to suggest otherwise. We also caution against over-interpretation of the posterior median in cases like the Lesotho example where the credible interval is relatively large.
Finally, across all country-years estimated, the median difference between the 90% interval width from our results and the corresponding interval width from the UN-IGME results is −1.7, −3.0, −3.7, −4.0, −4.6, and −5.5 deaths for the AR1, RW1, RW2, P-spline, N-spline, and linear models, respectively. This suggests our interval widths are slightly smaller than UN-IGME’s intervals.
6. Discussion.
In this paper we studied a set of latent temporal models, and used these models to estimate the national-level NMR for 41 countries in sub-Saharan Africa. We conducted a comprehensive validation exercise for these temporal models based on survey data, and supported our findings with an analysis of simulated full birth history data. Our results are evidence to support the use of a RW2 or a P-spline temporal model, over alternatives such as AR1, RW1, N-spline, or linear models. The AR1 and RW1 models return to the average linear trend from the observation years in the forecasts, and struggle when the trend at the tail has a slope which differs from the overall linear trend in the observations. The RW2 and P-spline models, by comparison, are more responsive to recent trends in the data. The N-spline model does not perform better than the other models, and it is volatile to user selection of knot quantity and placement. The linear model may be an okay approximation of the trend for some countries, but it typically has smaller uncertainty intervals than we would like, especially for forecasts. Additionally, we found poor coverage for both the N-spline and the linear models in our simulation study even for years with observations.
The comparison between the RW2 and the P-spline is more even. One notable difference is that the intervals tend to be larger for the RW2 model, especially in the forecasts. Larger intervals manifest in higher coverage for RW2, which we observed to a small degree in the simulation study and to a larger degree for the interior 1-year out-of-sample analysis using DHS and MICS data. However, the wider intervals for the RW2 are also related to a posterior distribution with density less concentrated around the mode, which manifests in a lower logCPO for the forecasted 5-year period for the RW2 model. A future study may focus on the RW2 and P-spline models and provide a deeper comparison of these two models.
We recognize that each country is different, and suggest that analysts take country-specific features and expert knowledge into account when choosing the latent temporal model; especially when choosing between the AR1 and RW1 on the one hand and RW2 and P-spline on the other hand. Based on the findings in this paper, it is our opinion that careful context-specific model selection is preferable to either a one-size-fits-all approach, or, something like Bayesian model averaging to combine multiple models (Yao et al., 2018).
Just as choosing the latent temporal model is important, choosing the hyperparameters for the priors is also important. In this paper we made the priors equivalent with respect to the implied marginal variance of the deviations from a linear temporal trend. This strategy allows us to make fair comparisons between the models. However, we note that in any application, the hyperparameters themselves should be carefully selected. We leave within-model comparisons across hyperparameter values as a possible extension to this study. An additional extension may be to utilize the PC prior for precision in the P-spline model proposed by Ventrucci and Rue (2016), where is based on the number of effective degrees of freedom. This prior has a consistent interpretation across changing number of knots (and estimation years).
One limitation of this study is that under our model for the analysis of NMR in sub-Saharan Africa, we could only realistically validate a 5-year forecast, since insufficient data from the most recent survey would remain to estimate the survey effect. Longer-term forecasts based on a modified model assuming there is no survey-specific bias could be validated by holding out the most recent survey entirely and comparing it to forecasted NMR. However, DHS and MICS surveys are often conducted every five years, so five-year short-term forecasts are still of interest when making estimates up to present day.
Our findings are relevant for short-term forecasts only, and we expect the estimates will be increasingly uninformative as they get farther from the last observation. For example, in our simulation exercise we observed that coverage for the RW2 and P-spline models is increasing over time in the forecasts, for countries with unbiased estimates, with coverage surpassing nominal coverage. With a lack of information to guide the models, they produce intervals which are too wide. Our results also highlight that inferring future trends using past trends alone, on the basis of a statistical model such as the AR1, RW1, RW2, P-spline, N-spline, or linear model studied in this paper, may be adequate for a few years, but that analysts should exercise caution when interpreting forecasts more than 5 years into the future.
Longer-term forecasts can incorporate additional structure or auxiliary information. For example, the B3 model uses a weighted combination of global and country-specific trends for projection years (Alkema and New, 2014). However, the B3 approach is ad-hoc, it is not consistent with the model used to fit the data, and it can lead to forecasts which do not appear to be supported by country-specific data. For the NMR, UN-IGME also utilizes the U5MR as a covariate. The GBD study uses additional covariates such as HIV and maternal education in their model for child mortality rates. Both UN-IGME and the GBD use global hierarchical models so that global information can inform country-specific estimates. While covariates and hierarchical modeling can be useful in cases where data are sparse, they can also impose strong trends in places where the data do not support them. For example, if under-five mortality is declining but the NMR is in truth stable, the covariate model may incorrectly forecast that the NMR is also declining. Future research should be conducted to explore alternative methods for leveraging information across locations and age groups. However, if possible, we would always prefer to use country-specific models. The comparison between our results and the UN-IGME estimates also demonstrates that decisions about data inclusions and exclusions has a large impact. Future research could establish a principled way to incorporate all data while addressing data quality.
For subnational estimation, the same survey data are often present, however, sampling variation is larger for subnational areas and some areas may not be sampled at all. One extension of this study could be to examine whether data sparsity and sampling variation impacts the conclusions drawn in this paper.
Finally, our forecasts to 2030 support the sobering findings by Hug et al. (2019) that most countries in sub-Saharan Africa are not on track to reach the SDG target for NMR. We estimated that due to stalled progress in the SDG era relative to the MDG era, only seven countries have posterior median NMR for 2030 lower than the SDG target of fewer than 12 deaths per 1000 live births according to at least one temporal model: Sao Tome and Principe, Niger, Comoros, Angola, Burundi, Ghana, and Rwanda. No temporal model estimates more than six countries will achieve the SDG target, and only Sao Tome and Principe has a posterior median below the SDG target for all six models.
Efforts to further reduce the NMR and accelerate progress towards the SDG target may be guided by the Every Newborn Action Plan, which details steps that can be taken to end preventable newborn deaths (World Health Organization (WHO) and United Nations Children’s Fund (UNICEF), 2020)). Such steps include established interventions related to antenatal care, nutrition, access to skilled birth attendants, and infrastructure, among other areas (World Health Organization (WHO), 2024; Hug et al., 2019). Interventions should also be informed by knowledge of the primary causes of deaths in the neonatal period, which are preterm birth, birth complications such as birth asphyxia, infections such as sepsis and meningitis, and congenital anomalies (World Health Organization (WHO), 2024; Hug et al., 2019). Since 2000, the most progress has been in reducing neonatal mortality from birth complications and infections (Hug et al., 2019; Paulson et al., 2021).
The set of countries projected to have posterior median NMR below the SDG target by 2030 is not uniform across all temporal models. Furthermore, four of the nine countries with a posterior median NMR below the SDG target by 2030 appear in the list of the ten countries with the largest differences across models, as presented in Figure 4 (Burundi, Comoros, Niger, and Rwanda). The disagreement across models reinforces the contributions of this paper, and we argue that the forecasts according to the RW2 and P-spline are more reliable because they are more sensitive to localized trends in the most recent observation years. We hope that by using the best statistical models to generate the most accurate estimates and short-term forecasts of the NMR, we can inform and motivate the global public health community to concentrate resources and efforts on progress towards these goals. However, we caution against over-interpretation of forecasts for countries with wide disagreement across models. In particular, these tend to be countries with highly non-linear historical trends, such as Democratic Republic of the Congo, Mozambique, and Madagascar.
Finally, we would like to acknowledge the uncertain future for household survey data on child mortality in low- and middle-income countries given the termination of the contract between the United States Agency for International Development and the Demographic and Health Survey series. These surveys are vital for tracking child mortality rates, in addition to many other indicators of health and well-being. As demonstrated in this paper, statistical models have limited ability to forecast trends in child mortality rates far into the future to compensate for a lack of data. This may be especially true if there are changes in the short-term trend in mortality rates in response to a withdrawal of global funding for health programs. In this context, additional work by the statistical and global health communities to make creative use of non-survey data or to target limited resources for new survey data collection may be critical.
Supplementary Material
(A) Results supplement: Validation
This supplement includes supplemental results related to the validation exercises based on simulated survey data and based on DHS and MICS data from sub-Saharan Africa.
(B) Methods supplement: Choice of
This supplement describes the selection of the parameter for tuning the hyperparameters for the priors on the structured temporal effects.
(C) Methods supplement: for simulations
Description of the parametric functions selected for baseline logit-NMR in the simulations, .
(D) Results supplement: Country rankings for AR1, RW1, RW2, N-spline, and linear models
Includes figures of posterior probability of receiving a given ranking of NMR within the 41 countries included, by country and year, according to the AR1, RW1, RW2, N-spline, or linear models. The analogous figure for the P-spline model is included in the main body the paper.
(E) Results supplement: Example model fits for simulation study
This supplement contains plots displaying one simulated 2020 survey for each hypothetical country (up-down, down-flat, down-down, and down-up), and the posterior median and 90 percent credible intervals for the NMR using an AR1, RW1, RW2, penalized spline, natural spline, or linear temporal model.
(F) Results supplement: Country-specific results for analysis of the NMR in sub-Saharan Africa
This supplement includes one figure for each country included in the analysis of data from countries in sub-Saharan Africa. The figures show posterior median and 90 precent credible intervals for the NMR using an AR1, RW1, RW2, penalized spline, natural spline, or linear model.
Funding.
Paulson, Li, and Wakefield were supported by grant R01HD112421-01 from the National Institutes of Health. Partial support for this research came from a Shanahan Endowment Fellowship and a Eunice Kennedy Shriver National Institute of Child Health and Human Development training grant, T32 HD101442, to the Center for Studies in Demography & Ecology at the University of Washington. The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health. Additional support for this research came from support to CSDE from the College of Arts & Sciences, the UW Provost, eSciences Institute, the Evans School of Public Policy & Governance, College of Built Environment, School of Public Health, the Foster School of Business, and the School of Social Work.
REFERENCES
- Adin A, Krainski ET, Lenzi A, Liu Z, Martínez-Minaya J and Rue H (2024). Automatic cross-validation in structured models: Is it time to leave out leave-one-out? Spatial Statistics 62 100843. [Google Scholar]
- United Nations Inter-agency Group for Child Mortality Estimation (UN IGME) (2024). Explanatory Notes: Child, adolescent and youth mortality trend series to 2022. [Google Scholar]
- Alexander M and Alkema L (2018). Global estimation of neonatal mortality using a Bayesian hierarchical splines regression model. Demographic Research 38 335–372. [Google Scholar]
- Alkema L and New JR (2014). Global estimation of child mortality using a Bayesian B-spline biasreduction model. The Annals of Applied Statistics 2122–2149. [Google Scholar]
- Breidt FJ and Opsomer JD (2017). Model-assisted survey estimation with modern prediction techniques. Statistical Science 32 190–205. Publisher: Institute of Mathematical Statistics. [Google Scholar]
- Bürkner P-C, Gabry J and Vehtari A (2020). Approximate leave-future-out cross-validation for Bayesian time series models. Journal of Statistical Computation and Simulation 90 2499–2523. Publisher: Taylor & Francis _eprint: 10.1080/00949655.2020.1783262. [DOI] [Google Scholar]
- Croft TN, Allen CK, Zachary BW et al. (2023). Guide to DHS Statistics. Rockville, Maryland, USA: ICF. [Google Scholar]
- Eilers PHC and Marx BD (1996). Flexible smoothing with B-splines and penalties. Statistical Science 11 89–121. Publisher: Institute of Mathematical Statistics. [Google Scholar]
- FINLAY JE, Özaltin E and Canning D (2011). The association of maternal age with infant mortality, child anthropometric failure, diarrhoea and anaemia for first births: evidence from 55 low- and middle-income countries. BMJ Open 1 e000226. Publisher: British Medical Journal Publishing Group Section: Global health. [Google Scholar]
- Foreman KJ, Marquez N, Dolgert A, Fukutaki K, Fullman N, McGaughey M, Pletcher MA, Smith AE, Tang K, Yuan C-W, Brown JC, Friedman J, He J, Heuton KR, Holmberg M, Patel DJ, Reidy P, Carter A, Cercy K, Chapin A, Douwes-Schultz D, Frank T, Goettsch F, Liu PY, Nandakumar V, Reitsma MB, Reuter V, Sadat N, Sorensen RJD, Srinivasan V, Updike RL, York H, Lopez AD, Lozano R, Lim SS, Mokdad AH, Vollset SE and Murray CJL (2018). Forecasting life expectancy, years of life lost, and all-cause and cause-specific mortality for 250 causes of death: reference and alternative scenarios for 2016–40 for 195 countries and territories. The Lancet 392 2052–2090. Publisher: Elsevier. [Google Scholar]
- Gneiting T and Raftery AE (2007). Strictly proper scoring rules, prediction, and estimation. Journal of the American Statistical Association 102 359–378. [Google Scholar]
- Godwin J and Wakefield J (2021). Space-time modeling of child mortality at the Admin-2 level in a low and middle income countries context. Statistics in medicine 40 1593–1638. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Gómez-Rubio V (2020). Bayesian Inference with INLA. Chapman Hall/CRC Press, Boca Raton, FL. [Google Scholar]
- Hájek J (1971). Discussion of “An Essay on the Logical Foundations of Survey Sampling, Part I,” by D. Basu. In Foundations of Statistical Inference (Proc. Sympos., Univ. Waterloo, Waterloo, Ont., 1970). Holt, Rinehart, and Winston, Toronto. [Google Scholar]
- Hug L, Alexander M, You D, Alkema L and Un Inter-agency Group for Child Mortality Estimation (2019). National, regional, and global levels and trends in neonatal mortality between 1990 and 2017, with scenario-based projections to 2030: a systematic analysis. The Lancet. Global Health 7 e710–e720. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kristensen K, Nielsen A, Berg CW, Skaug H and Bell BM (2016). TMB: Automatic differentiation and Laplace approximation. Journal of Statistical Software 70 1–21. [Google Scholar]
- Lang S and Brezger A (2004). Bayesian P-splines. Journal of computational and graphical statistics 13 183–212. [Google Scholar]
- Lee RD and Carter LR (1992). Modeling and forecasting U.S. mortality. Journal of the American Statistical Association 87 659–671. [Google Scholar]
- Liu Z and Rue H (2022). Leave-group-out cross-validation for latent Gaussian models. [Google Scholar]
- Lumley T (2010). Complex Surveys: A Guide to Analysis Using R: A Guide to Analysis Using R. John Wiley and Sons. [Google Scholar]
- Lumley T (2023). survey: analysis of complex survey samples. R package version 4.2. [Google Scholar]
- Mahy M (2003). Measuring child mortality in AIDS-affected countries Technical Report, United Nations, New York. [Google Scholar]
- Mercer LD, Wakefield J, Pantazis A, Lutambi AM, Masanja H and Clark S (2015). Space-time smoothing of complex survey data: Small area estimation for child mortality. The annals of applied statistics 9 1889–1905. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Msemburi W, Karlinsky A, Knutson V, Aleshin-Guendel S, Chatterji S and Wakefield J (2023). The WHO estimates of excess mortality associated with the COVID-19 pandemic. Nature 613 130–137. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Noori N, Proctor JL, Efevbera Y and Oron AP (2022). The effect of adolescent pregnancy on child mortality in 46 low- and middle-income countries. BMJ Global Health 7 e007681. Publisher: BMJ Specialist Journals Section: Original research. [Google Scholar]
- Federal Ministry of Health and Social Welfare of Nigeria (FMoHSW), National Population Commission (NPC) [Nigeria], And ICF (2024). Nigeria Demographic and Health Survey 2023–24: Key Indicators Report. Abuja, Nigeria, and Rockville, Maryland, USA: NPC and ICF. [Google Scholar]
- Paulson KR, Kamath AM, Alam T, Bienhoff K, Hay SI, Murray CJL, Wang H, Kassebaum NJ et al. (2021). Global, regional, and national progress towards Sustainable Development Goal 3.2 for neonatal and child health: all-cause and cause-specific mortality findings from the Global Burden of Disease Study 2019. The Lancet 398 870–905. [Google Scholar]
- Paulson KR, Fuglstad G-A, Li ZR and Wakefield J (2025). Supplement to “Temporal Models for Estimation and Short-term Forecasting of Neonatal Mortality Rates in Sub-Saharan Africa”.[provided by typesetter] [Google Scholar]
- Pedersen J and Liu J (2012). Child mortality estimation: appropriate time periods for child mortality estimates from full birth histories. PLOS Medicine 9 e1001289. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Prieto JR, Verhulst A and Guillot M (2021). Estimating the infant mortality rate from DHS birth histories in the presence of age heaping. PLOS ONE 16 e0259304. Publisher: Public Library of Science. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Rue H and Held L (2005). Gaussian Markov random fields: theory and applications. CRC press. [Google Scholar]
- Rue H, Martino S and Chopin N (2009). Approximate Bayesian inference for latent Gaussian models by using integrated nested Laplace approximations. Journal of the Royal Statistical Society Series B: Statistical Methodology 71 319–392. [Google Scholar]
- Rue H, Riebler A, Sørbye SH, Illian JB, Simpson DP and Lindgren FK (2017). Bayesian computing with INLA: A Review. Annual Review of Statistics and Its Application 4 395–421. Publisher: Annual Reviews. [Google Scholar]
- Schumacher AE, Kyu HH, Lim SS, Murray CJL et al. (2024). Global age-sex-specific mortality, life expectancy, and population estimates in 204 countries and territories and 811 subnational locations, 1950–2021, and the impact of the COVID-19 pandemic: a comprehensive demographic analysis for the Global Burden of Disease Study 2021. The Lancet 403 1989–2056. Publisher: Elsevier. [Google Scholar]
- Simpson D, Rue H, Riebler A, Martins TG and Sørbye SH (2017). Penalising Model Component Complexity: A Principled, Practical Approach to Constructing Priors. Statistical Science 32 1–28. [Google Scholar]
- Susmann H, Alexander M and Alkema L (2021). Temporal models for demographic and global health outcomes in multiple populations: Introducing a new framework to review and standardize documentation of model assumptions and facilitate model comparison. arXiv preprint arXiv:2102.10020. [Google Scholar]
- Sørbye SH and Rue H (2014). Scaling intrinsic Gaussian Markov random field priors in spatial modelling. Spatial Statistics 8 39–51. [Google Scholar]
- United Nations Children’s Fund (UNICEF) (2025). Levels & Trends in Child Mortality. [Google Scholar]
- Ventrucci M and Rue H (2016). Penalized complexity priors for degrees of freedom in Bayesian P-splines. Statistical Modelling 16 429–453. [Google Scholar]
- Wakefield J, Fuglstad G-A, Riebler A, Godwin J, Wilson K and Clark SJ (2019). Estimating under-five mortality in space and time in a developing world context. Statistical Methods in Medical Research 28 2614–2634. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Walker N, Hill K and Zhao F (2012). Child mortality estimation: methods used to adjust for bias due to AIDS in estimating trends in under-five mortality. PLOS Medicine 9 e1001298. [DOI] [PMC free article] [PubMed] [Google Scholar]
- World Health Organization (WHO) (2024). Newborn mortality fact sheet. https://www.who.int/news-room/fact-sheets/detail/newborn-mortality. [Google Scholar]
- World Health Organization (WHO) and United Nations Children’s Fund (UNiceF) (2020). Ending preventable newborn deaths and stillbirths by 2030. https://www.unicef.org/reports/ending-preventable-newborn-deaths-stillbirths-quality-health-coverage-2020-2025. [Google Scholar]
- Wilson K and Wakefield J (2021). Child mortality estimation incorporating summary birth history data. Biometrics 77 1456–1466. [DOI] [PubMed] [Google Scholar]
- Wu H, Zhao M, Liang Y, Liu F and Xi B (2021). Maternal age at birth and neonatal mortality: Associations from 67 low-income and middle-income countries. Paediatric and Perinatal Epidemiology 35 318–327. [DOI] [PubMed] [Google Scholar]
- Yao Y, Vehtari A, Simpson D and Gelman A (2018). Using stacking to average Bayesian predictive distributions (with discussion). Bayesian Analysis 13 917–1007. Publisher: International Society for Bayesian Analysis. [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
