Skip to main content
Wiley Open Access Collection logoLink to Wiley Open Access Collection
. 2026 Aug 22;68(5):e70162. doi: 10.1002/bimj.70162

Zero‐Inflated Promotion Cure‐Rate Models With Frailty and Flexible Activation Schemes

Danillo Assunção 1,2, Pedro L Ramos 3,, Vera Tomazella 2, Jeremias Leão 4, Lucas Osses 3, Agatha Rodrigues 5
PMCID: PMC13499460  PMID: 42631646

ABSTRACT

We propose a zero‐adjusted promotion cure‐rate regression model that jointly accommodates instantaneous failures (times at or near‐zero), a non‐susceptible cured fraction, competing risks, and latent heterogeneity modeled with a Gamma frailty. The number of latent causes follows a Poisson distribution and their activation times follow a Weibull distribution. We derive closed‐form survival and density functions under three activation schemes—minimum, maximum, and random—and allow covariates to act on the zero inflation and cure proportions as well as on the baseline time‐to‐event distribution. Parameters are estimated by maximum likelihood, and finite‐sample performance is assessed via Monte Carlo experiments. We illustrate the method with data on time to insulin initiation among 390 pregnant women with gestational diabetes, where the random activation scheme provides the best fit by Akaike Information Criterion and Bayesian Information Criterion and yields interpretable covariate effects. The proposed framework unifies and extends existing cure‐rate and zero‐adjusted survival models, providing a flexible tool for complex settings with cure fractions, competing risks, and frailty.

Keywords: activation schemes, competing risks, Gamma frailty, gestational diabetes, promotion cure‐rate, Weibull, zero‐adjusted survival

1. Introduction

Survival analysis is concerned with the time until an event of interest occurs. The foundational work by Kaplan and Meier (1958) introduced nonparametric estimators for survival functions, and (Cox 1972) proposed the semiparametric proportional hazards model. Both frameworks typically assume that every individual in the study is susceptible to the event and, given sufficiently long follow‐up, will eventually experience it. In practice, however, a subset of subjects may never undergo the event, either due to inherent immunity, effective treatment, or other factors, resulting in non‐susceptible, or cured, individuals.

To accommodate such long‐term survivors, (Berkson and Gage 1952) proposed the mixture cure‐rate model, which represents the population as a mixture of susceptible and cured subpopulations. The overall survival function then combines a point mass at infinity, the cured fraction, with a proper survival curve for the susceptible group. This approach not only estimates the cure proportion, but it also allows covariates to influence both the hazard of the event and the probability of cure. When no cure fraction is present, the model reduces to standard Kaplan–Meier or Cox regression methods (Kaplan and Meier 1958; Cox 1972). Applications of mixture cure‐rate models span many fields. In oncology, the outcome may be time to death or recurrence, recognizing that some patients achieve long‐term remission (Rahimzadeh et al. 2014). In criminology, researchers model time to recidivism for released offenders, allowing for those who never reoffend (Maller and Zhou 1996). In industrial reliability, the focus may be on time to failure of machinery components, where certain parts may effectively last indefinitely under normal conditions (Meeker et al. 2022). These examples demonstrate the versatility of cure‐rate models in contexts with a cured or immune fraction.

Despite their usefulness, standard cure‐rate models often assume a homogeneous susceptible population. In reality, unobserved factors, genetic, environmental, or behavioral, induce additional variability. Frailty models introduce an unobservable random effect into the hazard function to capture this latent heterogeneity (Vaupel et al. 1979). In the multiplicative frailty extension of the Cox model, each individual hazard is scaled by a frailty term, often assumed Gamma‐distributed, improving fit and interpretability when unexplained heterogeneity is present (Andersen et al. 1993; Hougaard 1995). Several authors have combined frailty with cure‐rate frameworks (Yu and Peng 2008; Calsavara et al. 2013; Leão et al. 2018). Moreover, in many applied settings, interventions or intrinsic processes impose a delay before failures become observable. For example, initiating medication may require a latent period before therapeutic benefits appear, or preventive maintenance may postpone mechanical breakdowns. Observations at or near time zero may, therefore, reflect this initial phase rather than true failure events. Zero‐adjusted, or zero‐inflated, cure‐rate models explicitly account for such instantaneous failures by allowing survival times equal to zero (de Oliveira et al. 2017; Calsavara, Rodrigues, Rocha, Tomazella, et al. 2019). While zeros frequently arise in real‐world survival data, the literature that simultaneously addresses zero inflation and cure fractions remains comparatively scarce.

The primary goal of this paper is to develop a comprehensive regression framework that simultaneously accommodates three key features often encountered in survival data, namely, a non‐negligible fraction of immediate failures, zero or near‐zero event times, a cure fraction of long‐term survivors, and unobserved heterogeneity among subjects. To this end, we extend the classical zero‐inflated mixture cure‐rate model by embedding a multiplicative frailty term into a promotion‐type hazard structure. Concretely, we assume that the number of latent competing causes follows a Poisson distribution, that their activation times follow a Weibull law, and that individual risk is further modulated by a Gamma‐distributed frailty. This unified approach yields closed‐form expressions for the survival and hazard functions under different activation schemes, minimum, maximum, and random, while accounting for latent variability. We estimate model parameters via maximum likelihood and investigate their finite‐sample behavior through Monte Carlo experiments. Finally, we illustrate the practical utility of the proposed framework by analyzing time to insulin initiation in a cohort of 390 pregnant women with gestational diabetes followed at a tertiary hospital in Brazil. These data exhibit both a non‐negligible proportion of very early insulin initiation times and a cured fraction corresponding to women who never require insulin during pregnancy, thereby providing a natural setting for a zero‐adjusted cure‐rate analysis. In this application, the proposed models offer improved flexibility and meaningful covariate interpretation relative to existing cure‐rate and zero‐adjusted survival formulations.

The remainder of this article is structured as follows. Section 2 reviews zero‐adjusted cure‐rate models under latent competing causes and activation schemes. Section 3 describes the formulation of the proposed frailty‐adjusted model. Section 4 outlines inference methods based on the likelihood function. Section 5 presents the simulation studies, including data generation, finite‐sample maximum likelihood estimates (MLE) performance, and misspecification and activation‐scheme identifiability. Section 6 illustrates the model using data on insulin use among pregnant women with gestational diabetes. Finally, Section 7 presents concluding remarks and future research directions.

2. Background

In this section, we present a brief description of the cure‐rate model proposed by Tsodikov et al. (1996) and by Tsodikov et al. (2003), later extended by Rodrigues et al. (2009), as well as a description of the zero‐adjusted promotion model and activation schemes.

2.1. Frailty Cure‐Rate Models

The statistical methods related to Tsodikov et al. (1996) and Chen et al. (1999) typically assume independence of individual survival times. Although this assumption is reasonable in many studies, it can be inappropriate when lifetimes are correlated, for example, when several recurrence times are observed for the same individual, characterizing multivariate survival data. In such cases, the independence assumption may fail. A common approach to incorporate this dependence is the frailty model, which has been widely used in the literature.

The frailty term is an unobservable, nonnegative random variable that is usually incorporated multiplicatively into the baseline hazard function. The multiplicative frailty model can be viewed as an extension of Cox (1972). Common choices for the frailty distribution include the Gamma and the log normal; the Gamma distribution is often preferred due to its algebraic tractability.

Following Elbers and Ridder (1982), identifiability requires that the frailty distribution has a finite mean. A convenient specification is a distribution with mean 1. In this paper, we assume a Gamma distribution with shape 1/α and scale α, so that the mean is 1 and α measures the amount of heterogeneity among individuals.

Let V denote the frailty variable. With multiplicative frailty, the hazard function for the ith individual at time t is

hi(t)=h(tvi)=vih0(t),i=1,,n, (1)

where v1,v2,,vn are the individual frailty terms and h0(t) is the baseline hazard function, common to all individuals. Note that in (1) the individual hazard increases if vi>1 and decreases if vi<1. Conditionally on the frailty term V, the model has a proportional hazards structure and may be interpreted as a frailty extension of the Cox proportional hazards model. The standard Cox model is recovered only in the degenerate case V1, that is, in the absence of unobserved heterogeneity. From (1), the conditional survival function is

S(tvi)=S0(t)vi, (2)

where S0(·) is the baseline survival function. To obtain the unconditional survival function, eliminating the unobserved frailty, we integrate out V. The marginal survival function is

S(t)=EVS(tV)=0eH0(t)vfV(v)dv=LVH0(t), (3)

where fV(·) is the density of V, H0(·) is the cumulative baseline hazard, and LV(·) denotes the Laplace transform. For a Gamma distribution with shape 1/α and scale α, the Laplace transform is

Lz(s)=(1+αs)1/α. (4)

Evaluating (4) at s=H0(t) yields the unconditional survival under Gamma frailty,

S(t)=1+αH0(t)1/α. (5)

Consequently, from (5), the corresponding density function is

f(t)=1+αH0(t)(α+1)/αh0(t), (6)

where h0(t) is the baseline hazard function.

2.2. Gamma Frailty Zero‐Adjusted Promotion Cure‐Rate Model

Cavenague de Souza (2020) studied a zero‐adjusted cure‐rate model with frailty, also known as zero‐inflated cure‐rate with frailty (ZICRF) , as an alternative for modeling survival data with both a cure fraction and a zero fraction. This model extends the ZICR framework proposed by Louzada et al. (2018). In Cavenague de Souza (2020), Gamma and log‐normal distributions were considered for the frailty and the survival time, respectively. For the analyzed dataset, the authors reported significant heterogeneity, indicating that a zero‐adjusted model with frailty is an attractive option for settings with unobserved heterogeneity.

To model survival data with a cure fraction, a zero fraction, and unobserved heterogeneity, the population survival function of the zero‐adjusted cure‐rate model with frailty, ZICR F, is given by Cavenague de Souza (2020)

Sp(tv)=p1+1p0p1S0(tv), (7)

where p0 is the proportion of survival times equal to zero, p1 is the proportion of cured individuals in the population, and S0(tv) is the survival function for the susceptible subgroup conditional on the frailty variable v.

Integrating out the frailty via the Laplace transform yields the unconditional population survival function,

Sp(t)=p1+1p0p1LVH0(t), (8)

where H0(t) is the cumulative baseline hazard and LV[·] denotes the Laplace transform associated with the frailty distribution.

Assuming that V follows a Gamma distribution with shape 1/α and scale α, the unconditional survival function of the ZICR F model is obtained by substituting (4) into (8),

Sp(t)=p1+1p0p11+αH0(t)1/α. (9)

As a consequence of (9), the corresponding density is

fp(t)=1p0p11+αH0(t)(α+1)/αh0(t).

These expressions allow heterogeneity to be modeled through frailty, enabling estimation of frailty parameters and covariate effects on survival. Furthermore, the model in (9) satisfies the following limiting cases:

  • If p0=0 and p1=0, one recovers the Gamma frailty survival function in (5).

  • If p0=0 and α0, the frailty degenerates at 1 and the model reduces to the standard Berkson and Gage mixture cure‐rate model (Berkson and Gage 1952).

2.3. Activation Schemes

In this section, we provide a brief description of activation schemes in the promotion cure‐rate model proposed by Yakovlev et al. (1993); Tsodikov et al. (1996), and Chen et al. (1999), where the modeling is grounded in a biological mechanism.

Following these studies, we assume that the subject under study has M tumor cells competing for metastasis and that the ith carcinogenic cell has a random activation time Zi, interpreted as an incubation time until a clinically detectable mass appears. The count M may also be interpreted as the number of cells affected by a bacterial or viral infection. The random variables Zi, i=1,2,, are assumed independent and identically distributed and independent of M, and M follows a Poisson distribution with parameter θ.

To accommodate individuals who are not susceptible to the disease under study, that is, subjects with initial tumor burden equal to zero and, theoretically, unlimited survival time, we adopt P(Z0=)=1. Let T denote the observed time to the event for exposed individuals. In line with (Cooner et al. 2007;;2006), we consider the following specifications:

  1. K=1: activation of the first latent factor leads to the event, hence
    T=Z(1,M)=min{Z1,,ZM}.
  2. K=q: activation occurs at the qth latent factor, hence
    T=Z(q,M).
  3. K=M: activation of all latent factors is required, hence
    T=Z(M,M)=max{Z1,,ZM}.

3. Frailty Zero‐Adjusted Cure‐Rate Model Under Different Activation Schemes

In this section, we formulate the model under three activation schemes. For each individual, there are multiple latent causes competing for the event of interest, with count MPoisson(θ). Let Zi denote the activation time of the ith latent cause. Following Section 2, we assume that Z1,Z2, are independent and identically distributed with distribution function F0(t)=1S0(t), independent of M. The observed time is defined according to the activation scheme: T=min{Z1,,ZM} (minimum), T=max{Z1,,ZM} (maximum), or T=Z(K) with K uniformly distributed on {1,,M} (random). These three activation mechanisms are summarized in Table 1.

TABLE 1.

Definition of the observed time T under each activation scheme.

Activation scheme
T
Minimum
T=,ifM=0,Z(1,M)=min{Z1,,ZM},ifM1
Maximum
T=,ifM=0,Z(M,M)=max{Z1,,ZM},ifM1
Random
T=Z(K,M),KMUniform{1,,M},k=1,,M

According to Chen et al. (1999), the survival function for the full population, including susceptible and non‐susceptible individuals, is

Sp(t)=P(T>t,M0)=P(M=0)+PT>t,M1=eθ+k=1P(T>tM=k)P(M=k)=eθ+k=1P(T>tM=k)θkk!eθ. (10)

The term P(T>tM) depends on the activation scheme:

P(T>tM)=S0(t)M,minimum,1F0(t)M,maximum,S0(t),random, sinceKMis uniform on{1,,M}.

After obtaining Sp(t), the density follows as fp(t)=ddtSp(t). The resulting survival and density functions (without frailty) are summarized in Table 2.

TABLE 2.

Population survival Sp(t) and density fp(t) under each activation scheme (no frailty).

Activation scheme
Sp(t)
fp(t)
Minimum
exp{θF0(t)}
θf0(t)exp{θF0(t)}
Maximum
1exp{θ[1F0(t)]}+eθ
θf0(t)exp{θ[1F0(t)]}
Random
eθ+(1eθ)S0(t)
(1eθ)f0(t)

Using F0(t)=1S0(t) and substituting S0(t) by the Gamma frailty survival in (5), we obtain the population densities and survivals with frailty shown in Table 3.

TABLE 3.

Population density fp(t) and survival Sp(t) with Gamma frailty for each activation scheme.

Activation scheme
fp(t)
Sp(t)
Minimum
θ(1+αH0(t))(α+1)/αh0(t)expθ1(1+αH0(t))1/α
expθ1(1+αH0(t))1/α
Maximum
θ(1+αH0(t))(α+1)/αh0(t)expθ(1+αH0(t))1/α
1expθ(1+αH0(t))1/α+eθ
Random
(1eθ)(1+αH0(t))(α+1)/αh0(t)
eθ+(1eθ)(1+αH0(t))1/α

The survival functions Sp(t) above are improper, with limtSp(t)=eθ>0, the cured fraction, irrespective of the activation scheme. Following Chen et al. (1999), we restrict attention to exposed individuals, M1, and define

Sp(t)=P(T>tM1).

For the ZICR model (no frailty), the survival and density for the non‐cured group are shown in Table 4.

TABLE 4.

Survival Sp(t) and density fp(t) for the non‐cured group (no frailty).

Activation scheme
Sp(t)
fp(t)
Minimum
exp{θF0(t)}eθ1eθ
exp{θF0(t)}1eθθf0(t)
Maximum
1exp{θ[1F0(t)]}1eθ
exp{θ[1F0(t)]}1eθθf0(t)
Random
S0(t)
f0(t)

Substituting S0(t) by (1+αH0(t))1/α yields Sp(t) and fp(t) for the Gamma frailty case, summarized in Table 5.

TABLE 5.

Conditional density fp(t) and survival Sp(t) with Gamma frailty under each activation scheme.

Activation scheme
fp(t)
Sp(t)
Minimum
expθ1(1+αH0(t))1/α1eθθ(1+αH0(t))(α+1)/αh0(t)
expθ1(1+αH0(t))1/αeθ1eθ
Maximum
expθ(1+αH0(t))1/α1eθθ(1+αH0(t))(α+1)/αh0(t)
1expθ(1+αH0(t))1/α1eθ
Random
(1+αH0(t))(α+1)/αh0(t)
(1+αH0(t))1/α

The functions in Table 5 are proper, satisfying Sp(0)=1 and Sp()=0. In what follows, we adopt the zero‐adjusted promotion cure‐rate specification

Sp(t)=p1+(1p0p1)Sp(t),t0, (11)

where Sp(t) and fp(t) depend on the activation scheme as above. Different choices for the latent distribution of Zj under each scheme generate different parametric families; in this paper, we assume a Weibull baseline for Zj.

4. Inference

Incorporating activation schemes for all competing causes into the zero‐adjusted promotion cure‐rate model allows us to decompose the population into three parts: the fraction that fails at the initial time (p0), the fraction that is not susceptible and therefore cured (p1), and the susceptible fraction 1p0p1 governed by a proper survival function Sp(t). We adopt the specification in (11), and model p0 and p1 via a multinomial logit so that p0,p1,1p0p1(0,1) (see Pereira et al. 2013; Hosmer et al. 2013; Louzada et al. 2018). Let x1i and x2i be covariate vectors linked to the zero mass and the cure fraction, with coefficients β1 and β2. Then

p0i=exp{x1iβ1}1+exp{x1iβ1}+exp{x2iβ2},p1i=exp{x2iβ2}1+exp{x1iβ1}+exp{x2iβ2}.

Equivalently, writing p0i=eνi and p1i=eθi with νi,θi>0 yields

νi=logexp{x1iβ1}1+exp{x1iβ1}+exp{x2iβ2},θi=logexp{x2iβ2}1+exp{x1iβ1}+exp{x2iβ2}.

To complete the specification of Sp(t) under each activation scheme, we take the baseline for the latent times Zj to be Weibull with shape α1i>0 and scale α2i>0,

F0(z;α1i,α2i)=1expzα2iα1i,f0(z;α1i,α2i)=α1iα2izα2iα1i1expzα2iα1i,

and link these parameters to covariates x3i and x4i through

α1i=exp{x3iβ3},α2i=exp{x4iβ4}.

When frailty is present, the Gamma variance parameter α>0 enters Sp(t) and fp(t) (see Table 5).

Finally, for observation i the mixture form is

Sp,i(t)=eθi+1eνieθiSp,i(t),

where Sp,i(t) and fp,i(t) depend on the chosen activation scheme, as summarized in Table 5.

4.1. Likelihood Function

We estimate the parameters by maximum likelihood. Each subject belongs to one of three subgroups: (i) zero time, ti=0, (ii) cured, and (iii) susceptible. Let the observed data be D={(ti,δi,x1i,x2i,x3i,x4i)}i=1n, where δi=1 denotes an observed event and δi=0 denotes right censoring. The individual likelihood contribution is

p0i,ifti=0,1p0ip1ifp,i(ti),ifti>0andδi=1,p1i+1p0ip1iSp,i(ti),ifti>0andδi=0, (12)

where Sp,i(t) and fp,i(t) are the proper survival and density for the susceptible group under the selected activation scheme.

Let ζ collect all free parameters. With a Weibull baseline and Gamma frailty, we write

ζ=β1,β2,β3,β4,α,

where the Weibull shape and scale are determined by β3 and β4, respectively, and α is the frailty variance parameter. The full likelihood is

L(ζ;D)i:ti=0p0ii:ti>0×(1p0ip1i)fp,i(ti)δip1i+(1p0ip1i)Sp,i(ti)1δi.

The MLE ζ^ solves the score equations U(ζ)=(ζ)/ζ=0. Since closed‐form solutions are not available, we use a quasi‐Newton method (Shanno 1970). In practice, all optimizations were carried out in R using the Broyden–Fletcher–Goldfarb–Shanno (BFGS) implementation in the optim function. In the simulation study, we initialized the algorithm at the true parameter values with small random perturbations, and in the insulin data application we used estimates from simpler submodels (without frailty and/or without zero inflation) as starting values. Convergence was monitored by the default optim criteria, requiring small relative changes in both the log‐likelihood and the parameter vector, and, when necessary, by checking that different starting points led to essentially identical solutions. For the models considered here, computing times were on the order of a few seconds on a standard laptop.

Under standard regularity and for sufficiently large n, Wald inference is based on the inverse observed information,

Var^(ζ^)=2(ζ)ζζζ=ζ^1

(see Migon et al. 2014; Ospina and Ferrari 2012). Thus, approximate (1α)×100% confidence intervals are

α^1±z1α/2se(α^1),α^2±z1α/2se(α^2),α^±z1α/2se(α^),β^·±z1α/2se(β^·),

where z1α/2 is the standard normal quantile and β^· denotes any component of β^1,β^2,β^3,β^4.

5. Simulation Studies

This section has three parts. First, we describe the common data‐generation mechanism used in all simulations. Second, we assess the finite‐sample performance of the MLE when the fitted model is correctly specified. Third, we evaluate misspecification and practical identifiability, including whether information criteria can distinguish the true activation scheme when the sample size changes.

5.1. Data Generation

In the first simulation experiment, we assess the finite‐sample performance of the MLE for the proposed models under different activation schemes and sample sizes n{100,250,500,750,1000}. To evaluate covariate effects, we generate a single binary covariate XBernoulli(0.5) indicating group membership. In these simulations, the four covariate vectors used in the linear predictors are scalar and take the same value, that is, x1i=x2i=x3i=x4i=xi, where xi{0,1}. In the correctly specified MLE experiment, the failure time follows a Weibull distribution with parameters (α1i,α2i). We link the zero inflation, cure fraction, and Weibull parameters to X by

νi=logexp{β10+xiβ11}1+exp{β10+xiβ11}+exp{β20+xiβ21},
θi=logexp{β20+xiβ21}1+exp{β10+xiβ11}+exp{β20+xiβ21}, (13)
α1i=exp{β30+xiβ31},α2i=exp{β40+xiβ41}.

Algorithm 1 generates samples from the zero‐adjusted promotion cure‐rate model under a chosen activation scheme and under the distributional specifications required by each simulation study. For the MLE performance study, G is Weibull and H is Gamma. For the misspecification study in Subsection 5.3, the same algorithm is used after replacing G by Gompertz or generalized exponential distributions and/or replacing H by an inverse Gaussian frailty.

ALGORITHM 1. Generator of random samples from the zero‐adjusted promotion cure‐rate model under a general activation scheme.

1: Choose the sample size n, the activation scheme A{minimum,maximum,random}, the latent activation‐time distribution G, the frailty distribution H, and the censoring distribution.
2: Set the parameters for p0i, p1i, G, and, when present, H. In the correctly specified MLE study, G is Weibull, and H is Gamma.
3: for
i=1,,n
do
4: Generate xiBernoulli(0.5) and compute p0i=eνi, p1i=eθi and the parameters of G and H for individual i.
5: Generate uiUniform(0,1).
6: if
uip0i
then
7: Set si=0.
8: else if
ui>1p1i
then
9: Set si=.
10: else
11: Set si=FA,G,H,i1{(uip0i)/(1p0ip1i)}, equivalently the root of Fi(si)ui=0, where Fi(t)=p0i+(1p0ip1i)FA,G,H,i(t).
12: end if
13: end for
14: Choose a finite censoring upper bound using the finite generated si values and generate wi from the censoring distribution.
15: Set ti=min(si,wi) and δi=I(ti<wi) for i=1,,n.

5.2. Finite‐Sample Performance of the MLE

We consider three parameter scenarios shown in Table 6.

TABLE 6.

Parameter values under different scenarios.

Parameters Scenarios
1 2 3
β10
2.00
0.50
0.50
β11
0.50 1.50 0.50
β20
1.00
0.75
1.00
β21
0.75 3.00 1.00
β30
0.50
2.00
0.75
β31
0.50 1.50 1.00
β40
1.50
1.25
1.25
β41
2.00 1.00 1.00
α
(0.25,0.50)
(0.25,0.50)
(0.25,0.50)

From Table 6, scenario 1 yields p09% for x=0 and 11% for x=1, while p124.47% and 38.90%, respectively. In scenario 2, p018.63% and 21.19%, with p130.72% and 57.61%. In scenario 3, p030.72% and 33.33%, with p118.63% and 33.33%. These values represent low, medium, and high levels of zero inflation and cure fraction. Scenario 1 has few early failures, Scenario 2 is intermediate, and Scenario 3 has many early failures.

For each combination of scenario, activation scheme, and sample size, we run 1000 Monte Carlo replications. We compute MLEs and standard errors and summarize performance by bias, root mean squared error (RMSE), and empirical coverage probability (CP) of nominal 95% Wald intervals. We use

RMSE(ζ)=Bias(ζ)2+Var(ζ),Bias(ζ)=ζ^ζ.

Due to Monte Carlo variability, CP will not be exactly 0.95; with 1000 replications, the expected 95% binomial interval is approximately [0.936,0.964].

Standard errors for p0 and p1 are obtained by the delta method with first‐order Taylor approximation. Optimization uses optim in R with the BFGS algorithm; see R Core Team RStudio Team (2020).

Figures 1, 2, 3 display bias, RMSE, and CP across sample sizes for the maximum, random, and minimum activation schemes, respectively. Main findings are as follows:

  1. Across schemes and scenarios, RMSEs of the MLEs shrink toward zero with n, indicating consistency. In scenario 1 under the maximum scheme (Figure 1), the parameter β21 needs about n500 for stable estimation. Under the minimum scheme (Figure 3), β11 and β21 in Scenarios 1 and 2 also require larger n.

  2. Biases tend to zero as n increases, and CPs are generally close to 95%. Under the random scheme (Figure 2), β31 approaches nominal CP more slowly than β40. Scenario 3 typically attains nominal CP from n250.

  3. Scenarios with larger cured and zero fractions (Scenarios 2 and 3) improve estimation of the regression parameters linked to p0=eν and p1=eθ, yielding smaller bias and better CP than scenario 1, due to more information in the zero and cured components.

  4. Conversely, in Scenario 1, with fewer zeros and cured individuals and thus longer observed times, estimation of the Weibull parameters (α1,α2) shows smaller bias and improved CP relative to Scenarios 2 and 3.

FIGURE 1.

FIGURE 1

Finite‐sample performance of the MLE under the maximum activation scheme. Each panel shows the bias, RMSE, and empirical coverage of nominal 95% Wald confidence intervals for the model parameters as a function of the sample size n. Colors identify the three simulation scenarios in Table 6: Scenario 1 (dark green), Scenario 2 (dark blue), and Scenario 3 (dark red); solid and dashed lines correspond to α=0.25 and α=0.50, respectively. All results are based on 1000 Monte Carlo replications for each combination of n, scenario, and frailty variance.

FIGURE 2.

FIGURE 2

Finite‐sample performance of the MLE under the random activation scheme. Each panel shows the bias, RMSE, and empirical coverage of nominal 95% Wald confidence intervals for the model parameters as a function of the sample size n. Colors identify the three simulation scenarios in Table 6: Scenario 1 (dark green), Scenario 2 (dark blue), and Scenario 3 (dark red); solid and dashed lines correspond to α=0.25 and α=0.50, respectively. All results are based on 1000 Monte Carlo replications for each combination of n, scenario, and frailty variance.

FIGURE 3.

FIGURE 3

Finite‐sample performance of the MLE under the minimum activation scheme. Each panel shows the bias, RMSE, and empirical coverage of nominal 95% Wald confidence intervals for the model parameters as a function of the sample size n. Colors identify the three simulation scenarios in Table 6: Scenario 1 (dark green), Scenario 2 (dark blue), and Scenario 3 (dark red); solid and dashed lines correspond to α=0.25 and α=0.50, respectively. All results are based on 1000 Monte Carlo replications for each combination of n, scenario, and frailty variance.

The conclusions from this MLE experiment concern estimation accuracy under correct specification. First, sample sizes in the range n250–500 are typically desirable for stable estimation of covariate effects on the zero‐inflation and cure components when the corresponding proportions are small. Second, moderate zero and cure fractions provide more information about the multinomial logit parameters, whereas datasets with fewer zeros and cured individuals are more informative about the Weibull baseline parameters. The separate question of whether the activation scheme itself can be distinguished from the data is addressed in Subsection 5.3.

5.3. Evaluating Misspecification and Activation‐Scheme Identifiability

This subsection addresses a different question from the MLE experiment: whether the model‐selection procedure can recover the correct distributional specification or activation mechanism. The proposed model includes several components: latent heterogeneity represented by a Gamma frailty, a Poisson number of latent causes, an activation scheme, and a Weibull distribution for the activation times, where all components were specified without covariates. We, therefore, considered misspecification scenarios, following the strategy of Louzada et al. (2025), in which the fitted model is compared with alternatives that use different activation‐time distributions and/or a different frailty distribution. The latent cause count was kept Poisson in all scenarios to preserve the competing‐risks interpretation.

Overall, six scenarios were considered:

  • Scenario I: Generate samples from the model with Gamma frailty (α=0.25) and a Weibull distribution (α1=2, α2=1) for the activation times and fit the data using the same frailty, but a Gompertz and a generalized exponential distribution for the activation times.

  • Scenario II: Generate samples from the model with Gamma frailty (α=0.25) and a Gompertz distribution (α1=2, α2=1) for the activation times and fit the data using the same frailty, but a Weibull and a generalized exponential distribution for the activation times.

  • Scenario III: Generate samples from the model with Gamma frailty (α=0.25) and a generalized exponential distribution (α1=2, α2=1) for the activation times and fit the data using the same frailty, but a Weibull and a Gompertz distribution for the activation times.

  • Scenario IV: Generate samples from the model with Gamma frailty (α=0.75) and a Weibull distribution (α1=2, α2=1) for the activation times and fit the data using the same activation times, but using an inverse Gaussian distribution for the frailty.

  • Scenario V: Generate samples from the model with Gamma frailty (α=0.75) and a Weibull distribution (α1=2, α2=1) for the activation times and fit the data using an inverse Gaussian distribution for the frailty and a generalized exponential distribution for the activation times.

  • Scenario VI: Generate samples from the model with Gamma frailty (α=0.75) and a Weibull distribution (α1=2, α2=1) for the activation times and fit the data using an inverse Gaussian distribution for the frailty and a Gompertz distribution for the activation times.

For each scenario 1000 samples were generated, and the Akaike Information Criterion (AIC) (Berkson and Gage 1952) was calculated for the correctly specified model and for the competing misspecified models. Table 7 reports the proportion of replications in which AIC selected the correctly specified model.

TABLE 7.

Proportion of samples in which AIC selected the correctly specified model in each misspecification scenario.

n
I II III IV V VI
50 0.390 0.728 0.398 0.561 0.596 0.851
150 0.597 0.798 0.554 0.650 0.583 0.995
300 0.743 0.865 0.611 0.777 0.624 1.000
500 0.841 0.936 0.642 0.860 0.649 1.000
1000 0.902 0.974 0.690 0.954 0.743 1.000

For activation‐time distributions, Scenario I shows that Weibull activation times are harder to recover in small samples, but the selection probability increases from 0.390 at n=50 to 0.902 at n=1000. Scenario II shows that a Gompertz activation‐time model is selected with probability 0.728 even at n=50, increasing to 0.974 at n=1000. Scenario III indicates that the generalized exponential distribution is the hardest to distinguish among the three alternatives, with a correct selection probability of 0.69 at n=1000. For frailty misspecification, Scenario IV shows that distinguishing Gamma from inverse Gaussian frailty can be difficult in small samples but improves markedly as n increases.

In Scenarios V and VI, both the frailty and activation‐time specifications are changed. Scenario V shows that when the activation time follows a generalized exponential distribution, the same difficulty seen in Scenario III persists. In contrast, Scenario VI shows that the original model is easily distinguished when the activation time follows a Gompertz distribution, with correct selection once the sample size exceeds 300 observations.

To provide direct evidence about activation‐scheme identifiability, we also generated data from the same general model but under the maximum activation scheme and compared it with the corresponding model under the minimum activation scheme using AIC. Figure 4 shows the proportion of replications in which the true activation scheme was selected as the sample size increases, based on 5000 replications. The correct‐selection rate increases substantially with n, but it is lower for small and moderate samples. This provides evidence for the practical finding that activation schemes may be weakly identifiable in finite samples when they induce similar survival patterns. Thus, the conclusions from this subsection are about robustness and identifiability rather than MLE accuracy. Information criteria are useful for comparing activation schemes and distributional assumptions, but their ability to select the correct specification depends on sample size and on how distinct the induced survival patterns are. In applications where competing schemes produce similar fitted curves, information criteria should therefore be complemented with substantive knowledge about the latent activation process and with diagnostic checks.

FIGURE 4.

FIGURE 4

Simulation results comparing activation schemes using 5000 replications. The x‐axis represents the sample size, while the y‐axis indicates the rate at which AIC selected the true activation scheme.

6. Application With Insulin Dataset

The dataset analyzed here, referred to as the insulin dataset, was introduced by Calsavara, Rodrigues, Rocha, Louzada, et al. (2019). It comprises retrospective times to initiation of insulin therapy for 390 pregnant women diagnosed with gestational diabetes mellitus (GDM) before 24 weeks of pregnancy. Follow‐up occurred during prenatal care at Hospital das Clínicas, University of São Paulo, Brazil, from 2012 to 2015. The GDM diagnosis was based on abnormal fasting plasma glucose in the first trimester (92–125 mg/dL), and insulin was started when glycemic targets were not achieved (fasting glucose > 95 mg/dL and 1 h postprandial glucose > 140 mg/dL), according to Souza et al. (2019). All women received nutritional and physical activity counseling and were instructed in self‐monitoring of blood glucose. Some women initiated insulin within the first 7 days, before the counseling program could take effect, so their survival times were recorded as zero. Conversely, a subset did not require insulin during pregnancy, characterizing a cured fraction.

This application matches a setting with a latent number of causal factors, as outlined in the Introduction. Accordingly, we consider three activation schemes. If a single causal factor is sufficient to trigger insulin use, the minimum scheme applies; if all factors must occur, the maximum scheme applies; otherwise, a random activation scheme is appropriate. The application uses the no‐frailty special case (α0); the Gamma frailty component is evaluated in the simulation studies.

We investigate whether the following covariates affect the time from the first obstetric visit to insulin initiation: family history of diabetes (0: no, n=143; 1: yes, n=247), prior history of GDM (0: primigravida, n=94; 1: no, n=249; 2: yes, n=47), smoking history (0: no, n=356; 1: yes, n=34), prior fetal macrosomia (0: no, n=356; 1: yes, n=34), prior chronic hypertension (0: no, n=282; 1: yes, n=108), and pre‐pregnancy body mass index (BMI) category (0: normal, BMI 25, n=84; 1: overweight, 25<BMI30, n=120; 2: obesity, BMI>30, n=186). Our goal is to assess whether these variables influence the probability of insulin use at the first visit and the probability of cure. The parameters p0 and p1 are linked to covariates by

p0=exp(xiβ1)1+exp(xiβ1)+exp(xiβ2),p1=exp(xiβ2)1+exp(xiβ1)+exp(xiβ2),

where xi=(x1i,x2i,x3i,x4i,x5i,x6i) for i=1,,390. Here, β1=(β10,β11,β12,β13,β14,β15,β16) are coefficients for the zero inflation component, and β2=(β20,β21,β22,β23,β24,β25,β26) are coefficients for the cure fraction.

We first fitted 18 models, one per covariate and activation scheme (six covariates by three schemes), using a Weibull baseline for time in all models. Results are reported in Tables A1 and A2. The three activation schemes yield very similar AIC and Bayesian Information Criterion (BIC) values, indicating that, for these data, the overall goodness‐of‐fit is not highly sensitive to the choice of activation mechanism. Nonetheless, the random activation scheme consistently provides the smallest AIC/BIC in almost all comparisons, and we, therefore, adopt it as our preferred specification. This choice is also clinically appealing, since insulin indication in gestational diabetes is known to be multifactorial and may require different combinations of factors across patients (Souza et al. 2019).

From the single covariate analyses, family history of diabetes, prior hypertension, and pre‐pregnancy BMI are important for explaining the cure fraction. In contrast, no individual risk factor shows a statistically significant association with insulin use at the first visit, as reflected in the β11 estimates. Tables A1 and A2 also report point estimates and 95% confidence intervals for p0 and p1. Standard errors for p0 and p1 were obtained via the delta method with first‐order Taylor approximation. For all activation schemes, both proportions are statistically relevant in the fitted models.

We then fitted a full model including all six covariates under each activation scheme. Estimates for these models are presented in Table A3. Prior fetal macrosomia is significant for explaining insulin use within the first 7 days, whereas family history of diabetes and smoking history primarily affect the cure component. For categorical covariates with three levels (e.g., pre‐pregnancy BMI and prior GDM), some activation schemes lead to slightly different patterns of statistical significance across categories. In these cases, we focus our substantive interpretation on the best‐fitting scheme (random activation) and on the direction and magnitude of the estimated effects, treating discrepancies across schemes as an indication of residual model uncertainty rather than as conflicting evidence.

Figure 5 compares, for selected covariates, the Kaplan–Meier estimates with the fitted population survival functions under the maximum, minimum, and random activation schemes. Table 8 reports the final selected model, with prior fetal macrosomia linked to p0 and history of diabetes and smoking linked to p1. The selected covariates are jointly relevant to describe early insulin use and the cure fraction across schemes. Prior fetal macrosomia has a positive effect (β11>0), increasing the probability of insulin at the first visit. History of diabetes shows a negative effect on the cure fraction (β22<0), indicating a lower chance of remaining insulin free during pregnancy. The QQ plot of the normalized randomized quantile residuals in Figure A1 supports the adequacy of the final specification.

FIGURE 5.

FIGURE 5

Kaplan–Meier estimates and fitted survival curves from the zero‐adjusted promotion cure‐rate model under the minimum, maximum, and random activation schemes, stratified by prior history of GDM, smoking, prior fetal macrosomia, prior chronic hypertension, pre‐pregnancy BMI, and family history of diabetes.

TABLE 8.

Maximum likelihood estimates (MLE), standard error (SE), 95% confidence interval (CI), and information criteria (AIC, BIC) for the final zero‐adjusted promotion cure‐rate model under the three activation mechanisms, insulin dataset.

Minimum Maximum Random
Parameter MLE SE MLE SE MLE SE
β10
0.770
0.197
0.769
0.197
0.769
0.197
β11 (Prior macrosomia) 1.118 0.428 1.111 0.428 1.115 0.428
β20
1.582 0.207 1.586 0.209 1.584 0.208
β21 (Smoking)
0.447
0.368
0.493
0.369
0.473
0.369
β22 (History diabetes)
0.675
0.235
0.674
0.235
0.674
0.235
AIC 1187.2404 1187.2572 1187.1498
BIC 1238.8003 1238.8171 1238.7097

Across all covariate patterns, the three fitted curves closely follow the empirical estimates and lie very near to each other. Thus, the conclusion from the application is that, for this dataset, the observable information is not sufficient to strongly favor one activation mechanism in terms of goodness‐of‐fit. The random‐activation model has the smallest AIC and BIC, but the differences are small; consequently, the activation schemes are best interpreted here as a sensitivity analysis for the latent competing‐risks structure. Importantly, the estimated covariate effects and the substantive conclusions about early insulin initiation and the cure fraction remain essentially unchanged across schemes. This application‐specific conclusion is distinct from the simulation conclusion in Subsection 5.3, which studies when activation schemes become distinguishable as sample size and distributional separation increase.

7. Concluding Remarks

This paper proposed a zero‐adjusted promotion cure‐rate regression framework with Gamma frailty under three activation schemes, minimum, maximum, and random. The construction unifies early failures and long‐term survivors within a single likelihood, and it nests relevant models as special cases, for example, the Berkson and Gage mixture cure model when p0=0 and α0 (Berkson and Gage 1952). Closed‐form expressions for survival and density functions were derived for each activation mechanism, which facilitates inference and interpretation, and the parameterization allows covariates to act on excess zeros, on the cure fraction, and on the failure time distribution.

The Monte Carlo study showed that maximum likelihood estimation performs well across activation schemes. Bias and RMSE decreased with sample size, and empirical coverage was close to the nominal level in most scenarios. The study also highlighted practical aspects of design; for instance, larger samples may be required for stable estimation of some covariate effects on p0 and p1 when the corresponding proportions are small.

The data application with time to insulin initiation in gestational diabetes illustrated the practical value of the approach. Information criteria favored the random activation scheme, which is consistent with a multifactorial clinical decision process. Prior fetal macrosomia increased the probability of insulin use at the first visit, while history of diabetes and smoking mainly explained the cure component. Randomized quantile residual diagnostics supported the adequacy of the final specification.

This work suggests several directions for further research. First, the Weibull baseline and the Poisson model for the number of latent causes are convenient and interpretable, yet alternative baselines, spline‐based hazards, or semiparametric specifications could improve robustness. Second, other frailty distributions, shared frailty structures for clustered data, or covariate‐dependent frailty could better capture unobserved heterogeneity. Third, model selection among activation schemes could be enhanced with predictive scoring rules or cross‐validation. Fourth, time‐varying covariates and dynamic zero inflation could accommodate evolving clinical states. Finally, extensions to handle misclassification of early events, interval censoring, or left truncation would broaden applicability.

An additional point concerns the assessment of model assumptions in practical applications. Although the quantile‐residual diagnostics in our case study indicate that the proposed model fits the insulin data well, this may not hold in other settings. For this reason, we emphasize that the use of our model should be accompanied by routine diagnostic checks, including the inspection of quantile residuals, comparison with simpler alternative specifications using information criteria, and examination of component‐specific estimates such as the zero probability and the cure fraction. These procedures provide a practical way to detect potential violations of the assumptions and to ensure that the chosen specification is supported by the empirical features of the data.

Although the model assumes a Gamma frailty distribution, this choice is widely used and captures a broad range of heterogeneity patterns through a single variance parameter. Alternative frailty specifications, such as log‐normal or inverse Gaussian frailties, may lead to similar population‐level behavior when the frailty variance is moderate. For this reason, the practical impact of the frailty distribution should be assessed empirically through diagnostic checks and model comparison rather than assumed a priori. The insulin application illustrates the no‐frailty special case, while the simulation studies evaluate estimation and misspecification of the frailty component.

In summary, the proposed family of zero‐adjusted promotion cure‐rate models with frailty provides a coherent and flexible framework for settings with instantaneous failures, cure fractions, and unobserved heterogeneity. Theoretical tractability, satisfactory finite‐sample behavior, and interpretability in the insulin application indicate that these models are a useful addition to the survival analysis toolkit.

Conflicts of Interest

The authors declare no conflicts of interest.

Open Research Badges

This article has earned an Open Data badge for making publicly available the digitally‐shareable data necessary to reproduce the reported results. The data is available in the Supporting Information section.

This article has earned an open data badge “Reproducible Research” for making publicly available the code necessary to reproduce the reported results. The results reported in this article could fully be reproduced.

Supporting information

Supporting File: bimj70162‐sup‐0001‐Datacode.zip.

Acknowledgments

Jeremias Leão is supported by the Brazilian agency CNPq (grant number 300393/2025‐3). Vera Tomazella is supported by the Brazilian agencies CNPq (grant number 301941/2025‐4) and FAPESP (grant number 2024/23079‐6)

Appendix A. Other Simulation Results

TABLE A1.

Maximum likelihood estimates (MLE), standard error (SE), 95% confidence interval (CI), and information criterion (AIC, BIC) values obtained by a zero‐adjusted promotion cure‐rate model under different activation mechanisms fitted for the insulin dataset

Minimum Maximum Random
Parameter MLE Lower Upper SE MLE Lower Upper SE MLE Lower Upper SE
β10 (intercept) −0.650 −1.348 0.048 0.356 −0.651 −1.349 0.047 0.356 −0.653 −1.351 0.045 0.356
β11 (family diabetes) 0.031 −0.782 0.844 0.415 0.032 −0.781 0.846 0.415 0.034 −0.779 0.848 0.415
β20 (intercept) 1.547 1.097 1.998 0.230 1.546 1.096 1.996 0.230 1.545 1.095 1.995 0.230
β21 (family diabetes) −0.682 −1.220 −0.144 0.275 −0.681 −1.219 −0.143 0.274 −0.680 −1.218 −0.142 0.274
AIC 1188.0048 1187.9497 1187.9162
BIC 1219.7339 1219.6788 1219.6454
p0(No)
0.084 0.038 0.129 0.023 0.084 0.038 0.129 0.023 0.084 0.038 0.129 0.023
p0(Yes)
0.138 0.095 0.181 0.022 0.138 0.095 0.181 0.022 0.138 0.095 0.181 0.022
p1(No)
0.755 0.685 0.826 0.036 0.755 0.685 0.826 0.036 0.755 0.685 0.826 0.036
p1(Yes)
0.607 0.546 0.668 0.031 0.607 0.546 0.668 0.031 0.607 0.546 0.668 0.031
β10 (intercept) −0.736 −1.128 −0.343 0.200 −0.735 −1.127 −0.343 0.200 −0.735 −1.127 −0.343 0.200
β11 (prior macrosomia) 0.736 −0.268 1.739 0.512 0.735 −0.268 1.739 0.512 0.735 −0.268 1.739 0.512
β20 (intercept) 1.143 0.886 1.400 0.131 1.142 0.885 1.399 0.131 1.142 0.885 1.399 0.131
β21 (prior macrosomia) −0.572 −1.428 0.284 0.437 −0.559 −1.414 0.297 0.436 −0.567 −1.423 0.290 0.437
AIC 1188.5019 1188.6249 1188.4775
BIC 1220.2311 1220.3540 1220.2067
p0 (No) 0.104 0.072 0.136 0.016 0.104 0.072 0.136 0.016 0.104 0.072 0.136 0.016
p0 (Yes) 0.265 0.117 0.414 0.076 0.264 0.116 0.412 0.075 0.265 0.116 0.413 0.076
p1 (No) 0.679 0.631 0.728 0.025 0.679 0.631 0.728 0.025 0.679 0.631 0.728 0.025
p1 (Yes) 0.469 0.302 0.637 0.086 0.473 0.305 0.640 0.086 0.471 0.303 0.638 0.086
β10 (intercept) −0.787 −1.251 −0.324 0.237 −0.787 −1.251 −0.323 0.237 −0.788 −1.251 −0.324 0.237
β11 (prior hypertension) 0.415 −0.320 1.150 0.375 0.413 −0.321 1.148 0.375 0.415 −0.319 1.150 0.375
β20 (intercept) 1.246 0.952 1.541 0.150 1.251 0.956 1.545 0.150 1.247 0.953 1.542 0.150
β21 (prior hypertension) −0.530 −1.063 0.003 0.272 −0.550 −1.083 −0.016 0.272 −0.538 −1.072 −0.005 0.272
AIC 1186.9748 1186.4715 1186.6558
BIC 1218.7040 1218.2006 1218.3850
p0 (No) 0.092 0.058 0.126 0.017 0.092 0.058 0.126 0.017 0.092 0.058 0.126 0.017
p0 (Yes) 0.184 0.111 0.257 0.037 0.186 0.112 0.259 0.037 0.185 0.112 0.258 0.037
p1 (No) 0.705 0.652 0.758 0.027 0.706 0.653 0.759 0.027 0.705 0.652 0.759 0.027
p1 (Yes) 0.548 0.454 0.642 0.048 0.544 0.450 0.638 0.048 0.546 0.452 0.640 0.048
β10 (intercept) −0.620 −0.999 −0.240 0.194 −0.619 −0.999 −0.240 0.194 −0.620 −1.000 −0.240 0.194
β11 (smoker) −0.074 −1.212 1.065 0.581 −0.074 −1.213 1.065 0.581 −0.071 −1.210 1.067 0.581
β20 (intercept) 1.143 0.884 1.401 0.132 1.144 0.885 1.402 0.132 1.143 0.884 1.401 0.132
β21 (smoker) −0.491 −1.298 0.315 0.411 −0.504 −1.311 0.304 0.412 −0.499 −1.308 0.309 0.412
AIC 1194.0061 1193.8797 1193.8464
BIC 1225.7352 1225.6088 1225.5756
p0 (No) 0.115 0.082 0.148 0.017 0.115 0.082 0.148 0.017 0.115 0.082 0.148 0.017
p0 (Yes) 0.146 0.028 0.265 0.060 0.147 0.028 0.266 0.061 0.147 0.028 0.266 0.061
p1 (No) 0.671 0.622 0.720 0.025 0.671 0.622 0.720 0.025 0.671 0.622 0.720 0.025
p1 (Yes) 0.561 0.395 0.727 0.085 0.558 0.392 0.725 0.085 0.559 0.392 0.726 0.085

TABLE A2.

Maximum likelihood estimates (MLE), standard error (SE), 95% confidence interval (CI), and information criterion (AIC, BIC) values obtained by zero‐adjusted promotion cure‐rate model under different activation mechanisms fitted for the insulin dataset.

Minimum Maximum Random
Parameter MLE Lower Upper SE MLE Lower Upper SE MLE Lower Upper SE
β10 (intercept) −0.533 −1.314 0.248 0.399 −0.532 −1.313 0.249 0.399 −0.532 −1.313 0.249 0.399
β11 (primigesta) 0.453 −0.654 1.560 0.565 0.452 −0.655 1.559 0.565 0.452 −0.655 1.559 0.565
β111 (No) −0.317 −1.233 0.599 0.467 −0.318 −1.234 0.598 0.467 −0.318 −1.234 0.598 0.467
β20 (intercept) 0.163 −0.484 0.810 0.330 0.169 −0.477 0.815 0.330 0.160 −0.488 0.807 0.330
β21 (primigesta) 1.506 0.629 2.383 0.448 1.502 0.625 2.378 0.447 1.510 0.632 2.387 0.448
β211 (No) 0.939 0.225 1.654 0.364 0.932 0.219 1.645 0.364 0.942 0.227 1.656 0.365
AIC 1188.0339 1188.1137 1187.9113
BIC 1235.6276 1235.7075 1235.5050
p0 (Primigesta) 0.128 0.060 0.195 0.034 0.128 0.060 0.195 0.034 0.128 0.060 0.195 0.034
p1 (Primigesta) 0.734 0.645 0.823 0.046 0.734 0.645 0.824 0.046 0.734 0.645 0.823 0.046
p0 (No) 0.096 0.060 0.133 0.019 0.096 0.060 0.133 0.019 0.096 0.060 0.133 0.019
p1 (No) 0.678 0.620 0.736 0.030 0.678 0.620 0.736 0.030 0.678 0.620 0.736 0.030
p0 (Yes) 0.212 0.096 0.329 0.060 0.212 0.095 0.329 0.059 0.213 0.096 0.330 0.060
p1 (Yes) 0.426 0.284 0.567 0.072 0.427 0.286 0.568 0.072 0.425 0.283 0.566 0.072
β10 (intercept) −0.548 −1.006 −0.091 0.233 −0.548 −1.005 −0.090 0.233 −0.548 −1.005 −0.090 0.233
β21 (Normal) −0.463 −1.696 0.769 0.629 −0.464 −1.697 0.768 0.629 −0.464 −1.696 0.769 0.629
β111 (Overweight) −0.108 −0.920 0.704 0.414 −0.108 −0.919 0.704 0.414 −0.108 −0.919 0.704 0.414
β20 (intercept) 0.758 0.421 1.095 0.172 0.755 0.418 1.091 0.172 0.756 0.420 1.093 0.172
β21 (Normal) 1.081 0.361 1.801 0.367 1.080 0.360 1.799 0.367 1.080 0.360 1.800 0.367
β211 (Overweight) 0.425 −0.135 0.986 0.286 0.434 −0.126 0.995 0.286 0.429 −0.131 0.990 0.286
AIC 1185.8927 1185.7000 1185.7147
BIC 1233.4865 1233.2937 1233.3084
p0 (Normal) 0.047 0.002 0.093 0.023 0.048 0.002 0.093 0.023 0.048 0.002 0.093 0.023
p1 (Normal) 0.822 0.740 0.904 0.042 0.821 0.739 0.903 0.042 0.821 0.740 0.903 0.042
p0 (Overweight) 0.108 0.053 0.164 0.028 0.108 0.053 0.164 0.028 0.108 0.053 0.164 0.028
p1 (Overweight) 0.683 0.599 0.766 0.043 0.684 0.601 0.767 0.042 0.683 0.600 0.766 0.043
p0 (Obesity) 0.156 0.104 0.208 0.027 0.156 0.104 0.208 0.027 0.156 0.104 0.208 0.027
p1 (Obesity) 0.575 0.504 0.646 0.036 0.574 0.503 0.645 0.036 0.574 0.503 0.646 0.036

TABLE A3.

Maximum likelihood estimates (MLE), standard error (SE), 95% confidence interval (CI), and information criterion (AIC, BIC) values obtained considering all risk factors by zero‐adjusted promotion cure‐rate model under different activation mechanisms fitted for the insulin dataset

Minimum Maximum Random
Parameter MLE Lower Upper SE MLE Lower Upper SE MLE Lower Upper SE
β10 (intercept) −1.405 −2.542 −0.268 0.580 −0.961 −2.087 0.166 0.575 −0.987 −2.103 0.129 0.569
β11 (History diabetes) 0.414 −0.421 1.250 0.426 0.036 −0.794 0.865 0.423 0.058 −0.768 0.884 0.421
β12 (smoking) 0.197 −0.980 1.374 0.601 −0.268 −1.410 0.873 0.582 −0.244 −1.381 0.892 0.580
β131 (primigesta) 1.081 −0.110 2.272 0.608 0.343 −0.899 1.585 0.634 0.245 −0.977 1.467 0.623
β132 (No) −0.221 −1.188 0.747 0.494 −0.405 −1.389 0.579 0.502 −0.474 −1.457 0.510 0.502
β141 (normal) −0.810 −2.103 0.483 0.660 −0.077 −1.381 1.228 0.666 −0.009 −1.316 1.299 0.667
β142 (Overweight) −0.126 −0.975 0.722 0.433 0.161 −0.698 1.020 0.438 0.170 −0.682 1.022 0.435
β15 (Prior fetal Macrosomia) 1.032 −0.033 2.097 0.543 1.047 −0.026 2.119 0.547 1.073 0.006 2.141 0.545
β16 (Chronic hypertension) 0.494 −0.300 1.288 0.405 0.208 −0.584 0.999 0.404 0.258 −0.532 1.048 0.403
β20 (Intercept) −0.350 −1.457 0.756 0.565 0.456 −0.438 1.351 0.456 0.382 −0.505 1.268 0.452
β21 (History diabetes) −0.197 −0.824 0.429 0.319 −0.799 −1.442 −0.156 0.328 −0.791 −1.429 −0.153 0.325
β22 (smoking) −0.341 −1.209 0.528 0.443 −1.298 −2.589 −0.007 0.659 −1.300 −2.578 −0.022 0.652
β231 (primigesta) 1.990 0.864 3.117 0.575 0.838 −0.353 2.030 0.608 0.718 −0.439 1.875 0.590
β232 (No) 1.195 0.251 2.139 0.482 0.747 −0.092 1.586 0.428 0.682 −0.166 1.529 0.432
β241 (normal) 0.368 −0.495 1.230 0.440 1.396 0.417 2.376 0.500 1.538 0.532 2.544 0.513
β242 (Overweight) 0.167 −0.458 0.792 0.319 0.635 −0.076 1.347 0.363 0.685 −0.026 1.396 0.363
β25 (Prior fetal Macrosomia) 0.172 −0.772 1.116 0.482 0.175 −0.825 1.175 0.510 0.275 −0.711 1.261 0.503
β26 (Chronic hypertension) −0.139 −0.752 0.474 0.313 −0.658 −1.405 0.089 0.381 −0.605 −1.335 0.124 0.372
AIC 1181.7417 1191.4407 1189.2752
BIC 1324.5230 1334.2220 1332.0565

FIGURE A1.

FIGURE A1

QQ plot of the normalized randomized quantile residuals with identity line for the final model.

Data Availability Statement

The data that support the findings of this study are available from the corresponding author upon reasonable request.

References

  1. Andersen, P. K. , Borgan Ø., Gill R. D., and Keiding N.. 1993. “Nonparametric Estimation.” In Statistical Models Based on Counting Processes, 176–331. Springer. [Google Scholar]
  2. Berkson, J. , and Gage R. P.. 1952. “Survival Curve for Cancer Patients Following Treatment.” Journal of the American Statistical Association 47, no. 259: 501–515. [Google Scholar]
  3. Calsavara, V. F. , Rodrigues A. S., Rocha R., F. Louzada, et al. 2019. “Zero‐Adjusted Defective Regression Models for Modeling Lifetime Data.” Journal of Applied Statistics 46, no. 13: 2434–2459. [Google Scholar]
  4. Calsavara, V. F. , Rodrigues A. S., Rocha R., Tomazella V., and Louzada F.. 2019. “Defective Regression Models for Cure Rate Modeling With Interval‐censored Data.” Biometrical Journal 61, no. 4: 841–859. [DOI] [PubMed] [Google Scholar]
  5. Calsavara, V. F. , Tomazella V. L., and Fogo J. C.. 2013. “The Effect of Frailty Term in the Standard Mixture Model.” Chilean Journal of Statistics 4, no. 2: 95–109. [Google Scholar]
  6. Cavenague de Souza, H. C. 2020. “Modelos de sobrevida para dados inflacionados de zeros aplicados a progressão do trabalho de parto.” PhD diss., Universidade de São Paulo .
  7. Chen, M.‐H. , Ibrahim J. G., and Sinha D.. 1999. “A New Bayesian Model for Survival Data With a Surviving Fraction.” Journal of the American Statistical Association 94, no. 447: 909–919. [DOI] [PMC free article] [PubMed] [Google Scholar]
  8. Cooner, F. , Banerjee S., Carlin B. P., and Sinha D.. 2007. “Flexible Cure Rate Modeling Under Latent Activation Schemes.” Journal of the American Statistical Association 102, no. 478: 560–572. [DOI] [PMC free article] [PubMed] [Google Scholar]
  9. Cooner, F. , Banerjee S., and McBean A. M.. 2006. “Modelling Geographically Referenced Survival Data With a Cure Fraction.” Statistical Methods in Medical Research 15, no. 4: 307–324. [DOI] [PMC free article] [PubMed] [Google Scholar]
  10. Cox, D. R. 1972. “Regression Models and Life‐Tables.” Journal of the Royal Statistical Society: Series B (Methodological) 34, no. 2: 187–202. [Google Scholar]
  11. de Oliveira, M. R. , Moreira F., and Louzada F.. 2017. “The Zero‐Inflated Promotion Cure Rate Model Applied to Financial Data on Time‐to‐Default.” Cogent Economics & Finance 5, no. 1: 1395950. [Google Scholar]
  12. Elbers, C. , and Ridder G.. 1982. “True and spurious duration dependence: The identifiability of the proportional hazard model.” The Review of Economic Studies 49, no. 3: 403–409. [Google Scholar]
  13. Hosmer, D. W. Jr , Lemeshow S., and Sturdivant R. X.. 2013. Applied Logistic Regression. Wiley Series in Probability and Statistics, Vol 398. John Wiley & Sons. [Google Scholar]
  14. Hougaard, P. 1995. “Frailty Models for Survival Data.” Lifetime Data Analysis 1, no. 3: 255–273. [DOI] [PubMed] [Google Scholar]
  15. Kaplan, E. L. , and Meier P.. 1958. “Nonparametric Estimation From Incomplete Observations.” Journal of the American Statistical Association 53, no. 282: 457–481. [Google Scholar]
  16. Leão, J. , Leiva V., Saulo H., and Tomazella V.. 2018. “Incorporation of Frailties Into a Cure Rate Regression Model and Its Diagnostics and Application to Melanoma Data.” Statistics in Medicine 37, no. 29: 4421–4440. [DOI] [PubMed] [Google Scholar]
  17. Louzada, F. , Moreira F. F., and de Oliveira M. R.. 2018. “A Zero‐Inflated Non Default Rate Regression Model for Credit Scoring Data.” Communications in Statistics‐Theory and Methods 47, no. 12: 3002–3021. [Google Scholar]
  18. Louzada, F. , Ramos P. L., de Souza H. C. C., Oyeneyin L., and da Silva Castro Perdoná G.. 2025. “On the Unification of Zero‐Adjusted Cure Survival Models.” Communications in Statistics‐Simulation and Computation . Ahead of print, April 26. 10.1080/03610918.2025.2496297. [DOI]
  19. Maller, R. A. , and Zhou X.. 1996. Survival Analysis With Long‐Term Survivors. John Wiley & Sons. [Google Scholar]
  20. Meeker, W. Q. , Escobar L. A., and Pascual F. G.. 2022. Statistical Methods for Reliability Data. John Wiley & Sons. [Google Scholar]
  21. Migon, H. S. , Gamerman D., and Louzada F.. 2014. Statistical Inference: An Integrated Approach. CRC Press. [Google Scholar]
  22. Ospina, R. , and Ferrari S. L.. 2012. “A General Class of Zero‐or‐One Inflated Beta Regression Models.” Computational Statistics & Data Analysis 56, no. 6: 1609–1623. [Google Scholar]
  23. Pereira, G. H. , Botter D. A., and Sandoval M. C.. 2013. “A Regression Model for special proportions.” Statistical Modelling 13, no. 2: 125–151. [Google Scholar]
  24. Rahimzadeh, M. , Baghestani A. R., Gohari M. R., and Pourhoseingholi M. A.. 2014. “Estimation of the Cure Rate in Iranian Breast Cancer Patients.” Asian Pacific Journal of Cancer Prevention 15, no. 12: 4839–4842. [DOI] [PubMed] [Google Scholar]
  25. Rodrigues, J. , Cancho V. G., de Castro M., and Louzada‐Neto F.. 2009. “On the Unification of Long‐Term Survival Models.” Statistics & Probability Letters 79, no. 6: 753–759. [Google Scholar]
  26. RStudio Team ., 2020. RStudio: Integrated Development Environment for R . Boston, MA: RStudio, PBC. [Google Scholar]
  27. Shanno, D. F. 1970. “Conditioning of Quasi‐Newton Methods for Function Minimization.” Mathematics of Computation 24, no. 111: 647–656. [Google Scholar]
  28. Souza, A. C. , Costa R. A., Paganoti C. F., et al. 2019. “Can We Stratify the Risk for Insulin Need in Women Diagnosed Early With Gestational Diabetes by Fasting Blood Glucose?” The Journal of Maternal‐Fetal & Neonatal Medicine 32, no. 12: 2036–2041. [DOI] [PubMed] [Google Scholar]
  29. Tsodikov, A. , Ibrahim J. G., and Yakovlev A.. 2003. “Estimating Cure Rates From Survival Data: An Alternative to Two‐Component Mixture Models.” Journal of the American Statistical Association 98, no. 464: 1063–1078. [DOI] [PMC free article] [PubMed] [Google Scholar]
  30. Tsodikov, A. D. , Yakovlev A. Y., and Asselain B.. 1996. Stochastic Models of Tumor Latency and Their Biostatistical Applications. Series in Mathematical Biology and Medicine, Vol 1. World Scientific. [Google Scholar]
  31. Vaupel, J. W. , Manton K. G., and Stallard E.. 1979. “The Impact of Heterogeneity in Individual Frailty on the Dynamics of Mortality.” Demography 16, no. 3: 439–454. [PubMed] [Google Scholar]
  32. Yakovlev, A. , Asselain B., Bardou V., et al. 1993. “A Simple Stochastic Model of Tumor Recurrence and Its Application to Data on Premenopausal Breast Cancer.” Biometrie et Analyse de Donnees Spatio‐Temporelles 12: 66–82. [Google Scholar]
  33. Yu, B. , and Peng Y.. 2008. “Mixture Cure Models for Multivariate Survival Data.” Computational Statistics & Data Analysis 52, no. 3: 1524–1532. [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supporting File: bimj70162‐sup‐0001‐Datacode.zip.

Data Availability Statement

The data that support the findings of this study are available from the corresponding author upon reasonable request.


Articles from Biometrical Journal. Biometrische Zeitschrift are provided here courtesy of Wiley

RESOURCES