Skip to main content
Wiley Open Access Collection logoLink to Wiley Open Access Collection
. 2026 Jul 12;25(4):e70110. doi: 10.1002/pst.70110

The Impact of Model Misspecification on the Individual Causal Association in Surrogate Endpoint Evaluation

Gokce Deliorman 1,, Florian Stijven 2, Wim Van der Elst 3, Maria del Carmen Pardo 1,4, Ariel Alonso 2
PMCID: PMC13356973  PMID: 42437657

ABSTRACT

Surrogate endpoints are often used in place of expensive, delayed, or rare clinical endpoints in clinical trials. However, regulatory authorities require thorough evaluation to accept these surrogate endpoints as reliable substitutes. One evaluation approach is the information‐theoretic causal inference framework, which quantifies surrogacy using the individual causal association (ICA). Like most causal inference methods, this approach relies on models that are only partially identifiable. For continuous outcomes, a normal model is often used. In this study, we explored the effects of model misspecification across various scenarios. We first considered true data‐generating mechanisms based on multivariate t and log‐normal distributions. We then used D‐vine copulas with Gaussian, Clayton, Gumbel, and Frank families to vary the unidentifiable copulas involving counterfactual pairs while preserving the observable bivariate margins, and considered intuitive restrictions on nonidentified correlations, including positivity and conditional independence. In all settings, the identifiability issue was addressed through sensitivity analysis. Finally, we illustrate the proposed sensitivity analyses using clinical‐trial data from schizophrenia studies, evaluating the ICA under several modeling assumptions. The results show that, in most scenarios considered, the impact of model misspecification is small; however, certain departures from the assumed model can materially affect the surrogacy assessment.

Keywords: D‐vine copula, individual causal association (ICA), model misspecification, nonnormality, surrogate endpoint

1. Introduction

Humanitarian and commercial considerations have sparked a relentless search for methods to accelerate the development of new therapies while also reducing the associated costs. In response to this pressing need, researchers and medical professionals have eagerly explored the potential of surrogate endpoints, a general strategy that has gathered considerable interest over the last decades. Surrogate endpoints are alternative outcomes used to replace or supplement the primary clinical outcomes, known as true endpoints, in the evaluation of experimental treatments or interventions.

The primary advantage of surrogate endpoints is their ability to be measured earlier, more conveniently, or more frequently than the true endpoint. This capability not only speeds up clinical trials but also streamlines the evaluation process and makes it more resource‐friendly. As a result, regulatory agencies across the globe, especially those in the United States, Europe, and Japan, began to introduce policies to regulate the use of surrogate endpoints in the evaluation of new treatments [1, 2].

However, the critical question is how to establish the adequacy of a surrogate endpoint. In other words, how can we ensure that the treatment's effect on the surrogate accurately predicts its impact on a more clinically meaningful true endpoint? To tackle this challenge, researchers have proposed various definitions of validity and formal criteria. These approaches can be categorized into two frameworks: methodologies based on single trial data (single trial setting or STS) and methodologies using data from multiple clinical trials. Prentice [3] introduced the first statistical framework for evaluating surrogate endpoints, defining surrogacy, and establishing corresponding operational criteria in the context of STS. Freedman et al. [4] highlighted the limitations of Prentice's criteria and proposed the proportion of treatment effect (PTE) as an alternative measure. Buyse and Molenberghs [5] further extended the methodology by introducing the relative effect (RE) and adjusted association (AA) for assessing surrogates in the STS. The meta‐analytic framework introduced later by Buyse et al. [6], based on expected causal treatment effects across multiple clinical trials, is considered one of the most general and effective methods for evaluating surrogate endpoints [7]. However, its practical implementation is hindered by stringent data requirements, which are often unavailable in the early stages of drug development when surrogate endpoints are most needed [1, 2, 8]. Consequently, developing new methods for the STS remains important in the field and, in this scenario, Alonso [9] introduced a general definition of surrogacy based on individual causal treatment effects and the notion of information gain (equivalently, reduction in uncertainty). To operationalize this definition, a general surrogacy metric, the so‐called individual causal association (ICA), was proposed. The ICA quantifies the association between the individual causal treatment effects on the surrogate and the true endpoint via their mutual information. Additionally, Deliorman et al. [10] extended the ICA using Rényi divergence, with the original definition as a special case of this broader formulation. This area of research is known as the information‐theoretic causal inference (ITCI) framework.

Assessing the ICA requires a joint causal model for the potential outcomes of both variables and a suitable rescaling of mutual information with desirable mathematical properties. For continuous outcomes, Alonso et al. [11] introduced a four‐dimensional normal model together with an associated ICA definition. Because some parameters of this model are not estimable from the observed data, the ICA cannot be identified without strong, unverifiable assumptions. To address this, the authors proposed a simulation‐based sensitivity analysis in which the ICA is evaluated over ranges of unidentifiable parameters of the potential‐outcomes distribution that are compatible with the data; the resulting distribution of ICA values is then summarized using frequency plots and descriptive statistics. We refer the interested reader to Alonso et al. [11] for details.

Because the model is only partially identifiable, misspecified parametric assumptions (e.g., normality) may bias the ICA and thus the assessment of the surrogate. To our knowledge, this issue has not been systematically studied despite its practical importance. We address it here through a combination of theoretical derivations and simulation studies.

The paper is organized as follows. Section 2 introduces a single‐trial framework for assessing surrogacy with two continuous outcomes. Section 3 revisits the normal causal model. Section 4 studies misspecification in two settings: Section 4.1 analyzes the case where the true data‐generating mechanism is multivariate t, drawing on theoretical results, and Section 4.2 considers a multivariate log‐normal mechanism via a Monte‐Carlo simulations. Section 5 introduces D‐vine copulas and employs them to assess misspecification when the data‐generating mechanism is indistinguishable from a normal model on the basis of the observed data. In Section 6, a real‐life case study is analyzed under different modeling assumptions. Finally, Section 7 provides concluding remarks.

2. General Theoretical Background

In the following, we outline the general framework introduced by Alonso et al. [11] and Alonso [9] to evaluate the validity of a putative surrogate endpoint S for a true endpoint T. This evaluation is carried out using data from a single randomized clinical trial with a well‐defined population, where only two treatments (Z=0/1) are being evaluated in a parallel study design.

The Neyman–Rubin potential outcome model assumes that each patient has a four‐dimensional vector of potential outcomes Y=T0,T1,S0,S1T with Sz and Tz representing the outcome for the surrogate and the true endpoint under treatment Z=z. The practical implementation of the model is based on several assumptions [12]. First, the stable unit treatment value assumption (SUTVA) links the observed outcomes to the potential outcomes as follows:

(S,T)T=ZS1,T1T+(1Z)S0,T0T.

Second, the full exchangeability assumption states that the potential outcomes are independent of the assigned treatment: T0,T1,S0,S1TZ. These two assumptions are typically met in randomized clinical trials and they will be used throughout the remainder of this paper.

The vector of individual causal treatment effects is defined as Δ=(ΔT,ΔS)T where ΔS=S1S0 and ΔT=T1T0. Alonso et al. [9] introduced the following definition of surrogacy in the STS.

Definition 1

In the STS, we shall say that S is a good surrogate for T if ΔS conveys a substantial amount of information on ΔT.

The concept of information has been rigorously defined in information theory [13]. The amount of “shared” information between ΔS and ΔT can be quantified using the mutual information between these individual causal treatment effects, denoted by I(ΔT,ΔS). Mutual information is always nonnegative, zero if and only if ΔS and ΔT are independent, symmetric, and invariant under bijective transformations. Despite its attractive mathematical properties, the interpretation of mutual information can be challenging in some scenarios. For instance, it lacks an upper bound when ΔS and ΔT are continuous. This issue is addressed by mapping mutual information onto the unit interval, ensuring that it takes a value of zero when ΔT and ΔS are independent and one when there is a nontrivial transformation ϕ such that ΔT=ϕ(ΔS) with probability one.

The mutual information between ΔS and ΔT is a functional of the joint distribution f(ΔT,ΔS), which is completely determined by the distribution of the vector of potential outcomes Y. When both the surrogate and true endpoints are continuous, a causal inference model for Y is often constructed using a four‐dimensional normal distribution. Subsequently, based on this model, the surrogate's validity is assessed. This approach will be detailed in the following section.

3. The Normal Causal Model

Alonso et al. [11] assumed that YN(μ,Σ) with μ=μT0,μT1,μS0,μS1T and,

Σ=σT0T0σT0T1σT0S0σT0S1σT0T1σT1T1σT1S0σT1S1σT0S0σT1S0σS0S0σS0S1σT0S1σT1S1σS0S1σS1S1. (1)

Under this assumption, Δ=(ΔT,ΔS)T=AYNμΔ,ΣΔ, where μΔ=(β,α)T, β=E(ΔT), α=E(ΔS) and ΣΔ=AΣAT with A the corresponding contrast matrix. Furthermore, these authors proposed to assess Definition 1 using the Squared Information Correlation Coefficient (SICC) [14, 15],

RH2=1e2I(ΔT,ΔS) (2)

where I(ΔT,ΔS) represents the mutual information between ΔT and ΔS, as defined by:

I(ΔT,ΔS)=f(ΔT,ΔS)lnf(ΔT,ΔS)f(ΔT)f(ΔS)dΔTdΔS

Basically, I(ΔT,ΔS) is the Kullback–Leibler (KL) divergence between the joint density function f(ΔT,ΔS) and the product of the marginal density functions f(ΔT)f(ΔS) [16]. For continuous outcomes the SICC satisfies the properties given in Section 2 and, therefore, one may argue that (2) is a suitable metric to assess Definition 1, that is, one may argue that it is a good metric of surrogacy. Alonso et al. [11] called this metric the individual causal association or ICA. Under the normality assumption, 2I(ΔT,ΔS)=log1ρΔ2 where ρΔ=corr(ΔT,ΔS) is given by,

ρΔ=σT0T0σS0S0ρT0S0+σT1T1σS1S1ρT1S1σT1T1σS0S0ρT1S0σT0T0σS1S1ρT0S1σT0T0+σT1T12σT0T0σT1T1ρT0T1σS0S0+σS1S12σS0S0σS1S1ρS0S1,

and ρXY denotes the correlation between the potential outcomes X and Y. Therefore, under this model, one has ICA=RH2=ρΔ2.

As noted by Holland [17], only two of the four potential outcomes are observable in practice; thus, while correlations such as ρT0S0 and ρT1S1 are estimable, correlations involving the counterfactual pairs Si,Tj, Si,Sj, and Ti,Tj with ij are not. Consequently, the ICA cannot be assessed without untestable assumptions. Alonso et al. [11] addressed this challenge via a simulation‐based sensitivity analysis that estimates ICA over ranges of unidentifiable parameters compatible with the observed data. This strategy, however, still assumes correct parametric specification of the distribution of Y, namely, that the vector of potential outcomes follows a four‐dimensional normal distribution. Because only the joint distributions of S0,T0 and S1,T1 are identifiable, this normality assumption for Y cannot be fully tested. In the next sections, we investigate the consequences of incorrectly assuming normality under several nonnormal data‐generating mechanisms.

4. Model Misspecification

As mentioned previously, when the surrogate and true endpoints are continuous outcomes, the ICA is quantified under the assumption that the vector of potential outcomes Y follows a four‐dimensional normal distribution, hereafter referred to as ICAN. This assumption is critical for deriving the theoretical properties of ICAN and for its implementation in the R package Surrogate [18, 19, 20]. A critical question is the sensitivity of ICAN to deviations from the assumed normal model. In following sections, we explore the impact of model misspecification on the assessment of surrogacy. Specifically, we will consider three scenarios in which the potential outcome vector Y is assumed to follow a four‐dimensional normal distribution but the actual data‐generating mechanism (DGM) departs from this assumption.

4.1. The t‐Distribution‐Based Data‐Generating Mechanism

The multivariate t‐distribution is a special case of a more general family, the so‐called elliptical distributions [21]. One way of defining a p‐dimensional multivariate t‐distribution is based on the fact that if yN(0,Σ) and uχν2 with yu then x=μ+y/u+ν follows a multivariate t‐distribution with density function,

f(x|μ,Σ,ν)=Γ[(ν+p)/2]Γ(ν/2)(νπ)p/2|Σ|1/21+1ν(xμ)TΣ1(xμ)(ν+p)/2.

The multivariate t‐distribution has parameters Σ, μ, ν and it is denoted as xtp(μ,Σ,ν). The expected value of x equals E(x)=μ for ν>1 (else undefined), however, Σ is not the covariance matrix of x since Var (x)=ν/(ν2)Σ for (ν>2). It has some interesting mathematical properties. For instance, let us consider the p‐dimensional vector,

x=x1x2tp(μ,Σ,ν) (3)

with xi a pi‐dimensional vector (p1+p2=p) and let us further consider the partitions,

μ=μ1μ2andΣ=Σ11Σ12Σ21Σ22. (4)

Ding [22] shows that,

x2x1tp2μ2|1,ν+d1ν+p1Σ22|1,ν+p1

with the following conditional mean μ21=μ2+Σ21Σ111x1μ1, conditional variance Σ221=Σ22Σ12Σ111Σ21 and d1=x1μ1TΣ111x1μ1.

The multivariate t‐distribution is also affine invariant. In other words, if xtp(μ,Σ,ν) then z=Ax+btpAμ+b,AΣAT,ν with A0 [23]. This is particularly appealing for the purposes of this work. In fact, affine invariance implies that if the vector of potential outcomes follows a four‐dimensional t‐distribution, then the bivariate vector of individual causal treatment effects will follow a bivariate t‐distribution and an analytical expression for the mutual information is available in this case. In fact, let us consider again the decomposition given in Equations ((3), (4)). Arellano‐Valle et al. [21] provided an expression for the mutual information between x1 and x2,

Itx1,x2=INx1,x2+ζ(ν)withζ(ν)=logΓ(ν/2)Γν+p1+p2/2Γν+p1/2Γν+p2/2+ν+p22ψν+p22+ν+p12ψν+p12ν+p1+p22ψν+p1+p22ν2ψν2 (5)

where ψ(x)=d/dx[Γ(x)] is the so‐called digamma function and,

INx1,x2=12log|Σ|Σ11Σ22

Notice that in Equation (5) INx1,x2 actually quantifies the mutual information between x1 and x2 under the normality assumption, which can be easily shown by considering Var(x)=ν/(ν2)Σ (for ν>2). Interestingly, all the information due to Σ arises only from INx1,x2 while the information due to ν comes from the remaining terms. It can also be shown that as ν increases, the t mutual information converges to the normal mutual information.

Let us now assume that the true DGM for the vector of potential outcomes is given by Yt4(μ,Σ,ν) with μ and Σ as before. The previous properties imply that Δt2μΔ,ΣΔ,ν and, hence, the true ICA value denoted by ICAt is given by:

ICAt=1e2It(ΔT,ΔS)=1e2IN(ΔT,ΔS)2ζ(ν)=11ρΔ2e2ζ(ν)ICAt=11ICANe2ζ(ν)

where

ζ(ν)=2logΓ(ν/2)Γ[(1+ν)/2]ν2+(1+ν)ψ1+ν2(1+ν)ψν22+νν. (6)

To derive the final equation for ζ(ν), we use the properties of the gamma and digamma functions: Γ(1+z)=zΓ(z) and ψ(1+z)=ψ(z)+1/z.

As illustrated in Figure 1, when ν approaches infinity, ζ(ν) converges to zero and ICAt=ICAN. This is expected as the multivariate t‐distribution converges to a multivariate normal distribution as the degrees of freedom ν increase. Furthermore, Figure 2 plots the pairs ICAt,ICAN (dashed line) for (ν=3,4,5,7), with the continuous line representing the identity function y=x for reference. The figure clearly shows that using the normal causal inference model, when the true DGM is characterized by a multivariate t‐distribution, will have only a mild impact on the ICA. Indeed, the effect of the misspecification is noticeable only when ICAt is small and ν=3. In this scenario, the ICAN is slightly larger than the true ICAt. When ν4, the effect of the misspecification becomes negligible. This discussion demonstrates that the normal causal model yields practically meaningful results when the t‐causal model is the true underlying DGM.

FIGURE 1.

FIGURE 1

ζ(ν) as a function of ν.

FIGURE 2.

FIGURE 2

ICAt versus ICAN. The black line is the identity line ICAN=ICAt.

Although encouraging, this finding should be taken with caution. The t‐distribution shares many properties with the normal distribution and converges to the normal rather quickly as the degrees of freedom increase. Therefore, the next section explores the effect of misspecification using a log‐normal distribution. Addressing this case purely from a mathematical standpoint is not feasible due to the complexity of the algebra involved. Consequently, a Monte Carlo procedure will be employed.

4.2. The Log‐Normal Data‐Generating Mechanism

In Section 4.1, we showed that assuming a normal causal inference model, when the true DGM is a multivariate t‐distribution, generally has a negligible impact on the ICA value that would be estimated. However, exploring this issue when the true DGM is based on other multivariate distributions, such as the multivariate log‐normal distribution, becomes mathematically challenging due to the lack of closed form expressions. Therefore, in order to dive deeper into this problem, we resort to a Monte Carlo procedure in this section. To that end, we now assume that the vector of potential outcomes Y follows a four‐dimensional log‐normal distribution with the density function,

fy|μy,Σ=12π|Σ|4i=14yie(ln(y)μ)TΣ1(ln(y)μ)2.

Under the log‐normal DGM, the distribution of the individual causal treatment effects Δ does not have a closed form and the true ICA, denoted hereafter as ICAL, cannot be computed analytically. To approximate the distribution of Δ and calculate the corresponding ICAL, we used Monte Carlo simulations. Specifically, we generated 200 different pairs of μ and Σ for the underlying log‐normal distribution, using a normal distribution for μ and a Wishart distribution for Σ. For each of these pairs (settings), My=2000 vectors of potential outcomes Y and the corresponding vectors of individual causal treatment effects Δ were generated. Using the previous My=2000 vectors of individual causal treatment effects Δ the ICAL value (as given in Equation (2)) can be numerically approximated.

Furthermore, we also computed the ICAN under the assumption that Y follows a four‐dimensional normal distribution, with the same mean and variance as the correct log‐normal distribution. This was done by calculating ρΔ2=corr(ΔT,ΔS)2 using the My=2000 vectors of individual causal treatment effects generated in each setting.

The main findings are summarized in Figure 3, which shows d=ICALICAN across the different settings. The x‐axis displays codes that correspond to the settings, arranged in the order they were generated without any reordering. It is evident from the figure that using a misspecified model can significantly impact the results in some cases. The maximum observed value for d was 0.662, indicating that ICAN can be substantially smaller than the true ICAL. Although the misspecified model generally yields smaller ICAs than the correct model (d>0), it can also produce larger values, with the minimum difference observed being d=0.042.

FIGURE 3.

FIGURE 3

Difference between ICAL and ICAN.

We further explored the settings where this difference was small and large, as defined in the Supporting Information Appendix (Section 1). We observed that in cases where the difference between ICAL and ICAN was substantial, say larger than 0.3, the underlying log‐normal distribution used to generate the potential outcomes was notably different from a normal distribution. This includes the distribution of the identifiable margins, suggesting that one will likely be able to detect the misspecification in those cases. To further investigate these differences, we computed the Kullback–Leibler (KL) divergence between the log‐normal and normal distributions to assess whether a larger KL corresponds to a greater difference between ICAL and ICAN. However, no clear pattern was observed.

The previous analyses have led to the following preliminary conclusions: When the true DGM is characterized by a four‐dimensional t‐distribution, the normal causal inference model will still deliver meaningful results. In addition, when the true DGM is characterized by a four‐dimensional log‐normal distribution, the normal causal inference model may be misleading but mostly in the scenarios where misspecification is detectableusing observable data. This emphasizes the importance of checking whether the identifiable parts of the model are compatible with the normality assumption. However, an intriguing question remains: Could a true DGM, which appears identical to the normal causal model based on observed data, lead to a different assessment of the surrogate? This issue will be examined in the following section through the use of D‐vine copulas.

5. Data Generating Mechanism Indistinguishable from the Normal Model

D‐vine copulas are a flexible class of copulas that allow for the modeling of complex dependence structures among multiple variables using bivariate copulas as building blocks. A D‐vine is a sequential tree‐like structure that allows any multivariate distribution to be factorized into a set of marginal and conditional bivariate distributions. These bivariate relationships are then described using copulas, making them easier to model and estimate individually, which is useful in the ITCI framework [24]. More details on (D‐vine) copulas and related concepts are provided in the Supporting Information Appendix (Section 2).

Let now fY be the joint density of Y=T0,S0,S1,T1T (note the reordering). The D‐vine density decomposition of fY is the product of four marginal densities and six bivariate copula densities:

fY=fT0fS0fS1fT1cT0S0cS0S1cS1T1cT0S0;S1cS0T1;S1cT0T1;S0S1, (7)

where (i) fT0, fS0, fS1, and fT1 are univariate density functions, (ii) cT0,S0, cS0,S1, and cS1,T1 are unconditional bivariate copula densities, and (iii) cT0,S1;S0, cS0,T1;S1, and cT0,T1;S0,S1 are conditional bivariate copula densities. The so‐called simplifying assumption, as detailed in Definition 3 in the Supporting Information Appendix and commonly used in practice Czado [25], will be adopted in this section. This assumption posits that the three conditional copula densities in (7) do not depend on the value of the conditioning variable and it is commonly used in practice because it simplifies the D‐vine copula model. The distribution of the observable bivariate margins follows immediately from the components in (7) as follows:

fS0T0(s,t)=fS0(s)fT0(t)cT0S0FT0(t),FS0(s)andfS1T1(s,t)=fS1(s)fT1(t)cS1T1FS1(s),FT1(t)

where FT0,FS0,FS1, and FT1 are the marginal distribution functions. If the marginal distributions are normal, and both cT0S0 and cS1T1 are Gaussian copulas, then the densities fS0T0 and fS1T1 will be bivariate normal. Consequently, the distribution of Y specified in (7) becomes indistinguishable from a four‐dimensional normal distribution based on observed data, regardless of the parametric choices for cS0S1cS0T1;S1cT0T1;S0S1.

As part of our objective to construct true DGMs that are indistinguishable from the normal causal model, we ensure that the observable bivariate margins, fS0T0 and fS1T1, are set as bivariate normal distributions for S0,T0T and S1,T1T. The parameters of these distributions were fixed to the estimated values obtained from two distinct case studies: age‐related macular degeneration (ARMD) and a data set from five randomized clinical trials focused on schizophrenia. It is important to emphasize that, in this section, ARMD and schizophrenia are not presented as two full applied case studies. Rather, they are used as two real‐data calibrations of the identifiable parts of the causal model. This allows us to evaluate the impact of alternative assumptions on the unidentifiable copulas under realistic observable margins. The full applied case study is presented separately in Section 6, where we focus on the schizophrenia data set.

Additionally, for the four unidentifiable copulas in (7), we explore four parametric choices: Gaussian, Clayton, Gumbel, and Frank.

We evaluated all combinations of these parametric copulas, resulting in 44=256 unique distributions for Y. A configuration consisting entirely of Gaussian copulas produces the four‐dimensional normal distribution. Regardless of the chosen copula types, the unidentifiable copulas are parameterized by Spearman's rho correlation parameters. Let Fk(ρ), for k=1,2,,256, denote one of these 256 distributions of Y, where ρ is the four‐dimensional vector of correlations that characterizes the association of the counterfactual pairs Si,Tj with ij. Following the methodology in Alonso [11], we sample 1000 ρ vectors under three distinct sampling schemes:

  1. No additional assumptions: The rho parameters are sampled from U(1,1).

  2. Positive restricted associations: The rho parameters are assumed to be positive and bounded away from zero and one.

  3. Conditional independence and positive restricted associations: In addition to positive restricted associations, we assume conditional independence: T0S1S0 and T1S0S1.

Basically, for each of the 256 distributions Fk(ρ) previously defined, we generate 1000 realizations in each of the scenarios (a)–(c). Notably, each distribution has normal observable margins, though only one of them is actually a four‐dimensional normal distribution. We analyze the differences d(ρ,F)=ICAF(ρ)ICAN(ρ) for each scenario, where ICAF(ρ) and ICAN(ρ) are the Individual Causal Associations calculated from the density Fk(ρ) and a four‐dimensional normal distribution, respectively, both derived from the same ρ vector. Since both distributions exhibit the same observable margins and identical values for the unidentifiable parameters ρ, the difference d(ρ,F) effectively highlights the impact of the parametric copulas that characterize the distribution of the counterfactual pairs on the ICA.

Figure 4 summarizes the main findings. The rows correspond to the two identifiable normal margins, fS0T0 and fS1T1, derived from both data sets. The columns represent the sampling schemes described in (a)–(c). The results show that although the impact of misspecification is generally mild (d(ρ,F)<0.13) in most cases, there are instances where distributions, indistinguishable from the normal based on observed data, can lead to significantly different assessments of surrogacy. It is important to highlight again that the differences d(ρ,F) quantify the impact of the parametric copulas used to characterize the distribution of the counterfactual pairs. Interestingly, when the association structure of these counterfactual pairs is restricted, as in setting (c), the effect of the misspecification substantially diminishes.

FIGURE 4.

FIGURE 4

D‐vine copula sensitivity analysis calibrated using identifiable margins estimated from two real data examples: ARMD and schizophrenia. The two rows correspond to the two choices of observable bivariate margins, fS0T0 and fS1T1, used to anchor the simulation at realistic parameter values. The columns correspond to the three sampling schemes for the unidentifiable associations: no additional assumptions, positive restricted associations (PA), and positive restricted associations plus conditional independence (CI). The plotted quantity is d=ICAFICAN, the difference between the ICA computed under the D‐vine copula density (ICAF) and the ICA obtained under the corresponding multivariate normal model (ICAN).

6. Applied Case Study: Schizophrenia Data

Schizophrenia is a psychiatric disorder characterized by psychotic manifestations such as delusions, hallucinations, and disorganized thinking and behavior [26]. We analyze a dataset comprising five double‐blind randomized clinical trials in patients with schizophrenia, available in R via the Surrogate package [27]. In this section we focus on the principal findings; implementation details and full R code are provided in Supporting Information Appendix (Section 3).

A total of 2123 patients participated, with 1589 allocated to Risperidone and 534 to control. Symptoms were assessed using the Positive and Negative Syndrome Scale (PANSS) and the Brief Psychiatric Rating Scale (BPRS). PANSS comprises 30 items, whereas BPRS is an 18‐item PANSS subscale [28]. We treat PANSS, a reliable but comparatively complex scale, as the true endpoint and BPRS, a simpler, easier‐to‐administer subscale, as the surrogate. The objective is to determine whether the BPRS subscale can reliably serve as a substitute for the more complex PANSS scale in evaluating the efficacy of Risperidone or similar antipsychotic treatments.

In prior work, both endpoints were treated as approximately normal [29]. We began by reassessing this assumption: density and Q–Q plots (Web Appendix C) suggested near normality for both endpoints, yet the Shapiro–Wilk test rejected normality for all except T0. With a large sample, this test has high power and may flag small, practically negligible departures from normality. Results from earlier sections indicate that mild deviations from normality typically have only a negligible or modest impact on surrogacy assessment; however, it is difficult to know a priori how substantial the impact may be in any concrete application. We therefore conducted the analysis under multiple modeling assumptions and compared results as a sensitivity analysis.

We first assessed the ICA under the assumption of normally distributed potential outcomes, using the closed‐form expression in Section 3. To address the identifiability issues described there, we implemented the sensitivity analysis of Alonso et al. [11], estimating the ICA over a wide range of unidentifiable correlations compatible with the observed data. Figure 5 displays the resulting frequency distribution of ICA values under normality, denoted ICAN. Although small ICAN values are compatible with the data, most values from the sensitivity analysis exceed 0.8. Specifically, ICAN ranged from 0.39 to 0.99, with a mean of about 0.92 and first and third quartiles of 0.91 and 0.95, respectively. Repeating the analysis with the t‐causal model produced essentially identical results, fully aligning with Section 4.1.

FIGURE 5.

FIGURE 5

Density plot of ICA values for the schizophrenia case study obtained from the three models as ICAN, ICAt, and ICAL.

For the multivariate log‐normal model, we proceeded analogously: we estimated the identifiable log‐normal parameters from the data, then gridded the remaining nonidentifiable correlations while holding ρS0,T0 and ρS1,T1 at their estimated values. We retained only grid points yielding a positive‐definite covariance matrix and computed the ICA for each. Figure 5 shows the corresponding distribution of ICAL values. Notably, diagnostics on the observable margins suggest that a log‐normal specification is a poor description of these data, so this constitutes a misspecified scenario. Even so, the log‐normal model yielded slightly larger ICA values with a narrower range (0.450.99), a mean and median of 0.96, and quartiles of 0.95 and 0.97. Thus, despite differing from the normal/t results, this analysis provides even stronger evidence in favor of the surrogate.

Finally, we analyzed the data using a D‐vine copula specification. In this setup, four pair‐copulas are unidentifiable and must be fixed by untestable parametric choices. We considered the Gaussian, Gumbel, Frank, and Clayton families for each of these four components (Section 5), yielding 44=256 candidate models. Figure 6 displays the frequency distributions for a random subset of four models under three sampling schemes for the unidentifiable associations: (i) no additional constraints, (ii) positive association, and (iii) positive association plus conditional independence. Three observations are in order. First, within each sampling scheme (i)–(iii), the frequency distributions across models look strikingly similar, indicating that the particular parametric choice for the unidentifiable pair‐copulas has little influence once the sampling is fixed. Second, imposing untestable constraints on the association structure can materially change the results; for example, restricting all associations to be positive has a pronounced effect on the distributional shape. The contrasts across sampling schemes illustrate the risks of relying on untestable association assumptions, a common practice in causal inference (e.g., PTE, RE), and underscore the value of our sensitivity analysis. Third, when all pair‐copulas are Gaussian and no restrictions are imposed on the association structure, the results closely match those obtained under the normal causal model, as expected.

FIGURE 6.

FIGURE 6

Frequency distributions of the ICA for the schizophrenia case study under various D‐vine copula models and sampling schemes.

As noted in Section 1, several methods have been proposed to assess surrogacy in the STS, including the Proportion of Treatment Effect (PTE) introduced by Freedman et al. [4] and refined by Parast et al. [30]. The PTE is defined in terms of expected causal treatment effects and, because there is no replication at this level in a single trial, its identification requires untestable assumptions. Moreover, as a ratio, the PTE can be numerically unstable when the overall expected treatment effect on the true endpoint is small. Stijven et al. [31] provide a comprehensive discussion of the PTE's definition, interpretation, and limitations. Using the robust approach of Parast et al. [30], the estimated PTE in our case study was 0.84 (95% CI: 0.66–1), which supports the surrogate's validity. We also estimated the adjusted association (AA) as 0.963 (95% CI: 0.9594–0.9667); the AA quantifies the association between both endpoints after adjusting for treatment. Note that, while a high AA may be necessary for surrogacy, it is not a sufficient condition because a high correlation alone is not enough to deem a surrogate valid.

Taken together, these results support BPRS as a viable surrogate for PANSS. Nonetheless, side‐by‐side comparisons of multiple metrics should be interpreted with care, because the estimands, assumptions, and sources of uncertainty differ across methods [32].

7. Discussion

In this study, we investigated the effects of model misspecification on the behavior of the ICA when the true distribution deviates from a four‐dimensional normal model. Through theoretical derivations and simulations we examined scenarios where the true DGM was characterized by a four‐dimensional t‐distribution and a log‐normal distribution. Our results show that the impact of model misspecification can vary, ranging from negligible when the underlying model follows a multivariate t‐distribution to significant under a log‐normal DGM. Notably, in the latter case, the substantial impact of misspecification was primarily observable when deviations from normality could be detected using the observed data.

Additionally, by employing D‐vine copulas, we explored scenarios where the true DGM was indistinguishable from the normal model based on observed data. In this setting, it is important to note that normality of the bivariate normal margins does not guarantee multivariate normality of the entire distribution. We found that in such cases, assuming normality might lead to misleading results, even though the observed data would support this assumption.

For any surrogacy metric, defining a threshold that guarantees good surrogacy is very challenging. The general intuition is that the further the true distribution deviates from the assumed model, such as from normality, the greater the potential impact on the results. In situations with substantial uncertainty regarding model assumptions, we recommend conducting sensitivity analyses under alternative plausible distributions to evaluate the robustness of the conclusions, as illustrated in the case study. Moreover, ICA results can be compared with those obtained from alternative surrogate evaluation measures. Although these methods may differ in methodology and computational implementation, it is desirable that they lead to conclusions in the same direction.

Our findings highlight the complexity of evaluating surrogates using causal inference models, which are only partially identifiable from the data. The unidentifiable components can substantially impact the assessment of surrogacy under certain conditions. Practically, our results underscore the importance of verifying that the identifiable components of the model are consistent with the observed data and suggest that considering different modeling assumptions for the counterfactual pairs is crucial to understanding how they influence the final conclusions.

Funding

This research was supported by the MCIN/AEI/10.13039/501100011033/ and by ERDF A way of making Europe (Grant/Award Number: PID2022‐137050NB‐I00) and Agentschap Innoveren & Ondernemen and Janssen Pharmaceutical Companies of Johnson & Johnson Innovative Medicine through a Baekeland Mandate (Grant/Award Number: HBC.2022.0145).

Ethics Statement

The authors have nothing to report.

Conflicts of Interest

The authors declare no conflicts of interest.

Supporting information

Figure S1: The density graphs of potential outcome vector when the difference is −0.00079.

Figure S2: The density graphs of potential outcome vector when the difference is 0.0058.

Figure S3: The density graphs of potential outcome vector when the difference is 0.0543.

Figure S4: The density graphs of potential outcome vector when the difference is 0.5705.

Figure S5: The density graphs of potential outcome vector when the difference is 0.3464.

Figure S6: The density graphs of potential outcome vector when the difference is 0.4398.

Table S1: Interpretation of the unidentifiable components of the D‐vine density fY in (3).

Table S2: Estimated bivariate normal distributions in the ARMD and Schizo data sets from the Surrogate R package.

Figure S7: Histograms of endpoints.

Figure S8: Q–Q plots of endpoints.

PST-25-0-s001.pdf (1.3MB, pdf)

Acknowledgments

This work is supported by grant PID2022‐137050NB‐I00 of the Spanish Ministry of Science and Innovation. Florian Stijven gratefully acknowledges funding from Agentschap Innoveren & Ondernemen and Janssen Pharmaceutical Companies of Johnson & Johnson Innovative Medicine through a Baekeland Mandate [grant number HBC.2022.0145]. We sincerely appreciate the meaningful suggestions and feedback provided by the editor which have substantially improved the quality of our study.

Data Availability Statement

The data that support the findings of this study are available in R Surrogate package. These data were derived from the following resources available in the public domain: https://cran.r‐project.org/web/packages/Surrogate/index.html.

References

  • 1. FDA U, CDER, CBER , Guidance for Industry: Expedited Programs for Serious Conditions – Drugs and Biologics (US FDA, 2014). [Google Scholar]
  • 2. Burzykowski T., Molenberghs G., and Buyse M., The Evaluation of Surrogate Endpoints (Springer‐Verlag, 2005). [Google Scholar]
  • 3. Prentice R. L., “Surrogate Endpoints in Clinical Trials: Definition and Operational Criteria,” Statistics in Medicine 8, no. 4 (1989): 431–440. [DOI] [PubMed] [Google Scholar]
  • 4. Freedman L. S., Graubard B. I., and Schatzkin A., “Statistical Validation of Intermediate Endpoints for Chronic Diseases,” Statistics in Medicine 11, no. 2 (1992): 167–178. [DOI] [PubMed] [Google Scholar]
  • 5. Buyse M. and Molenberghs G., “Criteria for the Validation of Surrogate Endpoints in Randomized Experiments,” Biometrics 54 (1998): 1014–1029. [PubMed] [Google Scholar]
  • 6. Buyse M., Molenberghs G., Burzykowski T., Renard D., and Geys H., “The Validation of Surrogate Endpoints in Meta‐Analyses of Randomized Experiments,” Biostatistics 1 (2000): 49–67. [DOI] [PubMed] [Google Scholar]
  • 7. Weir C. J. and Taylor R. S., “Informed Decision‐Making: Statistical Methodology for Surrogacy Evaluation and Its Role in Licensing and Reimbursement Assessments,” Pharmaceutical Statistics 21, no. 4 (2022): 740–756. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8. Alonso A., Bigirumurame T., Burzykowski T., et al., Applied Surrogate Endpoint Evaluation Methods With SAS and R (Chapman Hall/CRC, 2017). [Google Scholar]
  • 9. Alonso A., An Information‐Theoretic Approach for the Evaluation of Surrogate Endpoints (Wiley StatsRef: Statistics Reference Online, 2018). [Google Scholar]
  • 10. Deliorman G., Alonso A., and Pardo M. C., “A Rényi‐Divergence‐Based Family of Metrics for the Evaluation of Surrogate Endpoints in a Causal Inference Framework,” Statistics in Biopharmaceutical Research 18, no. 1 (2026): 31–38. [Google Scholar]
  • 11. Alonso A., Van Der Elst W., Molenberghs G., Buyse M., and Burzykowski T., “On the Relationship Between the Causal‐Inference and Meta‐Analytic Paradigms for the Validation of Surrogate Endpoints,” Biometrics 71, no. 1 (2015): 15–24. [DOI] [PubMed] [Google Scholar]
  • 12. Imbens G. W. and Rubin D. B., Causal Inference in Statistics, Social, and Biomedical Sciences (Cambridge University Press, 2015). [Google Scholar]
  • 13. Cover T. M. and Thomas J. A., “Information Theory and Statistics,” Elements of Information Theory 1, no. 1 (1991): 279–335. [Google Scholar]
  • 14. Linfoot E. H., “An Informational Measure of Correlation,” Information and Control 1, no. 1 (1957): 85–89. [Google Scholar]
  • 15. Joe H., “Relative Entropy Measures of Multivariate Dependence,” Journal of the American Statistical Association 84, no. 405 (1989): 157–164. [Google Scholar]
  • 16. Kullback S. and Leibler R. A., “On Information and Sufficiency,” Annals of Mathematical Statistics 22, no. 1 (1951): 79–86. [Google Scholar]
  • 17. Holland P. W., “Statistics and Causal Inference,” Journal of the American Statistical Association 81, no. 396 (1986): 945–960. [Google Scholar]
  • 18. Elst V. d W., Molenberghs G., and Alonso A., “Exploring the Relationship Between the Causal‐Inference and Meta‐Analytic Paradigms for the Evaluation of Surrogate Endpoints,” Statistics in Medicine 35, no. 8 (2016): 1281–1298. [DOI] [PubMed] [Google Scholar]
  • 19. Van Der Elst W., Alonso A., Coppenolle H., Meyvisch P., and Molenberghs G., “The Individual‐Level Surrogate Threshold Effect in a Causal Inference Setting With Normally Distributed Endpoints,” Pharmaceutical Statistics 20, no. 6 (2021): 1216–1231. [DOI] [PubMed] [Google Scholar]
  • 20. Alonso A., Van Der Elst W., Molenberghs G., and Florez A. J., “A Reflection on the Causal Interpretation of Individual‐Level Surrogacy,” Journal of Biopharmaceutical Statistics 29, no. 3 (2019): 529–540. [DOI] [PubMed] [Google Scholar]
  • 21. Arellano‐Valle R. B., Conreras‐Reyes J. E., and Genton M. G., “Shannon Entropy and Mutual Information for Multivariate Skew‐Elliptical Distributions,” Scandinavian Journal of Statistics 40, no. 1 (2013): 42–62. [Google Scholar]
  • 22. Ding P., “On the Conditional Distribution of the Multivariate t Distribution,” American Statistician 70, no. 3 (2016): 293–295. [Google Scholar]
  • 23. Kotz S. and Nadarajah S., Multivariate t‐Distributions and Their Applications (Cambridge University Press, 2004). [Google Scholar]
  • 24. Stijven F., Molenberghs G., V. Keilegom, I , Elst V. d W., and Alonso A., “Evaluating Time‐To‐Event Surrogates for Time‐To‐Event True Endpoints: An Information‐Theoretic Approach Based on Causal Inference,” Lifetime Data Analysis 31, no. 1 (2025): 1–23. [DOI] [PubMed] [Google Scholar]
  • 25. Czado C., Analyzing Dependent Data With Vine Copulas, vol. 222 (Springer, 2019). [Google Scholar]
  • 26. Offermanns S. and Rosenthal W., Encyclopedic Reference of Molecular Pharmacology (Springer Verlag, 2004). [Google Scholar]
  • 27. Van Der Elst W., Meyvisch P., Alonso A., et al., “Package Surrogate,” 2022.
  • 28. Peuskens J., “Risperidone in the Treatment of Patients With Chronic Schizophrenia: A Multi‐National, Multi‐Centre, Double‐Blind, Parallel‐Group Study Versus Haloperidol,” British Journal of Psychiatry 166, no. 6 (1995): 712–726. [DOI] [PubMed] [Google Scholar]
  • 29. Alonso A., Bigirumurame T., Burzykowski T., et al., Applied Surrogate Endpoint Evaluation Methods With Sas and r (CRC Press, 2016). [Google Scholar]
  • 30. Parast L., McDermott M. M., and Tian L., “Robust Estimation of the Proportion of Treatment Effect Explained by Surrogate Marker Information,” Statistics in Medicine 35, no. 10 (2016): 1637–1653. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31. Stijven F., Ariel A., and Molenberghs G., “Proportion of Treatment Effect Explained: An Overview of Interpretations,” Statistical Methods in Medical Research 33, no. 7 (2024): 1278–1296. [DOI] [PubMed] [Google Scholar]
  • 32. Joffe M. M. and Greene T., “Related Causal Frameworks for Surrogate Outcomes,” Biometrics 65, no. 2 (2009): 530–538. [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Figure S1: The density graphs of potential outcome vector when the difference is −0.00079.

Figure S2: The density graphs of potential outcome vector when the difference is 0.0058.

Figure S3: The density graphs of potential outcome vector when the difference is 0.0543.

Figure S4: The density graphs of potential outcome vector when the difference is 0.5705.

Figure S5: The density graphs of potential outcome vector when the difference is 0.3464.

Figure S6: The density graphs of potential outcome vector when the difference is 0.4398.

Table S1: Interpretation of the unidentifiable components of the D‐vine density fY in (3).

Table S2: Estimated bivariate normal distributions in the ARMD and Schizo data sets from the Surrogate R package.

Figure S7: Histograms of endpoints.

Figure S8: Q–Q plots of endpoints.

PST-25-0-s001.pdf (1.3MB, pdf)

Data Availability Statement

The data that support the findings of this study are available in R Surrogate package. These data were derived from the following resources available in the public domain: https://cran.r‐project.org/web/packages/Surrogate/index.html.


Articles from Pharmaceutical Statistics are provided here courtesy of Wiley

RESOURCES