ABSTRACT
Surrogate endpoints are often used in place of expensive, delayed, or rare clinical endpoints in clinical trials. However, regulatory authorities require thorough evaluation to accept these surrogate endpoints as reliable substitutes. One evaluation approach is the information‐theoretic causal inference framework, which quantifies surrogacy using the individual causal association (ICA). Like most causal inference methods, this approach relies on models that are only partially identifiable. For continuous outcomes, a normal model is often used. In this study, we explored the effects of model misspecification across various scenarios. We first considered true data‐generating mechanisms based on multivariate and log‐normal distributions. We then used D‐vine copulas with Gaussian, Clayton, Gumbel, and Frank families to vary the unidentifiable copulas involving counterfactual pairs while preserving the observable bivariate margins, and considered intuitive restrictions on nonidentified correlations, including positivity and conditional independence. In all settings, the identifiability issue was addressed through sensitivity analysis. Finally, we illustrate the proposed sensitivity analyses using clinical‐trial data from schizophrenia studies, evaluating the ICA under several modeling assumptions. The results show that, in most scenarios considered, the impact of model misspecification is small; however, certain departures from the assumed model can materially affect the surrogacy assessment.
Keywords: D‐vine copula, individual causal association (ICA), model misspecification, nonnormality, surrogate endpoint
1. Introduction
Humanitarian and commercial considerations have sparked a relentless search for methods to accelerate the development of new therapies while also reducing the associated costs. In response to this pressing need, researchers and medical professionals have eagerly explored the potential of surrogate endpoints, a general strategy that has gathered considerable interest over the last decades. Surrogate endpoints are alternative outcomes used to replace or supplement the primary clinical outcomes, known as true endpoints, in the evaluation of experimental treatments or interventions.
The primary advantage of surrogate endpoints is their ability to be measured earlier, more conveniently, or more frequently than the true endpoint. This capability not only speeds up clinical trials but also streamlines the evaluation process and makes it more resource‐friendly. As a result, regulatory agencies across the globe, especially those in the United States, Europe, and Japan, began to introduce policies to regulate the use of surrogate endpoints in the evaluation of new treatments [1, 2].
However, the critical question is how to establish the adequacy of a surrogate endpoint. In other words, how can we ensure that the treatment's effect on the surrogate accurately predicts its impact on a more clinically meaningful true endpoint? To tackle this challenge, researchers have proposed various definitions of validity and formal criteria. These approaches can be categorized into two frameworks: methodologies based on single trial data (single trial setting or STS) and methodologies using data from multiple clinical trials. Prentice [3] introduced the first statistical framework for evaluating surrogate endpoints, defining surrogacy, and establishing corresponding operational criteria in the context of STS. Freedman et al. [4] highlighted the limitations of Prentice's criteria and proposed the proportion of treatment effect (PTE) as an alternative measure. Buyse and Molenberghs [5] further extended the methodology by introducing the relative effect (RE) and adjusted association (AA) for assessing surrogates in the STS. The meta‐analytic framework introduced later by Buyse et al. [6], based on expected causal treatment effects across multiple clinical trials, is considered one of the most general and effective methods for evaluating surrogate endpoints [7]. However, its practical implementation is hindered by stringent data requirements, which are often unavailable in the early stages of drug development when surrogate endpoints are most needed [1, 2, 8]. Consequently, developing new methods for the STS remains important in the field and, in this scenario, Alonso [9] introduced a general definition of surrogacy based on individual causal treatment effects and the notion of information gain (equivalently, reduction in uncertainty). To operationalize this definition, a general surrogacy metric, the so‐called individual causal association (ICA), was proposed. The ICA quantifies the association between the individual causal treatment effects on the surrogate and the true endpoint via their mutual information. Additionally, Deliorman et al. [10] extended the ICA using Rényi divergence, with the original definition as a special case of this broader formulation. This area of research is known as the information‐theoretic causal inference (ITCI) framework.
Assessing the ICA requires a joint causal model for the potential outcomes of both variables and a suitable rescaling of mutual information with desirable mathematical properties. For continuous outcomes, Alonso et al. [11] introduced a four‐dimensional normal model together with an associated ICA definition. Because some parameters of this model are not estimable from the observed data, the ICA cannot be identified without strong, unverifiable assumptions. To address this, the authors proposed a simulation‐based sensitivity analysis in which the ICA is evaluated over ranges of unidentifiable parameters of the potential‐outcomes distribution that are compatible with the data; the resulting distribution of ICA values is then summarized using frequency plots and descriptive statistics. We refer the interested reader to Alonso et al. [11] for details.
Because the model is only partially identifiable, misspecified parametric assumptions (e.g., normality) may bias the ICA and thus the assessment of the surrogate. To our knowledge, this issue has not been systematically studied despite its practical importance. We address it here through a combination of theoretical derivations and simulation studies.
The paper is organized as follows. Section 2 introduces a single‐trial framework for assessing surrogacy with two continuous outcomes. Section 3 revisits the normal causal model. Section 4 studies misspecification in two settings: Section 4.1 analyzes the case where the true data‐generating mechanism is multivariate , drawing on theoretical results, and Section 4.2 considers a multivariate log‐normal mechanism via a Monte‐Carlo simulations. Section 5 introduces D‐vine copulas and employs them to assess misspecification when the data‐generating mechanism is indistinguishable from a normal model on the basis of the observed data. In Section 6, a real‐life case study is analyzed under different modeling assumptions. Finally, Section 7 provides concluding remarks.
2. General Theoretical Background
In the following, we outline the general framework introduced by Alonso et al. [11] and Alonso [9] to evaluate the validity of a putative surrogate endpoint for a true endpoint . This evaluation is carried out using data from a single randomized clinical trial with a well‐defined population, where only two treatments () are being evaluated in a parallel study design.
The Neyman–Rubin potential outcome model assumes that each patient has a four‐dimensional vector of potential outcomes with and representing the outcome for the surrogate and the true endpoint under treatment . The practical implementation of the model is based on several assumptions [12]. First, the stable unit treatment value assumption (SUTVA) links the observed outcomes to the potential outcomes as follows:
Second, the full exchangeability assumption states that the potential outcomes are independent of the assigned treatment: . These two assumptions are typically met in randomized clinical trials and they will be used throughout the remainder of this paper.
The vector of individual causal treatment effects is defined as where and . Alonso et al. [9] introduced the following definition of surrogacy in the STS.
Definition 1
In the STS, we shall say that is a good surrogate for if conveys a substantial amount of information on .
The concept of information has been rigorously defined in information theory [13]. The amount of “shared” information between and can be quantified using the mutual information between these individual causal treatment effects, denoted by . Mutual information is always nonnegative, zero if and only if and are independent, symmetric, and invariant under bijective transformations. Despite its attractive mathematical properties, the interpretation of mutual information can be challenging in some scenarios. For instance, it lacks an upper bound when and are continuous. This issue is addressed by mapping mutual information onto the unit interval, ensuring that it takes a value of zero when and are independent and one when there is a nontrivial transformation such that with probability one.
The mutual information between and is a functional of the joint distribution , which is completely determined by the distribution of the vector of potential outcomes . When both the surrogate and true endpoints are continuous, a causal inference model for is often constructed using a four‐dimensional normal distribution. Subsequently, based on this model, the surrogate's validity is assessed. This approach will be detailed in the following section.
3. The Normal Causal Model
Alonso et al. [11] assumed that with and,
| (1) |
Under this assumption, , where , , and with the corresponding contrast matrix. Furthermore, these authors proposed to assess Definition 1 using the Squared Information Correlation Coefficient (SICC) [14, 15],
| (2) |
where represents the mutual information between and , as defined by:
Basically, is the Kullback–Leibler (KL) divergence between the joint density function and the product of the marginal density functions [16]. For continuous outcomes the SICC satisfies the properties given in Section 2 and, therefore, one may argue that (2) is a suitable metric to assess Definition 1, that is, one may argue that it is a good metric of surrogacy. Alonso et al. [11] called this metric the individual causal association or ICA. Under the normality assumption, where is given by,
and denotes the correlation between the potential outcomes and . Therefore, under this model, one has .
As noted by Holland [17], only two of the four potential outcomes are observable in practice; thus, while correlations such as and are estimable, correlations involving the counterfactual pairs , , and with are not. Consequently, the cannot be assessed without untestable assumptions. Alonso et al. [11] addressed this challenge via a simulation‐based sensitivity analysis that estimates over ranges of unidentifiable parameters compatible with the observed data. This strategy, however, still assumes correct parametric specification of the distribution of , namely, that the vector of potential outcomes follows a four‐dimensional normal distribution. Because only the joint distributions of and are identifiable, this normality assumption for cannot be fully tested. In the next sections, we investigate the consequences of incorrectly assuming normality under several nonnormal data‐generating mechanisms.
4. Model Misspecification
As mentioned previously, when the surrogate and true endpoints are continuous outcomes, the ICA is quantified under the assumption that the vector of potential outcomes follows a four‐dimensional normal distribution, hereafter referred to as . This assumption is critical for deriving the theoretical properties of and for its implementation in the R package Surrogate [18, 19, 20]. A critical question is the sensitivity of to deviations from the assumed normal model. In following sections, we explore the impact of model misspecification on the assessment of surrogacy. Specifically, we will consider three scenarios in which the potential outcome vector is assumed to follow a four‐dimensional normal distribution but the actual data‐generating mechanism (DGM) departs from this assumption.
4.1. The t‐Distribution‐Based Data‐Generating Mechanism
The multivariate t‐distribution is a special case of a more general family, the so‐called elliptical distributions [21]. One way of defining a p‐dimensional multivariate t‐distribution is based on the fact that if and with then follows a multivariate t‐distribution with density function,
The multivariate t‐distribution has parameters , , and it is denoted as . The expected value of equals for (else undefined), however, is not the covariance matrix of since Var for . It has some interesting mathematical properties. For instance, let us consider the ‐dimensional vector,
| (3) |
with a ‐dimensional vector () and let us further consider the partitions,
| (4) |
Ding [22] shows that,
with the following conditional mean , conditional variance and .
The multivariate t‐distribution is also affine invariant. In other words, if then with [23]. This is particularly appealing for the purposes of this work. In fact, affine invariance implies that if the vector of potential outcomes follows a four‐dimensional t‐distribution, then the bivariate vector of individual causal treatment effects will follow a bivariate t‐distribution and an analytical expression for the mutual information is available in this case. In fact, let us consider again the decomposition given in Equations ((3), (4)). Arellano‐Valle et al. [21] provided an expression for the mutual information between and ,
| (5) |
where is the so‐called digamma function and,
Notice that in Equation (5) actually quantifies the mutual information between and under the normality assumption, which can be easily shown by considering (for ). Interestingly, all the information due to arises only from while the information due to comes from the remaining terms. It can also be shown that as increases, the t mutual information converges to the normal mutual information.
Let us now assume that the true DGM for the vector of potential outcomes is given by with and as before. The previous properties imply that and, hence, the true ICA value denoted by is given by:
where
| (6) |
To derive the final equation for , we use the properties of the gamma and digamma functions: and .
As illustrated in Figure 1, when approaches infinity, converges to zero and . This is expected as the multivariate t‐distribution converges to a multivariate normal distribution as the degrees of freedom increase. Furthermore, Figure 2 plots the pairs (dashed line) for , with the continuous line representing the identity function for reference. The figure clearly shows that using the normal causal inference model, when the true DGM is characterized by a multivariate t‐distribution, will have only a mild impact on the ICA. Indeed, the effect of the misspecification is noticeable only when is small and . In this scenario, the is slightly larger than the true . When , the effect of the misspecification becomes negligible. This discussion demonstrates that the normal causal model yields practically meaningful results when the t‐causal model is the true underlying DGM.
FIGURE 1.

as a function of .
FIGURE 2.

versus . The black line is the identity line .
Although encouraging, this finding should be taken with caution. The t‐distribution shares many properties with the normal distribution and converges to the normal rather quickly as the degrees of freedom increase. Therefore, the next section explores the effect of misspecification using a log‐normal distribution. Addressing this case purely from a mathematical standpoint is not feasible due to the complexity of the algebra involved. Consequently, a Monte Carlo procedure will be employed.
4.2. The Log‐Normal Data‐Generating Mechanism
In Section 4.1, we showed that assuming a normal causal inference model, when the true DGM is a multivariate t‐distribution, generally has a negligible impact on the ICA value that would be estimated. However, exploring this issue when the true DGM is based on other multivariate distributions, such as the multivariate log‐normal distribution, becomes mathematically challenging due to the lack of closed form expressions. Therefore, in order to dive deeper into this problem, we resort to a Monte Carlo procedure in this section. To that end, we now assume that the vector of potential outcomes follows a four‐dimensional log‐normal distribution with the density function,
Under the log‐normal DGM, the distribution of the individual causal treatment effects does not have a closed form and the true ICA, denoted hereafter as , cannot be computed analytically. To approximate the distribution of and calculate the corresponding , we used Monte Carlo simulations. Specifically, we generated 200 different pairs of and for the underlying log‐normal distribution, using a normal distribution for and a Wishart distribution for . For each of these pairs (settings), vectors of potential outcomes and the corresponding vectors of individual causal treatment effects were generated. Using the previous vectors of individual causal treatment effects the value (as given in Equation (2)) can be numerically approximated.
Furthermore, we also computed the under the assumption that follows a four‐dimensional normal distribution, with the same mean and variance as the correct log‐normal distribution. This was done by calculating using the vectors of individual causal treatment effects generated in each setting.
The main findings are summarized in Figure 3, which shows across the different settings. The x‐axis displays codes that correspond to the settings, arranged in the order they were generated without any reordering. It is evident from the figure that using a misspecified model can significantly impact the results in some cases. The maximum observed value for was , indicating that can be substantially smaller than the true . Although the misspecified model generally yields smaller ICAs than the correct model (), it can also produce larger values, with the minimum difference observed being .
FIGURE 3.

Difference between and .
We further explored the settings where this difference was small and large, as defined in the Supporting Information Appendix (Section 1). We observed that in cases where the difference between and was substantial, say larger than , the underlying log‐normal distribution used to generate the potential outcomes was notably different from a normal distribution. This includes the distribution of the identifiable margins, suggesting that one will likely be able to detect the misspecification in those cases. To further investigate these differences, we computed the Kullback–Leibler (KL) divergence between the log‐normal and normal distributions to assess whether a larger KL corresponds to a greater difference between and . However, no clear pattern was observed.
The previous analyses have led to the following preliminary conclusions: When the true DGM is characterized by a four‐dimensional t‐distribution, the normal causal inference model will still deliver meaningful results. In addition, when the true DGM is characterized by a four‐dimensional log‐normal distribution, the normal causal inference model may be misleading but mostly in the scenarios where misspecification is detectableusing observable data. This emphasizes the importance of checking whether the identifiable parts of the model are compatible with the normality assumption. However, an intriguing question remains: Could a true DGM, which appears identical to the normal causal model based on observed data, lead to a different assessment of the surrogate? This issue will be examined in the following section through the use of D‐vine copulas.
5. Data Generating Mechanism Indistinguishable from the Normal Model
D‐vine copulas are a flexible class of copulas that allow for the modeling of complex dependence structures among multiple variables using bivariate copulas as building blocks. A D‐vine is a sequential tree‐like structure that allows any multivariate distribution to be factorized into a set of marginal and conditional bivariate distributions. These bivariate relationships are then described using copulas, making them easier to model and estimate individually, which is useful in the ITCI framework [24]. More details on (D‐vine) copulas and related concepts are provided in the Supporting Information Appendix (Section 2).
Let now be the joint density of (note the reordering). The D‐vine density decomposition of is the product of four marginal densities and six bivariate copula densities:
| (7) |
where (i) , , , and are univariate density functions, (ii) , , and are unconditional bivariate copula densities, and (iii) , , and are conditional bivariate copula densities. The so‐called simplifying assumption, as detailed in Definition 3 in the Supporting Information Appendix and commonly used in practice Czado [25], will be adopted in this section. This assumption posits that the three conditional copula densities in (7) do not depend on the value of the conditioning variable and it is commonly used in practice because it simplifies the D‐vine copula model. The distribution of the observable bivariate margins follows immediately from the components in (7) as follows:
where , and are the marginal distribution functions. If the marginal distributions are normal, and both and are Gaussian copulas, then the densities and will be bivariate normal. Consequently, the distribution of specified in (7) becomes indistinguishable from a four‐dimensional normal distribution based on observed data, regardless of the parametric choices for .
As part of our objective to construct true DGMs that are indistinguishable from the normal causal model, we ensure that the observable bivariate margins, and , are set as bivariate normal distributions for and . The parameters of these distributions were fixed to the estimated values obtained from two distinct case studies: age‐related macular degeneration (ARMD) and a data set from five randomized clinical trials focused on schizophrenia. It is important to emphasize that, in this section, ARMD and schizophrenia are not presented as two full applied case studies. Rather, they are used as two real‐data calibrations of the identifiable parts of the causal model. This allows us to evaluate the impact of alternative assumptions on the unidentifiable copulas under realistic observable margins. The full applied case study is presented separately in Section 6, where we focus on the schizophrenia data set.
Additionally, for the four unidentifiable copulas in (7), we explore four parametric choices: Gaussian, Clayton, Gumbel, and Frank.
We evaluated all combinations of these parametric copulas, resulting in unique distributions for . A configuration consisting entirely of Gaussian copulas produces the four‐dimensional normal distribution. Regardless of the chosen copula types, the unidentifiable copulas are parameterized by Spearman's rho correlation parameters. Let , for , denote one of these 256 distributions of , where is the four‐dimensional vector of correlations that characterizes the association of the counterfactual pairs with . Following the methodology in Alonso [11], we sample 1000 vectors under three distinct sampling schemes:
No additional assumptions: The rho parameters are sampled from .
Positive restricted associations: The rho parameters are assumed to be positive and bounded away from zero and one.
Conditional independence and positive restricted associations: In addition to positive restricted associations, we assume conditional independence: and .
Basically, for each of the 256 distributions previously defined, we generate 1000 realizations in each of the scenarios (a)–(c). Notably, each distribution has normal observable margins, though only one of them is actually a four‐dimensional normal distribution. We analyze the differences for each scenario, where and are the Individual Causal Associations calculated from the density and a four‐dimensional normal distribution, respectively, both derived from the same vector. Since both distributions exhibit the same observable margins and identical values for the unidentifiable parameters , the difference effectively highlights the impact of the parametric copulas that characterize the distribution of the counterfactual pairs on the ICA.
Figure 4 summarizes the main findings. The rows correspond to the two identifiable normal margins, and , derived from both data sets. The columns represent the sampling schemes described in (a)–(c). The results show that although the impact of misspecification is generally mild () in most cases, there are instances where distributions, indistinguishable from the normal based on observed data, can lead to significantly different assessments of surrogacy. It is important to highlight again that the differences quantify the impact of the parametric copulas used to characterize the distribution of the counterfactual pairs. Interestingly, when the association structure of these counterfactual pairs is restricted, as in setting (c), the effect of the misspecification substantially diminishes.
FIGURE 4.

D‐vine copula sensitivity analysis calibrated using identifiable margins estimated from two real data examples: ARMD and schizophrenia. The two rows correspond to the two choices of observable bivariate margins, and , used to anchor the simulation at realistic parameter values. The columns correspond to the three sampling schemes for the unidentifiable associations: no additional assumptions, positive restricted associations (PA), and positive restricted associations plus conditional independence (CI). The plotted quantity is , the difference between the ICA computed under the D‐vine copula density () and the ICA obtained under the corresponding multivariate normal model ().
6. Applied Case Study: Schizophrenia Data
Schizophrenia is a psychiatric disorder characterized by psychotic manifestations such as delusions, hallucinations, and disorganized thinking and behavior [26]. We analyze a dataset comprising five double‐blind randomized clinical trials in patients with schizophrenia, available in R via the Surrogate package [27]. In this section we focus on the principal findings; implementation details and full R code are provided in Supporting Information Appendix (Section 3).
A total of 2123 patients participated, with 1589 allocated to Risperidone and 534 to control. Symptoms were assessed using the Positive and Negative Syndrome Scale (PANSS) and the Brief Psychiatric Rating Scale (BPRS). PANSS comprises 30 items, whereas BPRS is an 18‐item PANSS subscale [28]. We treat PANSS, a reliable but comparatively complex scale, as the true endpoint and BPRS, a simpler, easier‐to‐administer subscale, as the surrogate. The objective is to determine whether the BPRS subscale can reliably serve as a substitute for the more complex PANSS scale in evaluating the efficacy of Risperidone or similar antipsychotic treatments.
In prior work, both endpoints were treated as approximately normal [29]. We began by reassessing this assumption: density and Q–Q plots (Web Appendix C) suggested near normality for both endpoints, yet the Shapiro–Wilk test rejected normality for all except . With a large sample, this test has high power and may flag small, practically negligible departures from normality. Results from earlier sections indicate that mild deviations from normality typically have only a negligible or modest impact on surrogacy assessment; however, it is difficult to know a priori how substantial the impact may be in any concrete application. We therefore conducted the analysis under multiple modeling assumptions and compared results as a sensitivity analysis.
We first assessed the ICA under the assumption of normally distributed potential outcomes, using the closed‐form expression in Section 3. To address the identifiability issues described there, we implemented the sensitivity analysis of Alonso et al. [11], estimating the ICA over a wide range of unidentifiable correlations compatible with the observed data. Figure 5 displays the resulting frequency distribution of values under normality, denoted . Although small values are compatible with the data, most values from the sensitivity analysis exceed . Specifically, ranged from to , with a mean of about and first and third quartiles of and , respectively. Repeating the analysis with the ‐causal model produced essentially identical results, fully aligning with Section 4.1.
FIGURE 5.

Density plot of ICA values for the schizophrenia case study obtained from the three models as , , and .
For the multivariate log‐normal model, we proceeded analogously: we estimated the identifiable log‐normal parameters from the data, then gridded the remaining nonidentifiable correlations while holding and at their estimated values. We retained only grid points yielding a positive‐definite covariance matrix and computed the ICA for each. Figure 5 shows the corresponding distribution of values. Notably, diagnostics on the observable margins suggest that a log‐normal specification is a poor description of these data, so this constitutes a misspecified scenario. Even so, the log‐normal model yielded slightly larger ICA values with a narrower range (–), a mean and median of , and quartiles of and . Thus, despite differing from the normal/ results, this analysis provides even stronger evidence in favor of the surrogate.
Finally, we analyzed the data using a D‐vine copula specification. In this setup, four pair‐copulas are unidentifiable and must be fixed by untestable parametric choices. We considered the Gaussian, Gumbel, Frank, and Clayton families for each of these four components (Section 5), yielding candidate models. Figure 6 displays the frequency distributions for a random subset of four models under three sampling schemes for the unidentifiable associations: (i) no additional constraints, (ii) positive association, and (iii) positive association plus conditional independence. Three observations are in order. First, within each sampling scheme (i)–(iii), the frequency distributions across models look strikingly similar, indicating that the particular parametric choice for the unidentifiable pair‐copulas has little influence once the sampling is fixed. Second, imposing untestable constraints on the association structure can materially change the results; for example, restricting all associations to be positive has a pronounced effect on the distributional shape. The contrasts across sampling schemes illustrate the risks of relying on untestable association assumptions, a common practice in causal inference (e.g., PTE, RE), and underscore the value of our sensitivity analysis. Third, when all pair‐copulas are Gaussian and no restrictions are imposed on the association structure, the results closely match those obtained under the normal causal model, as expected.
FIGURE 6.

Frequency distributions of the ICA for the schizophrenia case study under various D‐vine copula models and sampling schemes.
As noted in Section 1, several methods have been proposed to assess surrogacy in the STS, including the Proportion of Treatment Effect (PTE) introduced by Freedman et al. [4] and refined by Parast et al. [30]. The PTE is defined in terms of expected causal treatment effects and, because there is no replication at this level in a single trial, its identification requires untestable assumptions. Moreover, as a ratio, the PTE can be numerically unstable when the overall expected treatment effect on the true endpoint is small. Stijven et al. [31] provide a comprehensive discussion of the PTE's definition, interpretation, and limitations. Using the robust approach of Parast et al. [30], the estimated PTE in our case study was ( CI: 0.66–1), which supports the surrogate's validity. We also estimated the adjusted association (AA) as ( CI: 0.9594–0.9667); the AA quantifies the association between both endpoints after adjusting for treatment. Note that, while a high AA may be necessary for surrogacy, it is not a sufficient condition because a high correlation alone is not enough to deem a surrogate valid.
Taken together, these results support BPRS as a viable surrogate for PANSS. Nonetheless, side‐by‐side comparisons of multiple metrics should be interpreted with care, because the estimands, assumptions, and sources of uncertainty differ across methods [32].
7. Discussion
In this study, we investigated the effects of model misspecification on the behavior of the ICA when the true distribution deviates from a four‐dimensional normal model. Through theoretical derivations and simulations we examined scenarios where the true DGM was characterized by a four‐dimensional t‐distribution and a log‐normal distribution. Our results show that the impact of model misspecification can vary, ranging from negligible when the underlying model follows a multivariate t‐distribution to significant under a log‐normal DGM. Notably, in the latter case, the substantial impact of misspecification was primarily observable when deviations from normality could be detected using the observed data.
Additionally, by employing D‐vine copulas, we explored scenarios where the true DGM was indistinguishable from the normal model based on observed data. In this setting, it is important to note that normality of the bivariate normal margins does not guarantee multivariate normality of the entire distribution. We found that in such cases, assuming normality might lead to misleading results, even though the observed data would support this assumption.
For any surrogacy metric, defining a threshold that guarantees good surrogacy is very challenging. The general intuition is that the further the true distribution deviates from the assumed model, such as from normality, the greater the potential impact on the results. In situations with substantial uncertainty regarding model assumptions, we recommend conducting sensitivity analyses under alternative plausible distributions to evaluate the robustness of the conclusions, as illustrated in the case study. Moreover, ICA results can be compared with those obtained from alternative surrogate evaluation measures. Although these methods may differ in methodology and computational implementation, it is desirable that they lead to conclusions in the same direction.
Our findings highlight the complexity of evaluating surrogates using causal inference models, which are only partially identifiable from the data. The unidentifiable components can substantially impact the assessment of surrogacy under certain conditions. Practically, our results underscore the importance of verifying that the identifiable components of the model are consistent with the observed data and suggest that considering different modeling assumptions for the counterfactual pairs is crucial to understanding how they influence the final conclusions.
Funding
This research was supported by the MCIN/AEI/10.13039/501100011033/ and by ERDF A way of making Europe (Grant/Award Number: PID2022‐137050NB‐I00) and Agentschap Innoveren & Ondernemen and Janssen Pharmaceutical Companies of Johnson & Johnson Innovative Medicine through a Baekeland Mandate (Grant/Award Number: HBC.2022.0145).
Ethics Statement
The authors have nothing to report.
Conflicts of Interest
The authors declare no conflicts of interest.
Supporting information
Figure S1: The density graphs of potential outcome vector when the difference is −0.00079.
Figure S2: The density graphs of potential outcome vector when the difference is 0.0058.
Figure S3: The density graphs of potential outcome vector when the difference is 0.0543.
Figure S4: The density graphs of potential outcome vector when the difference is 0.5705.
Figure S5: The density graphs of potential outcome vector when the difference is 0.3464.
Figure S6: The density graphs of potential outcome vector when the difference is 0.4398.
Table S1: Interpretation of the unidentifiable components of the D‐vine density in (3).
Table S2: Estimated bivariate normal distributions in the ARMD and Schizo data sets from the Surrogate R package.
Figure S7: Histograms of endpoints.
Figure S8: Q–Q plots of endpoints.
Acknowledgments
This work is supported by grant PID2022‐137050NB‐I00 of the Spanish Ministry of Science and Innovation. Florian Stijven gratefully acknowledges funding from Agentschap Innoveren & Ondernemen and Janssen Pharmaceutical Companies of Johnson & Johnson Innovative Medicine through a Baekeland Mandate [grant number HBC.2022.0145]. We sincerely appreciate the meaningful suggestions and feedback provided by the editor which have substantially improved the quality of our study.
Data Availability Statement
The data that support the findings of this study are available in R Surrogate package. These data were derived from the following resources available in the public domain: https://cran.r‐project.org/web/packages/Surrogate/index.html.
References
- 1. FDA U, CDER, CBER , Guidance for Industry: Expedited Programs for Serious Conditions – Drugs and Biologics (US FDA, 2014). [Google Scholar]
- 2. Burzykowski T., Molenberghs G., and Buyse M., The Evaluation of Surrogate Endpoints (Springer‐Verlag, 2005). [Google Scholar]
- 3. Prentice R. L., “Surrogate Endpoints in Clinical Trials: Definition and Operational Criteria,” Statistics in Medicine 8, no. 4 (1989): 431–440. [DOI] [PubMed] [Google Scholar]
- 4. Freedman L. S., Graubard B. I., and Schatzkin A., “Statistical Validation of Intermediate Endpoints for Chronic Diseases,” Statistics in Medicine 11, no. 2 (1992): 167–178. [DOI] [PubMed] [Google Scholar]
- 5. Buyse M. and Molenberghs G., “Criteria for the Validation of Surrogate Endpoints in Randomized Experiments,” Biometrics 54 (1998): 1014–1029. [PubMed] [Google Scholar]
- 6. Buyse M., Molenberghs G., Burzykowski T., Renard D., and Geys H., “The Validation of Surrogate Endpoints in Meta‐Analyses of Randomized Experiments,” Biostatistics 1 (2000): 49–67. [DOI] [PubMed] [Google Scholar]
- 7. Weir C. J. and Taylor R. S., “Informed Decision‐Making: Statistical Methodology for Surrogacy Evaluation and Its Role in Licensing and Reimbursement Assessments,” Pharmaceutical Statistics 21, no. 4 (2022): 740–756. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8. Alonso A., Bigirumurame T., Burzykowski T., et al., Applied Surrogate Endpoint Evaluation Methods With SAS and R (Chapman Hall/CRC, 2017). [Google Scholar]
- 9. Alonso A., An Information‐Theoretic Approach for the Evaluation of Surrogate Endpoints (Wiley StatsRef: Statistics Reference Online, 2018). [Google Scholar]
- 10. Deliorman G., Alonso A., and Pardo M. C., “A Rényi‐Divergence‐Based Family of Metrics for the Evaluation of Surrogate Endpoints in a Causal Inference Framework,” Statistics in Biopharmaceutical Research 18, no. 1 (2026): 31–38. [Google Scholar]
- 11. Alonso A., Van Der Elst W., Molenberghs G., Buyse M., and Burzykowski T., “On the Relationship Between the Causal‐Inference and Meta‐Analytic Paradigms for the Validation of Surrogate Endpoints,” Biometrics 71, no. 1 (2015): 15–24. [DOI] [PubMed] [Google Scholar]
- 12. Imbens G. W. and Rubin D. B., Causal Inference in Statistics, Social, and Biomedical Sciences (Cambridge University Press, 2015). [Google Scholar]
- 13. Cover T. M. and Thomas J. A., “Information Theory and Statistics,” Elements of Information Theory 1, no. 1 (1991): 279–335. [Google Scholar]
- 14. Linfoot E. H., “An Informational Measure of Correlation,” Information and Control 1, no. 1 (1957): 85–89. [Google Scholar]
- 15. Joe H., “Relative Entropy Measures of Multivariate Dependence,” Journal of the American Statistical Association 84, no. 405 (1989): 157–164. [Google Scholar]
- 16. Kullback S. and Leibler R. A., “On Information and Sufficiency,” Annals of Mathematical Statistics 22, no. 1 (1951): 79–86. [Google Scholar]
- 17. Holland P. W., “Statistics and Causal Inference,” Journal of the American Statistical Association 81, no. 396 (1986): 945–960. [Google Scholar]
- 18. Elst V. d W., Molenberghs G., and Alonso A., “Exploring the Relationship Between the Causal‐Inference and Meta‐Analytic Paradigms for the Evaluation of Surrogate Endpoints,” Statistics in Medicine 35, no. 8 (2016): 1281–1298. [DOI] [PubMed] [Google Scholar]
- 19. Van Der Elst W., Alonso A., Coppenolle H., Meyvisch P., and Molenberghs G., “The Individual‐Level Surrogate Threshold Effect in a Causal Inference Setting With Normally Distributed Endpoints,” Pharmaceutical Statistics 20, no. 6 (2021): 1216–1231. [DOI] [PubMed] [Google Scholar]
- 20. Alonso A., Van Der Elst W., Molenberghs G., and Florez A. J., “A Reflection on the Causal Interpretation of Individual‐Level Surrogacy,” Journal of Biopharmaceutical Statistics 29, no. 3 (2019): 529–540. [DOI] [PubMed] [Google Scholar]
- 21. Arellano‐Valle R. B., Conreras‐Reyes J. E., and Genton M. G., “Shannon Entropy and Mutual Information for Multivariate Skew‐Elliptical Distributions,” Scandinavian Journal of Statistics 40, no. 1 (2013): 42–62. [Google Scholar]
- 22. Ding P., “On the Conditional Distribution of the Multivariate t Distribution,” American Statistician 70, no. 3 (2016): 293–295. [Google Scholar]
- 23. Kotz S. and Nadarajah S., Multivariate t‐Distributions and Their Applications (Cambridge University Press, 2004). [Google Scholar]
- 24. Stijven F., Molenberghs G., V. Keilegom, I , Elst V. d W., and Alonso A., “Evaluating Time‐To‐Event Surrogates for Time‐To‐Event True Endpoints: An Information‐Theoretic Approach Based on Causal Inference,” Lifetime Data Analysis 31, no. 1 (2025): 1–23. [DOI] [PubMed] [Google Scholar]
- 25. Czado C., Analyzing Dependent Data With Vine Copulas, vol. 222 (Springer, 2019). [Google Scholar]
- 26. Offermanns S. and Rosenthal W., Encyclopedic Reference of Molecular Pharmacology (Springer Verlag, 2004). [Google Scholar]
- 27. Van Der Elst W., Meyvisch P., Alonso A., et al., “Package Surrogate,” 2022.
- 28. Peuskens J., “Risperidone in the Treatment of Patients With Chronic Schizophrenia: A Multi‐National, Multi‐Centre, Double‐Blind, Parallel‐Group Study Versus Haloperidol,” British Journal of Psychiatry 166, no. 6 (1995): 712–726. [DOI] [PubMed] [Google Scholar]
- 29. Alonso A., Bigirumurame T., Burzykowski T., et al., Applied Surrogate Endpoint Evaluation Methods With Sas and r (CRC Press, 2016). [Google Scholar]
- 30. Parast L., McDermott M. M., and Tian L., “Robust Estimation of the Proportion of Treatment Effect Explained by Surrogate Marker Information,” Statistics in Medicine 35, no. 10 (2016): 1637–1653. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31. Stijven F., Ariel A., and Molenberghs G., “Proportion of Treatment Effect Explained: An Overview of Interpretations,” Statistical Methods in Medical Research 33, no. 7 (2024): 1278–1296. [DOI] [PubMed] [Google Scholar]
- 32. Joffe M. M. and Greene T., “Related Causal Frameworks for Surrogate Outcomes,” Biometrics 65, no. 2 (2009): 530–538. [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Figure S1: The density graphs of potential outcome vector when the difference is −0.00079.
Figure S2: The density graphs of potential outcome vector when the difference is 0.0058.
Figure S3: The density graphs of potential outcome vector when the difference is 0.0543.
Figure S4: The density graphs of potential outcome vector when the difference is 0.5705.
Figure S5: The density graphs of potential outcome vector when the difference is 0.3464.
Figure S6: The density graphs of potential outcome vector when the difference is 0.4398.
Table S1: Interpretation of the unidentifiable components of the D‐vine density in (3).
Table S2: Estimated bivariate normal distributions in the ARMD and Schizo data sets from the Surrogate R package.
Figure S7: Histograms of endpoints.
Figure S8: Q–Q plots of endpoints.
Data Availability Statement
The data that support the findings of this study are available in R Surrogate package. These data were derived from the following resources available in the public domain: https://cran.r‐project.org/web/packages/Surrogate/index.html.
