Efficient Evaluation of Prediction Rules in Semi-Supervised Settings under Stratified Sampling

Jessica Gronsbell; Molei Liu; Lu Tian; Tianxi Cai

doi:10.1111/rssb.12502

. Author manuscript; available in PMC: 2023 Sep 1.

Published in final edited form as: J R Stat Soc Series B Stat Methodol. 2022 Apr 26;84(4):1353–1391. doi: 10.1111/rssb.12502

Efficient Evaluation of Prediction Rules in Semi-Supervised Settings under Stratified Sampling

Jessica Gronsbell ^1,², Molei Liu ^1,², Lu Tian ¹, Tianxi Cai ¹

PMCID: PMC9586151 NIHMSID: NIHMS1779685 PMID: 36275859

Abstract

In many contemporary applications, large amounts of unlabeled data are readily available while labeled examples are limited. There has been substantial interest in semi-supervised learning (SSL) which aims to leverage unlabeled data to improve estimation or prediction. However, current SSL literature focuses primarily on settings where labeled data is selected uniformly at random from the population of interest. Stratified sampling, while posing additional analytical challenges, is highly applicable to many real world problems. Moreover, no SSL methods currently exist for estimating the prediction performance of a fitted model when the labeled data is not selected uniformly at random. In this paper, we propose a two-step SSL procedure for evaluating a prediction rule derived from a working binary regression model based on the Brier score and overall misclassification rate under stratified sampling. In step I, we impute the missing labels via weighted regression with nonlinear basis functions to account for stratified sampling and to improve efficiency. In step II, we augment the initial imputations to ensure the consistency of the resulting estimators regardless of the specification of the prediction model or the imputation model. The final estimator is then obtained with the augmented imputations. We provide asymptotic theory and numerical studies illustrating that our proposals outperform their supervised counterparts in terms of efficiency gain. Our methods are motivated by electronic health record (EHR) research and validated with a real data analysis of an EHR-based study of diabetic neuropathy.

Keywords: Semi-Supervised Learning, Stratified Sampling, Model Evaluation, Risk Prediction

1. Introduction

Semi-supervised learning (SSL) has emerged as a powerful learning paradigm to address big data problems where the outcome is cumbersome to obtain and the predictors are readily available [Chapelle et al., 2009]. Formally, the SSL problem is characterized by two sources of data: (i) a relatively small sized labeled dataset $L$ with n observations on the outcome y and the predictors x and (ii) a much larger unlabeled dataset $U$ with $N ≫ n$ observations on only x. A promising application of SSL and the motivation for this work is in electronic health record (EHR) research. EHRs have immense potential to serve as a major data source for biomedical research as they have generated extensive information repositories on representative patient populations [Murphy et al., 2009, Kohane, 2011, Wilke et al., 2011]. Nonetheless, a primary bottleneck in recycling EHR data for secondary use is to accurately and efficiently extract patient level disease phenotype information [Sinnott et al., 2014, Liao et al., 2015]. Frequently, true phenotype status is not well characterized by disease-specific billing codes. For example, at Partner’s Healthcare, only 56% of patients with at least 3 International Classification of Diseases, Ninth Revision (ICD9) codes for rheumatoid arthritis (RA) have confirmed RA after manual chart review [Liao et al., 2010]. More accurate EHR phenotyping has been achieved by training a prediction model based on a number of features including billing codes, lab results and mentions of clinical terms in narrative notes extracted via natural language processing (NLP) [Liao et al., 2013, Xia et al., 2013, Ananthakrishnan et al., 2013, e.g.]. The model is traditionally trained and evaluated using a small amount of labeled data obtained from manual medical chart review by domain experts. SSL methods are particularly attractive for developing such models as they leverage unlabeled data to achieve higher estimation efficiency than their supervised counterparts. In practice, this increase in efficiency can be directly translated into requiring fewer chart reviews without a loss in estimation precision.

In the EHR phenotyping setting, it is often infeasible to select the labeled examples uniformly at random, either due to practical constraints or due to the nature of the application. For example, it may be necessary to oversample individuals for labeling with a particular billing code or diagnostic procedure for a rare disease to ensure that an adequate number of cases are available for model estimation. In some settings, training data may consist of a subset of individuals selected uniformly at random together with a set of registry patients whose phenotype status is confirmed through routine collection. Stratified sampling is also an effective strategy when interest lies in simultaneously characterizing multiple phenotypes. One may oversample patients with at least one billing code for several rare phenotypes and then perform chart review on the selected patients for all phenotypes of interest. Though it is critical to account for the aforementioned sampling mechanisms to make valid statistical inference, it is non-trivial in the context of SSL since the fraction of subjects being sampled for labeling is near zero. The challenge is further amplified when the fitted prediction models are potentially misspecified.

Existing SSL literature primarily concerns the setting in which $L$ is a uniform random sample from the underlying pool of data and thus the missing labels in $U$ are missing completely at random (MCAR) [Wasserman and Lafferty, 2008]. In this setting, a variety of methods for classification have been proposed including generative modeling [Castelli and Cover, 1996, Jaakkola et al., 1999], manifold regularization [Belkin et al., 2006, Niyogi, 2013] and graph-based regularization [Belkin and Niyogi, 2004]. While making use of both $U$ and $L$ can improve estimation, in many cases SSL is outperformed by supervised learning (SL) using only $U$ when the assumed models are incorrectly specified [Castelli and Cover, 1996, Jaakkola et al., 1999, Corduneanu, 2002, Cozman et al., 2002, 2003]. As model misspecification is nearly inevitable in practice, recent work has called for ‘safe’ SSL methods that are always at least as efficient as the SL counterparts. For example, several authors have considered safe SSL methods for discriminative models based on density ratio weighted maximum likelihood [Sokolovska et al., 2008, Kawakita and Kanamori, 2013, Kawakita and Takeuchi, 2014]. Though the true density ratio is 1 in the MCAR setting, the efficiency gain is achieved through estimation of density ratio weight, a statistical paradox previously observed in the missing data literature [Robins et al., 1992, 1994]. More recently, Krijthe and Loog [2016] introduced a SSL method for least squares classification that is guaranteed to outperform SL. Chakrabortty and Cai [2018] proposed an adaptive imputation-based SSL approach for linear regression that also outperforms supervised least squares estimation. It is unclear, however, whether these methods can be extended to accommodate additional loss functions. Moreover, none of the aforementioned methods are applicable to settings where the labeled data is not a uniform random sample from the underlying data such as the stratified sampling design. Additionally, the focus of existing work has been on the estimation of prediction models, rather than the estimation of model performance metrics. Gronsbell and Cai [2018] recently proposed a semi-supervised procedure for estimating the receiver operating characteristic parameters, but this method is similarly limited to the standard MCAR setting.

This paper addresses these limitations through the development of an efficient SS estimation method for model performance metrics that is robust to model misspecification in the presence of stratified sampling. Specifically, we develop an imputation based procedure to evaluate the prediction performance of a potentially misspecified binary regression model. To the best of our knowledge, the proposed method is the first SSL procedure that provides efficient and robust estimation of prediction performance measures under stratified sampling. We focus on two commonly used error measurements, the overall misclassification rate (OMR) and the Brier score. The proposed method involves two steps of estimation. In step I, the missing labels are imputed with a weighted regression with nonlinear basis functions to account for stratified sampling and to improve efficiency. In step II, the initial imputations are augmented to ensure the consistency of the resulting estimators regardless of the specification of the prediction model or the imputation model. Through theoretical results and numerical studies, we demonstrate that the SS estimators of prediction performance are (i) robust to the misspecification of the prediction or imputation model and (ii) substantially more efficient than their SL counterparts. We also develop an ensemble cross-validation (CV) procedure to adjust for overfitting and a perturbation resampling procedure for variance estimation.

The remainder of this paper is organized as follows. In Section 2, we specify the data structure and problem set-up. We then develop the estimation and bias correction procedure for the accuracy measures in Sections 3 and 4. Section 5 outlines the asymptotic properties of the estimators and section 6 introduces the perturbation resampling procedure for making inference. Our proposals are then validated through a simulation study in Section 7 and a real data analysis of an EHR-based study of diabetic neuropathy is presented in Section 8. We conclude with additional discussions in Section 9.

2. Preliminaries

2.1. Data Structure

Our interest lies in evaluating a prediction model for a binary phenotype y based on a predictor vector $x = {(1, x_{1}, \dots, x_{p})}^{⊤}$ for some fixed p. The underlying full data consists of $N = \sum_{s = 1}^{S} N_{s}$ independent and identically distributed random vectors

F = {F_{i} = {(y_{i}, u_{i}^{⊤})}^{⊤}}_{i = 1}^{N}

where $u_{i} = {(x_{i}^{⊤}, S_{i})}^{⊤}$ , $S_{i} \in {1, 2, \dots, S}$ is a discrete stratification variable that defines a fixed number of strata $S$ for sampling, and $N_{s} = \sum_{i = 1}^{N} I (S_{i} = s)$ is the sample size of stratum s. Throughout, we let $F_{0} = {(y_{0}, x_{0}^{⊤}, S_{0})}^{⊤}$ be a future realization of F.

Due to the difficulty in ascertaining y, a small uniform random sample is obtained from each stratum and labeled with outcome information. The observable data therefore consists of

D = {D_{i} = {(y_{i} V_{i}, u_{i}^{⊤}, V_{i})}^{⊤}}_{i = 1}^{N}

where $V_{i} \in {0, 1}$ indicates whether y_i is ascertained. We let

P (V_{i} = 1 ∣ F) = {\hat{π}}_{S_{i}}, {\hat{π}}_{s} = n_{s} / N_{s} and n_{s} = \sum_{j = 1}^{N} I (S_{j} = s) V_{j} .

Without loss of generality, we suppose that the first $n = \sum_{s = 1}^{S} n_{s}$ subjects are labeled and ${n_{s}, s = 1, \dots, S}$ are specified by design. We assume that

{\hat{ρ}}_{1 s} = n_{s} / n \overset{p}{\to} ρ_{1 s} \in (0, 1) and {\hat{ρ}}_{s} = N_{s} / N \overset{p}{\to} ρ_{s} \in (0, 1)

as n and $N \to \infty$ respectively. This ensures that

{\hat{π}}_{s} / {\hat{π}}_{t} \overset{p}{\to} (ρ_{1 s} ρ_{t}) / (ρ_{s} ρ_{1 t}) \in (0, 1)

as $n \to \infty$ for any pair of s and t [Mirakhmedov et al., 2014]. As in the standard SS setting, we further assume max_s ${\hat{π}}_{s} \overset{p}{\to} 0$ as $n \to \infty$ [Chakrabortty and Cai, 2018, Zhang et al., 2019]. This assumption distinguishes the current setting from (i) the familiar missing data setting where ${\hat{π}}_{s}$ is bounded above 0 and (ii) standard SSL under uniform random sampling as y is MCAR conditional on $S$ (i.e. $V ⊥ (y, x) ∣ S$ ) under stratified sampling.

2.2. Problem Set-Up

To predict y₀ based on x₀, we fit a working regression model

P (y = 1 ∣ x) = g (θ^{⊤} x)

(1)

where $θ = {(θ_{0}, θ_{1}, \dots, θ_{p})}^{⊤}$ is an unknown vector of regression parameters and $g (\cdot) : (- \infty, \infty) \to (0, 1)$ is a specified, smooth monotone function such as the expit function. The target model parameter, $\bar{θ}$ , is the solution to the estimating equation

U (θ) = E [x {y - g (θ^{⊤} x)}] = 0 .

We let the predicted value for y₀ be $Y ({\bar{θ}}^{⊤} x_{0})$ for some function $Y$ . In this paper, we aim to obtain SS estimators of the prediction performance of $Y ({\bar{θ}}^{⊤} x_{0})$ quantified by the Brier score

{\bar{D}}_{1} = E [{y_{0} - Y_{1} ({\bar{θ}}^{⊤} x_{0})}^{2}] with Y_{1} (x) = g (x),

and the overall misclassification rate (OMR) ${\bar{D}}_{2} = E [{y_{0} - Y_{2} ({\bar{θ}}^{⊤} x_{0})}^{2}]$ with $Y_{2} (x) = I {g (x) > c}$ for some specified constant c.

We focus on these two metrics as the convey distinct information about the performance of the prediction model. The OMR summarizes the overall discrimination capacity of the model while the Brier score summarizes the calibration of the model. More complete discussions regarding assessment of model performance can be found in Hand [1997, 2001], Gneiting and Raftery [2007] and Gerds et al. [2008].

To simplify presentation, we generically write

\bar{D} = D (\bar{θ}) with D (θ) = E [d {y_{0}, Y (θ^{⊤} x_{0})}],

where $d (y, z) = {(y - z)}^{2}, Y (\cdot) = Y_{1} (\cdot)$ for Brier score, and $Y (\cdot) = Y_{2} (\cdot)$ for the OMR. We will construct a SS estimator of $D (\bar{θ})$ to improve the statistical efficiency of its SL counterpart, ${\hat{D}}_{SL} ({\hat{θ}}_{SL})$ , where

{\hat{θ}}_{SL} {solves U}_{n} (θ) = \frac{1}{N} \sum_{i = 1}^{N} {\hat{w}}_{i} x_{i} {y_{i} - g (θ^{⊤} x_{i})} = 0, {\hat{D}}_{SL} (θ) = \frac{1}{N} \sum_{i = 1}^{N} {\hat{w}}_{i} d {y_{i}, Y (θ^{⊤} x_{i})},

and the weights ${\hat{w}}_{i} = V_{i} / {\hat{π}}_{S_{i}}$ account for the stratified sampling with $\sum_{i = 1}^{N} {\hat{w}}_{i} = N$ . Since ${\hat{w}}_{i} \overset{p}{\to} \infty$ for those with $V_{i} = 1$ , standard M-estimation theory cannot be directly applied to establish the asymptotic behavior of the SL estimators. We show in Appendix C that ${\hat{θ}}_{SL}$ is a root-n consistent estimator for $\bar{θ}$ and derive the asymptotic properties of ${\hat{D}}_{SL} ({\hat{θ}}_{SL})$ in Appendix D. We also note that throughout the article we use the subscripts n or N to index estimating equations to clarify if they are computed with the labeled or full data, respectively.

3. Estimation Procedure

Our approach to obtaining a SS estimator of $D (\bar{θ})$ proceeds in two steps. First, the missing outcomes are imputed with a flexible model to improve statistical efficiency. Next, the imputations are augmented so that the resulting estimators are consistent for $D (\bar{θ})$ regardless of the specification of the prediction model or the imputation model. The final estimator of the accuracy measure is then estimated using the full data and the augmented imputations. This estimation procedure is detailed in the subsequent sections.

We comment here that the initial imputation step allows for construction of a simple and efficient SS estimator of $\bar{θ}$ . As efficient estimation of $\bar{θ}$ may not be of practical utility in the prediction setting, we keep our focus on accuracy parameter estimation and defer some of the technical details of model parameter estimation to the Appendix. However, we do note that the estimator of $D (\bar{θ})$ inherently has two sources of estimation variability. The dominating source of variation is from estimating the accuracy measure itself while the second source is from the estimation of the regression parameter. Therefore, by leveraging a SS estimator of $\bar{θ}$ in estimating $D (\bar{θ})$ , we may further improve the efficiency of our SS estimator of the accuracy measure. These statements are elucidated by the influence function expansions of the SS and SL estimators of $D (\bar{θ})$ presented in Section 5.

3.1. Step 1: Flexible imputation

We propose to impute the missing y with an estimate of $m (u) = P (y = 1 ∣ u)$ . The purpose of the imputation step is to make use of $U$ as it essentially characterizes the covariate distribution due to its size. The accuracy metrics provide measures of agreement between the true and predicted outcomes and therefore depend on the covariate distribution. We thus expect to decrease estimation precision by incorporating $U$ into estimation. In taking an imputation-based approach, we rely on our estimate of m(u) to capture the dependency of y on u in order to glean information from $U$ . While a fully nonparametric method such as kernel smoothing allows for complete flexibility in estimating m(u), smoothing generally does not perform well with moderate p due to the curse of dimensionality [Kpotufe, 2010]. To overcome this challenge and allow for a rich model for m(u), we incorporate some parametric structure into the imputation step via basis function regression.

Let $Φ (u)$ be a finite set of basis functions with fixed dimension that includes x. We fit a working model

P (y = 1 ∣ u) = g {γ^{⊤} Φ (u)}

(2)

to $L$ and impute y as $g {{\tilde{γ}}^{⊤} Φ (u)}$ where $\tilde{γ}$ is the solution to

{\tilde{Q}}_{n} (γ) = \frac{1}{N} \sum_{i = 1}^{N} {\hat{w}}_{i} Φ_{i} [y_{i} - g {γ^{⊤} Φ_{i}}] - λ_{n} γ = 0

(3)

$Φ_{i} = Φ (u_{i})$ and $λ_{n} = o (n^{- \frac{1}{2}})$ is a tuning parameter to ensure stable fitting. Under Conditions 1–3 given in Section 5, we argue in Appendix D that $\tilde{γ}$ is a regular root-n consistent estimator for the unique solution, $\tilde{γ}$ , to $Q (γ) = E {Φ (u) [y - g {γ^{⊤} Φ (u)}]} = 0$ . We take our initial imputations as ${\tilde{m}}_{I} (u) = g {{\tilde{γ}}^{⊤} Φ (u)}$ .

With y imputed as ${\tilde{m}}_{I} (u)$ , we may also obtain a simple SS estimator for $\bar{θ}$ , ${\overset{ˇ}{θ}}_{SSL}$ , as the solution to

{\hat{U}}_{N} (θ) = \frac{1}{N} \sum_{i = 1}^{N} x_{i} {g ({\tilde{γ}}^{⊤} Φ_{i}) - g (θ^{⊤} x_{i})} = 0 .

The asymptotic behaviour of ${\overset{ˇ}{θ}}_{SSL}$ is presented and compared with ${\hat{θ}}_{SL}$ in the Appendix. When the working regression model (1) is correctly specified, it is shown that ${\hat{θ}}_{SL}$ is fully efficient and ${\overset{ˇ}{θ}}_{SSL}$ is asymptotically equivalent to ${\hat{θ}}_{SL}$ . When the outcome model in (1) is not correctly specified, but the imputation model in (2) is correctly specified, we show in Appendix E that ${\overset{ˇ}{θ}}_{SSL}$ is more efficient than ${\hat{θ}}_{SL}$ . When the imputation model is also misspecified, ${\overset{ˇ}{θ}}_{SSL}$ tends to be more efficient than ${\hat{θ}}_{SL}$ , but the efficiency gain is not theoretically guaranteed. We therefore obtain the final SS estimator, denoted as ${\hat{θ}}_{SSL} = {({\hat{θ}}_{SSL, 0}, \dots, {\hat{θ}}_{SSL, p})}^{⊤}$ , as a linear combination of ${\hat{θ}}_{SL}$ and ${\overset{ˇ}{θ}}_{SSL}$ to minimize the asymptotic variance. Details are provided in Appendices A and B.

3.2. Step 2: Robustness augmentation

To obtain an efficient SS estimator for $\bar{D} = D (\bar{θ})$ , we note that

d (y, Y) = y (1 - 2 Y) + Y^{2}

(4)

is linear in y when $y \in {0, 1}$ . With a given estimate of $m (\cdot)$ , denoted by $\tilde{m} (\cdot)$ , a SS estimate of $\bar{D}$ can be obtained as

\frac{1}{N} \sum_{i = 1}^{N} d {\tilde{m} (u_{i}), Y (θ^{⊤} x_{i})} .

However, $\tilde{m} (u)$ needs to be carefully constructed to ensure that the resulting estimator is consistent for $\bar{D}$ under possible misspecification of the imputation model. Using the expression in (4), a sufficient condition to guarantee consistency for $\bar{D}$ is that

E [{y_{0} - \tilde{m} (u_{0})} {1 - 2 Y ({\bar{θ}}^{⊤} x_{0})} ∣ \tilde{m} (\cdot)] \overset{p}{\to} 0 as n \to \infty .

(5)

This condition implies that $E [d {y_{0}, Y ({\bar{θ}}^{⊤} x_{0})} - d {\tilde{m} (u_{0}), Y ({\bar{θ}}^{⊤} x_{0})}] \overset{p}{\to} 0$ . Unfortunately, ${\tilde{m}}_{I} (u) = g {{\tilde{γ}}^{⊤} Φ (u)}$ does not satisfy (5) when (2) is misspecified. To ensure that (5) holds regardless of the adequacy of the imputation model used for estimating the regression parameters, we augment the initial imputation ${\tilde{m}}_{I} (u)$ as

{\tilde{m}}_{II} (u; θ) = g {{\tilde{γ}}^{⊤} Φ (u) + {\tilde{ν}}_{θ}^{⊤} z_{θ}}

where $z_{θ} = {[1, Y (θ^{⊤} x)]}^{⊤}$ and ${\tilde{ν}}_{θ}$ is the solution to the IPW estimating equation

P_{n} (ν, θ) = \frac{1}{N} \sum_{i = 1}^{N} {\hat{w}}_{i} {y_{i} - g ({\tilde{γ}}^{⊤} Φ_{i} + ν^{⊤} z_{i θ})} z_{i θ} = 0 for any given θ .

(6)

We let ${\bar{ν}}_{θ}$ be the limiting value of ${\tilde{ν}}_{θ}$ which solves the limiting estimation equation

R (ν ∣ θ) = E (z_{θ} [y - g {{\bar{γ}}^{⊤} Φ (u) + ν^{⊤} z_{θ}}]) = 0 .

(7)

This estimating equation is monotone in z_θ for any θ and thus ${\bar{ν}}_{θ}$ exists under mild regularity conditions. It also follows from (7) that (i) $E ([y - g {{\bar{γ}}^{⊤} Φ (u) + {\bar{ν}}_{θ}^{⊤} z_{θ}}]) = 0$ and (ii) $E (Y (θ^{⊤} x) [y - g {{\bar{γ}}^{⊤} Φ (u) + {\bar{ν}}_{θ}^{⊤} z_{θ}}]) = 0$ which ensure that the sufficiency condition in (5) is satisfied. We thus construct a SS estimator of D(θ) as

{\hat{D}}_{SSL} (θ) = \frac{1}{N} \sum_{i = 1}^{N} d {{\tilde{m}}_{II} (u_{i}; θ), Y (θ^{⊤} x_{i})} .

In Section 5.2, we present the asymptotic properties of ${\hat{D}}_{SSL} ({\hat{θ}}_{SSL})$ and ${\hat{D}}_{SSL} ({\overset{ˇ}{θ}}_{SSL})$ and compare ${\hat{D}}_{SSL} ({\overset{ˇ}{θ}}_{SSL})$ with its supervised counterpart, ${\hat{D}}_{SL} ({\hat{θ}}_{SL})$ . Similar to the SS estimation of $\bar{θ}$ , it is shown in Appendix F that the unlabeled data helps to reduce the asymptotic variance of ${\hat{D}}_{SSL} ({\overset{ˇ}{θ}}_{SSL})$ . Specifically, ${\hat{D}}_{SSL} ({\overset{ˇ}{θ}}_{SSL})$ is shown to be asymptotically more efficient than ${\hat{D}}_{SL} ({\hat{θ}}_{SL})$ when the imputation model is correct. In practice, however, we may want to use ${\hat{D}}_{SSL} ({\hat{θ}}_{SSL})$ instead of ${\hat{D}}_{SSL} ({\overset{ˇ}{θ}}_{SSL})$ to achieve improved finite sample performance.

4. Bias Correction via Ensemble Cross-Validation

Similar to the supervised estimators of the prediction performance measures, the proposed plug-in estimator uses the labeled data for both constructing and evaluating the prediction model and is therefore prone to overfitting bias [Efron, 1986]. K-fold cross-validation (CV) is a commonly used method to correct for such bias. However, it has been observed that CV tends to result in overly pessimistic estimates of accuracy measures, particularly when n is not very large relative to p [Jiang and Simon, 2007]. Bias correction methods such as the 0.632 bootstrap have been proposed to address this behavior [Efron, 1983, Efron and Tibshirani, 1997, Fu et al., 2005, Molinaro et al., 2005]. Here, we propose an alternative ensemble CV procedure that takes a weighted sum of the apparent and K-fold CV estimators.

We first construct a K-fold CV estimator by randomly partitioning $L$ into K disjoint folds of roughly equal size, denoted by ${L_{k}, k = 1, \dots, K}$ . Since N is assumed to be sufficiently large, no CV is necessary for projecting to the full data. For a given k, we use $L / L_{k}$ to estimate $\bar{γ}$ and $\bar{θ}$ , denoted as ${\tilde{γ}}_{(- k)}$ and ${\hat{θ}}_{(- k)}$ , respectively. The $n_{k}$ observations in $L_{k}$ are used in the augmentation step to obtain ${\tilde{ν}}_{{\hat{θ}}_{(- k)}, k}$ , the solution to $P_{n_{k}} (ν, {\hat{θ}}_{(- k)}) = 0$ . For the k^th fold, we estimate the accuracy measure as

{\hat{D}}_{k} ({\hat{θ}}_{(- k)}) = N^{- 1} \sum_{i = 1}^{N} d {{\tilde{m}}_{II, k} (u_{i}), Y ({\hat{θ}}_{(- k)}^{⊤} x_{i})},

where ${\tilde{m}}_{II, k} (u_{i}) = g ({\tilde{γ}}_{(- k)}^{⊤} Φ_{i} + {\tilde{ν}}_{{\hat{θ}}_{(- k)}, k}^{⊤} Z_{i {\hat{θ}}_{(- k)}})$ , and take the final CV estimator as ${\hat{D}}_{SSL}^{cv} = K^{- 1} \sum_{k = 1}^{K} {\hat{D}}_{k} ({\hat{θ}}_{(- k)})$ . In practice, we suggest averaging over several replications of CV to remove the variation due to the CV partition. We then obtain the weighted CV estimator with

{\hat{D}}_{SSL}^{ω} = ω {\hat{D}}_{SSL} + (1 - ω) {\hat{D}}_{SSL}^{cv}, where ω = K / (2 K - 1) .

We may similarly obtain a CV-based supervised estimator, denoted by ${\hat{D}}_{SL}^{cv}$ , as well as the corresponding weighted estimator, ${\hat{D}}_{SL}^{ω}$ . Note that the fraction of observations from stratum s in the kth fold, ${\hat{ρ}}_{1 s, k}$ , deviates from ${\hat{ρ}}_{1 s}$ in the order of $O (\sqrt{n_{s}} / n)$ . Although this deviation is asymptotically negligibile, it may be desirable to perform the K-fold partition within each strata to ensure that ${\hat{ρ}}_{1 s, k} = {\hat{ρ}}_{1 s}$ when n_s is small or moderate.

Using similar arguments as those given in Tian et al. [2007], it is not difficult to show that $n^{\frac{1}{2}} ({\hat{D}}_{SSL} - \bar{D})$ and $n^{\frac{1}{2}} ({\hat{D}}_{SSL}^{cv} - \bar{D})$ are first-order asymptotically equivalent. Thus, the ensemble CV estimator ${\hat{D}}_{SSL}^{ω}$ reduces the higher order bias of ${\hat{D}}_{SSL}$ and ${\hat{D}}_{SSL}^{cv}$ , but has the same asymptotic distribution. Although the empirical performance is promising, it is difficult to rigorously study the bias properties of ${\hat{D}}_{SSL}^{ω}$ as the regression parameter doesn’t necessarily minimize the loss function, $\hat{D} (θ)$ . We provide a heuristic justification of the ensemble CV method in Appendix H which assumes $\hat{θ}$ minimizes $\hat{D} (θ)$ .

5. Asymptotic Analysis

We next present the asymptotic properties of our proposed SS estimator of $\bar{D}$ . To facilitate our presentation, we first discuss the properties of ${\overset{ˇ}{θ}}_{SSL}$ as the accuracy parameter estimates inherently depend on the variability in estimating $\bar{θ}$ . We then present our main result high-lighting the efficiency gain of our proposed SS approach for accuracy parameter estimation. We conclude our theoretical analysis with two practical discussions of (i) intrinsic efficient estimation in the SS setting and (ii) optimal allocation in stratified sampling.

For our asymptotic analysis, we let $Σ_{1} ≻ Σ_{2}$ if $Σ_{1} - Σ_{2}$ is positive definite and $Σ_{1} ⪰ Σ_{2}$ if $Σ_{1} - Σ_{2}$ is positive semi-definite for any two symmetric matrices Σ₁ and Σ₂. For any matrix M and vectors v₁ and v₂, M_j· represents the j^th row vector, $v_{1}^{\otimes 2} = v_{1} v_{1}^{⊤}$ , and ${v_{1}, v_{2}} = {(v_{1}^{⊤}, v_{2}^{⊤})}^{⊤}$ is the vector concatenating v₁ and v₂. To establish our theoretical results, we recall that ${\hat{ρ}}_{1 s}$ and ${\hat{ρ}}_{s}$ converge to some fixed values ρ_1s and ρ_s in probability, as assumed in Section 2.1, and introduce the following three conditions.

Condition 1.

The basis Φ(u) contains x, has compact support, and is of fixed dimension. The density function for x, denoted by p(x), and P(y = 1 | u) are continuously differentiable in the continuous components of x and u, respectively. There is at least one continuous component of x with corresponding non-zero component in $\bar{θ}$ .

Condition 2.

The link function g(·) is continuously differentiable with derivative $\dot{g} (\cdot)$ .

Condition 3.

(A) There is no vector γ such that $P (γ^{⊤} Φ_{1} > γ^{⊤} Φ_{2} ∣ y_{1} > y_{2}) = 1$ and $E [Φ^{\otimes 2} \dot{g} {{\bar{γ}}^{⊤} Φ}] > 0$ . (B) There is a small neighborhood of $\bar{θ}, Θ = {θ : ∥ θ - \bar{θ} ∥_{2} < δ}$ for some δ > 0, such that for any θ ∈ Θ, there is no vector r such that $P (r^{⊤} {Φ_{1}, z_{1 θ}} > r^{⊤} {Φ_{2}, Z_{2 θ}} ∣ y_{1} > y_{2}) = 1$ and $E [z_{θ}^{\otimes 2} \dot{g} {{\bar{γ}}^{⊤} Φ + {\bar{ν}}_{θ}^{⊤} z_{θ}}] ≻ 0$ . (C) $E [x^{\otimes 2} \dot{g} ({\bar{θ}}^{⊤} x)] ≻ 0$ .

Remark 1.

Conditions 1–3 are commonly used regularity conditions in M-estimation theory and are satisfied in broad applications. Similar conditions can be found in Tian et al. [2007] and Section 5.3 of Van der Vaart [2000]. Condition 3(A) and 3(B) assume that there is no γ and ν such that $γ^{⊤} Φ + ν^{⊤} z_{θ}$ can perfectly separate the samples based on y. In our application of EHR data analysis, these conditions are typically satisfied as the outcomes of interest (i.e. disease status) do not perfectly depend on covariates such as billing codes, lab values, procedure codes, and other features extracted from free-text. Similar to Tian et al. [2007], Condition 3(A) ensures the existence and uniqueness of the limiting parameters $\bar{θ}$ and $\bar{γ}$ and Condition 3(B) ensures the existence and uniqueness of ${\bar{ν}}_{θ}$ .

5.1. Asymptotic Properties of ${\overset{ˇ}{θ}}_{SSL}$

The asymptotic properties of ${\overset{ˇ}{θ}}_{SSL}$ are summarized in Theorem 1 and the justification is provided in Appendix E.

Theorem 1.

Under Conditions 1–3, ${\overset{ˇ}{θ}}_{S S L} \overset{p}{\to} \bar{θ}$ , and

{\hat{W}}_{S S L} = n^{\frac{1}{2}} ({\overset{ˇ}{θ}}_{S S L} - \bar{θ}) = n^{\frac{1}{2}} \sum_{s = 1}^{S} ρ_{s} (n_{s}^{- 1} \sum_{i = 1}^{N} V_{i} I (S_{i} = s) e_{S S L i}) + o_{p} (1)

which weakly converges to $N (0, Σ_{S S L})$ where

Σ_{S S L} = \sum_{s = 1}^{S} ρ_{s}^{2} ρ_{1 s}^{- 1} E {e_{S S L i}^{\otimes 2} ∣ S_{i} = s}, e_{S S L i} = A^{- 1} x_{i} {y_{i} - g ({\bar{γ}}^{⊤} Φ_{i})}, and A = E {x_{i}^{\otimes 2} \dot{g} ({\bar{θ}}^{⊤} x_{i})} .

Remark 2.

To contrast with the supervised estimator ${\hat{θ}}_{S L}$ , we show in Appendix C that

{\hat{W}}_{S L} = n^{\frac{1}{2}} ({\hat{θ}}_{S L} - \bar{θ}) = n^{\frac{1}{2}} \sum_{s = 1}^{S} ρ_{s} (n_{s}^{- 1} \sum_{i = 1}^{N} V_{i} I (S_{i} = s) e_{S L i}) + o_{p} (1),

which weakly converges to $N (0, Σ_{S L})$ where

Σ_{S L} = \sum_{s = 1}^{S} ρ_{s}^{2} ρ_{1 s}^{- 1} E {e_{S L i}^{\otimes 2} ∣ S_{i} = s}, and e_{S L i} = A^{- 1} x_{i} {y_{i} - g ({\bar{θ}}^{⊤} x_{i})} .

It follows that when the imputation model $P (y = 1 ∣ u) = g {{\bar{γ}}^{⊤} Φ (u)}$ is correctly specified, $Σ_{S L} ⪰ Σ_{S S L}$ . When $P ({\bar{γ}}^{⊤} Φ (u) \neq {\bar{θ}}^{⊤} x) > 0$ , we have that $Σ_{S L} ≻ Σ_{S S L}$ .

5.2. Asymptotic Properties of ${\hat{D}}_{SSL} ({\overset{ˇ}{θ}}_{SSL})$ and ${\hat{D}}_{SSL} ({\hat{θ}}_{SSL})$

The asymptotic properties of ${\hat{D}}_{SSL} ({\overset{ˇ}{θ}}_{SSL})$ are summarized in Theorem 2 and the justification is provided in Appendix F.

Theorem 2.

Under Conditions 1–3, ${\hat{D}}_{S S L} ({\overset{ˇ}{θ}}_{S S L}) \overset{p}{\to} \bar{D}$ , and ${\overset{ˇ}{T}}_{S S L} = n^{\frac{1}{2}} {{\hat{D}}_{S S L} ({\overset{ˇ}{θ}}_{S S L}) - D (\bar{θ})}$ is asymptotically Gaussian with mean zero and variance $σ_{S S L}^{2}$ given in Appendix F. Also, ${\overset{ˇ}{T}}_{S S L}$ is asymptotically equivalent to

n^{\frac{1}{2}} \sum_{s = 1}^{S} ρ_{s} (n_{s}^{- 1} \sum_{i = 1}^{N} V_{i} I (S_{i} = s) [{d (y_{i}, {\bar{Y}}_{i}) - d (m_{I I, i}, {\bar{Y}}_{i})} + \dot{D} {(\bar{θ})}^{⊤} e_{S S L i}]),

where ${\bar{Y}}_{i} = Y ({\bar{θ}}^{⊤} x_{i})$ , $m_{I I, i} = g ({\bar{γ}}^{⊤} Φ_{i} + \bar{ν} \bar{θ} z_{i \bar{θ}})$ is the imputation model based approximation to $P (y = 1 ∣ u)$ and $\dot{D} (θ) = \partial D (θ) / \partial θ$ .

Remark 3.

We also show that ${\hat{T}}_{S S L} = n^{\frac{1}{2}} {{\hat{D}}_{S S L} ({\hat{θ}}_{S S L}) - D (\bar{θ})}$ is asymptotically equivalent to

n^{\frac{1}{2}} \sum_{s = 1}^{S} ρ_{s} (n_{s}^{- 1} \sum_{i = 1}^{N} V_{i} I (S_{i} = s) [{d (y_{i}, {\bar{Y}}_{i}) - d (m_{I I, i}, {\bar{Y}}_{i})} + \dot{D} {(\bar{θ})}^{⊤} {W e_{S S L i} + (I - W) e_{S L i}}]),

which is also asymptotically Gaussian with mean zero where W is a diagonal matrix defined in Appendix B.

Remark 4.

As shown in Appendix D, ${\hat{T}}_{S L} = n^{\frac{1}{2}} {{\hat{D}}_{S L} ({\hat{θ}}_{S L}) - D (\bar{θ})}$ is asymptotically Gaussian with mean zero and variance $σ_{S L}^{2}$ defined in Appendix D. It is equivalent to

n^{\frac{1}{2}} \sum_{s = 1}^{S} ρ_{s} (n_{s}^{- 1} \sum_{i = 1}^{N} V_{i} I (S_{i} = s) [{d (y_{i}, {\bar{Y}}_{i}) - D (\bar{θ})} + \dot{D} {(\bar{θ})}^{⊤} e_{S L i}]) .

We verify in Appendix F that when the imputation model is correctly specified, the asymptotic variance of ${\hat{D}}_{SSL} ({\overset{ˇ}{θ}}_{SSL})$ is smaller than that of ${\hat{D}}_{SL} ({\hat{θ}}_{SL})$ regardless of the specification of the working regression model in (1). This is because the accuracy measures always depend on the marginal distribution of x and the proposed SS approach leverages $U$ . Therefore ${\hat{D}}_{SSL} ({\overset{ˇ}{θ}}_{SSL})$ is asymptotically more efficient than ${\hat{D}}_{SL} ({\hat{θ}}_{SL})$ even when model (1) is correctly specified and ${\hat{θ}}_{SL}$ is fully efficient.

While we cannot theoretically guarantee that the SS estimator is more efficient than the supervised estimator under misspecification of the imputation model, the first and dominating term in the influence function expansion corresponds to the variability from estimating the accuracy measure. Even when the imputation model is misspecified, it may still provide a close approximation to P(y = 1 | u) and therefore result in reduced variability relative to the supervised approach. The second term of the influence function corresponds to the variability from estimation of the regression parameter. As the SS estimator of the regression parameter is more efficient than its supervised counterpart under model misspecification, we also expect this term to have smaller variation than its supervised counterpart. In our simulation studies, we evaluate the performance of our proposals under various model misspecifications to assess whether this heuristic justification holds up empirically. We also further study this limitation from a theoretical viewpoint in the next section where we introduce a SS estimator with the intrinsic efficiency property from the semiparametric inference literature for comparison.

5.3. Intrinsic Efficient SS Estimation

For simplicity, we begin our discussion of intrinsic efficient estimation focusing on estimation of the regression parameter. Recall that the idea in Section 3.1 is to (i) solve $N^{- 1} \sum_{i = 1}^{N} {\hat{w}}_{i} Φ_{i} [y_{i} - g {γ^{⊤} Φ_{i}}] - λ_{n} γ = 0$ to obtain estimated coefficients $\tilde{γ}$ for imputation and then (ii) solve $N^{- 1} \sum_{i = 1}^{N} x_{i} {g ({\tilde{γ}}^{⊤} Φ_{i}) - g (θ^{⊤} x_{i})} = 0$ to obtain the SS estimator, ${\hat{θ}}_{SSL}$ . By Theorem 1, for any $e \in ℝ^{p + 1} ∖ {0}$ , the asymptotic variance of $n^{\frac{1}{2}} (e^{⊤} {\hat{θ}}_{SSL} - e^{⊤} \bar{θ})$ can be expressed as:

\frac{1}{n} \sum_{i = 1}^{n} ζ_{i} {(e^{⊤} A^{- 1} x_{i})}^{2} {y_{i} - g ({\bar{γ}}^{⊤} Φ_{i})}^{2},

(8)

where $ζ_{i} = \sum_{s = 1}^{S} ρ_{s}^{2} ρ_{1 s}^{- 2} V_{i} I (S_{i} = s)$ for each $i \in {1, 2, \dots, N}$ . When the imputation model $P (y = 1 ∣ u) = g (γ^{⊤} Φ)$ is misspecified, an alternative estimating equation for γ may be used to directly reduce the asymptotic variance of the resulting SS estimator. Specifically, for a fixed Φ, we may find the estimating equation for γ that leads to the lowest asymptotic variance of the estimator for $e^{⊤} \bar{θ}$ , a property referred to as “intrinsic efficiency” in the semiparametric inference literature [Tan, 2010]. We briefly propose estimation procedures for an estimator achieving this property with potential to improve upon our original proposal under potential misspecification of the imputation model.

To directly minimize the asymptotic variance of the SS estimator of $e^{⊤} \bar{θ}$ given by (8), we obtain the estimated coefficients for the imputation model with

{\tilde{γ}}^{(1)} = \underset{γ}{argmin} \frac{1}{2 n} \sum_{i = 1}^{n} {\hat{ζ}}_{i} {(e^{⊤} {\hat{A}}^{- 1} x_{i})}^{2} {y_{i} - g (γ^{⊤} Φ_{i})}^{2} + λ_{n}^{(1)} ∥ γ ∥_{2}^{2}, s.t. \frac{1}{N} \sum_{i = 1}^{N} {\hat{w}}_{i} x_{i} {y_{i} - g (γ^{⊤} Φ_{i})} = 0,

(9)

where ${\hat{ζ}}_{i} = \sum_{s = 1}^{S} {\hat{ρ}}_{s}^{2} {\hat{ρ}}_{1 s}^{- 2} V_{i} I (S_{i} = s)$ , $\hat{A} = N^{- 1} \sum_{i = 1}^{N} x_{i}^{\otimes 2} \dot{g} ({\hat{θ}}_{SSL}^{⊤} x_{i})$ are the empirical estimates of $ζ_{i}$ and A, respectively, and $λ_{n}^{(1)} = O (n^{- \frac{1}{2}})$ is again a tuning parameter for stable fitting. We then solve $N^{- 1} \sum_{i = 1}^{N} x_{i} {g ({\tilde{γ}}^{(1) ⊤} Ψ_{i}) - g (θ^{⊤} x_{i})} = 0$ to obtain ${\hat{θ}}_{intri}$ , and return $e^{⊤} {\hat{θ}}_{intri}$ as the intrinsic efficient estimator for $e^{⊤} \bar{θ}$ . The moment condition in (9) is used for calibrating the potential bias from a misspecified imputation model and ensuring the consistency of ${\hat{θ}}_{intri}$ . This condition is explicitly imposed when constructing our original proposal.

To study the asymptotic properties of ${\hat{θ}}_{intri}$ and compare it with our original proposal, ${\hat{θ}}_{SSL}$ , we let

{\bar{γ}}^{(1)} = \underset{γ}{argmin} E [R {(e^{⊤} A^{- 1} x)}^{2} {y - g (γ^{⊤} Φ)}^{2}], s.t. E [x {y - g (γ^{⊤} Φ)}] = 0,

be the limit of ${\tilde{γ}}^{(1}$ , where $R = \sum_{s = 1}^{S} I (S = s) ρ_{s} / ρ_{1 s}$ . The proof of Theorem 3 is provided in Appendix G.2.

Theorem 3.

Under condition 1, and conditions A1 and A2 introduced in Appendix G.2, $n^{\frac{1}{2}} ({\hat{θ}}_{i n t r i} - \bar{θ})$ converges weakly to a mean zero normal distribution, and is asymptotically equivalent to $\hat{W} ({\bar{γ}}^{(1)})$ where

\hat{W} (γ) = n^{\frac{1}{2}} \sum_{s = 1}^{S} ρ_{s} [n_{s}^{- 1} \sum_{i = 1}^{N} V_{i} I (S_{i} = s) A^{- 1} x_{i} {y_{i} - g (γ^{⊤} Φ_{i})}] .

In addition: (i) when the imputation model $P (y = 1 ∣ u) = g (γ^{⊤} Φ)$ is correctly specified, ${\hat{θ}}_{i n t r i}$ is asymptotically equivalent to ${\hat{θ}}_{S S L}$ and (ii) the asymptotic variance of $n^{\frac{1}{2}} (e^{⊤} {\hat{θ}}_{i n t r i} - e^{⊤} \bar{θ})$ is minimized among estimators with ${e^{⊤} \hat{W} (γ) : E [x {y - g (γ^{⊤} Φ)}] = 0}$ . Consequently, the variance of the intrinsic efficient estimator is always less than or equal to the asymptotic variance of both $n^{\frac{1}{2}} (e^{⊤} {\hat{θ}}_{S L} - e^{⊤} \bar{θ})$ and $n^{\frac{1}{2}} (e^{⊤} {\hat{θ}}_{S S L} - e^{⊤} \bar{θ})$ .

The details and theoretical analysis of the intrinsic efficient estimation procedure of the accuracy measure $\bar{D}$ is presented Appendices G.1 and G.2. Similar to Theorem 3, we show that ${\hat{D}}_{intri}$ is asymptotically equivalent with our proposal, ${\hat{D}}_{SSL}$ , when the imputation model is correctly specified and has smaller asymptotic variance than ${\hat{D}}_{SSL}$ when the imputation model is misspecified. However, it is important to note that estimation based on intrinsic efficiency is a non-convex problem and one may encounter numerical optimization issues which may limit its use in practice. We provide simulation studies comparing the intrinsic efficient estimator to the proposed approach in Section S4 Supplement.

5.4. Optimal Allocation in Stratified Sampling

Another important practical issue is how to select the strata and the corresponding selection probabilities. Here we provide here a detailed assessment of the optimal (or Neyman) allocation of the labeled data across the strata. Specifically, the general form of the influence function for our estimators is

n^{1 / 2} \sum_{s = 1}^{S} ρ_{s} {n_{s}^{- 1} \sum_{i = 1}^{N} V_{i} I (S_{i} = s) f (F_{i})} + o_{p} (1)

for a function f with $σ_{s}^{2} = E {f^{2} (F_{i}) ∣ S_{i} = s}$ and $E {f (F_{i})} = 0$ . The asymptotic variance can then be expressed as

\sum_{s = 1}^{S} ρ_{s}^{2} {n_{s}^{- 2} \sum_{i = 1}^{N} V_{i} I (S_{i} = s) E {f^{2} (F_{i}) ∣ S_{i} = s}} = \sum_{s = 1}^{S} ρ_{s}^{2} \frac{σ_{s}^{2}}{n_{s}} = n^{- 1} \sum_{s = 1}^{S} n_{s} \sum_{s = 1}^{S} \frac{{(ρ_{s} σ_{s})}^{2}}{n_{s}} ⩾ n^{- 1} {(\sum_{s = 1}^{S} ρ_{s} σ_{s})}^{2},

by the Cauchy-Schwarz inequality, and equality holds if and only if

n_{s} = n \frac{ρ_{s} σ_{s}}{\sum_{s = 1}^{S} ρ_{s} σ_{s}} for s = 1, \dots, S .

(10)

The optimal sampling probabilities are therefore proportional to (i) the relative stratum size and (ii) the variability within the stratum, with greater weight placed on large stratum with high variability. Consequently, stratified sampling leads to a more efficient estimator than uniform random sampling when the allocation in (10) is used.

Remark 5.

There is a rich body of survey sampling literature concerning model-assisted approaches that address the practically important question of how to select the strata and the corresponding selection probabilities [Neyman, 1934, Särndal et al., 2003, Nedyalkova and Tillé, 2008]. The optimal allocation given by (10) is in similar spirit with the sampling schemes used in Cai and Zheng [2012], Liu et al. [2012]. It is particularly useful for HER-based phenotyping studies such as the diabetic neuropathy example in Section 8 as it is often straightforward for domain experts to define a filter variable(s) that yields relatively large stratum with increased prevalence of y (e.g. patients with notes containing terms related to the disease, relevant lab values, or specialist visits) and thus increased variability.

We provide additional numerical studies to illustrate Remark 5 in Section S5 of the Supplement.

6. Perturbation Resampling Procedure for Inference

We next propose a perturbation resampling procedure to construct standard error (SE) and confidence interval (CI) estimates in finite samples. Resampling procedures are particularly attractive for making inference about $\bar{D}$ when $Y = Y_{2}$ since ${\hat{D}}_{SSL} (θ)$ is not differentiable in θ. To this end, we generate a set of independent and identically distributed (i.i.d) non-negative random variables, $G = (G_{1}, \dots, G_{n})$ , independent of $D$ , from a known distribution with mean one and unit variance.

For each set of $G$ , we first obtain a perturbed version of ${\hat{θ}}_{SSL}$ as

{\hat{θ}}^{*} = {\hat{θ}}_{SSL} + {\hat{A}}^{- 1} \sum_{k = 1}^{K} \sum_{i \in L_{k}} \frac{{\hat{w}}_{i} (G_{i} - 1)}{\sum_{j = 1}^{n} {\hat{w}}_{j}} [x_{i} y_{i} - \hat{W} x_{i} g ({\tilde{γ}}_{(- k)}^{⊤} Φ_{i}) - (I - \hat{W}) x_{i} g ({\hat{θ}}_{(- k)}^{⊤} x_{i})],

where $\hat{A} = N^{- 1} \sum_{i = 1}^{N} x_{i}^{\otimes 2} \dot{g} ({\hat{θ}}_{SSL}^{⊤} x_{i})$ . We use CV to correct for variance underestimation due to overfitting. Next, we find the solution ${\tilde{γ}}^{*}$ to the perturbed objective function

{\tilde{Q}}_{n}^{*} (γ) = \frac{\sum_{i = 1}^{n} {\hat{w}}_{i} Φ_{i} [y_{i} - g (γ^{⊤} Φ_{i})] G_{i}}{\sum_{i = 1}^{n} {\hat{w}}_{i} G_{i}} - λ_{n} γ = 0

(11)

and the solution ${\tilde{ν}}^{*}$ that solves

P_{n}^{*} (ν, {\hat{θ}}^{*}) = \frac{\sum_{i = 1}^{n} {\hat{w}}_{i} {y_{i} - g ({\tilde{γ}}^{*} Φ_{i} + ν^{⊤} z_{i {\hat{θ}}^{*}})} z_{i {\hat{θ}}^{*}} G_{i}}{\sum_{i = 1}^{n} {\hat{w}}_{i} G_{i}} = 0

to obtain perturbed counterparts of $\tilde{γ}$ and $\tilde{ν}$ , respectively. We then compute ${\tilde{m}}_{11}^{*} (u_{i}) = g ({\tilde{γ}}^{* ⊤} Φ_{i} + {\tilde{ν}}^{* ⊤} z_{i {\hat{θ}}^{*}})$ and obtain the perturbed estimator of ${\hat{D}}_{SSL} ({\hat{θ}}_{SSL})$ as

{\hat{D}}_{SSL}^{*} ({\hat{θ}}^{*}) = N^{- 1} \sum_{i = 1}^{N} [{\tilde{m}}_{11}^{*} (u_{i}) {1 - 2 Y ({\hat{θ}}^{* ⊤} x_{i})} + Y^{2} ({\hat{θ}}^{* ⊤} x_{i})] .

Following arguments such as those in Tian et al. [2007], one may verify that $n^{\frac{1}{2}} {{\hat{D}}_{SSL}^{*} ({\hat{θ}}_{SSL}^{*}) - {\hat{D}}_{SSL} ({\hat{θ}}_{SSL})} ∣ D$ converges to the same limiting distribution as ${\hat{T}}_{SSL}$ . Additionally, it may be shown that ${\hat{T}}_{SSL}^{cv} = n^{\frac{1}{2}} {{\hat{D}}_{SSL}^{cv} - D (\bar{θ})}$ and hence ${\hat{T}}_{SSL}^{ω} = n^{\frac{1}{2}} {{\hat{D}}_{SSL}^{ω} - D (\bar{θ})}$ converge to the limiting distribution of ${\hat{T}}_{SSL}$ . We utilize these results to approximate the distribution of ${\hat{T}}_{SSL}^{ω}$ with the empirical distribution of a large number of perturbed estimates using the above resampling procedure to base inference for $D (\bar{θ})$ on the proposed bias-corrected estimator. The variance of ${\hat{D}}_{SSL}^{ω}$ can correspondingly be estimated with the sample variance and confidence intervals may be constructed accordingly.

7. Simulation Studies

We conducted extensive simulation studies to evaluate the performance of the proposed SSL procedures and to compare to existing methods. Throughout, we generated p = 10 dimensional covariates x from N(0, C) with $C_{k l} = 3 {(0.4)}^{| k - l |}$ . Stratified sampling was performed according to $S$ generated from the following two mechanisms:

$S \in {1, S = 2} with S = 1 + I (x_{1} + δ_{1} ⩽ 0.5) and δ_{1} \sim N (0, 1) .$
$S \in {1, 2, 3, S = 4} with S = 1 + I (x_{1} + δ_{1} ⩽ 0.5) + 2 I (x_{3} + δ_{2} ⩽ 0.5), δ_{1} \sim N (0, 1), δ_{2} \sim N (0, 1), and δ_{1} ⊥ δ_{2} .$

We let $S = {(I (S = 1), \dots, I (S = S - 1))}^{⊤}$ . For both settings, we sampled n_s = 100 or 200 observations from each stratum. Throughout, we let v₁ be the natural spline of x with 3 knots and v₂ be the interaction terms ${x_{1} : x_{- 1}, x_{2} : x_{- (1, 2)}}$ , where $x_{1} : x_{- 1}$ and $x_{2} : x_{- (1, 2)}$ represent interaction terms of x₁ with the remaining covariates and x₂ with covariates excluding x₁ and x₂, respectively. With $θ = {0, 1, 1, 0.5, 0.5, 0_{(p - 4) \times 1}}^{⊤}$ and $ϵ_{logistic}$ and $ϵ_{extreme}$ denoting noise generated from the logistic and extreme valuep−2,0.3q distributions, we simulated y from the following models:

$(M_{correct}, I_{correct})$ with correct outcome model and correct imputation model:
$y = I (θ^{⊤} x + ϵ_{logistic} > 2) and Φ = {(1, x^{⊤}, v_{1}^{⊤}, S^{⊤})}^{⊤};$
$(M_{incorrect}, I_{correct})$ with incorrect outcome model and correct imputation model:
$y = I [θ^{⊤} x + 0.5 {x_{1} x_{2} + x_{1} x_{5} - x_{2} x_{6} - I (S = 1)} + ϵ_{logistic} > 0] and Φ = {(1, x^{⊤}, v_{2}^{⊤}, S^{⊤})}^{⊤};$
$(M_{incorrect}, I_{incorrect})$ with incorrect outcome model and incorrect imputation model:
$y = I {θ^{⊤} x + x_{1}^{2} + x_{3}^{2} + \exp (- 2 - 3 x_{4} - 3 x_{6}) ϵ_{extreme} > 2} and Φ = {(1, x^{⊤}, v_{1}^{⊤}, S^{⊤})}^{⊤} .$

While the outcome model is misspecified in both (ii) and (iii), the misspecification is more severe in (iii) due to the higher magnitude of nonlinear effects. These configurations are chosen to mimic EHR settings where the signals are typically sparse and S is small. The covariate effects of 1 represent the strong signals from the main billing codes and free-text mentions of the disease of interest. The two weaker signals 0.5 characterize features such as related medications, signs, symptoms and lab results relevant to the disease of interest.

Across all settings, we compare our SS estimators to both the SL estimator and the alternative density ratio (DR) method [Kawakita and Kanamori, 2013, Kawakita and Takeuchi, 2014]. The basis function $φ (u)$ required in the DR method was chosen to be the same as $Φ (u)$ in our method for all settings. The details and theoretical properties of the DR method are further discussed in the Supplement. We employed the ensemble CV strategy to construct a bias corrected DR estimator for $\bar{D}$ , denoted as ${\hat{D}}_{DR}^{ω}$ , to ensure a fair comparison to our approach. The three settings of outcome and imputation models under (i), (ii), and (iii) allow us to verify the asymptotic efficiency of the proposed SSL procedures relative to the SL and DR methods under various scenarios of misspecification.

For each configuration, results are summarized with 500 independent data sets. The size of the unlabeled data was chosen to be 20,000 across all settings. For all our numerical studies including the real data application, CV was performed with either K = 3 or K = 6 and averaged over 20 replications. The estimated SEs were based on 500 perturbed realizations and the OMR was evaluated with c = 0.5. We let $λ_{n} = \log (2 p) / n^{1.5}$ when fitting the ridge penalized logistic regression. We focus primarily on results for S = 2 and K = 6, but include results for S = 4 and K = 3 in Section S3 of the Supplement as they show similar patterns. Additionally, our analyses concentrate on the performance of the accuracy metrics. Results for the regression parameter estimates can be found in Section S2 of the Supplement. The code to implement the proposed methods and run the simulation studies can be found at https://github.com/jlgrons/Stratified-SSL.

In Figure 1, we present the percent biases of the apparent, CV, and ensemble CV estimators of the accuracy parameters. Although all three estimators have negligible biases, the SSL exhibits slightly less bias than its supervised counterpart and the DR estimator under $(M_{incorrect}, I_{incorrect})$ . The ensemble CV method is effective in bias correction while the apparent estimators are optimistic and the standard CV estimator is pessimistic. For example, under $(M_{incorrect}, I_{incorrect})$ , n_s = 100 and S = 2, the percent bias of the SSL estimators for the OMR are −8.2%, 8.8% and −0.5% when we use the plug-in, 6-fold CV, and the ensemble CV methods, respectively. The efficiency of ${\hat{D}}_{SSL}^{w}$ and ${\hat{D}}_{DR}^{ω}$ relative to ${\hat{D}}_{SL}^{w}$ for both the Brier scoreSL and OMR are presented in Figure 2. Again, ${\hat{D}}_{SSL}^{w}$ is substantially more efficient than ${\hat{D}}_{SL}^{w}$ and ${\hat{D}}_{DR}^{ω}$ , with efficiency gains of approximately 15%–30% under $(M_{correct}, I_{correct})$ , 40% under $(M_{correct}, I_{correct})$ , and 40%–80% under $(M_{correct}, I_{correct})$ . Results for K = 3 have similar patterns and are presented in Figure S1 and S2 of the Supplement. In Table 1, we present the results for the interval estimation obtained from the perturbation resampling procedure for the SS estimator ${\hat{D}}_{SSL}^{ω}$ . The SEs are well approximated and the empirical coverage for the 95% CIs is close to the nominal level across all settings.

Figure 2: — Relative efficiency (RE) of ${\hat{D}}_{SSL}^{ω}$ (SSL) and ${\hat{D}}_{DR}^{ω}$ compared to ${\hat{D}}_{SL}^{ω}$ for the Brier score (BS) and OMR under (i) $(ℳ_{correct}, ℐ_{correct})$ , (ii) $(ℳ_{incorrect}, ℐ_{correct})$ , and (iii) $(ℳ_{incorrect}, ℐ_{incorrect})$ . Shown are the results obtained with K = 6 fold CV.

Table 1:

The 100×ESE of ${\hat{D}}_{SL}^{w}$ , ${\hat{D}}_{SSL}^{w}$ and ${\hat{D}}_{DR}^{w}$ under (i) $(ℳ_{correct}, ℐ_{correct})$ ; (ii) $(ℳ_{incorrect}, ℐ_{correct})$ , and (iii) $(ℳ_{incorrect}, ℐ_{incorrect})$ . For ${\hat{D}}_{SSL}^{w}$ , we also show the average of the 100 × ASE as well as the empirical coverage probability (CP) of the 95% confidence intervals constructed based on the resampling procedure.

(i)

(ℳ_{correct}, ℐ_{correct})

	Brier score				OMR

	${\hat{D}}_{SL}^{w}$	${\hat{D}}_{SSL}^{w}$		${\hat{D}}_{DR}^{w}$	${\hat{D}}_{SL}^{w}$	${\hat{D}}_{SSL}^{w}$		${\hat{D}}_{DR}^{w}$

	ESE	ESE_ASE	CP	ESE	ESE	ESE_ASE	CP	ESE

S = 2, n_s = 100	1.27	1.19_1.25	0.94	1.31	2.29	2.10_2.23	0.97	2.35
S = 2, n_s = 200	0.97	0.87_0.86	0.93	0.95	1.67	1.50_1.60	0.97	1.67
S = 4, n_s = 100	0.97	0.89_0.85	0.93	0.95	1.70	1.54_1.58	0.95	1.64
S = 4, n = 200	0.67	0.58_0.60	0.95	0.66	1.16	1.02_1.11	0.97	1.16

(ii)

(ℳ_{incorrect}, ℐ_{correct})

	Brier score				OMR

	${\hat{D}}_{SL}^{w}$	${\hat{D}}_{SSL}^{w}$		${\hat{D}}_{DR}^{w}$	${\hat{D}}_{SL}^{w}$	${\hat{D}}_{SSL}^{w}$		${\hat{D}}_{DR}^{w}$

	ESE	ESE_ASE	CP	ESE	ESE	ESE_ASE	CP	ESE

S = 2, n_s = 100	1.72	1.47_1.39	0.93	1.85	3.09	2.56_2.65	0.94	3.42
S = 2, n_s = 200	1.13	1.01_0.94	0.92	1.16	2.14	1.85_1.85	0.95	2.21
S = 4, n_s = 100	1.21	1.07_0.96	0.92	1.24	2.13	1.87_1.88	0.95	2.18
S = 4, n_s = 200	0.86	0.73_0.67	0.93	0.87	1.62	1.37_1.34	0.94	1.62

(iii)

(ℳ_{incorrect}, ℐ_{incorrect})

	Brier score				OMR

	${\hat{D}}_{SL}^{w}$	${\hat{D}}_{SSL}^{w}$		${\hat{D}}_{DR}^{w}$	${\hat{D}}_{SL}^{w}$	${\hat{D}}_{SSL}^{w}$		${\hat{D}}_{DR}^{w}$

	ESE	ESE_ASE	CP	ESE	ESE	ESE_ASE	CP	ESE

S = 2, n_s = 100	1.57	1.31_1.29	0.94	1.53	2.68	2.15_2.31	0.96	2.61
S = 2, n_s = 200	1.05	0.86_0.85	0.95	1.01	1.87	1.45_1.54	0.97	1.83
S = 4, n_s = 100	1.11	0.87_0.88	0.96	0.99	1.97	1.46_1.57	0.96	1.87
S = 4, n_s = 200	0.78	0.61_0.61	0.95	0.73	1.37	1.02_1.08	0.97	1.30

Open in a new tab

Remark 6.

While our simulation studies focus on the SSL estimators proposed in Section 3, we also investigated the finite sample performance of ${\hat{θ}}_{i n t r i}$ and ${\hat{D}}_{i n t r i}$ and compared them with our original proposals. The numerical studies are described in Section S4 of the Supplement and demonstrate that when the estimated coefficients for the imputation model of the original estimators are equal or close to those of the intrinsic efficient estimator, these two methods perform equivalently with respect to mean square errors (MSE). In contrast, under a setting where the coefficients for the imputation model differ across these two methods, ${\hat{θ}}_{i n t r i}$ and ${\hat{D}}_{i n t r i}$ have about 30% smaller MSE than ${\hat{θ}}_{S S L}$ and ${\hat{D}}_{S S L}$ , on average. The detailed results are presented in Table S6 of the Supplement.

Remark 7.

To illustrate the benefit of stratified sampling in both the supervised and SS settings, we provide numerical studies of the optimal allocation in S5 of the Supplement. Mimicking our example in Section 8, we let the risk of y significantly differ across the S = 2 sampling groups. The stratification variable is picked so that $P (y = 1 ∣ S = 1)$ is much lower than $P (y = 1 ∣ S = 2)$ . We consider two sampling strategies: (i) uniform random sampling of n subjects, and (ii) stratified sampling of n/2 subjects from each stratum. Since $P (y = 1 ∣ S = 1)$ is low and close to 0, the variability of this stratum $σ_{1}^{2} = E {f^{2} (F) ∣ S = 1}$ is smaller than that of $S$ = 2 with $P (y = 1 ∣ S = 2)$ not near 0 and 1. Connecting this with (10), the stratified sampling strategy oversamples within $S = 2$ so that its allocation of n₂ is more close to the optimal choice. Consistent with this observation, our simulation results indicate that stratified sampling is more efficient than uniform random sampling in both the supervised and SS settings with an average relative efficiency > 1.45 across different setups. We further inspect the supervised estimator of $D (\bar{θ})$ under setup (I) in Section S5, for which the optimal allocation is n₁ = 0.47n and n₂ = 0.53n and nearly coincides with our equal allocation of n across the two stratum.

8. Example: EHR Study of Diabetic Neuropathy

We applied the proposed SSL procedures to develop and evaluate an EHR phenotyping algorithm for classifying diabetic neuropathy (DN), a common and serious complication of diabetes resulting in nerve damage. The full study cohort consists of N = 16,826 patients over age 18 with one or more of 12 ICD9 codes relating to DN identified from Partners HealthCare EHR. An initial assessment of 100 charts by physicians revealed the prevalence of DN in the study cohort was approximately 17%. To obtain a labeled set with sufficient DN cases for model training and improve efficiency, the investigators decided to employ a stratified sampling scheme. To do so, a binary filter variable $S$ indicating whether a patient had a neurological exam and a neurology note with at least 1,000 characters was created. The prevalence of DN in the “enriched” stratum with $S = 1$ was expected to be higher than that in the stratum with $S = 0$ . As demonstrated in our theoretical analysis in Section 5.4 and our numerical studies in Section S5, oversampling within the enriched set can improve estimation efficiency relative to taking a uniform random sample and is a common approach taken in EHR-based analyses. For this study, the investigators sampled n₀ = 70 and n₁ = 538 patients from the N₀ = 13608 patients with $S = 0$ and the N₁ = 3218 patients with $S = 1$ , respectively, for developing the phenotyping algorithm.

To train the model for classifying DN, a set of 11 codified and NLP features related to DN were selected from an original list of 75 via an unsupervised screening as described in Yu et al. [2015]. The codified features included $S$ , diagnostic codes for diabetes, type 2 diabetes mellitus, diabetic neuropathy, other idiopathic peripheral autonomic neuropathy, and diabetes mellitus with neurological manifestation as well as normal glucose lab values and prescriptions for anti-diabetic medications. The NLP features included mentions of terms related to DN in the patient record including glycosylated hemoglobin (HgA1c), diabetic, and neuropathy. As all these features (with the exception of $S$ ) are count variables and tend to be highly skewed, we used the transformation $x \to \log (x + 1)$ to stabilize model fitting.

We developed DN classification models by fitting a logistic regression with the above features based on ${\hat{θ}}_{SL}, {\hat{θ}}_{SSL}$ and ${\hat{θ}}_{DR}$ obtained from density ratio weighted estimation. Since the proportion of observations with $S = 0$ is relatively low in the labeled data, we implemented 100 replications of 6-fold CV procedure by splitting the data randomly within each strata to improve the stability of the CV procedure. To construct the basis for ${\hat{θ}}_{SSL}$ and ${\hat{θ}}_{DR}$ , we used a natural spline with 3 knots on all covariates except $S$ . To improve training stability, we set the ridge tuning parameter λ_n = n⁻¹ when fitting the imputation model.

As shown in Table 2(a), the point estimates are reasonably similar which confirms the consistency and stability of SS estimator in a real data setting. As expected, we find that the two most influential predictors are the diagnostic code for diabetic neuropathy and anti-diabetic medications. Importantly, we note substantial efficiency gains of ${\hat{θ}}_{SSL}$ compared to ${\hat{θ}}_{SL}$ . The SSL estimates are > 50% more efficient than the SL estimates for several features including six diagnostic code features and one NLP feature for DN. Additionally, ${\hat{θ}}_{SSL}$ is the most efficient estimator among all three approaches for nearly all variables.

Table 2:

Results from the diabetic neuropathy EHR study: (a) estimates (Est.) of the regression parameters for both codified (COD) and NLP features based on ${\hat{θ}}_{SL}$ , ${\hat{θ}}_{SSL}$ and ${\hat{θ}}_{DR}$ along with their estimated SEs and the coordinate-wise relative efficiencies (RE) of ${\hat{θ}}_{SSL}$ compared to ${\hat{θ}}_{SL}$ and (b) ${\hat{D}}_{SL}^{ω}$ , ${\hat{D}}_{SSL}^{ω}$ and ${\hat{D}}_{DR}^{ω}$ along with their estimated SEs and relative efficiencies (RE) of ${\hat{D}}_{SSL}^{ω}$ and ${\hat{D}}_{DR}^{ω}$ compared to ${\hat{D}}_{SL}^{ω}$ .

(a) Estimates of the Regression Coefficients

		${\hat{θ}}_{SL}$		${\hat{θ}}_{SSL}$			${\hat{θ}}_{DR}$
		Est.	SE	Est.	SE	RE	Est.	SE	RE
	Intercept	−3.89	0.80	−3.85	0.73	1.21	−3.26	0.66	1.45

COD	Neurological exam & note	−1.09	0.81	−0.90	0.61	1.77	−2.38	1.16	0.48
	Diabetes	−0.24	0.57	0.02	0.43	1.79	−0.09	0.48	1.42
	Type 2 Diabetes Mellitus	−1.05	0.73	−0.52	0.60	1.62	−0.62	0.86	0.73
	Diabetic Neuropathy	1.99	0.86	1.79	0.68	1.58	1.66	1.08	0.63
	Anti-diabetic Meds	1.70	0.59	1.12	0.44	1.79	1.50	0.51	1.32
	Diabetes Mellitus with
	Neuro Manifestation	0.36	1.17	0.60	0.98	1.42	0.59	0.90	1.68
	Other Idiopathic Peripheral
	Autonomic Neuropathy	0.86	0.71	0.93	0.68	1.09	1.01	0.76	0.89
	Normal Glucose	−0.56	0.65	−0.20	0.51	1.64	−1.27	0.80	0.66

NLP	Diabetic	0.30	0.58	−0.57	0.49	1.39	0.14	0.65	0.80
	HgA1c	−0.52	0.75	−0.64	0.70	1.16	−0.67	0.85	0.79
	Neuropathy	0.27	0.58	0.30	0.47	1.55	0.37	0.54	1.15

(b) Estimates of the Accuracy Parameters (×100)

	${\hat{D}}_{SL}^{ω}$		${\hat{D}}_{SSL}^{ω}$			${\hat{D}}_{DR}^{ω}$
	Est.	SE	Est.	SE	RE	Est.	SE	RE
Brier score	8.97	2.09	9.59	1.68	1.55	9.60	1.87	1.26
OMR	12.87	3.25	14.01	2.55	1.63	12.29	3.04	1.14

Open in a new tab

In Table 2(b), we compare ${\hat{D}}_{SL}^{w}, {\hat{D}}_{SSL}^{w}$ and ${\hat{D}}_{DR}^{w}$ for the Brier score and OMR with c = 0.5. While the point estimates for the accuracy measures based on these different approaches are relatively similar, ${\hat{D}}_{SSL}^{w}$ is 55% more efficient than ${\hat{D}}_{SL}^{w}$ for the Brier score and 63% more efficient for the OMR. Again, ${\hat{D}}_{SSL}^{w}$ is substantially more efficient than the DR estimator ${\hat{D}}_{DR}^{w}$ . These results support the potential value of our method for EHR-based research as these gains in efficiency may be directly translated into requiring fewer labeled examples for model evaluation.

9. Discussion

In this paper, we focused on the evaluation of a classification rule derived from a working regression model under stratified sampling in the SS setting. In particular, we introduced a two-step imputation-based method for estimation of the Brier score and OMR that makes use of unlabeled data. Additionally, as a by-product of our procedure, we obtained an efficient SS estimator of the regression parameter. Through theoretical and numerical studies, we demonstrated the advantage of the SS estimator over the SL estimator with respect to efficiency. We also developed a weighted CV procedure to adjust for overfitting and a resampling procedure for making inference. Our numerical studies indicate that our proposed method outperforms the existing DR method for SSL in the finite sample studies and we provide further discussion of this finding in Section S6 of the Supplement. Importantly, this article is one of the first theoretical studies of labeling based on stratified sampling within the SSL literature. We focus on the stratified sampling scheme due to its direct application to a variety of EHR-based analyses, including the development of a phenotyping algorithm for diabetic neuropathy presented in the previous section.

In our numerical studies, we used spline functions with 3 or 4 knots and interaction terms for the imputation model. It would be possible to use more knots or add more features to the basis function for settings with a larger n. However, care must be taken to avoid overfitting and potential loss in the efficiency gain of the SS estimator in finite sample. Alternatively, other basis functions can be utilized provided that Φ(u) contains x in its components to ensure consistency of the regression parameter. In settings where nonlinear effects of x on y are present, it may be desirable to impose a more complex outcome model to improve the prediction performance. A potential approach is to explicitly include nonlinear basis functions in the outcome model. In Section S7 of the Supplementary Materials, we consider using the leading principal components (PCs) of x and Ψ(x) where Ψ(·) is a vector of nonlinear transformation functions under a variety of settings. This approach performs similarly or better than the commonly used random forest model with respect to predictive accuracy, suggesting the utility of our proposed methods in the presence of nonlinear effects. Our numerical results also illustrate the efficiency gain of the SS estimators of the Brier score and OMR relative to the SL and DR methods. It is important to note, however, that the dimensions of both x and Φ were assumed to be fixed in our asymptotic analysis. Accommodating more complex modeling with p not small relative to n requires extending the proposed SSL approach to settings where x and Φ are high dimensional.

For accuracy estimation, we proposed an ensemble CV estimator that eliminates first-order bias when the estimated regression parameter is the minimizer of the empirical performance measure. Though this condition may not hold when the outcome model is misspecified, we have found that the suggested weights perform well in our numerical studies. Such ensemble methods that accommodate the more general case in both the supervised and SS settings warrant further research. Additionally, an important setting where our proposed SSL procedure would be of great use is in drawing inferences about two competing regression models. As it is likely that at least one model is misspecified, we would expect to observe efficiency gains in estimating the difference in prediction error with the proposed method.

Lastly, while the present work focuses on the binary outcome along with the Brier score and OMR, the proposed SSL framework can potentially be extended to more general settings with continuous y and/or other accuracy parameters. In particular, for binary y and corresponding classification rule $Y_{2} = I {g (θ^{⊤} x) > c}$ , it would be of interest to consider estimation of the sensitivity, specificity, and weighted OMR with different threshold values to analyze the costs associated with the false positive and false negative errors.

Supplementary Material

supinfo

NIHMS1779685-supplement-supinfo.pdf^{(570.2KB, pdf)}

Acknowledgments

This research was supported by grants F31-GM119263, T32-NS048005, and R01HL089778 from the National Institutes of Health.

Appendix

Here we provide justifications for our main theoretical results. The following lemma confirming the existence and uniqueness of the limiting parameters $\bar{θ}, \bar{γ}$ , and ${\bar{ν}}_{θ}$ will be used in our subsequent derivations.

Lemma A1.

Under conditions 1–3, unique $\bar{θ}$ and $\bar{γ}$ exist. In addition, there exists δ > 0 such that a unique ${\bar{ν}}_{θ}$ exists for any θ satisfying $∥ θ - \bar{θ} ∥_{2} < δ$ .

Proof.

Conditions 3 (A) and (B) imply that there is no θ and γ such that with non-trivial probability,

I (Y_{1} > Y_{2}) = I (θ^{⊤} x_{1} > θ^{⊤} x_{2}),

and

I (Y_{1} > Y_{2}) = I (γ^{⊤} Φ_{1} > γ^{⊤} Φ_{2}) .

It follows directly from Appendix I of Tian et al. [2007] that there exist finite $\bar{θ}$ and $\bar{γ}$ that solve $U (\bar{θ})$ and $Q (\bar{γ})$ , respectively, and that $\bar{θ}$ and $\bar{γ}$ are unique. For $θ \in Θ$ , there exists no ν such that

I (Y_{1} > Y_{2}) = I ({\bar{γ}}^{⊤} Φ_{1} + ν^{⊤} z_{1 θ} > {\bar{γ}}^{⊤} Φ_{2} + ν^{⊤} z_{2 θ}),

I (Y_{1} > Y_{2}) = I ({\bar{γ}}^{⊤} Φ_{1} + ν^{⊤} z_{1 θ} < {\bar{γ}}^{⊤} Φ_{2} + ν^{⊤} z_{2 θ}),

which similarly implies there exists a finite ${\bar{ν}}_{θ}$ that is the solution to $R (ν ∣ θ) = 0$ . The solution is also unique as $E [z_{θ}^{\otimes 2} \dot{g} {{\bar{γ}}^{⊤} Φ + ν^{⊤} z_{θ}}] ≻ 0$ for any $ν$ . □

A. Estimation Procedure for ${\hat{θ}}_{SSL}$

We propose to obtain a simple SS estimator for $\bar{θ}, {\overset{ˇ}{θ}}_{SSL}$ , as the solution to

{\hat{U}}_{N} (θ) = \frac{1}{N} \sum_{i = 1}^{N} x_{i} {g ({\tilde{γ}}^{⊤} Φ_{i}) - g (θ^{⊤} x_{i})} = 0 .

(A.1)

Note that when u includes the stratum information $S$ and the imputation model (2) is correctly specified, there is no need to use the inverse probability weighted (IPW) estimating equation with ${\hat{w}}_{i}$ as in (3). However, under the general scenario where the imputation model may be misspecified, the unweighted estimating equation is not guaranteed to provide an asymptotically unbiased estimate of $\bar{γ}$ and necessitates the use of the IPW approach.

As detailed in the subsequent sections, when the imputation and outcome models are misspecified, the efficiency gain of ${\overset{ˇ}{θ}}_{SSL}$ relative to ${\hat{θ}}_{SL}$ , is not theoretically guaranteed. We therefore obtain the final SS estimator, denoted as ${\hat{θ}}_{SSL} = {({\hat{θ}}_{SSL, 0}, \dots, {\hat{θ}}_{SSL, p})}^{⊤}$ , as a linear combination of ${\hat{θ}}_{SL}$ and ${\overset{ˇ}{θ}}_{SSL}$ to minimize the asymptotic variance. For simplicity we consider here the component-wise optimal combination of the two estimators. That is, the jth component of $\bar{θ}, {\bar{θ}}_{j}$ , is estimated with

{\hat{θ}}_{SSL, j} = {\hat{W}}_{1 j} {\tilde{θ}}_{SSL, j} + (1 - {\hat{W}}_{1 j}) {\tilde{θ}}_{SL, j}

where ${\hat{W}}_{1 j}$ is the first component of the vector ${\hat{W}}_{j} = 1^{⊤} {\hat{Σ}}_{j}^{- 1} / (1^{⊤} {\hat{Σ}}_{j}^{- 1} 1)$ and ${\hat{Σ}}_{j}$ is a consistent estimator for $cov {{({\overset{ˇ}{θ}}_{SSL, j}, {\tilde{θ}}_{SL, j})}^{⊤}}$ . To estimate the variance of ${\hat{θ}}_{SSL}$ , one may rely on estimates of the influence functions of ${\overset{ˇ}{θ}}_{SSL}$ and ${\hat{θ}}_{SL}$ . To avoid under-estimation in a finite sample, we obtain bias-corrected estimates of the influence functions via K-fold CV. Details on the K-fold CV procedure, as well as the computations for the aforementioned estimation of ${\hat{Σ}}_{j}$ and ${\hat{W}}_{1 j}$ , are given in Appendix B.

B. Cross-validation Based Inference for ${\hat{θ}}_{SSL}$

Here we provide the details of the procedure to obtain ${\hat{θ}}_{SSL}$ as well as an estimate of its variance. We employ K-fold CV in the proposed procedure to adjust for overfitting in finite sample and denote each fold of $L$ as $L_{k}$ for $k = 1, \dots, K$ . First, we estimate ${\hat{Σ}}_{j}$ with

n^{- 1} \sum_{k = 1}^{K} \sum_{i \in L_{k}} {W_{j} ({\tilde{γ}}_{(- k)}, D_{i}), V_{j} ({\hat{θ}}_{(- k)}, D_{i})} {W_{j} ({\tilde{γ}}_{(- k)}, D_{i}), V_{j} ({\hat{θ}}_{(- k)}, D_{i})}^{⊤}

where ${\tilde{γ}}_{(- k)}$ is the estimator of $\bar{γ}$ based on $L / L_{k}, {\hat{θ}}_{(- k)}$ is the supervised estimator of $\bar{θ}$ based on $L / L_{k}$ ,

W ({\tilde{γ}}_{(- k)}, D_{i}) = {\hat{A}}^{- 1} [ϖ_{i} x_{i} {y_{i} - g ({\tilde{γ}}_{(- k)}^{⊤} Φ_{i})}], V ({\hat{θ}}_{(- k)}, D_{i}) = {\hat{A}}^{- 1} [ϖ_{i} x_{i} {y_{i} - g ({\hat{θ}}_{(- k)}^{⊤} x_{i})}], \hat{A} = N^{- 1} \sum_{i = 1}^{N} x_{i}^{\otimes 2} \dot{g} ({\overset{ˇ}{θ}}_{SSL}^{⊤} x_{i}) and ϖ_{i} = {\hat{w}}_{i} n / N

In practice, ${\hat{Σ}}_{j}$ may be unstable due to the high correlation of ${\overset{ˇ}{θ}}_{SSL}$ and ${\hat{θ}}_{SL}$ . One may use a regularized estimator ${({\hat{Σ}}_{j} + δ_{n} I)}^{- 1}$ with some $δ_{n} = O (n^{- \frac{1}{2}})$ to stabilize the estimation and obtain ${\hat{θ}}_{SSL}$ accordingly. The covariance of ${\hat{θ}}_{SSL}$ may then be consistently estimated with

n^{- 1} \sum_{k = 1}^{K} \sum_{i \in L_{k}} {Z ({\tilde{γ}}_{(- k)}, {\hat{θ}}_{(- k)}, D_{i})} {Z ({\tilde{γ}}_{(- k)}, {\hat{θ}}_{(- k)}, D_{i})}^{⊤} where

(B.1)

Z ({\tilde{γ}}_{(- k)}, {\hat{θ}}_{(- k)}, D_{i}) = \hat{W} {W ({\tilde{γ}}_{(- k)}, D_{i})} + (I - \hat{W}) {V ({\hat{θ}}_{(- k)}, D_{i})}

(B.2)

$\hat{W} = diag ({\hat{W}}_{10}, \dots, {\hat{W}}_{1 p})$ is an estimate of $W = diag (W_{10}, \dots, W_{1 p}), {\hat{W}}_{1 j}$ is the first component of ${\hat{W}}_{j} = 1^{⊤} {\hat{Σ}}_{j}^{- 1} / (1^{⊤} {\hat{Σ}}_{j}^{- 1} 1)$ and $W_{1 j}$ is the first component of $W_{j} = 1^{⊤} Σ_{j}^{- 1} / (1^{⊤} Σ_{j}^{- 1} 1)$ . The confidence intervals for the regression parameters can be constructed with the proposed variance estimates and the asymptotic normal distribution of the SS estimator.

C. Asymptotic Properties of ${\hat{θ}}_{SL}$

The main complication in deriving the asymptotic properties of the SL estimators arises from the fact that $P (V_{i} = 1 ∣ F) = {\hat{π}}_{S_{i}} \to 0$ as $n \to \infty$ and hence ${\hat{w}}_{i} = V_{i} / {\hat{π}}_{S_{i}}$ is an ill behaved random variable tending to infinity in the limit for those with V_i = 1. This substantially distinguishes the SS setting from the standard missing data literature. To overcome this complication, we note that for subjects in the labeled set, $N^{- 1} {\hat{w}}_{i} = \sum_{s = 1}^{S} I (S_{i} = s) {\hat{ρ}}_{s} n_{s}^{- 1} V_{i}$ and $V_{i} ∣ S_{i} = s$ , are independent identically distributed (i.i.d) random variables since the labeled observations are drawn randomly within each stratum. Also note that ${\hat{ρ}}_{s} \overset{p}{\to} ρ_{s}$ as assumed in Section 2.1. Hence for any function f with $var {f (F_{i}) ∣ S_{i} = s} < \infty$ ,

N^{- 1} \sum_{i = 1}^{N} {\hat{w}}_{i} f (F_{i}) = \sum_{s = 1}^{S} {\hat{ρ}}_{s} n_{s}^{- 1} \sum_{i = 1}^{N} V_{i} I (S_{i} = s) f (F_{i}) = \sum_{s = 1}^{S} {ρ_{s} + o_{p} (1)} {n_{s}^{- 1} \sum_{i = 1}^{N} V_{i} I (S_{i} = s) f (F_{i})} + o_{p} (1)

(C.1)

= \sum_{s = 1}^{S} ρ_{s} E {f (F_{i}) ∣ S_{i} = s} + o_{p} (1) = E {f (F_{i})} + o_{p} (1) .

(C.2)

We begin by verifying that ${\hat{θ}}_{SL}$ is consistent for $\bar{θ}$ . It suffices to show that (i) $\sup_{θ \in Θ} ∥ U_{n} (θ) - U (θ) ∥_{2} = o_{p} (1)$ and (ii) $\inf_{∥ θ - \bar{θ} ∥_{2} > ϵ} {‖ U (θ) ‖}_{2} > 0$ for any $ϵ > 0$ [Newey and McFadden, 1994, Lemma 2.8]. To this end, we write

U_{n} (θ) = \frac{1}{N} \sum_{i = 1}^{N} {\hat{w}}_{i} x_{i} {y_{i} - g (θ^{⊤} x_{i})} = \sum_{s = 1}^{S} ρ_{s} [n_{s}^{- 1} \sum_{V_{i} = 1} I (S_{i} = s) x_{i} {y_{i} - g (θ^{⊤} x_{i})}] + o_{p} (1) .

Under Conditions 1–3, x belongs to a compact set and $\dot{g} (θ^{⊤} x)$ is continuous and uniformly bounded for $θ \in Θ$ . From the uniform law of large numbers (ULLN) [Pollard, 1990, Theorem 8.2], $n_{s}^{- 1} \sum_{V_{i} = 1} I (S_{i} = s) x_{i} {y_{i} - g (θ^{⊤} x_{i})}$ converges to $E [x_{i} {y_{i} - g (θ^{⊤} x_{i})} ∣ S_{i} = s]$ in probability uniformly as $n \to \infty$ and $\sup_{θ \in Θ} {‖ U_{n} (θ) - U (θ) ‖}_{2} = o_{p} (1)$ . Furthermore, (ii) follows directly from Lemma A1 and consequently ${\hat{θ}}_{SL} \overset{p}{\to} \bar{θ}$ as $n \to \infty$ .

Next we consider the asymptotic normality of ${\hat{W}}_{SL} = n^{\frac{1}{2}} ({\hat{θ}}_{SL} - \bar{θ})$ . Noting that ${\hat{θ}}_{SL} \overset{p}{\to} \bar{θ}$ and ${\hat{ρ}}_{s} \overset{p}{\to} ρ_{s}$ , we apply Theorem 5.21 of Van der Vaart [2000] to obtain the Taylor expansion

{\hat{W}}_{SL} = n^{\frac{1}{2}} ({\hat{θ}}_{SL} - \bar{θ}) = n^{\frac{1}{2}} \frac{1}{N} \sum_{i = 1}^{N} {\hat{w}}_{i} A^{- 1} x_{i} {y_{i} - g ({\bar{θ}}^{⊤} x_{i})} + o_{p} (1) = n^{\frac{1}{2}} \sum_{s = 1}^{S} {ρ_{s} + o_{p} (1)} {n_{s}^{- 1} \sum_{V_{i} = 1} I (S_{i} = s) e_{LL i}} + o_{p} (1),

where $e_{SL i} = A^{- 1} x_{i} {y_{i} - g ({\bar{θ}}^{⊤} x_{i})}$ and $A = E {x_{i}^{\otimes 2} \dot{g} ({\bar{θ}}^{⊤} x_{i})}$ . It then follows by the classical Central Limit Theorem that ${\hat{W}}_{SL} \to N (0, Σ_{SL})$ in distribution and

{\hat{W}}_{SL} = n^{\frac{1}{2}} \sum_{s = 1}^{S} ρ_{s} {n_{s}^{- 1} \sum_{V_{i} = 1} I (S_{i} = s) e_{SL i}} + o_{p} (1),

where $Σ_{SL} = \sum_{s = 1}^{S} ρ_{s}^{2} ρ_{1 s}^{- 1} E {e_{SL i}^{\otimes 2} ∣ S_{i} = s}$ .

D. Asymptotic Properties of ${\hat{D}}_{SL} ({\hat{θ}}_{SL})$

We begin by showing that ${\hat{D}}_{SL} ({\hat{θ}}_{SL}) \overset{p}{\to} D (\bar{θ})$ as $n \to \infty$ . We first note that since ${\hat{ρ}}_{s} \overset{p}{\to} ρ_{s}$ ,

{\hat{D}}_{SL} (θ) = \sum_{s = 1}^{S} ρ_{s} [n_{s}^{- 1} \sum_{V_{i} = 1} I (S_{i} = s) d {y_{i}, Y (θ^{⊤} x_{i})}] + o_{p} (1) .

(D.1)

It follows by the ULLN that $\sup_{θ \in Θ} | {\hat{D}}_{SL} (θ) - D (θ) | = o_{p} (1)$ since $d {y, Y (θ^{⊤} x}$ is continuously differentiable in θ and uniformly bounded. The consistency of ${\hat{D}}_{SL} ({\hat{θ}}_{SL})$ for $D (\bar{θ})$ then follows from the fact that ${\hat{θ}}_{SL}$ converges in probability to $\bar{θ}$ as $n \to \infty$ .

To establish the asymptotic distribution of ${\hat{T}}_{SL} = n^{\frac{1}{2}} {{\hat{D}}_{SL} ({\hat{θ}}_{SL}) - D (\bar{θ})}$ , we first consider ${\tilde{T}}_{SL} (θ) = n^{\frac{1}{2}} {{\hat{D}}_{SL} (θ) - D (θ)}$ . We verify that there exists δ > 0, such that the classes of functions indexed by θ:

B_{1} = {I (S = s) | y - I {g (θ^{⊤} x) > c} | : ∥ θ - \bar{θ} ∥_{2} < δ}

and B_{2} = {I (S = s) {[y - g (θ^{⊤} x)]}^{2} : ∥ θ - \bar{θ} ∥_{2} < δ}

are Donsker classes. For $B_{1}$ , we note that ${I {g (θ^{⊤} x) > c} : ∥ θ - \bar{θ} ∥_{2} < δ}$ is a Vapnik-Chervonenkis class [Van der Vaart, 2000, Page 275] and thus

B_{1} = {I (s = s_{0}) [I (y = 0) I {g (θ^{⊤} x) > c} + I (y = 1) I {g (θ^{⊤} x) ⩽ c}] : ∥ θ - \bar{θ} ∥_{2} < δ}

is a Donsker class. For $B_{2}, {[y - g (θ^{⊤} x)]}^{2}$ is continuously differentiable in θ and uniformly bounded by a constant. It follows that $B_{2}$ is a Donsker class [Van der Vaart, 2000, Example 19.7]. By Theorem 19.5 of Van der Vaart [2000] we then have

n^{\frac{1}{2}} {{\hat{D}}_{SL} (θ) - D (θ)} = \sum_{s = 1}^{S} ρ_{s} ρ_{1 s}^{- \frac{1}{2}} n_{s}^{- \frac{1}{2}} \sum_{V_{i} = 1} I (S_{i} = s) [d {y_{i}, Y (θ^{⊤} x_{i})} - D (θ)] + o_{p} (1),

which converges weakly to a mean zero Gaussian process index by θ. Thus, $n^{\frac{1}{2}} {{\hat{D}}_{SL} (θ) - D (θ)}$ is stochastically equicontinuous at $\bar{θ}$ . In addition, note that D(θ) is continuously differentiable at $\bar{θ}$ and ${\hat{ρ}}_{s} \overset{p}{\to} ρ_{s}$ . It then follows that

{\hat{T}}_{SL} = n^{\frac{1}{2}} {{\hat{D}}_{SL} ({\hat{θ}}_{SL}) - D ({\hat{θ}}_{SL})} + n^{\frac{1}{2}} {D ({\hat{θ}}_{SL}) - D (\bar{θ})} = n^{\frac{1}{2}} \sum_{s = 1}^{S} ρ_{s} (n_{s}^{- 1} \sum_{V_{i} = 1} I (S_{i} = s) [d (y_{i}, {\bar{Y}}_{i}) - D (\bar{θ}) + \dot{D} {(\bar{θ})}^{⊤} e_{SL i}]) + o_{p} (1),

which converges in distribution to $N (0, σ_{SL}^{2})$ where

σ_{SL}^{2} = \sum_{s = 1}^{S} ρ_{s}^{2} ρ_{1 s}^{- 1} E [{d (y_{i}, {\bar{Y}}_{i}) - D (\bar{θ}) + \dot{D} {(\bar{θ})}^{⊤} e_{SL i}}^{2} ∣ S_{i} = s] .

E. Asymptotic Properties of ${\overset{ˇ}{θ}}_{SSL}$

We first consider the asymptotic properties of $\tilde{γ}$ . Under Conditions 1–3 and using that $λ_{n} = o (n^{- \frac{1}{2}})$ and ${\hat{ρ}}_{s} \overset{p}{\to} ρ_{s}$ , we can adapt the same procedure as in Appendix C to show that $\tilde{γ} \overset{p}{\to} \bar{γ}$ as $n \to \infty$ . We obtain the following Taylor series expansion

n^{\frac{1}{2}} (\tilde{γ} - \bar{γ}) = n^{\frac{1}{2}} \sum_{s = 1}^{S} ρ_{s} [n_{s}^{- 1} \sum_{i = 1}^{n} I (S_{i} = s) C^{- 1} Φ_{i} {y_{i} - g ({\bar{γ}}^{⊤} Φ_{i})}] + o_{p} (1)

where $C = E {Φ_{i}^{\otimes 2} \dot{g} ({\bar{γ}}^{⊤} Φ_{i})}$ . We then have that $n^{\frac{1}{2}} (\tilde{γ} - \bar{γ})$ converges to zero-mean Gaussian distribution. To verify that ${\overset{ˇ}{θ}}_{SSL}$ is consistent for $\bar{θ}$ , it suffices to show that (i) $\sup_{θ \in Θ} ∥ {\hat{U}}_{N} (θ) - U^{0} (θ) ∥_{2} = o_{p} (1)$ and (ii) $\inf_{∥ θ - \bar{θ} ∥_{2} ⩾ ϵ ∣} ∥ U (θ) ∥_{2} > 0$ , for any ϵ > 0 [Newey and McFadden, 1994, Lemma 2.8], where

U^{0} (θ) = E [x_{i} {g ({\bar{γ}}^{⊤} Φ_{i}) - g (θ^{⊤} x_{i})}] .

For (i), we first note that since $\tilde{γ} \overset{p}{\to} \bar{γ}, \dot{g} (\cdot)$ is continuous and Φ is bounded,

{\hat{U}}_{N} (θ) = N^{- 1} \sum_{i = 1}^{N} x_{i} {g ({\tilde{γ}}^{⊤} Φ_{i}) - g (θ^{⊤} x_{i})} = N^{- 1} \sum_{i = 1}^{N} x_{i} {g ({\bar{γ}}^{⊤} Φ_{i}) - g (θ^{⊤} x_{i})} + o_{p} (1) .

(E.1)

Note that $g ({\bar{γ}}^{⊤} Φ) - g (θ^{⊤} x)$ is bounded and continuously differentiable in θ. We then apply the ULLN to have that $\sup_{θ \in Θ} {‖ {\hat{U}}_{N} (θ) - U^{0} (θ) ‖}_{2} = o_{p} (1)$ . For (ii), we note

U^{0} (θ) = - E [x_{i} {y_{i} - g ({\bar{γ}}^{⊤} Φ_{i})}] + E [x_{i} {y_{i} - g (θ^{⊤} x_{i})}] = E [x_{i} {y_{i} - g (θ^{⊤} x_{i})}] .

Therefore, (ii) holds by Lemma A1 and ${\overset{ˇ}{θ}}_{SSL}$ is consistent for $\bar{θ}$ .

Now we consider the weak convergence of $n^{\frac{1}{2}} ({\overset{ˇ}{θ}}_{SSL} - \bar{θ})$ . Under Conditions 1–3, we have the Taylor expansion

n^{\frac{1}{2}} ({\overset{ˇ}{θ}}_{SSL} - \bar{θ}) = n^{\frac{1}{2}} A^{- 1} [N^{- 1} \sum_{i = 1}^{N} x_{i} {g ({\bar{γ}}^{⊤} Φ_{i}) - g ({\bar{θ}}^{⊤} x_{i})} + B (\tilde{γ} - \bar{γ})] + o_{p} (1),

where $B = E {x_{i} Φ_{i}^{⊤} \dot{g} ({\bar{γ}}^{⊤} Φ_{i})}$ . This expansion coupled with the fact that $N^{- 1} \sum_{i = 1}^{N} X_{i} {g ({\bar{γ}}^{⊤} Φ_{i}) - g ({\bar{θ}}^{⊤} x_{i})} = O_{p} (N^{- \frac{1}{2}}) = o_{p} (n^{- \frac{1}{2}})$ imply that

n^{\frac{1}{2}} ({\overset{ˇ}{θ}}_{SSL} - \bar{θ}) = n^{\frac{1}{2}} \sum_{s = 1}^{S} ρ_{s} [n_{s}^{- 1} \sum_{i = 1}^{n} A^{- 1} {BC}^{- 1} I (S_{i} = s) Φ_{i} {y_{i} - g ({\bar{γ}}^{⊤} Φ_{i})}] + o_{p} (1) .

Letting $x_{i} = {(x_{i 1}, \dots, x_{i p})}^{⊤}$ , we note that for $j = 1, \dots, p$ ,

{[{BC}^{- 1}]}_{j} = E {x_{i j} Φ_{i}^{⊤} \dot{g} ({\bar{γ}}^{⊤} Φ_{i})} E {[{Φ_{i}^{\otimes 2} \dot{g} ({\bar{γ}}^{⊤} Φ_{i})}]}^{- 1} = \underset{β}{argmin} E {\dot{g} ({\bar{γ}}^{⊤} Φ_{i}) {(x_{i j} - β^{⊤} Φ_{i})}^{2}} .

Since x_ij is a component of the vector $Φ (u)$ , the minimizer β can be chosen such that $x_{i j} - β^{⊤} Φ_{i} = 0$ for $i = 1, \dots, N$ which implies that $x_{i} = B C^{- 1} Φ_{i}$ . Thus

n^{\frac{1}{2}} ({\overset{ˇ}{θ}}_{SSL} - \bar{θ}) = n^{\frac{1}{2}} \sum_{s = 1}^{S} ρ_{s} {n_{s}^{- 1} \sum_{i = 1}^{n} I (S_{i} = s) e_{SSL i}} + o_{p} (1)

where $e_{SSL i} = A^{- 1} x_{i} {y_{i} - g ({\bar{γ}}^{⊤} Φ_{i})}$ . It then follows from the classical Central Limit Theorem that $n^{\frac{1}{2}} ({\overset{ˇ}{θ}}_{SSL} - \bar{θ}) \to N (0, Σ_{SSL})$ in distribution where

Σ_{SSL} = \sum_{s = 1}^{S} ρ_{s}^{2} ρ_{1 s}^{- 1} E {e_{SSL i}^{\otimes 2} ∣ S_{i} = s} .

We then see that

Σ_{SL} - Σ_{SSL} = \sum_{s = 1}^{S} ρ_{s}^{2} ρ_{1 s}^{- 1} [E {e_{SL i}^{\otimes 2} ∣ S_{i} = s} - E {e_{SSL i}^{\otimes 2} ∣ S_{i} = s}] = \sum_{s = 1}^{S} ρ_{s}^{2} ρ_{1 s}^{- 1} A^{- 1} E [x_{i}^{\otimes 2} {g ({\bar{γ}}^{⊤} Φ_{i}) - g ({\bar{θ}}^{⊤} x_{i})}^{2} + 2 x_{i}^{\otimes 2} {y_{i} - g ({\bar{γ}}^{⊤} Φ_{i})} {g ({\bar{γ}}^{⊤} Φ_{i}) - g ({\bar{θ}}^{⊤} x_{i})} ∣ S_{i} = s] .

Therefore, when the imputation model is correctly specified, it follows that

Σ_{SL} - Σ_{SSL} = \sum_{s = 1}^{S} ρ_{s}^{2} ρ_{1 s}^{- 1} E [A^{- 1} x_{i}^{\otimes 2} {g ({\bar{γ}}^{⊤} Φ_{i}) - g ({\bar{θ}}^{⊤} x_{i})}^{2} ∣ S_{i} = s] ⪰ 0 .

F. Asymptotic Properties of ${\hat{D}}_{SSL} ({\hat{θ}}_{SSL})$ and ${\hat{D}}_{sSL} ({\tilde{θ}}_{sSL})$

First note that by Lemma A1, there exists δ > 0 such that for all θ satisfying $∥ θ - \bar{θ} ∥_{2} < δ$ , so that ${\bar{ν}}_{θ}$ is unique. Then, similar to the derivations in Appendices C and E, we may show that ${\tilde{ν}}_{θ}$ is consistent for ${\bar{ν}}_{θ}$ and $n^{\frac{1}{2}} ({\tilde{ν}}_{θ} - {\bar{ν}}_{θ})$ is asymptotically Gaussian with mean zero.

Let $Θ_{δ} = Θ \cap {θ : ∥ θ - \bar{θ} ∥_{2} < δ}$ . For the consistency of ${\hat{D}}_{SSL} ({\hat{θ}}_{SSL})$ for $D (\bar{θ})$ , we note that the uniform consistency of ${\tilde{ν}}_{θ}$ for ${\bar{ν}}_{θ}$ and $\tilde{γ}$ for $\bar{γ}$ , together with the ULLN and under regularity Conditions 1–3 and ${\hat{ρ}}_{s} \overset{p}{\to} ρ_{s}$ , imply $\sup_{θ \in Θ_{δ}} | {\hat{D}}_{SSL} (θ) - D (θ) | = o_{p} (1)$ . It then follows from the consistency of ${\hat{θ}}_{SSL}$ and ${\overset{ˇ}{θ}}_{SSL}$ for $\bar{θ}$ that ${\hat{D}}_{SSL} ({\hat{θ}}_{SSL}) \overset{p}{\to} D (\bar{θ})$ and ${\hat{D}}_{SSL} ({\overset{ˇ}{θ}}_{SSL}) \overset{p}{\to} D (\bar{θ})$ as $n \to \infty$ .

To derive the asymptotic distribution for ${\hat{T}}_{SSL} = n^{\frac{1}{2}} {{\hat{D}}_{SSL} ({\hat{θ}}_{SSL}) - D (\bar{θ})}$ and ${\overset{˘}{T}}_{SSL} = n^{\frac{1}{2}} {{\hat{D}}_{SSL} ({\overset{ˇ}{θ}}_{SSL}) - D (\bar{θ})}$ , we first consider

{\tilde{T}}_{SSL} (θ) = n^{\frac{1}{2}} {{\hat{D}}_{SSL} (θ) - D (θ)} = n^{\frac{1}{2}} [N^{- 1} \sum_{i = 1}^{N} d {g ({\tilde{γ}}^{⊤} Φ_{i} + {\tilde{ν}}_{θ}), Y (θ^{⊤} x_{i})} - D (θ)]

Under Conditions 1–3 and by Taylor series expansion about $\bar{γ}$ and $\bar{ν}$ and the ULLN,

{\tilde{T}}_{SSL} (θ) = n^{\frac{1}{2}} [N^{- 1} \sum_{i = 1}^{N} d {g ({\bar{γ}}^{⊤} Φ_{i} + {\bar{ν}}_{θ}^{⊤} z_{i θ}), Y (θ^{⊤} x_{i})} - D (θ) + G_{θ} (\tilde{γ} - \bar{γ}) + H_{θ} ({\tilde{ν}}_{θ} - {\bar{ν}}_{θ})]

where $G_{θ} = E [Φ_{i}^{⊤} \dot{g} ({\bar{γ}}^{⊤} Φ_{i} + {\bar{ν}}_{θ}^{⊤} z_{i θ}) {1 - 2 Y (θ^{⊤} x_{i})}]$ and $H_{θ} = E [z_{i θ}^{⊤} \dot{g} ({\bar{γ}}^{⊤} Φ_{i} + {\bar{ν}}_{θ}^{⊤} z_{i θ}) {1 - 2 Y (θ^{⊤} x_{i})}]$ . From the previous section we have

n^{\frac{1}{2}} (\tilde{γ} - \bar{γ}) = n^{\frac{1}{2}} \sum_{s = 1}^{S} ρ_{s} [n_{s}^{- 1} \sum_{i = 1}^{n} C^{- 1} I (S_{i} = s) Φ_{i} {y_{i} - g ({\bar{γ}}^{⊤} Φ_{i})}] + o_{p} (1)

Similar arguments can be used to verify that

n^{\frac{1}{2}} {({\tilde{ν}}_{θ} - {\bar{ν}}_{θ})}^{⊤} = n^{\frac{1}{2}} J_{θ}^{- 1} [\sum_{s = 1}^{S} ρ_{s} n_{s}^{- 1} \sum_{i = 1}^{n} I (S_{i} = s) z_{i θ} {y_{i} - g ({\bar{γ}}^{⊤} Φ_{i} + {\bar{ν}}^{⊤} z_{i θ})} + K {(\tilde{γ} - \bar{γ})}^{⊤}] + o_{p} (1)

where $J_{θ} = E {z_{i θ}^{\otimes 2} \dot{g} ({\bar{γ}}^{⊤} Φ_{i} + {\bar{ν}}^{⊤} z_{i θ})}$ and $K_{θ} = - E {z_{i θ} Φ_{i}^{⊤} \dot{g} ({\bar{γ}}^{⊤} Φ_{i} + {\bar{ν}}^{⊤} z_{i θ})}$ . These results, together with the fact that $N^{- \frac{1}{2}} \sum_{i = 1}^{N} d {g ({\bar{γ}}^{⊤} Φ_{i} + {\bar{ν}}_{θ}^{⊤} z_{i θ}), Y (θ^{⊤} x_{i})} - D (θ))$ converges weakly to zero-mean Gaussian process in θ, imply that

{\tilde{T}}_{SSL} (θ) = n^{\frac{1}{2}} \sum_{s = 1}^{S} ρ_{s} (n_{s}^{- 1} \sum_{i = 1}^{n} I (S_{i} = s) [(G_{θ} + H_{θ} J_{θ}^{- 1} K_{θ}) C^{- 1} Φ_{i} {y_{i} - g ({\bar{γ}}^{⊤} Φ_{i})} + H_{θ} J_{θ}^{- 1} z_{i θ} {y_{i} - g ({\bar{γ}}^{⊤} Φ_{i} + {\bar{ν}}_{θ}^{⊤} z_{i θ})}]) + o_{p} (1) .

We may simplify the above expression by noting that ${1 - 2 Y (θ^{⊤} x_{i})}$ is a linear combination of z_iθ and hence $[H_{θ} J_{θ}^{- 1}] z_{i θ} = {1 - 2 Y (θ^{⊤} x_{i})}$ . Additionally, $H_{θ} J_{θ}^{- 1} K_{θ} = - G_{θ}$ which implies that $(G_{θ} + H_{θ} J_{θ}^{- 1} K_{θ}) C^{- 1} = 0$ . Thus,

{\tilde{T}}_{SSL} (θ) = n^{\frac{1}{2}} \sum_{s = 1}^{S} ρ_{s} [n_{s}^{- 1} \sum_{i = 1}^{n} I (S_{i} = s) {1 - 2 Y (θ^{⊤} x_{i})} {y_{i} - g ({\bar{γ}}^{⊤} Φ_{i} + {\bar{ν}}^{⊤} z_{i θ})}] .

This combined with the fact that D(θ) is continuously differentiable at $\bar{θ}, \hat{W}$ is consistent for its limiting value W introduced in Appendix B, and Conditions 1–3 then give that

{\hat{T}}_{SSL} = n^{\frac{1}{2}} {{\hat{D}}_{SSL} ({\hat{θ}}_{SSL}) - D ({\hat{θ}}_{SSL})} + n^{\frac{1}{2}} {D ({\hat{θ}}_{SSL}) - D (\bar{θ})}

= n^{\frac{1}{2}} \sum_{s = 1}^{S} ρ_{s} (n_{s}^{- 1} \sum_{i = 1}^{n} I (S_{i} = s) [{1 - 2 {\bar{Y}}_{i}} {y_{i} - g ({\bar{γ}}^{⊤} Φ_{i} + {\bar{ν}}_{\bar{θ}}^{⊤} z_{i \bar{θ}})} + \dot{D} {(\bar{θ})}^{⊤} {W e_{SSL i} + (I - W) e_{SL i}}]) + o_{p} (1)

= n^{\frac{1}{2}} \sum_{s = 1}^{S} ρ_{s} (n_{s}^{- 1} \sum_{i = 1}^{n} I (S_{i} = s) [{d (y_{i}, {\bar{Y}}_{i}) - d (m_{II, i}, {\bar{Y}}_{i})} + \dot{D} {(\bar{θ})}^{⊤} {W e_{SSL i} + (I - W) e_{SL i}}]) + o_{p} (1) .

Note that the existence of $\dot{D} (\bar{θ})$ is implied by Condition 1, namely, that the density function of ${\bar{θ}}^{⊤} x$ is continuously differentiable in ${\bar{θ}}^{⊤} x$ and $P (y = 1 ∣ u)$ is continuously differentiable in the continuous components of u. We then have that $n^{\frac{1}{2}} {{\hat{D}}_{SSL} ({\hat{θ}}_{SSL}) - D (\bar{θ})}$ converges to a zero mean normal random variable by the classical Central Limit Theorem. Using similar arguments as those for ${\overset{˘}{T}}_{SL}$ , we have that ${\overset{ˇ}{T}}_{SSL} = n^{\frac{1}{2}} {{\hat{D}}_{SSL} ({\overset{ˇ}{θ}}_{SSL}) - D (\bar{θ})}$ can be expanded as

n^{\frac{1}{2}} \sum_{s = 1}^{S} ρ_{s} (n_{s}^{- 1} \sum_{i = 1}^{n} I (S_{i} = s) [{d (y_{i}, {\bar{Y}}_{i}) - d (m_{II, i}, {\bar{Y}}_{i})} + \dot{D} {(\bar{θ})}^{⊤} e_{ssL i}]),

which also converges to a zero mean normal random variable.

Comparing the asymptotic variance of ${\overset{ˇ}{T}}_{SSL}$ with ${\hat{T}}_{SL}$ , first note that

{\hat{T}}_{sL} = n^{\frac{1}{2}} \sum_{s = 1}^{S} ρ_{s} (n_{s}^{- 1} \sum_{i = 1}^{n} I (S_{i} = s) [d (y_{i}, {\bar{Y}}_{i}) - D (\bar{θ}) + \dot{D} {(\bar{θ})}^{⊤} e_{SL i}]) + o_{p} (1)

= n^{\frac{1}{2}} \sum_{s = 1}^{S} ρ_{s} (n_{s}^{- 1} \sum_{i = 1}^{n} I (S_{i} = s) [d (y_{i}, {\bar{Y}}_{i}) - d (m_{II, i}, {\bar{Y}}_{i}) + d (m_{II, i}, {\bar{Y}}_{i}) - D (\bar{θ}) + \dot{D} {(\bar{θ})}^{⊤} A^{- 1} x_{i} {y_{i} - g ({\bar{θ}}^{⊤} x_{i})}]) + o_{p} (1)

= n^{\frac{1}{2}} \sum_{s = 1}^{S} ρ_{s} (n_{s}^{- 1} \sum_{i = 1}^{n} I (S_{i} = s) [(1 - 2 {\bar{Y}}_{i}) {y_{i} - g ({\bar{γ}}^{⊤} Φ_{i} + {\bar{ν}}_{\bar{θ}}^{⊤} z_{i \bar{θ}})} + d (m_{II, i}, {\bar{Y}}_{i}) - D (\bar{θ}) + \dot{D} {(\bar{θ})}^{⊤} A^{- 1} x_{i} {y_{i} - g ({\bar{γ}}^{⊤} Φ_{i}) + g ({\bar{γ}}^{⊤} Φ_{i}) - g ({\bar{θ}}^{⊤} x_{i})}]) + o_{p} (1) .

Letting $h_{1} (Φ_{i}) = 1 - 2 {\bar{Y}}_{i} + \dot{D} {(\bar{θ})}^{⊤} A^{- 1} x_{i}$ and

h_{2} (Φ_{i}) = d (m_{II, i}, {\bar{Y}}_{i}) - D (\bar{θ}) + \dot{D} {(\bar{θ})}^{⊤} A^{- 1} x_{i} {g ({\bar{γ}}^{⊤} Φ_{i}) - g ({\bar{θ}}^{⊤} x_{i})} .

we note that $h_{1} (Φ_{i})$ and $h_{2} (Φ_{i})$ are functions of Φ_i and do not depend on y_i. Thus, when $P (y = 1 ∣ u) = g ({\bar{γ}}^{⊤} Φ)$ , we have $\bar{ν} = 0$ and

σ_{SL}^{2} = \sum_{s = 1}^{S} ρ_{s}^{2} ρ_{1 s}^{- 1} E [h_{1}^{2} (Φ_{i}) {y_{i} - g ({\bar{γ}}^{⊤} Φ_{i})}^{2} + 2 h_{1} (Φ_{i}) h_{2} (Φ_{i}) {y_{i} - g ({\bar{γ}}^{⊤} Φ_{i})} + h_{2}^{2} (Φ_{i}) ∣ S_{i} = s] = \sum_{s = 1}^{S} ρ_{s}^{2} ρ_{1 s}^{- 1} E [h_{1} {(Φ_{i})}^{2} {y_{i} - g ({\bar{γ}}^{⊤} Φ_{i})}^{2} + h_{2} {(Φ_{i})}^{2} ∣ S_{i} = s]

while the asymptotic variance of ${\overset{ˇ}{T}}_{SSL}$ is

σ_{SSL}^{2} = \sum_{s = 1}^{S} ρ_{s}^{2} ρ_{1 s}^{- 1} E [h_{1}^{2} (Φ_{i}) {y_{i} - g ({\bar{γ}}^{⊤} Φ_{i})}^{2} ∣ S_{i} = s] .

Therefore, when $P (y = 1 ∣ u) = g ({\bar{γ}}^{⊤} Φ)$ , it follows that $Δ_{a V a r} : = σ_{SL}^{2} - σ_{SSL}^{2} > 0$ . Additionally, when model (1) is correct and $P (y = 1 ∣ u) = g ({\bar{γ}}^{⊤} Φ) = g ({\bar{θ}}^{⊤} x)$ , we have $h_{2} (Φ_{i}) = d (m_{II, i}, {\bar{Y}}_{i}) - D (\bar{θ})$ which is not equal to 0 with probability 1, so that again $Δ_{a V a r} > 0$ .

G. Intrinsic Efficient Estimation

G.1. Intrinsic Efficient Estimator for $\bar{D}$

We first introduce the intrinsic efficient estimator of the accuracy measures. Without loss of generality, we set the imputation basis for both θ and D(θ) as $Ψ_{θ i} = {[Φ_{i}^{⊤}, Y (θ^{⊤} x_{i})]}^{⊤}$ , where θ is plugged in with some preliminary estimator for θ, denoted as $\tilde{θ}$ . In practice, one may take $\tilde{θ}$ as either the simple SL estimator or the SSL estimator obtained following SectionT 3.1. We include $Y (θ^{⊤} x_{i})$ in the imputation basis to simplify the notation and presentation in this section. Although this distinguishes the following discussion from the proposal in Section 3.2, it is straightforward to extend our results to the original proposal.

Recall that for the original SSL estimator of the regression parameter, one first obtains ${\tilde{γ}}_{\hat{θ}}$ as the solution to

N^{- 1} \sum_{i = 1}^{N} {\hat{w}}_{i} Ψ_{{\tilde{θ}}_{i}} {y_{i} - g (γ^{⊤} Ψ_{{\tilde{θ}}_{i}})} - λ_{n} γ = 0

and then solves $N^{- 1} \sum_{i = 1}^{N} x_{i} {g ({\tilde{γ}}_{\tilde{θ}}^{⊤} Ψ_{{\tilde{θ}}_{i}}) - g (θ^{⊤} x_{i})} = 0$ to obtain the estimator of $\bar{θ}$ . Despite the change in basis, we still denote this estimator as ${\hat{θ}}_{SSL}$ with a slight abuse in the notation. Adapting the augmentation procedure in Section 3.2, we then find ${\tilde{γ}}_{{\hat{θ}}_{SSL}}$ as the solution to

N^{- 1} \sum_{i = 1}^{N} {\hat{w}}_{i} Ψ_{{\hat{θ}}_{SSL} i} {y_{i} - g (γ^{⊤} Ψ_{{\hat{θ}}_{SSL} i})} - λ_{n} γ = 0,

and estimate $\bar{D}$ with ${\hat{D}}_{SSL} = {\hat{D}}_{SSL} ({\hat{θ}}_{SSL})$ where ${\hat{D}}_{SSL} (θ) = N^{- 1} \sum_{i = 1}^{N} d {g ({\tilde{γ}}_{θ}^{⊤} Ψ_{θ i}), Y (θ^{⊤} x_{i})}$ . Extending Theorem 2, the asymptotic variance of $n^{\frac{1}{2}} {{\hat{D}}_{SSL} ({\hat{θ}}_{SSL}) - \bar{D}}$ may be expressed as

\frac{1}{n} \sum_{i = 1}^{n} E ζ_{i} {1 - 2 {\bar{Y}}_{i} + \dot{D} {(\bar{θ})}^{⊤} A^{- 1} x_{i}}^{2} {y_{i} - g ({\bar{γ}}_{\bar{θ}}^{⊤} {\bar{Ψ}}_{i})}^{2},

(G.1)

where ${\bar{γ}}_{\bar{θ}}$ represents the limits of ${\tilde{γ}}_{{\hat{θ}}_{SSL}} (or {\tilde{γ}}_{\tilde{θ}})$ , and ${\bar{Ψ}}_{i} = {\bar{Ψ}}_{{\bar{θ}}_{i}} = {[Φ_{i}^{⊤}, Y ({\bar{θ}}^{⊤} x_{i})]}^{⊤}$ . Analogous to the construction of $e^{⊤} {\hat{θ}}_{intri}$ , we consider minimizing the asymptotic variance given by (G.1) to estimate $\bar{D}$ . Specifically, we first solve for ${\tilde{γ}}_{\tilde{θ}}^{(2)}$ with

\underset{γ}{argmin} \frac{1}{2 n} \sum_{i = 1}^{n} {\hat{ζ}}_{i} {1 - 2 Y ({\tilde{θ}}^{⊤} x_{i}) + {\hat{\dot{D}}}^{⊤} {\hat{A}}^{- 1} x_{i}}^{2} {y_{i} - g (γ^{⊤} Ψ_{{\tilde{θ}}_{i}})}^{2} + λ_{n}^{(2)} ∥ γ ∥_{2}^{2}, s.t. \frac{1}{N} \sum_{i = 1}^{N} {\hat{w}}_{i} {[x_{i}^{⊤}, Y ({\tilde{θ}}^{⊤} x_{i})]}^{⊤} {y_{i} - g (γ^{⊤} Ψ_{{\tilde{θ}}_{i}})} = 0,

(G.2)

where $\hat{\dot{D}}$ is an estimation of $\dot{D} (\bar{θ})$ and the tuning parameter $λ^{(2)} = o (n^{- \frac{1}{2}})$ . Similar to (3) and (9), moment constraints in (G.2) calibrate potential bias of the estimators for $\bar{θ}$ and $\bar{D} (θ)$ .

Next, we present the construction of $\hat{\dot{D}}$ for the Brier score and OMR separately. For the Brier score, ${\bar{D}}_{1}$ , we take

\hat{\dot{D}} = {\hat{\dot{D}}}_{1} = \frac{1}{N} \sum_{i = 1}^{N} - 2 {\hat{w}}_{i} \dot{g} ({\tilde{θ}}^{⊤} x_{i}) {y_{i} - g ({\tilde{θ}}^{⊤} x_{i})} x_{i} .

For the the OMR, ${\bar{D}}_{2}$ , recall that a simple estimator is given by the empirical average

\frac{1}{N} \sum_{i = 1}^{N} {\hat{w}}_{i} {y_{i} + (1 - 2 y_{i}) I (g ({\tilde{θ}}^{⊤} x_{i}) > c)} .

Since $I (g (θ^{⊤} x_{i}) > c)$ is not a differentiable function of θ, we first smooth each $I (g ({\tilde{θ}}^{⊤} x_{i}) > c)$ as $\int_{c}^{+ \infty} K_{h} {g ({\tilde{θ}}^{⊤} x_{i}) - u} d u$ , where K(·) represents the Gaussian kernel function, and $K_{h} (a) : = h^{- 1} K (a / h)$ with some bandwidth h > 0. Then, $\dot{D} (\bar{θ})$ is estimated with

\hat{\dot{D}} = {\hat{\dot{D}}}_{2} = \frac{1}{N} \sum_{i = 1}^{N} {\hat{w}}_{i} (1 - 2 y_{i}) \dot{g} ({\tilde{θ}}^{⊤} x_{i}) K_{h} {g ({\tilde{θ}}^{⊤} x_{i}) - c} x_{i} .

With ${\tilde{γ}}_{\tilde{θ}}^{(2)}$ , we then solve

N^{- 1} \sum_{i = 1}^{N} x_{i} {g ({\tilde{γ}}_{\tilde{θ}}^{(2) ⊤} Ψ_{{\tilde{θ}}_{i}}) - g (θ^{⊤} x_{i})} = 0

to obtain ${\hat{θ}}_{intri}^{D}$ for estimation of $\bar{D}$ and employ the augmentation procedure in Section 3.2.

That is, we solve ${\tilde{γ}}_{{\hat{θ}}_{intri}^{D}}^{(2)}$ from

N^{- 1} \sum_{i = 1}^{N} {\hat{w}}_{i} Ψ_{{\hat{θ}}_{intri}^{D}} {y_{i} - g (γ^{⊤} Ψ_{{\hat{θ}}_{intri}^{D} i})} - λ_{n}^{(2)} γ = 0,

and estimate $\bar{D}$ by ${\hat{D}}_{intri} = {\hat{D}}_{intri} ({\hat{θ}}_{intri}^{D})$ where ${\hat{D}}_{intri} (θ) = N^{- 1} \sum_{i = 1}^{N} d {g ({\tilde{γ}}_{θ}^{(2)} Ψ_{θ i}), Y (θ^{⊤} x_{i})}$ .

To present the asymptotic properties of ${\hat{D}}_{intri}$ , we define

\begin{matrix} {\bar{γ}}_{\bar{θ}}^{(2)} = \underset{γ}{argmin} E [R {1 - 2 \bar{Y} + \dot{D} {(\bar{θ})}^{⊤} A^{- 1} x}^{2} {y - g (γ_{\bar{θ}}^{⊤} \bar{Ψ})}^{2}], \\ s.t. E {[x^{⊤}, Y ({\bar{θ}}^{⊤} x)]}^{⊤} {y - g (γ_{\bar{θ}}^{⊤} \bar{Ψ})} = 0. \end{matrix}

Theorem A1 provides the asymptotic expansion of ${\hat{D}}_{intri}$ and its proof, together with the proof of Theorem 3 from the main text, is detailed in Appendix G.2.

Theorem A1.

Under Conditions 1, Conditions A1 and A2 from Appendix G.2, and with the bandwidth $h ≍ n^{- \frac{1}{4}}$ , $n^{\frac{1}{2}} ({\hat{D}}_{intri} - \bar{D})$ weakly converges to a Gaussian distribution with mean zero, and is asymptotically equivalent to $\hat{T} ({\bar{γ}}_{\bar{θ}}^{(2)})$ where

\hat{T} (γ) = n^{\frac{1}{2}} \sum_{s = 1}^{S} ρ_{s} [n_{s}^{- 1} \sum_{i = 1}^{N} V_{i} I (S_{i} = s) {1 - 2 {\bar{Y}}_{i} + \dot{D} {(\bar{θ})}^{⊤} A^{- 1} x_{i}} {y_{i} - g (γ^{⊤} {\bar{Ψ}}_{i})}] .

This implies that (i) ${\hat{D}}_{i n t r i}$ is asymptotically equivalent to ${\hat{D}}_{S S L}$ when the imputation model $P (y = 1 ∣ u) = g (γ^{⊤} \bar{Ψ})$ is correctly specified and (ii) the asymptotic variance of $n^{\frac{1}{2}} ({\hat{D}}_{i n t r i} - \bar{D})$ is minimized among ${\hat{T} (γ) : E {[x^{⊤}, Y ({\bar{θ}}^{⊤} x)]}^{⊤} {y - g (γ^{⊤} \bar{Ψ})} = 0}$ . Consequently, the asymptotic variance of the intrinsic efficient estimator is always less than or equal to the asymptotic variance of $n^{\frac{1}{2}} ({\hat{D}}_{S S L} - \bar{D})$ and $n^{\frac{1}{2}} ({\hat{D}}_{S L} - \bar{D})$ .

G.2. Asymptotic Properties of ${\hat{θ}}_{intri}$ and ${\hat{D}}_{intri}$

We first introduce the smoothness condition on the link function g(·), which is stronger than Condition 2, but still holds for the most commonly used link functions such as the logit and probit functions.

Condition A1.

The link function $g (\cdot) \in (0, 1)$ is continuously twice differentiable with derivative $\dot{g} (\cdot)$ and the second order derivative $\ddot{g} (\cdot)$ .

Given Condition A1, we let ${\bar{γ}}^{(2)} = {\bar{γ}}_{\bar{θ}}^{(2)}$ and define

A_{1} = E [R {(e^{⊤} A^{- 1} x)}^{2} Φ^{\otimes 2} {{\dot{g}}^{2} ({\bar{γ}}^{(1) T} Φ) + \ddot{g} ({\bar{γ}}^{(1) ⊤} Φ) [y - g ({\bar{γ}}^{(1) ⊤} Φ)]}],

A_{2} = E [R {1 - 2 \bar{Y} + \dot{D} {(\bar{θ})}^{⊤} A^{- 1} x}^{2} {\bar{Ψ}}^{\otimes 2} {{\dot{g}}^{2} ({\bar{γ}}^{(2) ⊤} \bar{Ψ}) + \ddot{g} ({\bar{γ}}^{(2) ⊤} \bar{Ψ}) [y - g ({\bar{γ}}^{(2) T} \bar{Ψ})]}],

$B_{1} = E [Φ x^{⊤} \dot{g} ({\bar{γ}}^{(1) ⊤} Φ)]$ and $B_{2} = E [\bar{Ψ} {x^{⊤}, Y ({\bar{θ}}^{⊤} x)} \dot{g} ({\bar{γ}}^{(2) ⊤} \bar{Ψ})]$ . We next present the regularity condition on the covariates and regression coefficients required by Theorem 3.

Condition A2.

There exists $Θ^{'} = {θ : ∥ θ - \bar{θ} ∥_{2} < δ^{'}}$ for some δ^′ > 0, such that for any $θ \in Θ^{'}$ , there is no γ such that $P (γ^{⊤} Φ_{1} > γ^{⊤} Φ_{2} ∣ y_{1} > y_{2}) = 1$ or $P (γ^{⊤} Ψ_{θ 1} > γ^{⊤} Ψ_{θ 2} ∣ y_{1} > y_{2}) = 1$ . It is also the case that $A ≻ 0, A_{1} ≻ 0, A_{2} ≻ 0, B_{1}^{⊤} A_{1}^{- 1} B_{1} ≻ 0$ and $B_{2}^{⊤} A_{2}^{- 1} B_{2} ≻ 0$ .

Remark A1.

Condition A2 is analog to Condition 3. It assumes there is no linear combination of $Φ$ or $\bar{Ψ}$ perfectly separating the samples based on y, and the Hessian matrices of the constrained least square problems for ${\bar{γ}}^{(1)}$ and ${\bar{γ}}^{(2)}$ are positive definite. Again, these assumptions are mild and common in the M-estimation literature [Van der Vaart, 2000].

Under these regularity conditions, we present the proofs of Theorem 3 and A1. In our development, we take $\tilde{θ}$ as the SSL estimator for θ introduced in Section 3.1, but note that proof remains basically unchanged when taking $\tilde{θ}$ as the SL estimator. We first derive the consistency (error rates) of ${\hat{\dot{D}}}_{1}$ and ${\hat{\dot{D}}}_{2}$ . For the Brier score, letr ${\bar{\dot{D}}}_{1}$ denote the derivative the of $D_{1} (θ)$ evaluated at $\bar{θ}$ . We then use ${\hat{ρ}}_{s} \overset{p}{\to} ρ_{s}$ , Theorem 1, Conditions 1 and A1, and the classical Central Limit Theorem to derive that

\begin{array}{l} {\hat{\dot{D}}}_{1} - {\bar{\dot{D}}}_{1} = \frac{1}{N} \sum_{i = 1}^{N} - 2 {\hat{w}}_{i} \dot{g} ({\tilde{θ}}^{⊤} x_{i}) {y_{i} - g ({\tilde{θ}}^{⊤} x_{i})} x_{i} - \frac{1}{N} \sum_{i = 1}^{N} - 2 w_{i} \dot{g} ({\bar{θ}}^{⊤} x_{i}) {y_{i} - g ({\bar{θ}}^{⊤} x_{i})} x_{i} \\ + \frac{1}{N} \sum_{i = 1}^{N} - 2 w_{i} \dot{g} ({\bar{θ}}^{⊤} x_{i}) {y_{i} - g ({\bar{θ}}^{⊤} x_{i})} x_{i} - E [2 \dot{g} ({\bar{θ}}^{⊤} x) {y - g ({\bar{θ}}^{⊤} x)} x] \\ = O_{p} (∥ \tilde{θ} - \bar{θ} ∥_{2}) + O_{p} (n^{- \frac{1}{2}}) = O_{p} (n^{- \frac{1}{2}}) . \end{array}

For the estimator of the derivative for the OMR, let ${\bar{\dot{D}}}_{2}$ be the limiting value of ${\hat{\dot{D}}}_{2}$ . We then have

{\hat{\dot{D}}}_{2} - {\bar{\dot{D}}}_{2} = \frac{1}{N} \sum_{i = 1}^{N} (1 - 2 y_{i}) [{\hat{w}}_{i} \dot{g} ({\tilde{θ}}^{⊤} x_{i}) K_{h} {g ({\tilde{θ}}^{⊤} x_{i}) - c} - w_{i} \dot{g} ({\bar{θ}}^{⊤} x_{i}) K_{h} {g ({\bar{θ}}^{⊤} x_{i}) - c}] x_{i} + \frac{1}{N} \sum_{i = 1}^{N} w_{i} (1 - 2 y_{i}) \dot{g} ({\bar{θ}}^{⊤} x_{i}) K_{h} {g ({\bar{θ}}^{⊤} x_{i}) - c} x_{i} - E [(1 - 2 y) \dot{g} ({\bar{θ}}^{⊤} x) x ∣ g ({\bar{θ}}^{⊤} x) = c] f_{g} (c) = : Δ_{1} + Δ_{2}

where $f_{g} (c)$ represent the density function of $g ({\bar{θ}}^{⊤} x)$ evaluated at c. This follows from the fact that ${\bar{\dot{D}}}_{2} = E [(1 - 2 y) \dot{g} ({\bar{θ}}^{⊤} x) x ∣ g ({\bar{θ}}^{⊤} x) = c] f_{g} (c)$ . Since the Gaussian kernel K(·) is continuously differentiable and by Theorem 1, Conditions 1 and A1, we have

{‖ Δ_{1} ‖}_{2} = h^{- 1} O_{p} (∥ \tilde{θ} - \bar{θ} ∥_{2}) = O_{p} (n^{- \frac{1}{2}} h^{- 1}) .

For Δ₂, Condition 1 and the classical Central Limit Theorem imply that

\frac{1}{N} \sum_{i = 1}^{N} w_{i} (1 - 2 y_{i}) \dot{g} ({\bar{θ}}^{⊤} x_{i}) K_{h} {g ({\bar{θ}}^{⊤} x_{i}) - c} x_{i} - E (1 - 2 y) \dot{g} ({\bar{θ}}^{⊤} x) K_{h} {g ({\bar{θ}}^{⊤} x) - c} x = O_{p} {{(n h)}^{- \frac{1}{2}}},

and from Condition 1,

E (1 - 2 y) \dot{g} ({\bar{θ}}^{⊤} x) K_{h} {g ({\bar{θ}}^{⊤} x) - c} x - E [(1 - 2 y) \dot{g} ({\bar{θ}}^{⊤} x) x ∣ g ({\bar{θ}}^{⊤} x) = c] f_{g} (c) = \int_{0}^{1} {r (u) f_{g} (u) K_{h} (u - c) - r (c) f_{g} (c)} d u = \int_{- c / h}^{(1 - c) / h} {r (c + h v) f_{g} (c + h v) K (v) - r (c) f_{g} (c)} d v = O (h),

where $r (u) = \int_{{x : g ({\bar{θ}}^{⊤} x) = u}} {1 - 2 P (y = 1 ∣ x)} \dot{g} ({\bar{θ}}^{⊤} x) x f_{x ∣ g} (x ∣ u) d x$ , and $f_{x ∣ g} (\cdot ∣ u)$ represent the density of x given that $g ({\bar{θ}}^{⊤} x) = u$ . By Condition 1, there exists C > 0 such that $∥ r (a) - r (b) ∥_{2} ⩽ C | a - b |$ for any $a, b \in ℝ$ . Thus, we have ${‖ Δ_{2} ‖}_{2} = O_{p} {{(n h)}^{- \frac{1}{2}} + h}$ and with $h ≍ n^{- \frac{1}{4}}$ , we obtain ${‖ {\hat{\dot{D}}}_{2} - {\bar{\dot{D}}}_{2} ‖}_{2} = O_{p} (n^{- \frac{1}{4}})$ . It then follows that for both Brier score and OMR, $∥ \hat{\dot{D}} - \bar{\dot{D}} ∥_{2} = O_{p} (n^{- \frac{1}{4}}) = o_{p} (1)$ .

Leveraging these results, we establish the asymptotic normality of ${\tilde{γ}}^{(1)}$ and ${\tilde{γ}}_{\tilde{θ}}^{(2)}$ . Similar to Appendices C and E, we apply the ULLN [Pollard, 1990], together with Conditions 1, A1, and A2, and the facts that ${\hat{ρ}}_{s} \overset{p}{\to} ρ_{s}, {\hat{ρ}}_{1 s} \overset{p}{\to} ρ_{1 s}$ , and that $\tilde{θ}, {\hat{A}}^{- 1}$ and $\hat{\dot{D}}$ are consistent for their respective limits, to obtain

\sup_{γ \in Γ^{(1)}} | \frac{1}{n} \sum_{i = 1}^{n} {\hat{ζ}}_{i} {(e^{⊤} {\hat{A}}^{- 1} x_{i})}^{2} {y_{i} - g (γ^{⊤} Φ_{i})}^{2} - E [R {(e^{⊤} A^{- 1} x)}^{2} {y - g (γ^{⊤} Φ)}^{2}] | = o_{p} (1); \sup_{γ \in Γ^{(1)}} {‖ \frac{1}{N} \sum_{i = 1}^{N} {\hat{w}}_{i} x_{i} {y_{i} - g (γ^{⊤} Φ_{i})} - E [x {y - g (γ^{⊤} Φ)}] ‖}_{2} = o_{p} (1);

\sup_{γ \in Γ^{(2)}} ∣ \frac{1}{n} \sum_{i = 1}^{n} {\hat{ζ}}_{i} {1 - 2 Y ({\tilde{θ}}^{⊤} x_{i}) + {\hat{\dot{D}}}^{⊤} {\hat{A}}^{- 1} x_{i}}^{2} {y_{i} - g (γ^{⊤} Ψ_{{\tilde{θ}}_{i}})}^{2} - E [R {1 - 2 \bar{Y} + \dot{D} {(\bar{θ})}^{⊤} A^{- 1} x}^{2} {y - g (γ^{⊤} \bar{Ψ})}^{2}] ∣ = o_{p} (1);

\sup_{γ \in Γ^{(2)}} \frac{1}{N} {‖ \sum_{i = 1}^{N} {\hat{w}}_{i} {[x_{i}^{⊤}, Y ({\tilde{θ}}^{⊤} x_{i})]}^{⊤} {y_{i} - g (γ^{⊤} Ψ_{\tilde{θ} i})} - E {[x^{⊤}, Y ({\bar{θ}}^{⊤} x)]}^{⊤} {y - g (γ^{⊤} \bar{Ψ})} ‖}_{2} = o_{p} (1),

where $Γ^{(1)}$ and $Γ^{(2)}$ are two compact sets containing ${\bar{γ}}^{(1)}$ and ${\bar{γ}}^{(2)}$ , respectively. This implies that ${‖ {\tilde{γ}}^{(1)} - {\bar{γ}}^{(1)} ‖}_{2} = o_{p} (1)$ and ${‖ {\tilde{γ}}_{\tilde{θ}}^{(2)} - {\bar{γ}}^{(2)} ‖}_{2} = O_{p} (1)$ . We then expand (9) and (G.2) to derive that

{\tilde{γ}}^{(1)} = \underset{γ}{argmin} {(γ - {\bar{γ}}^{(1)})}^{⊤} [A_{1} (γ - {\bar{γ}}^{(1)}) + 2 {1 + o_{p} (1)} Ξ_{11} + o_{p} ({‖ {\tilde{γ}}^{(1)} - {\bar{γ}}^{(1)} ‖}_{2} + n^{- \frac{1}{2}})], s.t. B_{1}^{⊤} (γ - {\bar{γ}}^{(1)}) - {1 + o_{p} (1)} Ξ_{12} + o_{p} ({‖ {\tilde{γ}}^{(1)} - {\bar{γ}}^{(1)} ‖}_{2} + n^{- \frac{1}{2}}) = 0;

{\tilde{γ}}_{\tilde{θ}}^{(2)} = \underset{γ}{argmin} {(γ - {\bar{γ}}^{(2)})}^{⊤} [A_{2} (γ - {\bar{γ}}^{(2)}) + 2 {1 + o_{p} (1)} Ξ_{21} + o_{p} ({‖ {\tilde{γ}}^{(2)} - {\bar{γ}}^{(2)} ‖}_{2} + n^{- \frac{1}{2}})], s.t. B_{2}^{⊤} (γ - {\bar{γ}}^{(2)}) - {1 + o_{p} (1)} Ξ_{22} + o_{p} ({‖ {\tilde{γ}}^{(2)} - {\bar{γ}}^{(2)} ‖}_{2} + n^{- \frac{1}{2}}),

where Ξ_{11} = \frac{1}{n} \sum_{i = 1}^{n} ζ_{i} {(e^{⊤} A^{- 1} x_{i})}^{2} \dot{g} ({\bar{γ}}^{(1) ⊤} Φ_{i}) Φ_{i} {y_{i} - g ({\bar{γ}}^{(1) ⊤} Φ_{i})};

Ξ_{12} = \frac{1}{N} \sum_{i = 1}^{N} w_{i} x_{i} {y_{i} - g ({\bar{γ}}^{(1) ⊤} Φ_{i})};

Ξ_{21} = \frac{1}{n} \sum_{i = 1}^{n} ζ_{i} {1 - 2 Y ({\bar{θ}}^{⊤} x_{i}) + {\dot{D}}^{⊤} A^{- 1} x_{i}}^{2} \dot{g} ({\bar{γ}}^{(2) ⊤} {\bar{Ψ}}_{i}) {\bar{Ψ}}_{i} {y_{i} - g ({\bar{γ}}^{(2) ⊤} {\bar{Ψ}}_{i})} + {\dot{Ξ}}_{θ, 21}^{⊤} (\tilde{θ} - \bar{θ});

Ξ_{22} = \frac{1}{N} \sum_{i = 1}^{N} w_{i} {[x_{i}^{⊤}, Y ({\bar{θ}}^{⊤} x_{i})]}^{⊤} {y_{i} - g ({\bar{γ}}^{(2) ⊤} {\bar{Ψ}}_{i})} + {\dot{Ξ}}_{θ, 22}^{⊤} (\tilde{θ} - \bar{θ}),

and ${\dot{Ξ}}_{θ, 21}$ , ${\dot{Ξ}}_{θ, 22}$ are two fixed loading matrices of the order O(1). By Condition 1 and the classical Central Limit Theorem, $n^{\frac{1}{2}} {(Ξ_{11}^{⊤}, Ξ_{12}^{⊤}, Ξ_{21}^{⊤}, Ξ_{22}^{⊤})}^{⊤}$ converges to a Gaussian distribution with mean 0. By Theorem $n^{\frac{1}{2}} (\tilde{θ} - \bar{θ})$ also converges to a mean-zero Gaussian distribution. Analogous to the proof of Theorem 5.21 of Van der Vaart [2000], we then obtain

{\tilde{γ}}^{(1)} - {\bar{γ}}^{(1)} = [A_{1}^{- 1} - A_{1}^{- 1} B_{1} {(B_{1}^{⊤} A_{1}^{- 1} B_{1})}^{- 1} B_{1}^{⊤} A_{1}^{- 1}] Ξ_{11} + A_{1}^{- 1} B_{1} {(B_{1}^{⊤} A_{1}^{- 1} B_{1})}^{- 1} Ξ_{12} = O_{p} (n^{- \frac{1}{2}});

{\tilde{γ}}_{\tilde{θ}}^{(2)} - {\bar{γ}}^{(2)} = [A_{2}^{- 1} - A_{2}^{- 1} B_{2} {(B_{2}^{⊤} A_{2}^{- 1} B_{2})}^{- 1} B_{2}^{⊤} A_{2}^{- 1}] Ξ_{21} + A_{2}^{- 1} B_{2} {(B_{2}^{⊤} A_{2}^{- 1} B_{2})}^{- 1} Ξ_{22} = O_{p} (n^{- \frac{1}{2}}) .

(G.3)

By Conditions 1, A1 and A2, the consistency of ${\hat{ρ}}_{1}$ for its limit, and the asymptotic expansion of ${\tilde{γ}}^{(1)} - {\bar{γ}}^{(1)}$ derived above, we can use the argument of Appendix E to show that ${\hat{θ}}_{intri} \overset{p}{\to} \bar{θ}$ and obtain the expansion

n^{\frac{1}{2}} ({\hat{θ}}_{intri} - \bar{θ}) = n^{\frac{1}{2}} A^{- 1} [N^{- 1} \sum_{i = 1}^{N} x_{i} {g ({\bar{γ}}^{(1) ⊤} Φ_{i}) - g ({\bar{θ}}^{⊤} x_{i})} + B_{1}^{⊤} ({\tilde{γ}}^{(1)} - {\bar{γ}}^{(1)})] + o_{p} (1), = n^{\frac{1}{2}} \sum_{s = 1}^{S} ρ_{s} [n_{s}^{- 1} \sum_{i = 1}^{n} I (S_{i} = s) A^{- 1} x_{i} {y_{i} - g ({\bar{γ}}^{(1) ⊤} Φ_{i})}] + o_{p} (1) = \hat{W} ({\bar{γ}}^{(1)}) + o_{p} (1) .

The second equality follows from the fact that

B_{1}^{⊤} ({\tilde{γ}}^{(1)} - {\bar{γ}}^{(1)}) = 0 + Ξ_{12} .

Thus, the asymptotic variance of $n^{\frac{1}{2}} (e^{⊤} {\hat{θ}}_{intri} - e^{⊤} \bar{θ}) is E [R {(e^{⊤} A^{- 1} x)}^{2} {y - g ({\bar{γ}}^{(1) ⊤} Φ)}^{2}]$ , which is minimized among those of ${e^{⊤} \hat{W} (γ) : E [x {y - g (γ^{⊤} Φ)}] = 0}$ . From Theorem $n^{\frac{1}{2}} ({\hat{θ}}_{SSL} - \bar{θ})$ is asymptotically equivalent with $W (\bar{γ})$ . Therefore, when the imputation model is correctly specified, that is, there exists γ₀ such that $P (y = 1 ∣ u) = g (Φ^{⊤} γ_{0}), \bar{γ} = {\bar{γ}}^{(1)} = γ_{0}$ , it follows that $n^{\frac{1}{2}} ({\hat{θ}}_{intri} - \bar{θ})$ is asymptotically equivalent to $n^{\frac{1}{2}} ({\hat{θ}}_{SSL} - \bar{θ})$ . This completes the proof of Theorem 3.

Using our previous arguments, we next establish Theorem A1. Similar to (G.3), we expand

n^{\frac{1}{2}} ({\hat{θ}}_{intri}^{D} - \bar{θ}) as

n^{\frac{1}{2}} ({\hat{θ}}_{intri}^{D} - \bar{θ}) = n^{\frac{1}{2}} A^{- 1} [N^{- 1} \sum_{i = 1}^{N} x_{i} {g ({\bar{γ}}^{(2) ⊤} Ψ_{{\tilde{θ}}_{i}}) - g ({\bar{θ}}^{⊤} x_{i})} + B_{1}^{⊤} ({\tilde{γ}}_{\tilde{θ}}^{(2)} - {\bar{γ}}^{(2)})] + o_{p} (1), = n^{\frac{1}{2}} \sum_{s = 1}^{S} ρ_{s} [n_{s}^{- 1} \sum_{i = 1}^{n} I (S_{i} = s) A^{- 1} x_{i} {y_{i} - g ({\bar{γ}}^{(2) ⊤} {\bar{Ψ}}_{i})}] + n^{\frac{1}{2}} A^{- 1} [N^{- 1} \sum_{i = 1}^{N} x_{i} {g ({\bar{γ}}^{(2) ⊤} Ψ_{\tilde{θ} i}) - g ({\bar{γ}}^{(2) ⊤} {\bar{Ψ}}_{i})} + {\dot{Ξ}}_{θ, 22}^{⊤} (\tilde{θ} - \bar{θ})] + o_{p} (1) = n^{\frac{1}{2}} \sum_{s = 1}^{S} ρ_{s} [n_{s}^{- 1} \sum_{i = 1}^{n} I (S_{i} = s) A^{- 1} x_{i} {y_{i} - g ({\bar{γ}}^{(2) T} {\bar{Ψ}}_{i})}] + o_{p} (1),

The third equality follows from the fact that $n^{\frac{1}{2}} (\tilde{θ} - \bar{θ}) = O_{p} (1)$ and

\partial (N^{- 1} \sum_{i = 1}^{N} x_{i} {g ({\bar{γ}}^{(2) ⊤} Ψ_{{\tilde{θ}}_{i}}) - g ({\bar{γ}}^{(2) ⊤} {\bar{Ψ}}_{i})}) / \partial θ + {\dot{Ξ}}_{θ, 22} = o_{p} (1) .

Using this result, and applying similar arguments as those used for ${\tilde{γ}}_{\tilde{θ}}^{(2)}$ , we have that

{\tilde{γ}}_{{\hat{θ}}_{intri}^{D}}^{(2)} - {\bar{γ}}^{(2)} = [A_{2}^{- 1} - A_{2}^{- 1} B_{2} {(B_{2}^{⊤} A_{2}^{- 1} B_{2})}^{- 1} B_{2}^{⊤} A_{2}^{- 1}] Ξ_{21}^{'} + A_{2}^{- 1} B_{2} {(B_{2}^{⊤} A_{2}^{- 1} B_{2})}^{- 1} Ξ_{22}^{'} = O_{p} (n^{- \frac{1}{2}}), where Ξ_{21}^{'} = \frac{1}{n} \sum_{i = 1}^{n} ζ_{i} {1 - 2 Y ({\bar{θ}}^{⊤} x_{i}) + {\dot{D}}^{⊤} A^{- 1} x_{i}}^{2} \dot{g} ({\bar{γ}}^{(2) ⊤} {\bar{Ψ}}_{i}) {\bar{Ψ}}_{i} {y_{i} - g ({\bar{γ}}^{(2) ⊤} {\bar{Ψ}}_{i})} + {\dot{Ξ}}_{θ, 21}^{⊤} ({\hat{θ}}_{intri}^{D} - \bar{θ}); Ξ_{22}^{'} = \frac{1}{N} \sum_{i = 1}^{N} w_{i} {[x_{i}^{⊤}, Y ({\bar{θ}}^{⊤} x_{i})]}^{⊤} {y_{i} - g ({\bar{γ}}^{(2) ⊤} {\bar{Ψ}}_{i})} + {\dot{Ξ}}_{θ, 22}^{⊤} ({\hat{θ}}_{intri}^{D} - \bar{θ}) .

We then follow the same procedure as in Appendix F (specifically, noting that ${\tilde{γ}}_{{\hat{θ}}_{intri}^{D}}^{(2)}$ corresponds to the ${\hat{θ}}_{intri}^{D}$ plugged into ${\hat{D}}_{intri} (θ)$ , the derivation for the augmentation approach in Section 3.2 can be used directly) to derive that ${\hat{D}}_{intri} \overset{p}{\to} \bar{D}$ and

n^{\frac{1}{2}} ({\hat{D}}_{intri} - \bar{D}) = n^{\frac{1}{2}} [N^{- 1} \sum_{i = 1}^{N} {1 - 2 {\bar{Y}}_{i} + {\bar{D}}^{⊤} A^{- 1} x_{i}} {y_{i} - g ({\bar{γ}}^{(2) ⊤} {\bar{Ψ}}_{i})}] + o_{p} (1) = \hat{T} ({\bar{γ}}_{\bar{θ}}^{(2)}) + o_{p} (1) .

By the definition of ${\bar{γ}}_{\bar{θ}}^{(2)}$ , the asymptotic variance of $\hat{T} ({\bar{γ}}_{\bar{θ}}^{(2)})$ is minimized among those of ${\hat{T} (γ) : E {[x^{⊤}, Y ({\bar{θ}}^{⊤} x)]}^{⊤} {y - g (γ^{⊤} \bar{Ψ})} = 0}$ . Additionally, we may use a similar procedure as that in Appendix F to derive that

n^{\frac{1}{2}} ({\hat{D}}_{SSL} - \bar{D}) = \hat{T} ({\bar{γ}}_{\bar{θ}}) + o_{p} (1) .

Thus, when the imputation model for estimating D, i.e. $P (y = 1 ∣ u) = g (γ^{⊤} \bar{Ψ})$ is correct, we have ${\bar{γ}}_{\bar{θ}} = {\bar{γ}}_{\bar{θ}}^{(2)}$ and that $n^{\frac{1}{2}} ({\hat{D}}_{intri} - \bar{D})$ is asymptotically equivalent to $n^{\frac{1}{2}} ({\hat{D}}_{SSL} - \bar{D})$ . These arguments establish Theorem A1.

H. Justification for Weighted CV Procedure

To provide a heuristic justification for the weights for our ensemble CV method, consider an arbitrary smooth loss function d(·,·) and let $D (θ) = E [d {y_{0}, Y (θ^{⊤} x_{0})}]$ . Let $\hat{D} (θ)$ denote the empirical unbiased estimate of $D (θ)$ and suppose that $\hat{θ}$ minimizes $\hat{D} (θ)$ (i.e. $\dot{\hat{D}} (\hat{θ}) = 0$ ). Suppose that $n^{\frac{1}{2}} (\hat{θ} - \bar{θ}) \to N (0, Σ)$ in distribution. Then by a Taylor series expansion of $\hat{D} (\bar{θ}) at \hat{θ}$ ,

\hat{D} (\hat{θ}) = \hat{D} (\bar{θ}) - \frac{1}{2} {(\hat{θ} - \bar{θ})}^{⊤} \ddot{\hat{D}} (\hat{θ}) (\hat{θ} - \bar{θ}) + o_{p} (n^{- 1}) and

E {\hat{D} (\hat{θ})} = D (\bar{θ}) - \frac{1}{2} n^{- 1} Tr {\ddot{D} (\bar{θ}) Σ} + o_{p} (n^{- 1})

where $\ddot{D} (\bar{θ}) = \partial D (θ) / \partial θ \partial θ^{⊤}$ . For the K-fold CV estimator, ${\hat{D}}_{c v} = K^{- 1} \sum_{k = 1}^{K} {\hat{D}}_{k} ({\hat{θ}}_{(- k)})$ , we note that since ${\hat{D}}_{k} (θ)$ is independent of ${\hat{θ}}_{(- k)}$

E ({\hat{D}}_{c v}) = D (\bar{θ}) + K^{- 1} \sum_{k = 1}^{K} E {{\dot{\hat{D}}}_{k} (\bar{θ})} E ({\hat{θ}}_{(- k)} - \bar{θ}) + \frac{1}{2} \frac{K}{K - 1} n^{- 1} Tr {\ddot{D} (\bar{θ}) Σ} + o_{p} (n^{- 1}) = D (\bar{θ}) + \frac{1}{2} \frac{K}{K - 1} n^{- 1} Tr {\ddot{D} (\bar{θ}) Σ} + o_{p} (n^{- 1}),

where the second equality follows from the fact that $E {{\dot{\hat{D}}}_{k} (\bar{θ})} = \dot{D} (\bar{θ}) = 0$ when $\bar{θ}$ minimizes $D (θ)$ . Letting ${\hat{D}}_{ω} = ω \hat{D} (\hat{θ}) + (1 - ω) {\hat{D}}_{c v}$ with $ω = K / (2 K - 1)$ , it follows that $ω n^{- 1} - (1 - ω) K n^{- 1} / (K - 1) = 0$ and thus

E ({\hat{D}}_{ω}) = D (\bar{θ}) + o_{p} (n^{- 1}) .

REFERENCES

Ananthakrishnan A, Cai T, Savova G, Cheng S, Chen P, Perez R, Gainer V, Murphy S, Szolovits P, Xia Z, et al. Improving case definition of crohn’s disease and ulcerative colitis in electronic medical records using natural language processing: a novel informatics approach. Inflammatory bowel diseases, 19(7):1411–1420, 2013. [DOI] [PMC free article] [PubMed] [Google Scholar]
Belkin M. and Niyogi P. Semi-supervised learning on riemannian manifolds. Machine learning, 56(1–3):209–239, 2004. [Google Scholar]
Belkin M, Niyogi P, and Sindhwani V. Manifold regularization: A geometric framework for learning from labeled and unlabeled examples. The Journal of Machine Learning Research, 7:2399–2434, 2006. [Google Scholar]
Cai T. and Zheng Y. Evaluating prognostic accuracy of biomarkers in nested case–control studies. Biostatistics, 13(1):89–100, 2012. [DOI] [PMC free article] [PubMed] [Google Scholar]
Castelli V. and Cover TM The relative value of labeled and unlabeled samples in pattern recognition with an unknown mixing parameter. Information Theory, IEEE Transactions on, 42(6):2102–2117, 1996. [Google Scholar]
Chakrabortty A. and Cai T. Efficient and adaptive linear regression in semi-supervised settings. The Annals of Statistics, 46(4):1541–1572, 2018. [Google Scholar]
Chapelle O, Scholkopf B, and Zien A. Semi-supervised learning (chapelle O. et al. , eds.; 2006)[book reviews]. IEEE Transactions on Neural Networks, 20(3):542–542, 2009. [Google Scholar]
Corduneanu AAD Stable mixing of complete and incomplete information. PhD thesis, Massachusetts Institute of Technology, 2002. [Google Scholar]
Cozman FG, Cohen I, and Cirelo M. Unlabeled data can degrade classification performance of generative classifiers. In FLAIRS Conference, pages 327–331, 2002. [Google Scholar]
Cozman FG, Cohen I, and Cirelo MC Semi-supervised learning of mixture models. In Proceedings of the 20th International Conference on Machine Learning (ICML-03), pages 99–106, 2003. [Google Scholar]
Efron B. Estimating the error rate of a prediction rule: improvement on cross-validation. Journal of the American Statistical Association, 78(382):316–331, 1983. [Google Scholar]
Efron B. How biased is the apparent error rate of a prediction rule? Journal of the American Statistical Association, 81(394):461–470, 1986. [Google Scholar]
Efron B. and Tibshirani R. Improvements on cross-validation: the 632+ bootstrap method. Journal of the American Statistical Association, 92(438):548–560, 1997. [Google Scholar]
Fu WJ, Carroll RJ, and Wang S. Estimating misclassification error with small samples via bootstrap cross-validation. Bioinformatics, 21(9):1979–1986, 2005. [DOI] [PubMed] [Google Scholar]
Gerds TA, Cai T, and Schumacher M. The performance of risk prediction models. Biometrical Journal: Journal of Mathematical Methods in Biosciences, 50(4):457–479, 2008. [DOI] [PubMed] [Google Scholar]
Gneiting T. and Raftery AE Strictly proper scoring rules, prediction, and estimation. Journal of the American statistical Association, 102(477):359–378, 2007. [Google Scholar]
Gronsbell JL and Cai T. Semi-supervised approaches to efficient evaluation of model prediction performance. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 80(3):579–594, 2018. [DOI] [PMC free article] [PubMed] [Google Scholar]
Hand DJ Construction and assessment of classification rules. Wiley, 1997. [Google Scholar]
Hand DJ Measuring diagnostic accuracy of statistical prediction rules. Statistica Neerlandica, 55(1):3–16, 2001. [Google Scholar]
Jaakkola T, Haussler D, et al. Exploiting generative models in discriminative classifiers. Advances in neural information processing systems, pages 487–493, 1999.
Jiang W. and Simon R. A comparison of bootstrap methods and an adjusted bootstrap approach for estimating the prediction error in microarray classification. Statistics in medicine, 26(29):5320–5334, 2007. [DOI] [PubMed] [Google Scholar]
Kawakita M. and Kanamori T. Semi-supervised learning with density-ratio estimation. Machine learning, 91(2):189–209, 2013. [Google Scholar]
Kawakita M. and Takeuchi J. Safe semi-supervised learning based on weighted likelihood. Neural Networks, 53:146–164, 2014. [DOI] [PubMed] [Google Scholar]
Kohane IS Using electronic health records to drive discovery in disease genomics. Nature Reviews Genetics, 12(6):417–428, 2011. [DOI] [PubMed] [Google Scholar]
Kpotufe S. The curse of dimension in nonparametric regression. PhD thesis, UC San Diego, 2010. [Google Scholar]
Krijthe J. and Loog M. Projected estimators for robust semi-supervised classification. arXiv preprint arXiv:1602.07865, 2016. [Google Scholar]
Liao KP, Cai T, Gainer V, Goryachev S, Zeng-treitler Q, Raychaudhuri S, Szolovits P,Churchill S, Murphy S, Kohane I, et al. Electronic medical records for discovery research in rheumatoid arthritis. Arthritis care & research, 62(8):1120–1127, 2010. [DOI] [PMC free article] [PubMed] [Google Scholar]
Liao KP, Kurreeman F, Li G, Duclos G, Murphy S, Guzman PR, Cai T, Gupta N, Gainer V, Schur P, et al. Autoantibodies, autoimmune risk alleles and clinical associations in rheumatoid arthritis cases and non-ra controls in the electronic medical records. Arthritis and rheumatism, 65(3):571, 2013. [DOI] [PMC free article] [PubMed] [Google Scholar]
Liao KP, Cai T, Savova GK, Murphy SN, Karlson EW, Ananthakrishnan AN, Gainer VS, Shaw SY, Xia Z, Szolovits P, et al. Development of phenotype algorithms using electronic medical records and incorporating natural language processing. bmj, 350: h1885, 2015. [DOI] [PMC free article] [PubMed] [Google Scholar]
Liu D, Cai T, and Zheng Y. Evaluating the predictive value of biomarkers with stratified case-cohort design. Biometrics, 68(4):1219–1227, 2012. [DOI] [PMC free article] [PubMed] [Google Scholar]
Mirakhmedov SM, Jammalamadaka SR, and Mohamed IB On edgeworth expansions in generalized urn models. Journal of Theoretical Probability, 27(3):725–753, 2014. [Google Scholar]
Molinaro AM, Simon R, and Pfeiffer RM Prediction error estimation: a comparison of resampling methods. Bioinformatics, 21(15):3301–3307, 2005. [DOI] [PubMed] [Google Scholar]
Murphy S, Churchill S, Bry L, Chueh H, Weiss S, Lazarus R, Zeng Q, Dubey A, Gainer V, Mendis M, et al. Instrumenting the health care enterprise for discovery research in the genomic era. Genome research, 19(9):1675–1681, 2009. [DOI] [PMC free article] [PubMed] [Google Scholar]
Nedyalkova D. and Tillé Y. Optimal sampling and estimation strategies under the linear model. Biometrika, 95(3):521–537, 2008. [Google Scholar]
Newey WK and McFadden D. Large sample estimation and hypothesis testing. Handbook of econometrics, 4:2111–2245, 1994. [Google Scholar]
Neyman J. On the two different aspects of the representative method: The method of stratified sampling and the method of purposive selection. Journal of the Royal Statistical Society, 97(4):558–606, 1934. [Google Scholar]
Niyogi P. Manifold regularization and semi-supervised learning: Some theoretical analyses. The Journal of Machine Learning Research, 14(1):1229–1250, 2013. [Google Scholar]
Pollard D. Empirical processes: theory and applications. In NSF-CBMS regional conference series in probability and statistics, pages i–86. JSTOR, 1990. [Google Scholar]
Robins JM, Mark SD, and Newey WK Estimating exposure effects by modelling the expectation of exposure conditional on confounders. Biometrics, pages 479–495, 1992. [PubMed]
Robins JM, Rotnitzky A, and Zhao LP Estimation of regression coefficients when some regressors are not always observed. Journal of the American statistical Association, 89 (427):846–866, 1994. [Google Scholar]
Särndal C-E, Swensson B, and Wretman J. Model assisted survey sampling. Springer Science & Business Media, 2003. [Google Scholar]
Sinnott JA, Dai W, Liao KP, Shaw SY, Ananthakrishnan AN, Gainer VS, Karlson EW, Churchill S, Szolovits P, Murphy S, et al. Improving the power of genetic association tests with imperfect phenotype derived from electronic medical records. Human genetics, 133(11):1369–1382, 2014. [DOI] [PMC free article] [PubMed] [Google Scholar]
Sokolovska N, Cappé O, and Yvon F. The asymptotics of semi-supervised learning in discriminative probabilistic models. In Proceedings of the 25th international conference on Machine learning, pages 984–991. ACM, 2008. [Google Scholar]
Tan Z. Bounded, efficient and doubly robust estimation with inverse weighting. Biometrika, 97(3):661–682, 2010. [Google Scholar]
Tian L, Cai T, Goetghebeur E, and Wei L. Model evaluation based on the sampling distribution of estimated absolute prediction error. Biometrika, 94(2):297–311, 2007. [Google Scholar]
Van der Vaart AW Asymptotic statistics, volume 3. Cambridge University Press, 2000. [Google Scholar]
Wasserman L. and Lafferty JD Statistical analysis of semi-supervised regression. In Advances in Neural Information Processing Systems, pages 801–808, 2008.
Wilke R, Xu H, Denny J, Roden D, Krauss R, McCarty C, Davis R, Skaar T, Lamba J, and Savova G. The emerging role of electronic medical records in pharmacogenomics. Clinical Pharmacology & Therapeutics, 89(3):379–386, 2011. [DOI] [PMC free article] [PubMed] [Google Scholar]
Xia Z, Secor E, Chibnik LB, Bove RM, Cheng S, Chitnis T, Cagan A, Gainer VS, Chen PJ, Liao KP, et al. Modeling disease severity in multiple sclerosis using electronic health records. PloS one, 8(11):e78927, 2013. [DOI] [PMC free article] [PubMed] [Google Scholar]
Yu S, Liao KP, Shaw SY, Gainer VS, Churchill SE, Szolovits P, Murphy SN, Kohane IS, and Cai T. Toward high-throughput phenotyping: unbiased automated feature extraction and selection from knowledge sources. Journal of the American Medical Informatics Association, 22(5):993–1000, 2015. [DOI] [PMC free article] [PubMed] [Google Scholar]
Zhang A, Brown LD, Cai TT, et al. Semi-supervised inference: General theory and estimation of means. Annals of Statistics, 47(5):2538–2566, 2019. [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

supinfo

NIHMS1779685-supplement-supinfo.pdf^{(570.2KB, pdf)}

[R1] Ananthakrishnan A, Cai T, Savova G, Cheng S, Chen P, Perez R, Gainer V, Murphy S, Szolovits P, Xia Z, et al. Improving case definition of crohn’s disease and ulcerative colitis in electronic medical records using natural language processing: a novel informatics approach. Inflammatory bowel diseases, 19(7):1411–1420, 2013. [DOI] [PMC free article] [PubMed] [Google Scholar]

[R2] Belkin M. and Niyogi P. Semi-supervised learning on riemannian manifolds. Machine learning, 56(1–3):209–239, 2004. [Google Scholar]

[R3] Belkin M, Niyogi P, and Sindhwani V. Manifold regularization: A geometric framework for learning from labeled and unlabeled examples. The Journal of Machine Learning Research, 7:2399–2434, 2006. [Google Scholar]

[R4] Cai T. and Zheng Y. Evaluating prognostic accuracy of biomarkers in nested case–control studies. Biostatistics, 13(1):89–100, 2012. [DOI] [PMC free article] [PubMed] [Google Scholar]

[R5] Castelli V. and Cover TM The relative value of labeled and unlabeled samples in pattern recognition with an unknown mixing parameter. Information Theory, IEEE Transactions on, 42(6):2102–2117, 1996. [Google Scholar]

[R6] Chakrabortty A. and Cai T. Efficient and adaptive linear regression in semi-supervised settings. The Annals of Statistics, 46(4):1541–1572, 2018. [Google Scholar]

[R7] Chapelle O, Scholkopf B, and Zien A. Semi-supervised learning (chapelle O. et al. , eds.; 2006)[book reviews]. IEEE Transactions on Neural Networks, 20(3):542–542, 2009. [Google Scholar]

[R8] Corduneanu AAD Stable mixing of complete and incomplete information. PhD thesis, Massachusetts Institute of Technology, 2002. [Google Scholar]

[R9] Cozman FG, Cohen I, and Cirelo M. Unlabeled data can degrade classification performance of generative classifiers. In FLAIRS Conference, pages 327–331, 2002. [Google Scholar]

[R10] Cozman FG, Cohen I, and Cirelo MC Semi-supervised learning of mixture models. In Proceedings of the 20th International Conference on Machine Learning (ICML-03), pages 99–106, 2003. [Google Scholar]

[R11] Efron B. Estimating the error rate of a prediction rule: improvement on cross-validation. Journal of the American Statistical Association, 78(382):316–331, 1983. [Google Scholar]

[R12] Efron B. How biased is the apparent error rate of a prediction rule? Journal of the American Statistical Association, 81(394):461–470, 1986. [Google Scholar]

[R13] Efron B. and Tibshirani R. Improvements on cross-validation: the 632+ bootstrap method. Journal of the American Statistical Association, 92(438):548–560, 1997. [Google Scholar]

[R14] Fu WJ, Carroll RJ, and Wang S. Estimating misclassification error with small samples via bootstrap cross-validation. Bioinformatics, 21(9):1979–1986, 2005. [DOI] [PubMed] [Google Scholar]

[R15] Gerds TA, Cai T, and Schumacher M. The performance of risk prediction models. Biometrical Journal: Journal of Mathematical Methods in Biosciences, 50(4):457–479, 2008. [DOI] [PubMed] [Google Scholar]

[R16] Gneiting T. and Raftery AE Strictly proper scoring rules, prediction, and estimation. Journal of the American statistical Association, 102(477):359–378, 2007. [Google Scholar]

[R17] Gronsbell JL and Cai T. Semi-supervised approaches to efficient evaluation of model prediction performance. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 80(3):579–594, 2018. [DOI] [PMC free article] [PubMed] [Google Scholar]

[R18] Hand DJ Construction and assessment of classification rules. Wiley, 1997. [Google Scholar]

[R19] Hand DJ Measuring diagnostic accuracy of statistical prediction rules. Statistica Neerlandica, 55(1):3–16, 2001. [Google Scholar]

[R20] Jaakkola T, Haussler D, et al. Exploiting generative models in discriminative classifiers. Advances in neural information processing systems, pages 487–493, 1999.

[R21] Jiang W. and Simon R. A comparison of bootstrap methods and an adjusted bootstrap approach for estimating the prediction error in microarray classification. Statistics in medicine, 26(29):5320–5334, 2007. [DOI] [PubMed] [Google Scholar]

[R22] Kawakita M. and Kanamori T. Semi-supervised learning with density-ratio estimation. Machine learning, 91(2):189–209, 2013. [Google Scholar]

[R23] Kawakita M. and Takeuchi J. Safe semi-supervised learning based on weighted likelihood. Neural Networks, 53:146–164, 2014. [DOI] [PubMed] [Google Scholar]

[R24] Kohane IS Using electronic health records to drive discovery in disease genomics. Nature Reviews Genetics, 12(6):417–428, 2011. [DOI] [PubMed] [Google Scholar]

[R25] Kpotufe S. The curse of dimension in nonparametric regression. PhD thesis, UC San Diego, 2010. [Google Scholar]

[R26] Krijthe J. and Loog M. Projected estimators for robust semi-supervised classification. arXiv preprint arXiv:1602.07865, 2016. [Google Scholar]

[R27] Liao KP, Cai T, Gainer V, Goryachev S, Zeng-treitler Q, Raychaudhuri S, Szolovits P,Churchill S, Murphy S, Kohane I, et al. Electronic medical records for discovery research in rheumatoid arthritis. Arthritis care & research, 62(8):1120–1127, 2010. [DOI] [PMC free article] [PubMed] [Google Scholar]

[R28] Liao KP, Kurreeman F, Li G, Duclos G, Murphy S, Guzman PR, Cai T, Gupta N, Gainer V, Schur P, et al. Autoantibodies, autoimmune risk alleles and clinical associations in rheumatoid arthritis cases and non-ra controls in the electronic medical records. Arthritis and rheumatism, 65(3):571, 2013. [DOI] [PMC free article] [PubMed] [Google Scholar]

[R29] Liao KP, Cai T, Savova GK, Murphy SN, Karlson EW, Ananthakrishnan AN, Gainer VS, Shaw SY, Xia Z, Szolovits P, et al. Development of phenotype algorithms using electronic medical records and incorporating natural language processing. bmj, 350: h1885, 2015. [DOI] [PMC free article] [PubMed] [Google Scholar]

[R30] Liu D, Cai T, and Zheng Y. Evaluating the predictive value of biomarkers with stratified case-cohort design. Biometrics, 68(4):1219–1227, 2012. [DOI] [PMC free article] [PubMed] [Google Scholar]

[R31] Mirakhmedov SM, Jammalamadaka SR, and Mohamed IB On edgeworth expansions in generalized urn models. Journal of Theoretical Probability, 27(3):725–753, 2014. [Google Scholar]

[R32] Molinaro AM, Simon R, and Pfeiffer RM Prediction error estimation: a comparison of resampling methods. Bioinformatics, 21(15):3301–3307, 2005. [DOI] [PubMed] [Google Scholar]

[R33] Murphy S, Churchill S, Bry L, Chueh H, Weiss S, Lazarus R, Zeng Q, Dubey A, Gainer V, Mendis M, et al. Instrumenting the health care enterprise for discovery research in the genomic era. Genome research, 19(9):1675–1681, 2009. [DOI] [PMC free article] [PubMed] [Google Scholar]

[R34] Nedyalkova D. and Tillé Y. Optimal sampling and estimation strategies under the linear model. Biometrika, 95(3):521–537, 2008. [Google Scholar]

[R35] Newey WK and McFadden D. Large sample estimation and hypothesis testing. Handbook of econometrics, 4:2111–2245, 1994. [Google Scholar]

[R36] Neyman J. On the two different aspects of the representative method: The method of stratified sampling and the method of purposive selection. Journal of the Royal Statistical Society, 97(4):558–606, 1934. [Google Scholar]

[R37] Niyogi P. Manifold regularization and semi-supervised learning: Some theoretical analyses. The Journal of Machine Learning Research, 14(1):1229–1250, 2013. [Google Scholar]

[R38] Pollard D. Empirical processes: theory and applications. In NSF-CBMS regional conference series in probability and statistics, pages i–86. JSTOR, 1990. [Google Scholar]

[R39] Robins JM, Mark SD, and Newey WK Estimating exposure effects by modelling the expectation of exposure conditional on confounders. Biometrics, pages 479–495, 1992. [PubMed]

[R40] Robins JM, Rotnitzky A, and Zhao LP Estimation of regression coefficients when some regressors are not always observed. Journal of the American statistical Association, 89 (427):846–866, 1994. [Google Scholar]

[R41] Särndal C-E, Swensson B, and Wretman J. Model assisted survey sampling. Springer Science & Business Media, 2003. [Google Scholar]

[R42] Sinnott JA, Dai W, Liao KP, Shaw SY, Ananthakrishnan AN, Gainer VS, Karlson EW, Churchill S, Szolovits P, Murphy S, et al. Improving the power of genetic association tests with imperfect phenotype derived from electronic medical records. Human genetics, 133(11):1369–1382, 2014. [DOI] [PMC free article] [PubMed] [Google Scholar]

[R43] Sokolovska N, Cappé O, and Yvon F. The asymptotics of semi-supervised learning in discriminative probabilistic models. In Proceedings of the 25th international conference on Machine learning, pages 984–991. ACM, 2008. [Google Scholar]

[R44] Tan Z. Bounded, efficient and doubly robust estimation with inverse weighting. Biometrika, 97(3):661–682, 2010. [Google Scholar]

[R45] Tian L, Cai T, Goetghebeur E, and Wei L. Model evaluation based on the sampling distribution of estimated absolute prediction error. Biometrika, 94(2):297–311, 2007. [Google Scholar]

[R46] Van der Vaart AW Asymptotic statistics, volume 3. Cambridge University Press, 2000. [Google Scholar]

[R47] Wasserman L. and Lafferty JD Statistical analysis of semi-supervised regression. In Advances in Neural Information Processing Systems, pages 801–808, 2008.

[R48] Wilke R, Xu H, Denny J, Roden D, Krauss R, McCarty C, Davis R, Skaar T, Lamba J, and Savova G. The emerging role of electronic medical records in pharmacogenomics. Clinical Pharmacology & Therapeutics, 89(3):379–386, 2011. [DOI] [PMC free article] [PubMed] [Google Scholar]

[R49] Xia Z, Secor E, Chibnik LB, Bove RM, Cheng S, Chitnis T, Cagan A, Gainer VS, Chen PJ, Liao KP, et al. Modeling disease severity in multiple sclerosis using electronic health records. PloS one, 8(11):e78927, 2013. [DOI] [PMC free article] [PubMed] [Google Scholar]

[R50] Yu S, Liao KP, Shaw SY, Gainer VS, Churchill SE, Szolovits P, Murphy SN, Kohane IS, and Cai T. Toward high-throughput phenotyping: unbiased automated feature extraction and selection from knowledge sources. Journal of the American Medical Informatics Association, 22(5):993–1000, 2015. [DOI] [PMC free article] [PubMed] [Google Scholar]

[R51] Zhang A, Brown LD, Cai TT, et al. Semi-supervised inference: General theory and estimation of means. Annals of Statistics, 47(5):2538–2566, 2019. [Google Scholar]

PERMALINK

Efficient Evaluation of Prediction Rules in Semi-Supervised Settings under Stratified Sampling

Jessica Gronsbell

Molei Liu

Lu Tian

Tianxi Cai

Abstract

1. Introduction

2. Preliminaries

2.1. Data Structure

2.2. Problem Set-Up

3. Estimation Procedure

3.1. Step 1: Flexible imputation

3.2. Step 2: Robustness augmentation

4. Bias Correction via Ensemble Cross-Validation

5. Asymptotic Analysis

Condition 1.

Condition 2.

Condition 3.

Remark 1.

5.1. Asymptotic Properties of θˇSSL

Theorem 1.

Remark 2.

5.2. Asymptotic Properties of D^SSL(θˇSSL) and D^SSL(θ^SSL)

Theorem 2.

Remark 3.

Remark 4.

5.3. Intrinsic Efficient SS Estimation

Theorem 3.

5.4. Optimal Allocation in Stratified Sampling

Remark 5.

6. Perturbation Resampling Procedure for Inference

7. Simulation Studies

Figure 1:

Figure 2:

Table 1:

Remark 6.

Remark 7.

8. Example: EHR Study of Diabetic Neuropathy

Table 2:

9. Discussion

Supplementary Material

Acknowledgments

Appendix

Lemma A1.

Proof.

A. Estimation Procedure for θ^SSL

B. Cross-validation Based Inference for θ^SSL

C. Asymptotic Properties of θ^SL

D. Asymptotic Properties of D^SL(θ^SL)

E. Asymptotic Properties of θˇSSL

F. Asymptotic Properties of D^SSL(θ^SSL) and D^sSL(θ˜sSL)

G. Intrinsic Efficient Estimation

G.1. Intrinsic Efficient Estimator for D¯

Theorem A1.

G.2. Asymptotic Properties of θ^intri and D^intri

Condition A1.

Condition A2.

Remark A1.

H. Justification for Weighted CV Procedure

REFERENCES

Associated Data

Supplementary Materials

ACTIONS

PERMALINK

RESOURCES

Similar articles

Cited by other articles

Links to NCBI Databases

5.1. Asymptotic Properties of ${\overset{ˇ}{θ}}_{SSL}$

5.2. Asymptotic Properties of ${\hat{D}}_{SSL} ({\overset{ˇ}{θ}}_{SSL})$ and ${\hat{D}}_{SSL} ({\hat{θ}}_{SSL})$

A. Estimation Procedure for ${\hat{θ}}_{SSL}$

B. Cross-validation Based Inference for ${\hat{θ}}_{SSL}$

C. Asymptotic Properties of ${\hat{θ}}_{SL}$

D. Asymptotic Properties of ${\hat{D}}_{SL} ({\hat{θ}}_{SL})$

E. Asymptotic Properties of ${\overset{ˇ}{θ}}_{SSL}$

F. Asymptotic Properties of ${\hat{D}}_{SSL} ({\hat{θ}}_{SSL})$ and ${\hat{D}}_{sSL} ({\tilde{θ}}_{sSL})$

G.1. Intrinsic Efficient Estimator for $\bar{D}$

G.2. Asymptotic Properties of ${\hat{θ}}_{intri}$ and ${\hat{D}}_{intri}$