ABSTRACT
High‐dimensional regression problems, for example with genomic or drug exposure data, typically involve automated selection of a sparse set of regressors. Penalized regression methods like the LASSO can deliver a family of candidate sparse models. To select one, there are criteria balancing log‐likelihood and model size, the most common being AIC and BIC. These two methods do not take into account the implicit multiple testing performed when selecting variables in a high‐dimensional regression, which makes them too liberal. We propose the extended AIC (EAIC), a new information criterion for sparse model selection in high‐dimensional regressions. It allows for asymptotic FWER control when the candidate regressors are independent. It is based on a simple formula involving model log‐likelihood, model size, the total number of candidate regressors, and the FWER target. In a simulation study over a wide range of linear and logistic regression settings, we assessed the variable selection performance of the EAIC and of other information criteria (including some that also use the number of candidate regressors: mBIC, mAIC, and EBIC) in conjunction with the LASSO. Our method controls the FWER in nearly all settings, in contrast to the AIC and BIC, which produce many false positives. We also illustrate it for the automated signal detection of adverse drug reactions on the French pharmacovigilance spontaneous reporting database.
Keywords: FWER control, high‐dimensional regression, information criterion, LASSO, pharmacovigilance, variable selection
1. Introduction
This work focuses on Information Criteria (IC) in model selection in the context of high‐dimensional regressions, such as those encountered in genomic and drug safety studies. IC are model selection criteria that balance goodness of fit against model complexity (see [1] for the underlying statistical concepts). In practice, an IC selects one model among a family of parametric models by minimizing a function of the number of parameters of the model , denoted as , its log‐likelihood on the data , and the number of observations . The function has the following general form:
| (1) |
where is the penalty function. Very often, the penalty is a linear function of the model size:
| (2) |
where is the weight that penalizes the addition of one new parameter into . The most popular IC are the Akaike Information Criterion (AIC) [2] and the Bayesian Information Criterion (BIC) [3]. They follow the linear form with the following values of :
AIC:
BIC:
In high‐dimensional variable selection, we have a regression model with a large number of covariates, , and we aim to select a sparse sub‐model of it. For combinatorial reasons, the number of sub‐models is too large to allow the computation of each sub‐model's IC. We therefore use a preliminary variable selection algorithm, which yields a smaller family of candidate sub‐models. The LASSO and other penalized regressions with a regularization parameter are examples of such algorithms. Then we select the model with the lower IC among this family. The combination of BIC and preliminary LASSO has been used and compared with other methods, for example in Sabourin et al. [4] and Courtois et al. [5].
However, in high‐dimensional variable selection, the AIC and BIC can yield a large number of false positives [6, 7, 8]. They are not associated with control of multiple testing statistical criteria such as the family‐wise error rate (FWER) or the false discovery rate (FDR).
To address this issue, several authors have proposed information criteria that account for the high‐dimensional nature of variable selection settings by incorporating in their formula, making the addition of a new parameter into the model more penalized when is large to account for the large number of candidate models. An overview of them can be found in Bogdan and Frommlet [9].
Foster and George [10] proposed the Risk Inflation Criterion (RIC), which follows the linear form (2) with . Bogdan and Frommlet (2021) introduced the modified AIC (mAIC), where a positive constant is added to the RIC's linear weight: where is a constant for which the authors recommend the value .
Other proposals are modifications of the BIC which add terms to the BIC's ‐dependent penalty. Bogdan et al. [11] proposed the modified BIC (mBIC) which—in its version stated in Bogdan and Frommlet—also follows (2), with and a constant that the authors suggest be chosen at . Chen and Chen [6] presented the extended BIC (EBIC), which, in its version stated in Bogdan and Frommlet (2021), follows the general form of an IC (1) with the non‐linear penalty depending on a constant . There exists also another version of the EBIC, introduced in Chen and Chen [7] which has a more conservative linear penalty, and variations of the mBIC and mAIC with less conservative non‐linear penalties. The various IC may coincide for some values of their constant parameters.
In this work, we propose our own extended version of the AIC, which we call the Extended Akaike Information Criterion (EAIC). This criterion follows the linear form (2). Its is independent from (as in the classical AIC) and dependent on (as in the mBIC and EBIC); it shares these two features with the RIC and the mAIC. We have built the EAIC so that when the variables are independent, it controls the FWER at a user‐specified level when both and are large. The control relies on a property of the asymptotic distribution of the likelihood ratio between two nested models. Section 2 introduces this property, whose detailed proof is in the supporting information, and its meaning in terms of multiple likelihood ratio tests.
Section 3 focuses on variable selection with information criteria. It describes the existing criteria in terms of likelihood‐ratio tests, defines the EAIC based on the results of Section 2, and describes its combination with the LASSO [12]. We conducted a simulation study to assess the performances of the different information criteria, including the EAIC, in both linear and logistic regression models in a variety of high‐dimensional settings, which include a wide range of values of , , signal strength, and correlation between the variables. This is described in Section 4. In Section 5, we illustrate our approach by an application in the field of signal detection in pharmacovigilance. This application is based on data from the French national spontaneous reporting database and compares the performances of the various IC using a drug reference set pertaining to drug‐induced liver injuries (DILI) [13].
2. FWER Control Bound in Multiple Independent Likelihood Ratio Tests
This section presents a property of likelihood‐ratio tests of the insertion of one more parameter to a parametric model. This property makes it possible to approximately control the FWER in multiple independent tests by setting the test's critical value at a simple function of the approximate number of tests and of the desired FWER.
2.1. Rationale and Notations
Consider a statistical model with log‐likelihood . Let be a subspace of the parameter space having dimension and a subspace having dimension , containing . Let be the maximum of log‐likelihood over , the maximal log‐likelihood over , and . Let be the hypothesis that the true parameter vector is in and be the alternative hypothesis that it is in .
Wilks' theorem [14] asserts that under , has asymptotically a distribution with one degree of freedom, that is, that follows asymptotically a gamma distribution of parameters . Let be the probability density function of this distribution, and its cumulative distribution function. In a likelihood‐ratio test, the statistic is used to test against . A procedure that rejects in favor of if , for some bound , has an asymptotic probability of incorrectly rejecting when it holds.
Suppose now that there are subspaces of parameters of dimension , all containing . Let be the hypothesis that the true parameter vector is in , the maximum of the log‐likelihood over , and .
is tested against each , for with a procedure that rejects in favour of if for some bound which does not depend on . These tests are a tool of model selection: if is rejected for none of the , then the model where the parameters are in is selected. If it is rejected at least once, then a model with more parameters will be selected. We want to control the probability of this event under .
We assume that the likelihood‐ratio tests are independent, that is, the are independently distributed. Then, under , the asymptotic probability that the correct model is selected (i.e., that none of the tests reject ) is . is the asymptotic cumulative distribution function of . We will prove an asymptotic property of for large , which translates into FWER control.
2.2. A Simple Case
For all , the sequence with:
| (3) |
has the property that:
In other words, when performing a large number of independent likelihood‐ratio tests of against for , if the follow their asymptotic distribution, then rejecting in favor of if controls the family‐wise error rate at level .
Given the definition (3), the result has its simplest form when . The bound controls the family‐wise error rate at level .
2.3. General Case
In practice, the critical value must not vary with the exact number of tests . This is because, as detailed in Section 3, using an information criterion amounts to doing several multiple testing procedures where varies a little but stays the same across the families of tests. Therefore, we replace by a proxy quantity approximately equal to it, , in the definition of . We will prove that:
For all , for every sequence (denoted just in the following) such that , the sequence defined in (3) satisfies:
This means that when performing a large number of independent likelihood‐ratio tests where the likelihood ratios follow their asymptotic distribution, for all values of approximately equal to , the bound controls the family‐wise error rate at level .
We prove this property (and consequently its sub‐case stated in Section 2.2) in subsection 1 of the supporting information.
Therefore, for any family of independent tests when both the number of observations and of tests are large, we can reasonably suppose that the family‐wise error rate is controlled at this level (although the theorem above does not deal with the asymptotics in : it assumes that is large enough to use the asymptotic distribution of the ).
3. Variable Selection in High‐Dimensional Regressions
We use the result of Section 2 to quantify the FWER of information‐criteria‐based sparse model selection methods in the simple case of comparing a model with all those having one more parameter, and then to build an FWER‐controlling information criterion, the EAIC. We focus on generalized linear regression models, although the same reasoning would work for any large parametric model.
3.1. AIC and BIC
Consider a generalized linear model:
is assumed to be a large integer, the are assumed to be independent, and the are assumed to be zero when is not part of some unknown sparse subset of . The goal of variable selection is to determine based on observations of and the . For every subset of , it is possible to compute the log‐likelihood of the following sub‐model (also called model ):
Then one model is selected by minimizing an information criterion depending on and on the number of variables in , such as the AIC. It may be interpreted in terms of a likelihood‐ratio test. Let be a subset containing and having one more variable than ; then
Therefore, when minimizing the AIC over and , is selected if , which has an asymptotic probability of if . The AIC is therefore a way of setting the type 1 error rate at a fixed level.
Nonetheless, in this high‐dimensional regression problem, the use of the AIC (or any other similar criterion) is akin to multiple likelihood‐ratio testing. Since variables are not included in A, there are subsets that contain and have one more variable than A. When we compare the AIC of all these subsets, the probability of selecting one of them rather than if is the true model (family‐wise error rate) is (assuming that each follows the cumulative distribution function and that they are independent of each other), which approaches 1 when is large.
The test independence condition is seldom exactly verified. It is closer to being verified when the are not, or are weakly, correlated with each other. A nonzero correlation between the would typically create a positive correlation between the corresponding . The probability of wrongly selecting a model larger than when is the true model would then be lower than in the independent case, but it would still converge to 1.
If the BIC is used rather than the AIC, then if has one more variable than :
Assuming again that each follows and that the tests are independent, the family‐wise error rate is:
If the models that we are comparing are sufficiently sparse that , then:
If goes to infinity while is of the order of or smaller, the BIC's family‐wise error rate converges to 0. However, as this is a high‐dimensional regression, we cannot make this hypothesis. For example, if is of the order of for some larger than 1, the BIC's family‐wise error rate converges to 1. Variable selection based on the BIC does not asymptotically bound the family‐wise error rate in general.
3.2. Extended AIC
The results of Section 2 suggest a way of controlling the family‐wise error rate while using a criterion similar to the AIC. For every , for every subset of , let us define the extended AIC:
where the weight is, as defined in (3):
Unlike the AIC and BIC, the EAIC is not a function of the intrinsic characteristics of the model only. It also takes into account the fact that is considered along with all the other sub‐models of a ‐dimensional regression model.
Like the AIC and BIC, the definition translates into the following difference for two nested models A and B, where B has one more parameter:
B is chosen over A if and only if . Therefore, by minimizing the EAIC among all the subsets containing the true active subset and having one more variable than , we have the situation in Section 2.3 where and . Since the setting is high‐dimensional (large and large ) and the subsets that we are comparing are relatively sparse (), the asymptotic setting of Section 2.3 ( and ) is well approached, and we make the assumption that the follow their asymptotic distribution with cumulative distribution function . Therefore, the family‐wise error rate approaches
3.2.1. Coincidence With Other Information Criteria
For some values of or , the EAIC coincides with other information criteria. It coincides with the RIC (of weight ) when:
that is,
The term (which is equivalent to at large ) has been noted in Bogdan and Frommlet [9] to be interpretable as the RIC's family‐wise test level. decreases at a very slow rate, with , , and .
The EAIC also coincides with the mAIC (of weight ) when:
that is,
or
decreases in at the same slow rate as , asymptotically multiplied by the constant . At the recommended , , , and .
3.2.2. Simplification and Approximation of the EAIC
We noted in Section 2.2 that simplifies when . Therefore, the EAIC has the following simpler version, whose family‐wise error rate approaches :
Furthermore, owing to the asymptotics of when , the EAIC has the following approximation at small values of :
3.3. Model Selection Procedure With the EAIC
We have defined a new information criterion that can be computed for any given submodel. Like the others IC, it could in principle be used to select a sparse submodel of interest by computing it for every submodel and choosing the one with the lowest IC. Although the FWER properties of Sections 3.1 and 3.2 do not automatically translate into equivalent results for this procedure, they are a good indication of the tendency of information criteria to produce false positives.
In practice, as a consequence of the high dimension of the regression problem, the number of subsets of variables is actually too large to compute the likelihood of each of them. A solution is to use a pre‐selection procedure that yields a smaller family of sparse candidate models that explain the outcome particularly well and then to minimize the IC among these models. This is a computationally feasible proxy for minimizing the IC among all the sub‐models. The pre‐selection procedure we used is the LASSO penalized regression [12], where the variations of the regularization parameter deliver a list of sparse models, the LASSO path.
Note that the log‐likelihood used in the information criteria is non‐penalized. This is because Wilks' theorem describes the distributions of the maximal loglikelihoods of nested models. Comparing their maximal penalized log‐likelihoods, or the log‐likelihoods at the parameters where the penalized log‐likelihoods are maximal, would introduce an extra term in the due to the difference in penalization; the sparser, higher‐ models shrink their LASSO estimates of more, which puts their log‐likelihood further from its maximal value. Therefore, to combine the LASSO and an IC, it is necessary to fit each of the models selected by the LASSO without penalty, and not just to reuse the models fitted with the LASSO's penalty.
4. Simulation Study
We conducted two simulation studies. The first one aimed to assess the FWER control given by the EAIC in a simple framework, where it tests all the models having one active variable against the null one (with no active variables) under the null hypothesis. The second one explored the performance of the LASSO‐based complete model selection procedure of Section 3.3 in settings with active variables. The R code of our simulation studies is provided in our supporting information.
4.1. FWER Control Under the Null Hypothesis
We used the EAIC to compare the null model with the models with one active variable . We simulated data under the null model, computed each model's EAIC, and selected the model with one active variable having the lowest EAIC if it was also lower than the null model's EAIC. This amounts to multiple likelihood‐ratio tests where the EAIC performs multiple testing corrections. For a given setting, we assessed the control of the FWER by computing the estimated FWER, which is the proportion of simulated datasets where a non‐null model is selected.
This can be viewed as the first step of the complete model selection procedure: we minimize the EAIC among the models of size only 0 or 1 instead of a whole sequence of models. Unlike in Section 3.3, the minimization is done across all these candidate models without using the LASSO or any other pre‐selection method. For computational reasons, the exception to this were the settings where , where we computed the EAIC of the model including only for the where was at least of its maximal value across all .
4.1.1. Data Simulation
We simulated 1000 datasets from each of 30 settings, described by the following parameters:
The regression model family: linear or logistic.
The number of observations is , or .
The number of regressors , or (only in non‐correlated settings) .
The correlation matrix used to simulate the regressors is a Toeplitz matrix , with or . For computational reasons, settings where were excluded.
The regressors were drawn following a standard normal distribution. The outcome was drawn independently from the , following a standard normal distribution in the linear settings and a Bernoulli distribution with probability 0.5 in the logistic settings. We used the EAIC with .
4.1.2. Simulation Results
As shown in Table 1, the FWERs that we observed in the 30 settings ranged between 0.029 and 0.055, in agreement with the targeted level.
TABLE 1.
Null hypothesis simulation study: observed FWER by setting, averaged on 1000 simulations.
| Parameters |
|
|
|
|||||
|---|---|---|---|---|---|---|---|---|
| Linear |
|
|
.046 | .045 | .055 | |||
|
|
.030 | .030 | .039 | |||||
|
|
.043 | .040 | .049 | |||||
|
|
|
.039 | .049 | |||||
|
|
.033 | .050 | ||||||
|
|
.032 | .048 | ||||||
| Logistic |
|
|
.047 | .045 | .049 | |||
|
|
.037 | .037 | .036 | |||||
|
|
.029 | .046 | .039 | |||||
|
|
|
.039 | .050 | |||||
|
|
.036 | .044 | ||||||
|
|
.029 | .050 | ||||||
4.2. Simulation of the Full Procedure
4.2.1. Design
4.2.1.1. Data Simulation
We compared the performance of the EAIC with that of other Information Criteria‐based variable selection methods on 308 different settings of the generalized linear model . The settings were determined by:
The regression model family: linear or logistic.
The number of observations is , or .
The number of regressors , or (only in non‐correlated settings where ) .
The correlation matrix used to simulate the regressors is a Toeplitz matrix , with or (only where ) .
- The empirical signal‐to‐noise ratio. This quantity was based on the signal‐to‐noise ratio used in Sabourin et al. 2015, although unlike them we used an empirical version, which is defined for every generalized linear model. It is the ratio of the empirical variance of the signals () over the empirical mean of the noises' variances ():
A larger SNR means a more easily observable impact of each active on . Variable selection is expected to be more efficient in settings with large SNR. We set .
We simulated 1000 datasets in each of these settings. In each simulation, we first drew the regressor matrix following a standard normal distribution (with or without correlation depending on the setting). Then we drew 10 active regressors uniformly among the regressors. We determined their coefficients with the following algorithm:
Draw non‐normalized coefficients at random with and following a uniform distribution on , each independently from one another and from their sign;
Define and as the normalized coefficients: where we computed and via numerical equation solving (function multiroot in the R package rootSolve version 1.8.2.3) [15] so that is the desired value.
We then simulated the simulation's outcome vector following the generalized linear model based on the simulated and . In the linear model, .
Due to computational limitations in the simulation of , the settings where or were not included.
4.2.1.2. Method Comparison
For each simulated dataset, we ran the LASSO algorithm using R's glmnet package version 4.1‐2 [16] with the default parameters and at most 100 degrees of freedom (parameter dfmax). We applied seven different variable selection methods, which all consist in choosing one of the models appearing on the LASSO path. Six methods select the model that minimizes an IC:
the AIC
the BIC
the EBIC at
the mAIC at
the mBIC at
the EAIC at
and the EAIC at .
These methods require fitting the non‐penalized version of each of the models on the LASSO path and then using the resulting log‐likelihood to compute the criteria. The seventh method, which we call “oracle”, uses the information that there are 10 active variables. It consists in selecting the first model with at least 10 variables on the LASSO path. Its purpose is to illustrate how accurate variable selection can be expected for each setting.
We computed estimates of the family‐wise error rate (FWER), the false discovery rate (FDR), and the sensitivity of each of these methods for each of the 308 settings by averaging over the 1 000 simulated datasets. The sensitivity is the expected proportion of selected variables among the active variables; since all settings have 10 active variables, we indicate the average number of those true positives, which is proportional to the estimated sensitivity.
4.2.2. Results
For each combination of , , , and model family, we plotted the graph of each method's FWER (Figures 1, 2, 3, 4), average number of true positives (Figures S1–S4) and FDR (Figures S5–S8) varying with the SNR.
FIGURE 1.

Full procedure simulation study: FWER by setting, averaged on 1000 simulations. Linear model, .
FIGURE 2.

Full procedure simulation study: FWER by setting, averaged on 1000 simulations. Logistic model, .
FIGURE 3.

Full procedure simulation study: FWER by setting, averaged on 1000 simulations. Linear model, .
FIGURE 4.

Full procedure simulation study: FWER by setting, averaged on 1000 simulations. Logistic model, .
In Figures 1, 2, 3, 4, the oracle's FWER indicates the models in which it is possible to capture the exact 10 active variables by stopping at a node of the LASSO path. When or the SNR is low, it is almost never possible since the oracle's FWER is at or close to 1. In non‐correlated settings with and sufficiently high SNR (Figures 1 and 2), the oracle's FWER approaches 0, which shows that the correct set of 10 variables is on the LASSO path.
The AIC's FWER, almost always at 1, shows that this method almost always produces false positives.
In non‐correlated settings (Figures 1 and 2), the BIC's FWER tends not to depend on the SNR (except in high‐SNR, high‐ , logistic models where there is a decrease) but is highly dependent on both and . In linear settings with few observations () or many variables (), it is at or close to 1, meaning that the BIC also almost always produces false positives. Only at large and small does the BIC have a moderate FWER.
Figures 1 and 2 show that in accordance with its theoretical basis, the extended AIC approximately controls the FWER at the desired level in most non‐correlated settings. The only exceptions are the setting with few observations (), high SNR, and moderate number of variables ( in the linear model, in the logistic one). In those settings, the high SNR and moderate imply that a large proportion of the active variables are captured (as seen on the sensitivity curves of Figures S1 and S2). Therefore, the EAIC compares models of size of the order of 10. At this size, the convergence in distribution guaranteed by Wilks' theorem is slower than when comparing models of very small size. Therefore, at , the asymptotic regime is not reached, so the theoretical basis for FWER control by the EAIC does not hold. Moreover, in non‐correlated settings, the EAIC's FWER does not drop much below its nominal value except in high‐SNR logistic models (Figure 2).
The other information criteria that penalize for —mBIC, mAIC, and EBIC—generally have a low FWER, except in the same settings where the EAICs' FWER control fails: , high SNR, and moderate number of variables. They exhibit patterns similar to the EAIC at 0.05's FWER, although the mBIC's and the EBIC's FWER show more variation, being sometimes close to 0 and having higher maximal values. The mAIC has a FWER always slightly below that of the EAIC at with a bit more divergence at large , in accordance with their coincidence described in Section 3.2.1.
In correlated settings (Figures 3 and 4), the FWER varies somewhat more with the SNR for all methods except the AIC (which has a FWER always close to 1). When , the FWERs have a similar pattern to those observed in non‐correlated settings, with an even higher spike at high SNR and a moderate number of variables. When , there is some deviation from the FWER control, which was observed in non‐correlated settings. Over all the settings, the EAIC at has a maximal FWER of 0.612, the EAIC at has 0.251, the EBIC has 0.288, the mAIC has 0.247, and the mBIC has 0.224. In the most favorable correlated settings, at large and SNR (when the correct model is almost always on the LASSO path, as shown by the oracle's FWER approaching 0), the EAICs control the FWER at or below its nominal level.
Figures S1–S4 show the number of true positives captured by each method in each setting, which is 10‐fold their sensitivities since all settings are simulated with 10 active variables. The sensitivities increase with the SNR and with but decrease with . The most conservative methods in terms of FWER are also the least sensitive. The EAIC at , mAIC, mBIC, and EAIC have sensitivities close to each other's. However, there are two types of settings where two groups differ in sensitivity:
In the large , low SNR settings, EAIC and mAIC have a higher sensitivity than EBIC and mBIC. They also have a slightly higher FWER, but the EAIC still controls its FWER at the nominal level.
On the contrary, in the small , small , high SNR settings, EBIC and mBIC have a higher sensitivity than EAIC and mAIC. This comes with the price of higher FWERs, which reach their highest values in these settings.
The curves of FDR (Figures S5–S8 in the supporting information) show the same hierarchy of methods, with the AIC being the least conservative and the EBIC being on average the most conservative. While the mAIC, EAIC at , and mBIC have a higher FDR than the EBIC averaged across the settings, their FDR is more stable, being respectively at most 0.075, 0.083, and 0.107 in non‐correlated settings and respectively at most 0.125, 0.126, and 0.141 in correlated settings, while the EBIC reaches 0.128 in a non‐correlated setting and 0.171 in a correlated setting. Among those four, the mBIC has the highest FDR when is small. The EAIC at , although generally not conservative, has a FDR that never approaches 1 (maximal FDR equal to 0.623), as opposed to the BIC, which has a FDR close to 1 in small , large , low SNR settings.
5. Application to Pharmacovigilance Data
To illustrate the behavior of the variable selection methods on real data, we applied them to the French pharmacovigilance database (BNPV). We used the same data preprocessing as described in [5], yielding a database of spontaneous reports of adverse drug reactions from 1 January 2000 to 29 December 2017 with 6617 different adverse events (coded according to the Preferred Term level of the Medical Dictionary for Regulatory Activities, MedDRA) and different drugs (coded with the 5th level of the Anatomical Therapeutic Chemical hierarchy) reported at least 10 times. We focused on a binary outcome, the adverse event Drug‐Induced Liver Injury (DILI), and we used a logistic regression model based on drugs as binary covariates. A pharmacovigilance signal was defined as a variable that is selected with a positive estimated coefficient.
To assess the performances of the methods, we used the same reference set of pharmacovigilance signals as [5] pertaining to DILI [13]. It includes 203 negative controls (drugs known not to be associated with DILI) and 133 positives (drugs known to be associated with DILI).
We implemented six methods that minimize an information criterion on the LASSO path: the AIC, the BIC, the EAIC at various levels of , the EBIC, the mBIC, and the mAIC. Table 2 shows the results of these methods, including the EAIC at , , and . The False Discovery Proportion (FDP), specificity, and sensitivity were computed on drugs with known status. The results are shown in Table 2. As in most simulation settings, the AIC generated the most signals. The mBIC generated the least.
TABLE 2.
Performance of each method on the BNPV dataset in terms of number of pharmacovigilance signals (variables positively associated with DILI), False Discovery Proportion (FDP), specificity and sensitivity.
| Methods | Signals | Signals with known status | False positives | FDP (%) | Specificity (%) | Sensitivity (%) |
|---|---|---|---|---|---|---|
| AIC | 187 | 69 | 5 | 7.2 | 97.5 | 48.1 |
| EAIC at , BIC | 170 | 65 | 5 | 7.7 | 97.5 | 45.1 |
| EAIC at mAIC | 150 | 59 | 4 | 6.8 | 98.0 | 41.4 |
| EAIC at , EBIC | 142 | 55 | 2 | 3.6 | 99.0 | 39.8 |
| mBIC | 118 | 48 | 2 | 4.2 | 99.0 | 34.6 |
The sensitivities of the methods reflect this hierarchy, ranging from for the AIC to for the mBIC. Several pairs of criteria selected the same model and therefore had the same characteristics on this dataset. This is the case of the EAIC at with the BIC (the level at which the EAIC and BIC were mathematically equivalent being for these values of and ), the EAIC at with the mAIC, and the EAIC at with the EBIC. With only two known false positives, these last two had the best specificity together with mBIC (), and the single best FDP ().
6. Discussion
This work focuses on high‐dimensional sparse model selection using information criteria. Compared to variable selection based on univariate models, multivariate model selection is necessary to avoid producing false positives or false negatives when the regressors are correlated. Consequently, although an independence hypothesis was involved in the design of the method we propose, it was important to also measure its properties in settings with correlated regressors.
The AIC and BIC, which are classical IC, are often used for high‐dimensional sparse model selection. In high dimensions, they do not have consistency properties and tend to select models with many false positives, as confirmed by our extensive simulation study. However, we show that limiting the number of false positives while using an IC in high dimension is possible. This is done by choosing a criterion that adjusts for the dimensionality: either the extended BIC, the modified BIC, the modified AIC, or one of their variants; or our proposal, the extended AIC.
Unlike the other information criteria, the EAIC is designed to control the family‐wise error rate at a specified level. Our mathematical result suggests that it achieves this when the candidate variables are not correlated and and are large enough for asymptotic approximations to hold. In practice, the simulations show that the EAIC's FWER is indeed close to its specified level in nearly all non‐correlated settings (across wide ranges of and , in both linear and logistic models), the only exceptions being when both and are small in conjunction with the active variables having a strong effect.
Moreover, the “modified” or “extended” criteria achieve satisfactory FDR performance in contrast with the classical criteria. Although we do not provide mathematical results supporting this claim, our simulations show that while both the AIC and BIC have complete failure modes (settings where their FDR approaches 1), this is not the case of the criteria that penalize for . Even in settings where the candidate variables are correlated, their FDR is at most about 0.6 (EAIC at ) or always below 0.2 (all the other IC which correct for ).
The various IC used in combination with the LASSO are deterministic, simple, and fast methods. They only require fitting the LASSO once and then fitting a sequence of low‐dimensional, non‐penalized regressions without having to re‐sample or simulate data as opposed to (among others) cross‐validation. Although our simulations focused on the LASSO, another pre‐selection procedure can be used in its place as long as it provides a small family of sparse candidate models. This is the case of other penalized regressions such as SCAD [17] and elasticnet [18]. It has been observed that the LASSO path can include avoidable false positives early on [19], which may induce to combine the IC with another pre‐selection method.
Compared to the other information criteria, the EAIC's parametrization is directly interpretable, as is the FWER that one obtains under asymptotic and test independence assumptions—which is in practice the FWER observed in most non‐correlated settings. While the mAIC controls the FWER at known values, which slowly decrease with [9], the EAIC is the only information criterion that controls the FWER at any level that the user may input.
The EAIC uses the number of covariates by interpreting it as an approximation of the number of likelihood‐ratio tests that IC optimization implicitly performs and by making the strong assumption that those tests are independent. Since the tests are in practice non‐independent, mostly because of the correlation between covariates, it can be desirable to replace with a measure of the effective number of independent tests to which the non‐independent tests are approximately equivalent when correcting for multiple testing. The notion is used in multiple testing in genomic analysis, with varying definitions [20, 21]. In the context of genomics, Bogdan et al. [22] have proposed an ‐dependent effective number of tests, which they use in the mBIC, and might also work with the EAIC.
Therefore, we recommend that statistical practitioners who need a simple method for high‐dimensional variable selection based on the LASSO or another pre‐selection method and who want to make sure that they do not mostly select false positives use one of the criteria that adjust for the dimension—mBIC, EBIC, mAIC, or EAIC—possibly with a measure of the effective number of tests, and preferentially the EAIC if a specific level of FWER is targeted.
Conflicts of Interest
The authors declare no conflicts of interest.
Supporting information
Supporting information.
Data S1. Supporting information.
Acknowledgements
The authors thank the regional pharmacovigilance centers and the Agence Nationale de Sécurité du Médicament et des Produits de Santé (ANSM) for providing the pharmacovigilance database. They also thank the reviewers of Statistics in Medicine for their in‐depth commentaries, which allowed them to supplement this work.
Pascale Tubert‐Bitter and Ismaïl Ahmed, Contributed equally.
Funding: This work was supported by the National Natural Science Foundation of China under grant numbers 62373358, 62173177, and 62276119, the Key Program of the National Natural Science Foundation of China under grant number 62333016, and the SuQian Sci ∖& Tech Program undger grant numbers K202225 and Z2023129.
Data Availability Statement
We share our simulation code in the supporting information. The pharmacovigilance data provided by the ANSM are not freely available for privacy reasons.
References
- 1. Emiliano P. C., Vivanco M. J. F., and de Menezes F. S., “Information Criteria: How Do They Behave in Different Models?,” Computational Statistics & Data Analysis 69 (2014): 141–153, 10.1016/j.csda.2013.07.032. [DOI] [Google Scholar]
- 2. Akaike H., “Information Theory and an Extention of the Maximum Likelihood Principle,” in 2nd International Symposium on Information Theory (Budapest: Akademiai Kiado, 1973), 267–281. [Google Scholar]
- 3. Schwarz G., “Estimating the Dimension of a Model,” Annals of Statistics 6, no. 2 (1978): 461–464, 10.1214/aos/1176344136. [DOI] [Google Scholar]
- 4. Sabourin J. A., Valdar W., and Nobel A. B., “A Permutation Approach for Selecting the Penalty Parameter in Penalized Model Selection,” Biometrics 71, no. 4 (2015): 1185–1194, 10.1111/biom.12359. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5. Courtois E., Tubert‐Bitter P., and Ahmed I., “New Adaptive lasso Approaches for Variable Selection in Automated Pharmacovigilance Signal Detection,” BMC Medical Research Methodology 21, no. 1 (2021): 271, 10.1186/s12874-021-01450-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6. Chen J. and Chen Z., “Extended Bayesian Information Criteria for Model Selection With Large Model Spaces,” Biometrika 95, no. 3 (2008): 759–771, 10.1093/biomet/asn034. [DOI] [Google Scholar]
- 7. Chen J. and Chen Z., “Extended BIC for Small‐n‐Large‐P Sparse GLM,” Statistica Sinica 22, no. 2 (2012): 555–574. [Google Scholar]
- 8. Broman K. W. and Speed T. P., “A Model Selection Approach for the Identification of Quantitative Trait Loci in Experimental Crosses,” Journal of the Royal Statistical Society, Series B: Statistical Methodology 64, no. 4 (2002): 641–656, 10.1111/1467-9868.00354. [DOI] [Google Scholar]
- 9. Bogdan M. and Frommlet F., “Identifying Important Predictors in Large Data Bases ‐ Multiple Testing and Model Selection,” in Handbook of Multiple Comparisons (New York: Chapman and Hall/CRC, 2021), 44. [Google Scholar]
- 10. Foster D. P. and George E. I., “The Risk Inflation Criterion for Multiple Regression,” Annals of Statistics 22, no. 4 (1994): 1947–1975, 10.1214/aos/1176325766. [DOI] [Google Scholar]
- 11. Bogdan M., Ghosh J. K., and Doerge R. W., “Modifying the Schwarz Bayesian Information Criterion to Locate Multiple Interacting Quantitative Trait Loci,” Genetics 167, no. 2 (2004): 989–999, 10.1534/genetics.103.021683. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12. Tibshirani R., “Regression Shrinkage and Selection via the Lasso,” Journal of the Royal Statistical Society: Series B: Methodological 58, no. 1 (1996): 267–288. [Google Scholar]
- 13. Chen M., Suzuki A., Thakkar S., Yu K., Hu C., and Tong W., “DILIrank: The Largest Reference Drug List Ranked by the Risk for Developing Drug‐Induced Liver Injury in Humans,” Drug Discovery Today 21, no. 4 (2016): 648–653, 10.1016/j.drudis.2016.02.015. [DOI] [PubMed] [Google Scholar]
- 14. Wilks S. S., “The Large‐Sample Distribution of the Likelihood Ratio for Testing Composite Hypotheses,” Annals of Mathematical Statistics 9, no. 1 (1938): 60–62, 10.1214/aoms/1177732360. [DOI] [Google Scholar]
- 15. Soetaert K. and Herman P. M., A Practical Guide to Ecological Modelling. Using R as a Simulation Platform (Dordrecht: Springer, 2009). [Google Scholar]
- 16. Friedman J., Hastie T., and Tibshirani R., “Regularization Paths for Generalized Linear Models via Coordinate Descent,” Journal of Statistical Software 33, no. 1 (2010): 1–22, 10.18637/jss.v033.i01. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17. Fan J. and Li R., “Variable Selection via Nonconcave Penalized Likelihood and its Oracle Properties,” Journal of the American Statistical Association 96, no. 456 (2001): 1348–1360, 10.1198/016214501753382273. [DOI] [Google Scholar]
- 18. Zou H. and Hastie T., “Regularization and Variable Selection via the Elastic net,” Journal of the Royal Statistical Society, Series B: Statistical Methodology 67, no. 2 (2005): 301–320, 10.1111/j.1467-9868.2005.00503.x. [DOI] [Google Scholar]
- 19. Su W., Bogdan M., and Candès E., “False Discoveries Occur Early on the Lasso Path,” Annals of Statistics 45, no. 5 (2017): 2133–2150. [Google Scholar]
- 20. Galwey N. W., “A New Measure of the Effective Number of Tests, A Practical Tool for Comparing Families of Non‐Independent Significance Tests,” Genetic Epidemiology 33, no. 7 (2009): 559–568, 10.1002/gepi.20408. [DOI] [PubMed] [Google Scholar]
- 21. Li M. X., Yeung J. M. Y., Cherny S. S., and Sham P. C., “Evaluating the Effective Numbers of Independent Tests and Significant p‐Value Thresholds in Commercial Genotyping Arrays and Public Imputation Reference Datasets,” Human Genetics 131, no. 5 (2012): 747, 10.1007/s00439-011-1118-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22. Bogdan M., Frommlet F., Biecek P., Cheng R., Ghosh J. K., and Doerge R. W., “Extending the Modified Bayesian Information Criterion (mBIC) to Dense Markers and Multiple Interval Mapping,” Biometrics 64, no. 4 (2008): 1162–1169, 10.1111/j.1541-0420.2008.00989.x. [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Supporting information.
Data S1. Supporting information.
Data Availability Statement
We share our simulation code in the supporting information. The pharmacovigilance data provided by the ANSM are not freely available for privacy reasons.
