Abstract
In order to solve large-scale lasso problems, screening algorithms have been developed that discard features with zero coefficients based on a computationally efficient screening rule. Most existing screening rules were developed from a spherical constraint and half-space constraints on a dual optimal solution. However, existing rules admit at most two half-space constraints due to the computational cost incurred by the half-spaces, even though additional constraints may be useful to discard more features. In this paper, we present AdaScreen, an adaptive lasso screening rule ensemble, which allows to combine any one sphere with multiple half-space constraints on a dual optimal solution. Thanks to geometrical considerations that lead to a simple closed form solution for AdaScreen, we can incorporate multiple half-space constraints at small computational cost. In our experiments, we show that AdaScreen with multiple half-space constraints simultaneously improves screening performance and speeds up lasso solvers.
Index Terms—: Lasso, Screening Rule, Ensemble
1. Introduction
In modern applications of machine learning, data sets with a large number of features are common. Examples include human genomes [3], gene expression measurements [6], spam data [28], and social media data [13]. Given such data, one of the fundamental challenges in machine learning is feature selection, that is, given input features (e.g., genetic variants) and an output variable to predict (e.g., disease status), select a subset of the input features that is relevant to the output. Among the many feature selection techniques, the lasso [22] has been very successful at solving the feature selection problem under the assumption that most features are irrelevant to the output [22]. However, solving the lasso problems on very high dimensional data (e.g., billions of features) still remains a computational challenge. In other words, either it takes too long to solve such big lasso problems, or the input data are too large to store in memory.
To address the computational problem, researchers have developed screening algorithms that can be used as a preprocessing step to efficiently discard features irrelevant to the output from the lasso model. The key idea is that after the screening step the input feature set is reduced, and we can use any lasso solver to solve the lasso problem on the reduced set. If the screening can be performed efficiently, and the resulting reduced feature set is small enough, we can obtain a lasso solution efficiently. Existing screening rules include the safe feature elimination (SAFE) rule [7], the sphere tests [31], the strong rule [23], the dome test [26], [29], efficient dual polytope projection (EDPP) rule [25], the two hyperplane test [27], the Sasvi rule [16], safe rules for the lasso [11], and GAP Safe screening rules for sparse multi-task and multi-class models [19]. Except for the strong rule, all screening rules are safe (i.e., features discarded by screening are guaranteed to be excluded in an optimal lasso solution). Recently, screening rules for graphical lasso [18] and L1 logistic regression have been developed [24]; further, Bonnefoy et al. [1] proposed a dynamic screening, where features are discarded in every iteration of an optimization algorithm if a certain condition is met. In this paper, we limit ourselves to lasso screening algorithms used before applying the optimization algorithms. A summary of existing lasso screening rules is shown in Table 1.
Table 1.
Summary of lasso screening algorithms with their sphere and half-space constraints (Seq. column shows whether sequential screening is supported or not).
| Algorithm | Sphere Const. (center) | Sphere Const. (radius) | Half-space Const. | Reference | Seq. |
|---|---|---|---|---|---|
| Sasvi | [16] | Yes | |||
| Two Hyperplane Test | adaptive | adaptive | adaptive up to two half-space const. | [30] | Yes |
| EDPP | (see Sec. 2.1) | None | [25] | Yes | |
| DOME | θ* (λmax) | [26], [29] | Yes | ||
| Strong rule | θ* (λ0) | None | [23] | Yes | |
| Sphere test | [31] | No | |||
| SAFE rule | θ* (λmax) | None | [7] | No | |
| AdaScreen | adaptive | adaptive | adaptive up to K half-space const. | ours | Yes |
We note that the screening algorithms described in this paper are different from sure-screening algorithms [8], [9], [10], which, instead of discarding features with zero coefficients in a global optimal solution of the lasso problem, aims at discarding features irrelevant to the output (features uninvolved in the true model). Also, consistency of the lasso model [33] is beyond the scope of this paper.
While screening techniques have been successful in enabling us to solve big lasso problems efficiently, they become ineffective in discarding irrelevant features as lasso optimal solutions become non-sparse. The reason for this degradation is that most screening methods rely on finding a region containing the optimal solution that is as small as possible. For a non-sparse optimal solution such a region typically is difficult to estimate, resulting in low screening efficiency. One can incorporate multiple half-space constraints into a screening rule to mitigate the problem, but for existing screening algorithms the advantage of this is substantially reduced due the increased amount of computation needed to evaluate the screening rule.
In this paper, we present an adaptive screening rule ensemble for the lasso, referred to as AdaScreen. We first propose an adaptive screening rule ensemble that can include any one sphere and multiple half-space constraints on a dual optimal solution (it is adaptive in the sense that any constraints can be chosen). Then we derive a closed-form solution for AdaScreen, which allows us to efficiently evaluate multiple half-space constraints. The contributions of this paper are as follows: (1) We provide an adaptive screening rule ensemble with an efficient closed-form solution; and (2) we provide several instances of AdaScreen with different choices of constraints on a dual optimal solution. We note that AdaScreen offers a framework to combine multiple screening rules, rather than a novel screening rule. In our experiments, we confirm that AdaScreen can incorporate any sphere and any multiple half-space constraints, each resulting in novel screening rules. Furthermore, across a range of datasets we show that AdaScreen with multiple half-spaces can reject significantly more features with zero coefficients than existing screening rules. Note that AdaScreen is the first screening framework that allows us to improve screening performance, taking advantage of multiple half-space constraints. While in principle any half-space constraints could be used to achieve a computational speed-up, AdaScreen leads to safe screening iff all constraints are derived from safe screening rules. We further experimentally confirmed that the proposed screening rules are safe.
Notations In the following, upper case letters denote matrices, boldface lowercase letters denote vectors, and italic lowercase letters denote scalars. Subscript denotes column index of matrices or element index of vectors. A superscript asterisk (*) denotes an optimal solution.
2. Background: Lasso Screening
The lasso [22] is defined as a linear regression from a set of J feature vectors to a target vector , where N is the sample size, with a squared loss and L1 norm penalty:
| (1) |
where is a design matrix, where each column represents a feature, is a regression coefficient vector, and the regularization parameter λ controls the sparsity of β. Lasso screening rules aim at discarding features with zero coefficients at a global optimal solution by computing simple, computationally inexpensive tests for each feature.
We can obtain a screening rule from a Lagrangian of the lasso problem in (1) [25], given by
| (2) |
where θ is a dual variable. By introducing an auxiliary variable z = y – Xβ, we can obtain (2) from (1) using the method of Lagrange multipliers [2]. Taking a subgradient of (2) with respect to βj, we get , where sj is a subgradient of the L1 norm, defined by: if βj ≠ 0, sj = sign(βj); otherwise, sj ∈ [−1, 1]. Based on this, we can derive a screening rule:
| (3) |
However, use of (3) is impractical because θ*(λ) is unknown. Instead, we estimate the region where θ*(λ) exists, denoted by θ*(λ) ∈ Θ, and then, use the following rule to identify zero coefficients:
| (4) |
From (4), we can see that screening rules can be developed by finding Θ and a solution for . Furthermore, screening efficiency (i.e., the proportion of rejected features to all features) increases as the size of Θ decreases, making the left-hand side in (4) smaller. Screening time increases as the time to solve increases; thus, a closed-form solution is desirable.
2.1. Screening via Dual Polytope Projection
So far, we have shown how screening rules are developed in general. In this section, we review dual polytope projection (DPP) screening [25]. Screening rules are derived from a dual form of the lasso (See [25]):
| (5) |
where θ is the dual variable. The optimal parameters in the primal (1) and the dual (5) are related via [25], where we have made the dependence on the regularization parameter explicit. Furthermore, in (5), a dual optimal solution θ* can be found by projecting onto the constraints . In other words, θ* is the vector closest to , which satisfies C because it minimizes the objective function in (5) while satisfying all the constraints. We represent this projection as , where the projection operator is defined as
In DPP, the Θ in (4) is obtained using the “nonexpansiveness” property of the projection operator PC (·) [17]:
| (6) |
Using (6), a range of θ*(λ) ∈ Θ can be estimated by
| (7) |
where we used the fact that and . Let us now assume that λ < λ0 and θ*(λ0) are given, which shall be discussed in Section 2.2.
This provides us a region that contains θ*(λ), that is, Θ = B(o,ρ), which represents a sphere centered at o with radius ρ, where o = θ*(λ0) and from (7). Further, using the spherical region Θ = B(o,ρ), θ*(λ) can be represented as:
| (8) |
where ∥v∥2 ≤ ρ. Using (3) and (8), we can derive a basic DPP lasso screening rule as follows:
| (9) |
| (10) |
where we used by the Cauchy-Schwarz inequality.
Different screening rules use differently estimated regions for Θ. Based on DPP, Wang et al. [25] developed enhanced DPP (EDPP) that achieves a smaller Θ than that of DPP via projections of rays and the “firmly nonexpansiveness property” of the projection operator PC(·). Specifically, EDPP uses a spherical region Θ = B(o,ρ), where
| (11) |
where , and v1(λ0) and v2(λ, λ0) are defined by
| (12) |
| (13) |
where . In comparison to DPP, EDPP uses smaller ρ, yielding more efficient screening. We refer readers to [25] for the details of EDPP.
2.2. One-shot and Sequential Screening
Here we describe how screening rules are used in two different scenarios: the first is one-shot screening, and the second is sequential screening. One-shot screening is for when we want to solve lasso problems for a single λ parameter. In sequential screening, we want to solve lasso problems for a given sequence of λ parameters in descending order, Λ = {λ1, …, λT}. Typically, λ1 = λmax, where λmax is the smallest λ that sets β*(λmax) = 0. In this setting, we discard features based on a screening rule for λt given and then run a lasso solver on the unscreened features, resulting in . Subsequently, based on , we keep iterating screening and the lasso solver until the lasso solution for desired λ is achieved. It is called sequential screening due to its sequential nature. Throughout this paper, we assume a sequential screening setting; thus we denote λ0 by the previous λ parameter (i.e., λ0 ≡ λt−1) and λ by the current one (i.e., λ ≡ λt). Note that one-shot screening can be easily obtained by setting λ0 ≡ λmax.
It is well known that screening efficiency degrades as the gap between λ and λ0 increases [25]. Thus, when desired λ is small, sequential screening is more appropriate than the one-shot version because it may include intermediate parameters between λmax and λ. However, even sequential screening becomes more ineffective as the lasso solution becomes dense. This phenomenon can be explained as follows. Suppose that we solve lasso problems with a sequence of geometrically spaced λ parameters, that is, λ = αλ0, where 0 < α < 1. Then for both DPP and EDPP, as λ0 decreases, ρ increases. For DTP, ; for EDPP, , where . Obviously, as the sphere radius of ρ increases, the screening efficiency degrades. The same argument can be applied to linearly or logarithmically spaced λ sequences. Note that other screening rules have the same issue because they also rely on the spherical region containing θ*(λ) with the radius of ρ (see Table 1).
3. AdaScreen: Adaptive Lasso Screening Rule Ensemble
In this section, to improve the screening efficiency for small λ parameters, we present AdaScreen, an adaptive form of lasso screening that can take any one sphere constraint, and any multiple half-space constraints on θ*(λ). Note that AdaScreen can also be viewed as a general lasso screening framework, which can be instantiated with any sphere and half-space constraints; the whole space of AdaScreen remains to be explored.
Recall that lasso screening rules are derived based on the estimation of the region Θ that includes a dual optimal solution θ*(λ). Given a sphere constraint ∥θ*(λ) − o∥ ≤ ρ, one can represent θ*(λ) = o + v, where o is a fixed center of a sphere, and v is a free vector with ∥v∥ ≤ ρ. Furthermore, K half-space constraints on θ*(λ) are represented by or , k = 1, … K. Now we propose AdaScreen:
| (14) |
where and are defined by
| (15) |
Here is a set of user-defined half-space constraints on θ*(λ). One can easily see that AdaScreen is derived as follows:
Note that and can be solved in the same way except that we replace xj by −xj. Thus, without loss of generality, we restrict the presentation to a closed-form solution for only.
AdaScreen has the following unique properties: first, one can incorporate any sphere or half-space constraints into (15), resulting in a new lasso screening rule. Second, AdaScreen directly estimates (or ) for achieving a better bound without resorting to only the Cauchy Schwarz inequality. Both help us achieve a better screening rule. The former makes Θ smaller by allowing us to add more constraints; the latter gives us a smaller feasible region for a half-space constraint considering the angle between xj and v. We start with a closed-form solution for AdaScreen with a single half-space constraint in and a single sphere constraint. Then, we will extend it to the case for multiple half-space constraints, followed by the discussion on the choice of the constraints. AdaScreen is summarized in Algorithm 2.
3.1. Closed-form Solution with One Sphere and One Half-space Constraint
To efficiently compute , we first derive a closed-form solution given one sphere and one half-space constraint. In fact, Sasvi [16] derived a closed-form solution for the same problem when one sphere and one half-space come from variational inequality constraints. Here we seek a different closed-form solution based on geometry, which satisfies the following requirements. First, we need a general solution that allows any sphere and any half-space constraints; second, we need a solution that admits additional half-space constraints with negligible computational overhead.
In , is a constant, thus solving boils down to optimizing the following problem:
| (16) |
Here we assume that (16) includes a non-empty feasible region. A sphere constraint can be provided by any existing screening methods such as EDPP, and the candidate half-space constraints shall be discussed in Section 3.3.
One can view (16) as the problem of finding a hyperplane with the maximum of z such that two constraints in (16) are satisfied. Since is linear in v, its maximum can be obtained in the following two extreme cases: (1) is tangent to the sphere ∥v∥2 ≤ ρ at the contact; (2) meets the intersection between the half-space and the sphere. Below, we show that for all cases, a solution is given by the following closed-form:
| (17) |
where ρ′ is a constant, determined by each case.
We define , where α is the angle between xj and ak, and , i.e., the distance between and the origin. Hereafter, let us consider the case when 0 ≤ α ≤ π (if π < α ≤ 2π, the same derivation can be applied by setting α ← 2π − α).
In the first case, meets the sphere constraint at its boundary, and thus we disregard the half-space constraint. Using Cauchy Schwarz inequality, the maximum of is given by . Thus, ρ′ = ρ. The geometry in Figure 1(a) shows that if , we use this case. Note that if , ; thus, for the next case, the range of α should be .
Fig. 1.
Graphical illustration for three different cases, where z is maximized in . (a): the case where is tangent to the sphere at its contact, and meets the half-space constraint; (b): the case where , and the hyperplane meets the intersection between the sphere and the half-space constraint; (c): the case where , and the hyperplane meets the intersection between the sphere and the half-space constraint.

In the second case, meets the intersection between the half-space and the sphere constraint. We define , which is non-negative because for this case. Let us consider the following two sub-cases.
- Sub-case with . The geometry of this case is depicted in Figure 1(b), where v′ is the vector on the hyperplane with the direction and the length ρ′, given by
Plugging into , we get the maximum of the objective in (17).(18) - Sub-case with . The geometry of this case is depicted in Figure 1(c), where v′ is the vector on the hyperplane with the direction and the length ρ′, given by
The maximum z is obtained by plugging ρ′ into (17).(19)
Summarizing the above results, we have the following condition for ρ′:
| (20) |
3.2. AdaScreen with Multiple Half-space Constraints
Solving (16), we showed how to compute the maximum of with one sphere and one half-space constraint. However, we want to use multiple half-space constraints, denoted by . Here we show that an upper bound on with multiple half-space constraints can be found by solving (16) multiple times ( times) and taking the minimum among them, as follows:
| (21) |
| (22) |

In Algorithm 1, we summarize a closed-form solution for with a sphere and multiple half-space constraints. We note that AdaScreen is safe given a sphere and half-space constraints from safe screening rules, as shown below.
Proposition 1. AdaScreen never discards non-zeros in β*(λ), given sphere and half-space constraints obtained from safe screening rules.
Proof 1. Suppose that we are given sphere and half-space constraints from safe screening rules. From (14), a safe lasso screening rule is if , where . Here, we show that AdaScreen computes an upper bound on , leading to a safe screening rule. can be shown in a similar fashion. Since the first term is a constant, let us consider the second term. When a single half-space constraint is given, AdaScreen takes a closed-form solution for the second term, given by (17); when multiple half-space constraints are given, AdaScreen takes an upper bound on , given by (22). Therefore, AdaScreen takes an upper bound on , and it is a safe lasso screening rule.
3.3. Selection of Half-space Constraints
So far, we presented how to solve AdaScreen, assuming that one sphere and K half-space constraints are given by users. Here we discuss how to select half-space constraints that may increase the screening efficiency.
The following set of half-space constraints can be useful to increase the screening efficiency:
| (23) |
| (24) |
The set of constraints in (23) is originated from dual lasso constraints; (24) stems from variational inequality [16], proven to be useful. We propose that useful global and half-space constraints can be selected in the pool of constraints in (23) and (24).
Here we claim that the half-spaces in (23) can be useful for screening. To achieve a small ρ′ < ρ in (20), we need , and thus small r is desirable. Let us denote θ*(λ) = o + v* (v* is a fixed unknown vector). By dual lasso constraints, , ∀k. Thus, and , ∀k, that lead to the following: , where equality holds when . Thus, for each feature, r is smallest when . Furthermore, Proposition 2 shows that can never be discarded by an independent screening rule if ; thus, (23) give us nontrivial constraints.
Proposition 2. Suppose that we are given λ0 and λ (λ0 > λ). For all j = 1, …, J, if , then xj can never be discarded by any independent screening rules with λ.
Proof 2. We assume that , that is,
| (25) |
Suppose that there exists an independent screening rule which can discard xk at λ, that is,
| (26) |
Note that a screening rule estimates the region of θ*(λ) based on θ*(λ0) as follows: θ*(λ) = θ*(λ0) + v, ∥v∥ ≤ ρ. By plugging it into (26), we get
| (27) |
| (28) |
However, (28) contradicts to the assumption (25), and thus the proof is completed.
For half-space constraints, one may use (24) by setting λ0 = λmax and/or a subset of constraints in (23) using the following procedure. Following the λ sequence, we start solving a lasso problem using screening with only one sphere constraint (lasso can be solved efficiently with large λ). When the number of nonzero coefficients is larger than a user-defined threshold at λt, we generate half-space constraints based on (23). From λt+1, we use the half-space constraints for screening.
We note that combination of (23) and (24) generates a feasible set for θ*(λ) because the former is a feasible set for θ*(λ) by the dual form of the lasso (5) and it has been proven that the latter is a feasible set for θ*(λ) [16]. Therefore, AdaScreen is safe if combination of (23) and (24) and a sphere constraint from a safe screening rule are used.
3.4. Analysis of Screening Time Complexity
The half-space constraints can be divided into global and local constraints. The global half-space constraints (Hglobal) refer to the half-space constraints which are used for all features and all λ parameters. The local half-space constraints (Hlocal) are generated for each λ, but applicable to all features.
Here we analyze the screening time complexity under two different scenarios: AdaScreen with (1) K global half-space constraints, and (2) K local half-space constraints. Let us first consider the screening complexity of AdaScreen when K global half-spaces are used. In such a case, the time complexity of AdaScreen is , where stems from a sphere constraint, and stems from K half-space constraints. Thus, if K << ∣Λ∣ and K << N, K half-space constraints barely affects the screening speed. When using K local half-space constraints, the time complexity of AdaScreen is . Therefore, approximately, use of local half-space constraints is K times more expensive than use of global half-space constraints. For sparse data, depending on screening rules, the time complexity can be reduced. For example, for each feature vector, given DPP and a local half-space in Eq. (23), the time complexity of AdaScreen is O(max(N′, Q′)K ∣Λ∣), where N′ is the number of nonzeros in the feature vector of interest and Q′ is the number of nonzeros in the normal vector of the local half-space constraint. In Algorithm 2, we summarize AdaScreen with one sphere constraint, L global half-space constraints, and K − L local half-space constraints. In practice, a small number of local half-space constraints (e.g., 5 half-spaces) is sufficient to significantly benefit from the constraints.
4. Experiments
We systematically evaluate the screening performance in simulations and several real world datasets. To get a broad comparison of existing screening rules, we chose DPP [25], DOME [29], SAFE [7], strong rule [23], EDPP [25], and Sasvi [16], covering both classical as well as recent methods. For the classical ones including DOME and SAFE rules, we used one-shot screening, while for the others, we used a sequential screening setting. We settled for three instances of AdaScreen that incorporate constraints from the advanced techniques EDPP and Sasvi. Each instance uses EDPP sphere constraint with (a) Sasvi half-space in (24) as global half-space constraint (EDPP+Sasvi); (b) 100 half-spaces in (23) as local half-space constraints (EDPP + DualLasso(100)); and (c) all half-space constraints in (a) and (b) (EDPP+Sasvi+DualLasso(100)).
We evaluated screening performances based on the feature rejection ratio, defined by
and the runtime used to solve the lasso problems over a sequence of λ parameters. Throughout the experiments, we use the standard coordinate descent algorithm [12], implemented in scikit-learn [20], with tolerance of 10−4. We use a geometrically spaced sequence of 65 λ values and step length 0.9, starting from λmax to ~ 10−3λmax. Experiments were run on a single machine with 48 2.1GHz AMD cores and 384GB RAM.
4.1. Simulation Study
We first investigate screening performance over different degrees of feature correlations in the simulated data. Note that as more features are correlated with each other, the lasso takes longer to converge. This can be seen by inspecition of the coordinate descent update rule [12]: , where S(βj, λ) ≡ sign(β) (∣β∣ − λ). If , ∀j ≠ k (i.e., all features are completely independent), the update rule becomes , and then the optimal j-th coefficient can be found by a single update. When features are highly correlated, multiple iterations are needed to converge the lasso objective because the update of one coefficient affects optimal values of other coefficients. We generated simulation data with 250 samples and 10,000 dimensions as follows: x1 ← r and xj ← cxj−1 + (1 − c)r for j ≥ 2, where is a random vector drawn from a uniform distribution unif(0, 1), and c represents the degree of feature correlations in the simulated data.
4.1.1. Effects of Feature Correlations on Screening
Figure 2 shows the rejection ratio and runtime of the screening methods on simulated datasets with the degree of feature correlation (c) set to 0.1, 0.5 and 0.9. Here we did not have the case c = 0 (all features are independent from each other) because it is a trivial case for lasso optimization and existing sequential screening methods (e.g., EDPP) and AdaScreen perform equally well. In all settings, AdaScreen with local half-space constraints achieved better rejection ratio and runtime than all the other methods. Notably, the performance gap between AdaScreen with local half-spaces and others was more substantial when the degree of feature correlations was high (e.g., c=0.5 & c=0.9); further, AdaScreen’s rejection ratio was barely affected by the degree of feature correlations. It demonstrates that when features are highly correlated, local half-spaces effectively reduce the region Θ, resulting in high rejection ratio. We also observed that recent methods (AdaScreen, Sasvi, EDPP) dramatically outperformed the classical methods (DPP, DOME, SAFE). In Figure 2(second and third column), AdaScreen’s rejection ratio dropped when λ = 0.9λmax decreased from λmax because there exist no local half-spaces to be used (i.e., (23) is the empty set because ); after that, the rejection ratio increased taking advantage of local half-spaces.
Fig. 2.
Rejection ratio and runtime of various screening methods on simulated datasets with the feature correlations of 0.1 (first column), 0.5 (second column), and 0.9 (third column) (see text for details) given a sequence of λ parameters. Two AdaScreen instances with 100 dual lasso half-space constraints outperformed the other methods in terms of both rejection ratio and runtime.
4.1.2. Effects of Local Half-space Constraints on Screening
Figure 3 (third column) shows the impact of increasing the number of local half-space constraints in (23), in terms of the rejection ratio. As expected, the more half-space constraints are taken into account, the higher the rejection ratios. Interestingly, AdaScreen with five local half-spaces showed substantially better rejection ratio than AdaScreen with one local half-space, and the benefits were quickly saturated when we allowed ≥ 5 local half-space constraints.
Fig. 3.
Rejection ratio (first row) and runtime (second row) for comparison between AdaScreen and BagScreen (first column); comparison between AdaScreen and strong rule (second column); and demonstration of the effects of local half-spaces on screening (third column) on simulated datasets with the feature correlation of 0.5 given a sequence of λ parameters.
4.1.3. Comparison between AdaScreen and Trivial Combination of Screening Rules (BagScreen)
Here we show that AdaScreen provides a non-trivial way of combining multiple constraints for screening. To this end, let us compare AdaScreen with BagScreen. BagScreen simply tests multiple screening rules, and then discards the j-th feature if any of them discards it. To ensure fair comparison, for AdaScreen, we used the EDPP spherical constraint, and Sasvi and DOME half-spaces as global half-space constraints; for BagScreen, we combined EDPP, Sasvi, and DOME rules. Figure 3 (first column) shows the rejection ratios by AdaScreen and BagScreen given the simulated data with feature correlation c = 0.5 and step length 0.85 for the geometric sequence. AdaScreen maintained > 0.1 rejection ratios throughout all λs, but BagScreen discarded no features for most λs. It is because AdaScreen generated an efficient screening rule with a sphere and multiple half-space constraints, while BagScreen is a set of weaker screening rules with a sphere and a half-space constraint. Here high rejection ratios were unachievable by AdaScreen because we did not use local half-spaces. The experimental results clearly demonstrate that AdaScreen combines multiple constraints in a non-trivial way, resulting in high screening performance.
4.1.4. Comparison between AdaScreen and Strong Rule
The strong rule has been developed as a screening algorithm for lasso-type problems such as lasso and L1 logistic regression. It is developed based on an assumption that given λ and λ0, correlation between xj and the residual (i.e., ) does not change more than ∣λ − λ0∣. Therefore, the strong rule is not safe, meaning that it may discard features whose coefficients are nonzero in a global optimal solution. In Figure 3 (second column) we compare the rejection ratios by strong rule, AdaScreen, Sasvi, and EDPP in the simulated data with the feature correlation c = 0.5 for 1,000 samples; we used step length 0.75 for the geometric sequence. Notably, AdaScreen substantially outperformed the strong rule, and the strong rule was significantly better than Sasvi and EDPP. This result shows that with multiple half-spaces, we can obtain safe screening rules that are even superior to the strong rule.
4.2. Experiments on Real Datasets
We performed experiments on real world datasets including PEMS [4], [15], Alzheimer’s disease [32] and PIE [21] datasets. The PEMS input data contains 440 samples (daily records) and 138672 features (963 sensors by 144 timestamps) describing the occupancy rate of multiple car lanes in San Francisco; the PEMS output data consists of the day of the week. The input from the Alzheimer’s disease dataset contains 540 samples (270 disease and 270 healthy individuals) and 511997 features (genetic variants); its output data contains the expression levels of a randomly selected gene. The PIE dataset (11554 examples and 1024 features) is a face recognition dataset that contains 11554 gray face images of 68 people under various conditions and expressions. For the output of PIE dataset, we randomly chose one feature in PIE images, and then generated input data by concatenating all features except the selected feature for output. We note that PEMS and Alzheimer’s disease are large-scale datasets for a single machine setting; for example, it took more than 3 hours for coordinate descent lasso solver implemented in scikit-learn with the DPP screening rule to solve the lasso problem on the full Alzheimer’s disease dataset.
4.2.1. Screening Efficiency and Runtimes against Baseline Competitors
Figure 4 shows our main result on real-world data sets, namely rejection ratio and runtime by AdaScreen, Sasvi, EDPP, DPP, DOME, and SAFE rules on PEMS, PIE, and Alzheimer’s disease datasets. For PEMS dataset, AdaScreen with 100 local half-space constraints maintained very high rejection ratios (> 0.9) throughout all λ parameters, and it also showed significantly better runtime than all the other methods. For Alzheimer’s disease and PIE datasets, AdaScreen with local half-spaces also showed the best screening performance for all λ parameters. These results confirm that multiple local half-spaces employed by AdaScreen are truly useful to improve the rejection ratio at a low computational cost; as a result, we achieved a speedup in runtime compared to the screening rules without local half-space constraints.
Fig. 4.
Rejection ratio and runtime on PEMS (first column), Alzheimer’s disease (second column), and PIE image (third column) datasets by three instances of AdaScreen, Sasvi, EDPP, DPP, DOME, and SAFE rules given a sequence of λ parameters.
4.2.2. Speed-up Comparison over Path-solver without AdaScreen
So far, experiments showed screening rejection ratios and accumulated runtime behavior. While runtime behavior is a very informative measure to know exactly how much time it takes to get to this λ, it obfuscates the driving reason on where the optimization benefits the most from screening, i.e. where the difference in computation time is the highest. Fig. 5 shows the expected accumulated speedup when comparing the path-solver with screening against the path-solver without screening. As can be seen in the figures, all three experiments gain the most in the beginning of the λ sequence. This is partially due to high screening ratios as well as longer distances between consecutive λ’s where the path solver needs more iterations to reach a sufficient optimum. Experiments have been repeated 10 times and mean speedup values are reported.
Fig. 5.
Speed-up comparison of a variety of screening algorithms on PEMS (left), Alzheimer’s disease (center), and PIE image (right) dataset.
4.2.3. Accuracy Assessment and Impact of Regularization
We are interested in solving a specific problem and hence, interested in reaching a specific λ for which the problem is solved sufficiently accurate. To assess the accuracy of lasso with λs tested in our experiments, we divided the data into training (80%) and test sets (remaining 20%). Fig. 6 shows the accuracy on the test sets in mean squared error (MSE) for the three real-world datasets PIE, PEMS, and Alzheimer. The experiment was repeated 10 times and mean accuracies are reported. For PEMS and PIE, accuracies seem to reach a plateau for decreasing λ, as more and more features get activated. For Alzheimer on the other hand, a minimum error is reached for a low number of features early on in the λ sequence. It demonstrates that the range of λ parameters tested includes practically useful λ. We also note that in applications for feature selection [14], lasso solution is useful even when the test error is not minimal because the goal is to find a small feature set associated with outputs.
Fig. 6.
Accuracy in mean squared error (MSE) for varying regularization parameter λ on PEMS (left), Alzheimer’s disease (center), and PIE image (right) dataset.
4.2.4. Solver Comparison
Fig. 7 shows the speedup performance on PIE dataset comparing various path-solver using AdaScreen against the corresponding path-solver without screening. We chose 5 distinct solvers demonstrating the usefulness of our approach among different types of optimization techniques. Among those is our standard method used throughout the experiments, a coordinate descent solver (implementation used from scikit-learn), a LARS solver, a coordinate descent solver with active set selection (based on scikit-learn coordinate descent solver), proximal gradient descent solver, and accelerated proximal gradient descent solver. We repeated the experiments 10 times and report the means of the measured speedup. Fig. 7 demonstrates that AdaScreen is able to achieve significant speed-up (> 100×) for a wide range of λs with various solvers with different optimization algorithms.
Fig. 7.
The speed-up for various solvers using AdaScreen when compared against the corresponding solver without screening.

5. Software & Practical Considerations
Previous sections dealt with the principal properties of AdaScreen for ensemble screening. In this section, we will introduce our Python software package which is easily extendable for other solvers, screening rules, and settings.
5.1. Implementation
We developed our screening framework in the Python programming language and it can be freely downloaded at http://nicococo.github.io/AdaScreen/ or conveniently downloaded and installed automatically, using the Python pip command. An UML (Unified Modeling Language) diagram of our implementation is shown in Figure 8. It is designed to efficiently implement various screening rules without changing the lasso path solver (e.g., scikit-learn lasso solver [20], cf. Table 2). Even though different screening rules require different constraints and equations, they all share common data structures; thus, we wrap all of them into a single framework. An advantage of this approach is that the lasso path solver needs to interact with only one abstract class for screening rules.
Fig. 8.
UML diagram of our screening implementation: the modular design allows us to easily implement screening rules by integrating any sphere and any multiple half-space constraints.
Table 2.
List of screening rules and their properties implemented in our screening software package.
| Name | Sequential | One-shot | Safe | Strong | Ensemble |
|---|---|---|---|---|---|
| DPP | x | x | x | - | - |
| EDPP | x | x | x | - | - |
| Sasvi | x | x | x | - | - |
| DOME | x | x | x | - | - |
| ST3 | x | x | x | - | - |
| SAFE | - | x | x | - | - |
| Seq-SAFE | x | x | x | - | - |
| Strong Rule | x | x | - | x | - |
| This Work | |||||
| AdaScreen | x | x | x | x | x |
To systematically manage data structures involved in screening, we divide them into Globals and Locals. Globals refer to variables that do not change over the lambda path (e.g., the inputs X, y, λmax). In contrast, Locals refer to variables that change over the lambda path (e.g., the previous regularization parameter λ0 or the solution β*(λ0) from the previous λ0).
Furthermore, we designed our screening framework in such a way that all screening rules can be derived from the abstract base class. Screening rules with a single sphere constraint can be implemented by overloading the getSphere function. For more advanced rules, corresponding functions need to be overloaded. For example, to implement AdaScreen with EDPP sphere constraint and Sasvi local half-space constraint, we first instantiate EDPP, Sasvi, and AdaScreen. Then in AdaScreen, we simply call setSphereRule(EDPP) and addHalfspace(Sasvi).
6. Discussions
We presented an adaptive lasso screening rule ensemble, AdaScreen, which can include any sphere and multiple half-space constraints. AdaScreen takes advantage of multiple half-spaces based on a simple, computationally efficient closed-form solution. We experimentally validated that AdaScreen benefits from multiple half-spaces and compared the rejection ratio and the runtime performance against a set of state-of-the-art competitors as well as a naïve implementation of the screening ensemble (BagScreen). Further, we provide a Python software package, which includes various screening methods, solvers, and settings.
Some datasets consist of a large number of categorical or binary features with only few entries set to non-zero values. In those cases, it would suffice to use single sphere constraint based screening rules (e.g. EDPP). Naturally, those datasets only maintain a weak correlation structure with features or examples being orthogonal (i.e. 〈xi, xj〉 = 0, i ≠ j) frequently. Figure 2 shows that the gain of using AdaScreen with multiple halfspace constraints decreases when the correlation structure is less prominent.
It is worthwhile to mention that we considered lasso problems with a sequence of λ parameters. However, we often solve a lasso problem with a fixed λ. To this end, data adaptive sequential screening (DASS) has been developed whose key idea is to choose a feedback-controlled sequence of λ parameters to reach λ, which attempts to increase screening efficiency across a λ sequence. A promising direction of future research for AdaScreen would be to use a λ sequence suggested by DASS, rather than a fixed one.
Furthermore, one interesting direction of research is to develop distributed AdaScreen with a parallel lasso algorithm to solve very large problems. For such a parallel screening, MapReduce [5] would be an appropriate framework because screening rules are embarrassingly parallel. Furthermore, extensions of AdaScreen that deal with different loss functions such as logistic loss or hinge loss would be an interesting research direction. We are also interested in incorporating AdaScreen in the lasso optimization procedure, along the lines of Bonnefoy et al. [1] to further improve rejection ratio and runtime.
7. Acknowledgments
Part of the work was done while SL, NG, and CL were with Microsoft Research, Los Angeles. NG was supported by BMBF ALICE II grant 01IB15001B, and EPX was supported by NIH R01GM114311 and NIH P30DA035778. Seunghak Lee, Nico Görnitz, and Christoph Lippert contributed equally to this work.
Biographies

Seunghak Lee is a Data Scientist at Human Longevity, Inc (HLI) in Mountain View, CA. His research interests include machine learning and computational biology with a focus on integrative approaches to the analysis of genetic and biomedical datasets, genome-wide association studies, distributed optimization, and large-scale machine learning algorithms and systems. Prior to HLI, he was a project scientist in Machine Learning Department, Carnegie Mellon University. He has received his Ph.D. in Computer Science from Carnegie Mellon University in 2015.

Nico Görnitz is a research associate in the machine learning group at the TU Berlin (Berlin Institute of Technology, Germany) headed by Klaus-Robert Müller. He is interested in machine learning in general and in one-class learning based anomaly detection for data with dependency structure, large-margin structured output learning and corresponding optimization techniques in specific. Applications that Nico has been working on cover computational biology, computer security, computational sustainability, brain-computer-interfaces, natural language processing and porosity estimation for geosciences.

Eric P. Xing is a Professor of Machine Learning in the School of Computer Science at Carnegie Mellon University, and the director of the CMU Center for Machine Learning and Health. His principal research interests lie in the development of machine learning and statistical methodology; especially for solving problems involving automated learning, reasoning, and decision-making in high-dimensional, multimodal, and dynamic possible worlds in social and biological systems. Professor Xing received a Ph.D. in Molecular Biology from Rutgers University, and another Ph.D. in Computer Science from UC Berkeley. His current work involves, 1) foundations of statistical learning, including theory and algorithms for estimating time/space varying-coefficient models, sparse structured input/output models, and nonparametric Bayesian models; 2) framework for parallel machine learning on big data with big model in distributed systems or in the cloud; 3) computational and statistical analysis of gene regulation, genetic variation, and disease associations; and 4) application of machine learning in social networks, natural language processing, and computer vision. He is an associate editor of the Annals of Applied Statistics (AOAS), the Journal of American Statistical Association (JASA), the IEEE Transaction of Pattern Analysis and Machine Intelligence (PAMI), the PLoS Journal of Computational Biology, and an Action Editor of the Machine Learning Journal (MLJ), the Journal of Machine Learning Research (JMLR). He is a member of the DARPA Information Science and Technology (ISAT) Advisory Group, and a Program Chair of ICML 2014.

David Heckerman is a Distinguished Scientist and Director of Microsoft Genomics at Microsoft. In his current scientific work, he is developing machine-learning and statistical approaches for biological and medical applications including genomics (see https://github.com/microsoftgenomics) and HIV vaccine design. In his early work, he demonstrated the importance of probability theory in Artificial Intelligence, and developed methods to learn graphical models from data, including methods for causal discovery. At Microsoft, he has developed numerous applications including the junk-mail filters in Outlook, Exchange, and Hotmail, machine-learning tools in SQL Server and Commerce Server, handwriting recognition in the Tablet PC, text mining software in Sharepoint Portal Server, troubleshooters in Windows, and the Answer Wizard in Office. David received his Ph.D. (1990) and M.D. (1992) from Stanford University, and is an ACM and AAAI Fellow.

Christoph Lippert is a Data Scientist at Human Longevity, Inc (HLI) in Mountain View, CA. It is a focus of his research to provide statistical and computational methodology for advanced genetic analyses of heritable traits and diseases. His research spans the fields of Machine Learning, Statistical Genetics and Bioinformatics. Before joining HLI, he has held research positions at Microsoft Research, the Max Planck Institute (MPI) for Developmental Biology, the MPI for Intelligent Systems, and Siemens Corporate Technology. He has received his Ph.D. in Bioinformatics from University of Tubingen, Germany in 2014.
Contributor Information
Seunghak Lee, Human Longevity Inc., Mountain View, CA 94041, USA..
Nico Görnitz, Machine Learning Group, Department of Software Engineering and Theoretical Computer Science, Berlin Institute of Technology, Berlin, 10578, Germany..
Eric P. Xing, School of Computer Science, Carnegie Mellon University, Pittsburgh, PA 15213, USA.
David Heckerman, Microsoft Research, Los Angeles, CA 90024, USA..
Christoph Lippert, Human Longevity Inc., Mountain View, CA 94041, USA..
References
- [1].Bonnefoy A, Emiya V, Ralaivola L, Rémi G, et al. A dynamic screening test principle for the lasso. In European Signal Processing Conference, 2014. [Google Scholar]
- [2].Boyd S and Vandenberghe L. Convex optimization. Cambridge university press, 2004. [Google Scholar]
- [3].Collins FS, Morgan M, and Patrinos A. The human genome project: lessons from large-scale biology. Science, 300(5617):286–290, 2003. [DOI] [PubMed] [Google Scholar]
- [4].Cuturi M. Fast global alignment kernels. In International Conference on Machine Learning, pages 929–936, 2011. [Google Scholar]
- [5].Dean J and Ghemawat S. MapReduce: simplified data processing on large clusters. Communications of the ACM, 51(1):107–113, 2008. [Google Scholar]
- [6].Dixon AL et al. A genome-wide association study of global gene expression. Nature genetics, 39(10):1202–1207, 2007. [DOI] [PubMed] [Google Scholar]
- [7].El Ghaoui L, Viallon V, and Rabbani T. Safe feature elimination in sparse supervised learning. Pacific Journal of Optimization, 8(4):667–698, 2012. [Google Scholar]
- [8].Fan J, Feng Y, and Song R. Nonparametric independence screening in sparse ultra-high-dimensional additive models. Journal of the American Statistical Association, 106(494), 2011. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [9].Fan J and Lv J. Sure independence screening for ultrahigh dimensional feature space. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 70(5):849–911, 2008. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [10].Fan J, Song R, et al. Sure independence screening in generalized linear models with NP-dimensionality. The Annals of Statistics, 38(6):3567–3604, 2010. [Google Scholar]
- [11].Fercoq O, Gramfort A, and Salmon J. Mind the duality gap: safer rules for the lasso. In International Conference on Machine Learning, 2015. [Google Scholar]
- [12].Friedman J, Hastie T, Hofling H, and Tibshirani R. Pathwise coordinate optimization. Annals of Applied Statistics, 1(2):302–332, 2007. [Google Scholar]
- [13].Lang K. Newsweeder: Learning to filter netnews. In International Conference on Machine Learning, 1995. [Google Scholar]
- [14].Lee S, Kong S, and Xing EP. A network-driven approach for genome-wide association mapping. Bioinformatics, 32(12):i164–i173, 2016. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [15].Lichman M. UCI machine learning repository, 2013. [Google Scholar]
- [16].Liu J, Zhao Z, Wang J, and Ye J. Safe screening with variational inequalities and its applicaiton to lasso. In International Conference on Machine Learning, 2014. [Google Scholar]
- [17].Luenberger DG. Optimization by vector space methods. John Wiley & Sons, 1969. [Google Scholar]
- [18].Mazumder R and Hastie T. Exact covariance thresholding into connected components for large-scale graphical lasso. The Journal of Machine Learning Research, 13(1):781–794, 2012. [PMC free article] [PubMed] [Google Scholar]
- [19].Ndiaye E, Fercoq O, Gramfort A, and Salmon J. Gap safe screening rules for sparse multi-task and multi-class models. In Advances in Neural Information Processing Systems, pages 811–819, 2015. [Google Scholar]
- [20].Pedregosa F et al. Scikit-learn: Machine learning in python. The Journal of Machine Learning Research, 12:2825–2830, 2011. [Google Scholar]
- [21].Sim T, Baker S, and Bsat M. The CMU pose, illumination, and expression (PIE) database. In IEEE International Conference on Automatic Face and Gesture Recognition, pages 46–51. IEEE, 2002. [Google Scholar]
- [22].Tibshirani R. Regression shrinkage and selection via the lasso. Journal of the Royal Statistical Society. Series B (Methodological), 58(1):267–288, 1996. [Google Scholar]
- [23].Tibshirani R, Bien J, Friedman J, Hastie T, Simon N, Taylor J, and Tibshirani RJ. Strong rules for discarding predictors in lasso-type problems. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 74(2):245–266, 2012. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [24].Wang J, Zhou J, Liu J, Wonka P, and Ye J. A safe screening rule for sparse logistic regression. In Advances in Neural Information Processing Systems, 2014. [Google Scholar]
- [25].Wang J, Zhou J, Wonka P, and Ye J. Lasso screening rules via dual polytope projection. In Advances in Neural Information Processing Systems, 2013. [Google Scholar]
- [26].Wang Y, Xiang ZJ, and Ramadge PJ. Lasso screening with a small regularization parameter. In IEEE International Conference on Acoustics, Speech and Signal Processing, pages 3342–3346. IEEE, 2013. [Google Scholar]
- [27].Wang Y, Xiang ZJ, and Ramadge PJ. Tradeoffs in improved screening of lasso problems. In IEEE International Conference on Acoustics, Speech and Signal Processing, pages 3297–3301. IEEE, 2013. [Google Scholar]
- [28].Webb S, Caverlee J, and Pu C. Introducing the webb spam corpus: Using email spam to identify web spam automatically. In Conference on Email and Anti-Spam, 2006. [Google Scholar]
- [29].Xiang ZJ and Ramadge PJ. Fast lasso screening tests based on correlations. In IEEE International Conference on Acoustics, Speech and Signal Processing, pages 2137–2140. IEEE, 2012. [Google Scholar]
- [30].Xiang ZJ, Wang Y, and Ramadge PJ. Screening tests for lasso problems. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2016. [DOI] [PubMed] [Google Scholar]
- [31].Xiang ZJ, Xu H, and Ramadge PJ. Learning sparse representations of high dimensional data on large scale dictionaries. In Advances in Neural Information Processing Systems, 2011. [Google Scholar]
- [32].Zhang B et al. Integrated systems approach identifies genetic nodes and networks in late-onset Alzheimers disease. Cell, 153(3):707–720, 2013. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [33].Zhao P and Yu B. On model selection consistency of lasso. The Journal of Machine Learning Research, 7:2541–2563, 2006. [Google Scholar]








