ABSTRACT
In this paper, we propose a flexible Bayesian inference to identify significantly correlated high‐dimensional functions with the response variable, which is challenging because the relationship between the response variable and high‐dimensional functions is unknown and complex due to the dependence among high‐dimensional functions. For example, in genetics pathway‐based analysis, a pathway is a set of genes that serve a particular cellular or physiological function. A pathway is a high‐dimensional function of genes. A pathway‐based analysis can detect subtle changes in expression levels that are not detectable using a gene‐based analysis. However, these pathways are not independent of each other. Because the clinical outcome is affected by multiple pathway sets, it is inappropriate to model sets using marginal analysis, such as a single‐pathway analysis. Estimating set effects based on a single set ignores the fact that sets interact with each other and, thus, result in false positives or false negatives. In this paper, we propose a generalized fused kernel machine regression to test significantly correlated high‐dimensional functions with the response variable, which can be either continuous or binary variables. We develop a data‐driven, flexible Bayesian inference for adjusting multiple tests using the Bayes factor that accommodates dependence through a simple yet flexible structure. The benefits of our method are illustrated through a simulation study and our motivating data on genetic pathway analysis related to type II diabetes.
Keywords: Bayes factor, fused model, kernel machine regression, multiple testing
1. Introduction
Analyzing correlated high‐dimensional data is a challenging problem in genomics, proteomics, and other related areas. For example, it is important to identify significant genetic pathway effects associated with biomarkers or disease status. A gene pathway is a set of genes that functionally work together to regulate a certain biological process. Pathway‐based analysis can consider the dependency structures among genes and the possibility that several moderately regulated genes may significantly impact the clinical outcomes. As pathways are sets of genes that serve a particular cellular or physiological function, we refer to a pathway as a set and a gene as an element. Pathway‐based analyses have attracted extensive interest [1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12] because it is possible to explore multiple pathways with profound impact on human biology and drug development. However, it is challenging to quantify the overall set effect and test which sets are highly associated with the outcome among several multiple pathways due to unknown dependence structures between pathways.
One possible way to estimate the overall set effect is via the kernel machine method, as it is a powerful nonparametric statistical learning model for learning unknown function spaces, especially for high‐dimensional data. A number of methods have been developed for testing the overall set effect [4, 13, 14]. Liu et al. [13] proposed a flexible framework by connecting the kernel machine with a linear fixed model that simultaneously estimates fixed effects and set effects. Liu et al. [13] and Pang et al. [8] developed score tests to identify the overall set effects. This method was extended to predict disease in survival analyses [15, 16].
On the other hand, Stingo et al. [17] and Cheng et al. [2] developed a method to test set effects under the Bayesian framework. Kim et al. [6, 7] considered Bayes factor‐based and resampling‐based methods to identify significant sets. Fang et al. [3, 18] further developed a kernel machine‐based approach for evaluating interactions between pathways and covariates. Fang et al. [18] also developed a nonparametric variable selection procedure for single‐set analysis. Testing for the overall set effect has been well‐studied. However, the standard formulation for estimating the set effect is based on a myopic strategy that assumes only one set at a time. This single‐set‐based test has several limitations. The score test [13, 14] is obtained from an asymptotic distribution of test statistics. The Bayesian survival kernel procedure [16] is developed for a single set analysis. Second, because the outcome is affected by multiple sets, it is inappropriate to model sets using marginal analysis. Estimating set effects based on a single set ignores the fact that sets interact with each other and, thus, result in false positives or false negatives.
Hence, to overcome the limitation of the single‐set‐based test, we propose a fused kernel machine approach to identify significant multiple‐set‐based analyses and develop a data‐driven flexible Bayesian inference to adjust multiple tests using a Bayes factor. We model the unknown high‐dimensional functions of multisets via fused Gaussian kernel machines to consider the possibility that elements within the same set interact and that sets are dependent. This fused structure enables information sharing across related sets, stabilizes inference in high‐dimensional settings, and improves statistical efficiency without requiring explicit knowledge of inter‐set connectivity.
We provide several novel contributions: (1) we develop a generalized fused multi‐kernel machine regression framework that identifies the correlated signal sets associated with continuous or binary response variables; (2) we introduce a Bayesian inference to detect multiple significant set effects simultaneously under a parsimonious dependence structure; and (3) we develop a data‐driven flexible Bayesian inference to adjust multiple tests using the Bayes factor.
The rest of the article is organized as follows. In Section 2, we first describe the problem setup and notation. We introduce the generalized fused kernel machine regression. Section 2.3 explains Bayesian inference to estimate parameters and describe multiple hypothesis tests. Section 3.3 provides how to adjust multiple comparisons using the Bayes factor. In Section 4, we conduct a simulation study to investigate the performance of our methods in terms of true positive rate (TPR), false positive rate (FPR), accuracy, and precision. In Section 5, we demonstrate the advantage of our approach using genetic pathway‐based analysis of type II diabetes [1]. Finally, Section 6 contains concluding remarks.
2. Problem Setup
In this section, we first define some notations, nonparametric models, and test hypotheses.
2.1. Notations and Model Setting
Let be the response variable of the sample, , where can be a continuous or categorical variable. Define is a vector. Let be the matrix ‐element of the sample in the set, , that is, . Let be the ‐covariate of the sample, that is, . Define is a matrix with , which is a matrix, and is a matrix.
Let , where and , which has the group lasso structure, where represents the kernel matrix for a set with the component . Let , which is the norm of the difference between two vectors which has a fused lasso structure.
Rationale for group lasso and fused lasso: Here, a group structure is imposed because the variables within each set act collectively rather than individually. For instance, a gene pathway represents a collection of genes that jointly regulate a specific biological process, and the contribution of individual genes cannot be considered in isolation. The group structure framework accommodates this by specifying a kernel function that models the joint effect of multiple genes within a pathway. This approach allows for nonlinear effects of individual gene expressions and captures the complex interactions that may occur among genes within the same pathway.
We also impose a fused structure to account for the fact that the sets are not independent. Some pathways share genes, while others may not overlap but still display strong correlations through their constituent genes. The fused framework with autocorrelation enables us to capture such dependencies among pathways.
We emphasize that the fused‐lasso structure is not intended to represent a biological ordering or a mechanistic sequence among pathways. Instead, it is introduced as a statistical coupling prior across pathway‐specific latent models, with the goal of stabilizing inference and encouraging parsimony in high‐dimensional settings. From a biological perspective, pathways are not isolated entities but overlapping functional modules that often share genes, regulators, and signaling components. Encouraging similarity across pathway‐specific effects is therefore biologically plausible even in the absence of a known or reliable inter‐pathway network. Statistically, the fused‐lasso penalty corresponds to a first‐order Gaussian Markov random field prior and has been widely used as a structured shrinkage and smoothing device [19, 20]. From a modeling standpoint, this structure can be viewed as a path‐graph Laplacian prior, representing one of the simplest forms of graph‐regularized dependence. While more complex graph‐structured priors could in principle be adopted, they would require reliable prior knowledge of inter‐pathway connectivity, which is often unavailable or highly uncertain in practice. Moreover, much of the existing literature [21, 22, 23] on graph‐ or network‐based approaches in pathway genetic analysis has focused primarily on network estimation, rather than on formal hypothesis testing under dependence. Therefore, the proposed fused structure does not aim to recover or represent the true biological pathway network. Rather, it provides a robust and parsimonious dependence structure that enables stable inference and hypothesis testing in the absence of reliable network information. The primary goal of this paper is to develop a flexible Bayesian testing framework under a computationally tractable regularization scheme, rather than to model biological pathway interactions mechanistically. Our proposed Bayesian framework enables hypothesis testing under a simple yet flexible dependence structure.
For example, consider three pathways where pathway 1 shares genes with both pathways 2 and 3, while pathways 2 and 3 do not overlap. This means that the fused lasso component shrinks differences in elements of pathway 1 across neighboring pathway 2 or 3, and smooths adjacent estimates towards one another. Alternatively, if the three pathways are highly correlated or the three pathways have no shared genes but are highly correlated, or the correlation structure among pathways is unknown, all possible distinct orderings of pathways may be considered. The fused structure captures both direct overlaps and indirect correlations between pathways, enabling more accurate detection of their joint association with the response.
Hence, by jointly modeling multiple sets through the group structure and incorporating inter‐pathway correlations through the fused structure, our method improves the accuracy of identifying pathways significantly associated with the response. In contrast, traditional pathway‐based analyses typically adopt a myopic strategy that tests one pathway at a time under an independence assumption, increasing the risk of biased or misleading inference.
2.2. Generalized Fused Multi‐Kernel Machine Regression
By denoting as a link function, as the mean of , the distribution of given as , and the conditional prior distribution of as , which has the group‐fused structure, we consider the following generalized fused multi‐kernel machine regression model (GFKM):
| (1) |
where is a vector of regression coefficients for the covariate effects, is an unknown function in reproducing kernel Hilbert space (RKHS), variance is an unknown parameter of for a continuous response but for binary response with probit link, parameters are associated with group lasso, and parameters are associated with a fused lasso, both of which parameterize , that is, , is controlled by , and is controlled by the . Note that is the tuning parameter for the group lasso, and is the tuning parameter for the fused lasso.
We simply denote as . The form of is derived from representing the Laplace (double exponential) conditional prior of as a scale mixture of a normal distribution combined with an exponential density [24]. The complete form of can be expressed as follows:
where 0 is a matrix with all zero elements and and are block matrices of size .
This matrix depends on parameters , , and . The form of , and are
for ,
for ,
for ,
where the kernel structure assigned to each leads to the diagonal block of size , while the off‐diagonal components involving the parameters work to shrink random effects that are adjacent sets. We consider that the prior distribution for is controlled by , and 's prior distribution is controlled by . Thus, is the tuning parameter for the group prior, and is the tuning parameter for the fused prior. The detailed prior specification is in Section 2.3.
2.3. Bayesian Sampling for Parameter Estimation
This section describes how to estimate parameters in GFKM (1). We explain the prior specification and illustrate the full conditional distribution for continuous and binary variable cases. For the binary response variable, we use the probit link function. The specification of prior distributions is as follows:
| (2) |
where the hyperparameters (, ) for the prior on the regression coefficients () are set to (), is inverse gamma distribution with shape parameter and scale parameter . The full form of is described in Section 2.1, respectively. The full conditional distributions for all parameters except for have closed forms. So, we draw the MCMC sample using Metropolis‐Hastings (MH) for and Gibbs sampling for all parameters except for . The detailed forms of the conditional distributions for continuous and binary responses are summarized in Appendix A.
The rationale for defining the regularization parameters as and is as follows. The conditional prior reflects the Bayesian group and fused lasso formulation introduced in Section 2.2, Equation (1). This conditional prior can be expressed as a mixture of Gaussians, as shown in the two equations described in Section S1 of Supporting Information. The squared form of appears naturally in the parameters of the Gamma distribution. By defining the regularization parameter in squared form, the resulting integral becomes tractable and yields a closed‐form expression. In this setup, the Gamma distribution acts as a mixing distribution for the variance term () of the Gaussian. The same reasoning applies to the parameters and , which play an analogous role. Hyperparameters are selected via grid search, and the optimal values are those that maximize the likelihood function.
3. Hypothesis Test for Identifying Correlated Multiple Sets of Variables
Our main question of interest is to identify correlated multiple sets of variables associated with the response. We conduct the statistical inference based on the Bayes factor [25].
To identify which set is associated with the response, one may consider the following null and alternative hypothesis, : vs. : . Let denote the Bayes factor (BF) in favor of alternative hypothesis against null hypothesis . That is,
where is the marginal likelihood of data y under ( or 1) given the model‐specific parameter vector , . is the density function under given , and is the prior distribution density under . Because the integral for marginal likelihood under does not have a closed form, we approximate this integration using the following method proposed by Newton and Raftery [26],
where are MCMC draws of size from the posterior distribution . Larger values of suggest that the data favor the model under the alternative hypothesis . Large values of BF favor . This means that the data indicate that is more strongly supported by the data than . Kass and Rafter [25] suggested how to interpret the value of BF. We interpret the value of BF as not worth more than a bare mention if , positive if , strong if , and very strong if . However, this hypothesis is formulated under the assumption that functions are independent, and the interpretation of does not incorporate adjustments for multiple comparisons. Thus, we suggest a data‐driven flexible Bayesian inference based on to adjust multiple tests in Section 3.3.
3.1. Multiple Hypothesis Tests
As we specified in GFKM(1), the set depends on , , or both and , where shrinks random effects of and . The is used to evaluate the existence of the overall set effect associated with the response variable. We also use or both and to evaluate the existence of fused kernel structures between sets associated with the response variable. As a simple illustration, we consider . When the number of sets is three, the depends on the overall set effect and . The has , , and since the set has a correlation structure with and . The has , and .
3.2. Null and Alternative Hypotheses
The corresponding null and alternative hypothesis are denoted as and , :
is , which is the same as , and is , which is the same as .
is , which is the same as , and is , which is the same as .
is , which is the same as , and is , which is the same as .
Using three sets, we can think of ordering of three sets. However, the model estimations of GFKM using some of the orderings of the three sets are the same. For example, the GFKM estimation based on the ordering (, , ) is the identical to that based on the reverse ordering (, , ), because the correlation structure between and is the same as that between and . Also, the model estimation using the following order (, , ) is also same as that of (, , ). Similarly, we can have the same model estimation between the models using the order (, , ) and the order (, , ). Figure 1 shows which orderings of three sets are identical or not. Hence, we only need to consider three models: (1) GFKM with (, , ); (2) GFKM with (, , ); and (3) GFKM with (, , ). For each model, we conduct the following hypothesis test: versus , .
FIGURE 1.

The orderings of three sets. The GFKM with the following order of sets (,,) is the same as that using the order (,,); The GFKM with the following order of sets (,,) is equal to that of (,,); the GFKM based on the following ordering sets (,,) matches that of ordering (,,).
Using our approach, we can identify which three pathways are associated with the response variable by accounting for the possibility that one pathway may be strongly connected to the other two, through the execution of three distinct tests. These connections reflect inter‐pathway relationships that cannot be detected under the assumption of pathway independence. Therefore, our method captures both direct overlaps and indirect correlations among pathways, enabling more accurate detection of their joint association with the response.
3.3. Data‐Driven Adjustment of Bayes Factor for Multiple Tests
The multiplicity arises when multiple hypothesis tests are conducted, leading to an increased likelihood of false positives. In the sets, failing to account for multiplicity can lead to incorrect identification of sets because multiple sets have overlapping elements and element interactions with other sets. Thus, the decision of cutoff values for BF to account for multiplicity is essential for reliable set identification. Traditional Bayesian approaches decide an arbitrary threshold for BF, such as 1, 3, 20, 150 [25] to determine the strength of evidence. However, this approach may not consider the effect of multiplicity. Zhang et al. [16] suggested a method to determine a threshold of BF for multiple comparisons under a single set test. We extend this method to adjust BF for the fused multi‐kernel sets test that accounts for the complexity of multiplicity, thereby providing more reliable criteria for evaluating BF using a data‐driven procedure.
Consider sets with combinations and sets in the model. We built GFKM on sets with . Consider hypotheses tests and , . Let , where is the Bayes factor for the ‐ set (, ). For the ‐ set, are computed multiple times (e.g., 100, 500, 1000 times) with different MCMC samples of the same sample size and burn‐in. The final is determined by averaging the that fall within the first and third quantiles of the collected . We remove outliers for because they highly affect the nonsignificant decision. Outliers of were defined as those with Bayes factor values greater than 100 000. These removed outliers were directly interpreted as significant results. Since the Bayes factor of this magnitude provides overwhelming evidence in favor of the alternative hypothesis, no additional multiple comparison adjustment was applied to them.
Let be an arbitrary number of classes (), and represents a vector of class centers. Define as a label vector ‐ set (, ), where the if ‐ set belongs to the ‐ class, and for all . To decide the number of classes for BF (called ), the steps of our algorithm are summarized as follows:
- Step 1: standardization. We standardize (, ) by using following formula,
where and are mean and standard deviation of , respectively. Step 2: Determine the objective function. To calculate the distance or similarity of two points, we use Euclidean distance. To minimize variance within class, our objective function is .
Step 3: Optimize the objective function. First, we candidate initial values of each class (). Second, update using objective function (), which means that assign each to nearest initial value (assign ‐ set to ‐ class). Third, we update each such that . Lastly, update and until , where is previous value of for iterations, and is criterion, which can be smaller values.
We get the range of (), , where is the lower bound of class and is the upper bound of the class. We decide on cutoff values for BF within any values in the range. When determining the number of classes for BFs , there are four cases as follows:
When , if is included in the higher class, then the ‐ set is considered to be a significant set (), whereas if is included in the lower class, then the ‐ set is considered to be a insignificant set. To establish the cutoff values for determining a significant set and an insignificant set, use the range within which falls, specifically .
When , if is close to the center of the highest class, then the ‐ set is regarded to be a strongly significant set (). If belongs to the second highest class, then the ‐ set is considered to be a significant set (). To determine cutoff values for identifying a strongly significant set and a significant set, use the range in which lies, namely . The ‐ set is considered to be a insignificant set, when is close to the center of the last class. To establish the cutoff values for determining a significant set and an insignificant set, use the range within which falls, specifically .
When , if is neighboring the center of the highest class, then the ‐ set is considered to be a very strongly significant set. We consider the ‐ set to be a strongly significant set with in the second highest class. To set cutoff values for classifying a very strongly significant set () and a strongly significant set (), use the range within which falls, specifically . When is involved in the third‐highest class, the ‐ set is regarded to be a significant set (). To determine cutoff values for identifying a strongly significant set and a significant set, use the range in which lies, namely . The ‐ set is considered to be a insignificant set, when is adjacent to the center of the last class. To establish the cutoff values for determining a significant set and an insignificant set, use the range within which falls, specifically .
The entire procedure of a data‐driven adjustment of the BF‐based test (denoted adj‐BF) for multiple hypothesis tests is summarized in Algorithm 1, and the flowchart of adj‐BF is displayed in Figure 2.
ALGORITHM 1. Data‐driven adjustment of BF method for multiple hypothesis tests (adj‐BF).

FIGURE 2.

The flowchart of adjustment of BF‐based test (adj‐BF) for multiple hypothesis tests: can be any value in the range [, ], .
4. Simulation
We conducted a simulation study to investigate the performance of our approach. Our test is a BF based on a GFKM, denoted as BF‐GFKM.
There are two comparison methods. One is a frequent test based on a semiparametric additive kernel machine regression (FT‐SAKM) [27], which is a random effect model‐based approach for continuous response data that jointly considers multiple kernel functions. FT‐SAKM's test is based on the score test. The other is a BF based on the generalized additive kernel machine model (BF‐GAKM), in which we could build our model (1) with . BF‐GAKM can be the Bayesian version of Schweiger et al. [27].
When the response variable is continuous, we compared our BF‐GFKM with FT‐SAKM and BF‐GAKM. We also compare them using the adj‐BF threshold with those using traditional BF cut thresholds such as 1, 3, 5, and 10 for a continuous response. For the binary response variable, we compared our BF‐GFKM with BF‐GAKM. We also compare them using the adj‐BF threshold with those using traditional BF cut thresholds such as 1, 1.5, and 2 for binary response cases.
We considered three simulation scenarios: (i) correlated sets without shared elements, (ii) independent sets without shared elements, and (iii) correlated sets with shared elements. In addition, we conducted an additional simulation study under a more complex strong dependence structure than AR(1), which is summarized in Section S3.4 of the Supporting Information. The simulated dependence structure in S3.4 does not satisfy the defining properties of an AR(1) model, as correlations are neither determined by a single decay parameter nor monotone in index distance, and do not rely on an ordering of locations. The correlation structure represents a meaningful misspecification relative to the assumed model. We also performed a sensitivity analysis by varying the hyperparameters of the regularization terms, as summarized in Section S4 of the Supporting Information.
We compared the performance of our approach with the alternative in terms of TPR, FPR, accuracy, and precision. TP, FP, FN, and TN are defined as follows:
Evaluation metrics, which are TR rate (TPR), FP rate (FPR), Accuracy, and Precision, are defined as follows,
These criteria are evaluated by testing hypotheses using adj‐BF. We considered cutoff values for the significant set and the insignificant set. The significant sets are estimated BF greater than cutoff values at simulation (), where is denoted as the cutoff value of the BF obtained from our adj‐BF method for multiple testings, and is any value in the range [, ].
4.1. Correlated Sets Without Shared Elements
We considered four simulation settings for continuous response, as well as two simulation settings for binary response in Section 4.1.1. For each case, we conducted three hypotheses ( and , ) when . We compared the performance of our approach with the alternative in terms of TPR, FPR, accuracy, and precision.
4.1.1. Setting
We set (three sets) and (two covariates). Because each unknown function has a different number of elements, we set , which is a vector of the different numbers of elements for each unknown function. We varied the values of and p to represent the situations in which for . The simulation settings 1–2 are for continuous response, and 3 is for binary response. For each combination of , three different hypothesis tests were conducted, and 100 simulations were run. We ran 10 000 MCMC iterations with 2000 burn‐in times. We kept one sample for every five draws to reduce autocorrelation in MCMC samples. The computational complexity for fitting our model is and the computational time for the hypothesis across different simulation cases and sample sizes is summarized in S2 of the Supporting Information.
We first generated the covariates , and set . We considered , , which contains two parts: shared common () and nonshared () parts because unknown functions have shared elements and nonshared elements. Because can be varied by an unknown or misspecified mechanism, is differently defined, as described in Cases 1–3, which are explained in brief. We also generated from with the correlation matrix , where , is a matrix with all‐ones components, and
which has a AR(1) correlation for . Each was generated from N(0,1), and then were transformed to Uniform(−2,2). A more complex dependence structure than AR(1) is considered in Section S3.4.
To generate a continuous response variable, we considered the following Cases 1–2 with . To generate a binary response variable, we used the probit link and considered Case 3 with :
Case 1: with , , where , , , , and .
Case 2: , , where , , , , and .
Case 3: if or otherwise with , where , , , , and .
Case 1 was considered because the nonparametric function has a complex form with nonlinear functions of 's and interactions of 's. Cases 2–3 were considered to investigate how our methods with a Gaussian kernel can capture misspecified functional forms, such as quadratic functions or cubic functions.
We compare the performance of BF‐GFKM with BF‐GAKM and FT‐SAKM for Cases 1–2 and with BF‐GAKM for Case 3.
4.1.2. Simulation Result
Table 1 summarized the ranges of obtained from adj‐BF method for three hypothesis tests ( vs. , ) of BF‐GFKM and BF‐GAKM in all cases. The simulation results for Case 1, Case 2, and Case 3 are summarized in Tables 2 and 3, Tables 4 and S1, and Tables S2 and 5, respectively. Note that the results for Case 2 with are presented in Table S1 of the Supporting Information, and the results for Case 3 with are presented in Table S2 of the Supporting Information. Since the results in Tables 4 and S1 are similar, as are those in Tables 5 and S2, we provide Tables S1 and S2 in the Supporting Information.
TABLE 1.
The ranges of , which is the cutoff values of for determining a significant set and an insignificant set from adj‐BF method for three hypothesis tests ( vs. , ) of BF‐GFKM and BF‐GAKM in all simulation Cases.
| (n,) | Method | vs. | vs. | vs. | ||
|---|---|---|---|---|---|---|
| ( ) | ( ) | ( ) | ||||
| Continuous | Case 1 | (50,60,55,50) | BF‐GFKM | (10.326, 32.217) | (19,719, 26.923) | (5.634, 41.742) |
| BF‐GAKM | (3.144, 16.672) | (9.326, 312.672) | (2.507, 6188.328) | |||
| (100,110,105,100) | BF‐GFKM | (0.001, 120.373) | (0.093, 15.559) | (0.006, 1656.376) | ||
| BF‐GAKM | (0.523, 5.581) | (1.355, 2.855) | (, 4210.605) | |||
| Case 2 | (50,60,55,50) | BF‐GFKM | (8.173, 12.480) | (1.408, 611.794) | (9.748, 75.762) | |
| BF‐GAKM | (1.450, 2.166) | (0.056, 1133.734) | (0.679, 52.728) | |||
| (100,110,105,100) | BF‐GFKM | (, 6329.68) | (0.015, ) | (0.021, ) | ||
| BF‐GAKM | (0.765, 45.074) | (, 4379.843) | (, 104.695) | |||
| Binary | Case 3 | (80,90,85,80) | BF‐GFKM | (3.562, 3.909) | (2.059, 2.303) | (6.729, 7.390) |
| BF‐GAKM | (6.211, 6.453) | (5.007, 5.276) | (7.480, 7.815) | |||
| (100,110,105,100) | BF‐GFKM | (6.730, 7.200) | (1.489, 1.578) | (5.362, 5.704) | ||
| BF‐GAKM | (7.800, 8.583) | (8.067, 9.490) | (6.777, 8.275) |
TABLE 2.
The results of three hypothesis tests ( vs. , ) for Case 1 with .
| Decision Rule | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Setting | Hypothesis | Method | Criteria |
|
|
|
|
|
|||||
| , , , | vs. | BF‐GFKM | TPR | 1 | 1 | 1 | 1 | 1 | |||||
| BF‐GAKM | 1 | 0.990 | 0.990 | 0.990 | 0.990 | ||||||||
| FT‐SAKM | 0.840 | ||||||||||||
| BF‐GFKM | FPR | 0.350 | 0.140 | 0.080 | 0.050 | 0.040 | |||||||
| BF‐GAKM | 0.030 | 0.020 | 0.010 | 0.010 | 0.010 | ||||||||
| FT‐SAKM | 0.040 | ||||||||||||
| BF‐GFKM | Accuracy | 0.825 | 0.930 | 0.960 | 0.975 | 0.980 | |||||||
| BF‐GAKM | 0.985 | 0.985 | 0.990 | 0.990 | 0.990 | ||||||||
| FT‐SAKM | 0.900 | ||||||||||||
| BF‐GFKM | Precision | 0.741 | 0.877 | 0.926 | 0.952 | 0.962 | |||||||
| BF‐GAKM | 0.971 | 0.980 | 0.990 | 0.990 | 0.990 | ||||||||
| FT‐SAKM | 0.955 | ||||||||||||
| vs. | BF‐GFKM | TPR | 1 | 1 | 1 | 1 | 1 | ||||||
| BF‐GAKM | 1 | 1 | 1 | 1 | 1 | ||||||||
| FT‐SAKM | 0.840 | ||||||||||||
| BF‐GFKM | FPR | 0.360 | 0.170 | 0.150 | 0.100 | 0.060 | |||||||
| BF‐GAKM | 0.070 | 0.020 | 0.010 | 0 | 0 | ||||||||
| FT‐SAKM | 0.190 | ||||||||||||
| BF‐GFKM | Accuracy | 0.820 | 0.915 | 0.925 | 0.950 | 0.970 | |||||||
| BF‐GAKM | 0.965 | 0.990 | 0.995 | 1 | 1 | ||||||||
| FT‐SAKM | 0.825 | ||||||||||||
| BF‐GFKM | Precision | 0.735 | 0.855 | 0.870 | 0.909 | 0.943 | |||||||
| BF‐GAKM | 0.935 | 0.980 | 0.990 | 1 | 1 | ||||||||
| FT‐SAKM | 0.816 | ||||||||||||
| vs. | BF‐GFKM | TPR | 1 | 1 | 1 | 1 | 1 | ||||||
| BF‐GAKM | 1 | 1 | 1 | 1 | 1 | ||||||||
| FT‐SAKM | 0.840 | ||||||||||||
| BF‐GFKM | FPR | 0.240 | 0.100 | 0.050 | 0.020 | 0.020 | |||||||
| BF‐GAKM | 0.030 | 0 | 0 | 0 | 0 | ||||||||
| FT‐SAKM | 0.040 | ||||||||||||
| BF‐GFKM | Accuracy | 0.880 | 0.950 | 0.975 | 0.990 | 0.990 | |||||||
| BF‐GAKM | 0.985 | 1 | 1 | 1 | 1 | ||||||||
| FT‐SAKM | 0.905 | ||||||||||||
| BF‐GFKM | Precision | 0.806 | 0.909 | 0.952 | 0.980 | 0.980 | |||||||
| BF‐GAKM | 0.971 | 1 | 1 | 1 | 1 | ||||||||
| FT‐SAKM | 0.966 | ||||||||||||
Note: The significance level of FT‐SAKM. The ranges of , which is the cutoff values of to determine a significant set and an insignificant set for BF‐GFKM across three hypothesis tests, are (10.326, 32.217), (19.719, 26.923), and (5.634, 41.742), while those for BF‐GAKM are (3.144, 16.672), (9.326, 312.672), and (2.507, 6188.328).
TABLE 3.
The results of three hypothesis tests ( vs. , ) for Case 1 with .
| Decision Rule | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Setting | Hypothesis | Method | Criteria |
|
|
|
|
|
|||||
| , , , | vs. | BF‐GFKM | TPR | 1 | 1 | 1 | 1 | 1 | |||||
| BF‐GAKM | 0.980 | 0.980 | 0.980 | 0.970 | 0.980 | ||||||||
| FT‐SAKM | 0.950 | ||||||||||||
| BF‐GFKM | FPR | 0 | 0 | 0 | 0 | 0 | |||||||
| BF‐GAKM | 0 | 0 | 0 | 0 | 0 | ||||||||
| FT‐SAKM | 0.060 | ||||||||||||
| BF‐GFKM | Accuracy | 1 | 1 | 1 | 1 | 1 | |||||||
| BF‐GAKM | 0.990 | 0.990 | 0.990 | 0.985 | 0.990 | ||||||||
| FT‐SAKM | 0.945 | ||||||||||||
| BF‐GFKM | Precision | 1 | 1 | 1 | 1 | 1 | |||||||
| BF‐GAKM | 1 | 1 | 1 | 1 | 1 | ||||||||
| FT‐SAKM | 0.941 | ||||||||||||
| vs. | BF‐GFKM | TPR | 1 | 1 | 1 | 1 | 1 | ||||||
| BF‐GAKM | 0.960 | 0.940 | 0.920 | 0.890 | 0.950 | ||||||||
| FT‐SAKM | 0.950 | ||||||||||||
| BF‐GFKM | FPR | 0 | 0 | 0 | 0 | 0 | |||||||
| BF‐GAKM | 0 | 0 | 0 | 0 | 0 | ||||||||
| FT‐SAKM | 0.100 | ||||||||||||
| BF‐GFKM | Accuracy | 1 | 1 | 1 | 1 | 1 | |||||||
| BF‐GAKM | 0.980 | 0.970 | 0.960 | 0.945 | 0.975 | ||||||||
| FT‐SAKM | 0.950 | ||||||||||||
| BF‐GFKM | Precision | 1 | 1 | 1 | 1 | 1 | |||||||
| BF‐GAKM | 1 | 1 | 1 | 1 | 1 | ||||||||
| FT‐SAKM | 0.909 | ||||||||||||
| vs. | BF‐GFKM | TPR | 1 | 1 | 1 | 1 | 1 | ||||||
| BF‐GAKM | 1 | 1 | 1 | 1 | 1 | ||||||||
| FT‐SAKM | 0.950 | ||||||||||||
| BF‐GFKM | FPR | 0 | 0 | 0 | 0 | 0 | |||||||
| BF‐GAKM | 0 | 0 | 0 | 0 | 0 | ||||||||
| FT‐SAKM | 0.030 | ||||||||||||
| BF‐GFKM | Accuracy | 1 | 1 | 1 | 1 | 1 | |||||||
| BF‐GAKM | 1 | 1 | 1 | 1 | 1 | ||||||||
| FT‐SAKM | 0.960 | ||||||||||||
| BF‐GFKM | Precision | 1 | 1 | 1 | 1 | 1 | |||||||
| BF‐GAKM | 1 | 1 | 1 | 1 | 1 | ||||||||
| FT‐SAKM | 0.969 | ||||||||||||
Note: The significance level of FT‐SAKM. The ranges of , which is the cutoff values of to determine a significant set and an insignificant set for BF‐GFKM across three hypothesis tests, are (0.001, 120.373), (0.093, 15.559), and (0.006, 1656.376), while those for BF‐GAKM are (0.523, 5.581), (1.355, 2.855), and (, 4210.605).
TABLE 4.
The results of three hypothesis tests ( vs. , ) for Case 2 with .
| Decision Rule | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Setting | Hypothesis | Method | Criteria |
|
|
|
|
|
|||||
| , , , | vs. | BF‐GFKM | TPR | 1 | 1 | 1 | 1 | 1 | |||||
| BF‐GAKM | 1 | 0.980 | 0.970 | 0.960 | 0.990 | ||||||||
| FT‐SAKM | 0.960 | ||||||||||||
| BF‐GFKM | FPR | 0.320 | 0.110 | 0.060 | 0.020 | 0.020 | |||||||
| BF‐GAKM | 0.010 | 0 | 0 | 0 | 0 | ||||||||
| FT‐SAKM | 0.080 | ||||||||||||
| BF‐GFKM | Accuracy | 0.840 | 0.945 | 0.970 | 0.990 | 0.990 | |||||||
| BF‐GAKM | 0.995 | 0.990 | 0.985 | 0.980 | 0.995 | ||||||||
| FT‐SAKM | 0.940 | ||||||||||||
| BF‐GFKM | Precision | 0.758 | 0.901 | 0.943 | 0.980 | 0.980 | |||||||
| BF‐GAKM | 0.990 | 1 | 1 | 1 | 1 | ||||||||
| FT‐SAKM | 0.923 | ||||||||||||
| vs. | BF‐GFKM | TPR | 1 | 1 | 1 | 1 | 1 | ||||||
| BF‐GAKM | 1 | 1 | 1 | 1 | 1 | ||||||||
| FT‐SAKM | 0.960 | ||||||||||||
| BF‐GFKM | FPR | 0.010 | 0 | 0 | 0 | 0 | |||||||
| BF‐GAKM | 0 | 0 | 0 | 0 | 0 | ||||||||
| FT‐SAKM | 0.140 | ||||||||||||
| BF‐GFKM | Accuracy | 0.995 | 1 | 1 | 1 | 1 | |||||||
| BF‐GAKM | 1 | 1 | 1 | 1 | 1 | ||||||||
| FT‐SAKM | 0.910 | ||||||||||||
| BF‐GFKM | Precision | 0.990 | 1 | 1 | 1 | 1 | |||||||
| BF‐GAKM | 1 | 1 | 1 | 1 | 1 | ||||||||
| FT‐SAKM | 0.873 | ||||||||||||
| vs. | BF‐GFKM | TPR | 1 | 1 | 1 | 1 | 1 | ||||||
| BF‐GAKM | 1 | 1 | 1 | 1 | 1 | ||||||||
| FT‐SAKM | 0.960 | ||||||||||||
| BF‐GFKM | FPR | 0.240 | 0.080 | 0.050 | 0.010 | 0.010 | |||||||
| BF‐GAKM | 0 | 0 | 0 | 0 | 0 | ||||||||
| FT‐SAKM | 0.100 | ||||||||||||
| BF‐GFKM | Accuracy | 0.880 | 0.960 | 0.975 | 0.995 | 0.995 | |||||||
| BF‐GAKM | 1 | 1 | 1 | 1 | 1 | ||||||||
| FT‐SAKM | 0.930 | ||||||||||||
| BF‐GFKM | Precision | 0.806 | 0.926 | 0.952 | 0.990 | 0.990 | |||||||
| BF‐GAKM | 1 | 1 | 1 | 1 | 1 | ||||||||
| FT‐SAKM | 0.906 | ||||||||||||
Note: The significance level of FT‐SAKM. The ranges of , which is the cutoff values of to determine a significant set and an insignificant set for BF‐GFKM across three hypothesis tests, are (8.173, 12.480), (1.408, 611.794), and (9.748, 75.762), while those for BF‐GAKM are (1.450, 2.166), (0.056, 1133.734), and (9.748, 75.762).
TABLE 5.
The results of three hypothesis tests ( vs. , ) for Case 3 with .
| Decision Rule | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Setting | Method | Criteria |
|
|
|
|
|||||
| , , , | vs. | BF‐GFKM | TPR | 0.910 | 0.880 | 0.870 | 0.670 | ||||
| BF‐GAKM | 0.900 | 0.880 | 0.840 | 0.660 | |||||||
| BF‐GFKM | FPR | 0.540 | 0.440 | 0.430 | 0.190 | ||||||
| BF‐GAKM | 0.720 | 0.620 | 0.470 | 0.170 | |||||||
| BF‐GFKM | Accuracy | 0.685 | 0.720 | 0.720 | 0.740 | ||||||
| BF‐GAKM | 0.590 | 0.630 | 0.685 | 0.745 | |||||||
| BF‐GFKM | Precision | 0.628 | 0.667 | 0.669 | 0.779 | ||||||
| BF‐GAKM | 0.556 | 0.587 | 0.641 | 0.795 | |||||||
| vs. | BF‐GFKM | TPR | 0.750 | 0.700 | 0.670 | 0.700 | |||||
| BF‐GAKM | 0.940 | 0.930 | 0.880 | 0.630 | |||||||
| BF‐GFKM | FPR | 0.320 | 0.250 | 0.200 | 0.250 | ||||||
| BF‐GAKM | 0.660 | 0.590 | 0.530 | 0.280 | |||||||
| BF‐GFKM | Accuracy | 0.715 | 0.725 | 0.735 | 0.725 | ||||||
| BF‐GAKM | 0.640 | 0.670 | 0.675 | 0.675 | |||||||
| BF‐GFKM | Precision | 0.701 | 0.737 | 0.770 | 0.737 | ||||||
| BF‐GAKM | 0.588 | 0.612 | 0.624 | 0.692 | |||||||
| vs. | BF‐GFKM | TPR | 0.920 | 0.910 | 0.900 | 0.750 | |||||
| BF‐GAKM | 0.970 | 0.950 | 0.940 | 0.780 | |||||||
| BF‐GFKM | FPR | 0.460 | 0.390 | 0.310 | 0.140 | ||||||
| BF‐GAKM | 0.610 | 0.520 | 0.460 | 0.130 | |||||||
| BF‐GFKM | Accuracy | 0.730 | 0.760 | 0.795 | 0.805 | ||||||
| BF‐GAKM | 0.680 | 0.715 | 0.740 | 0.825 | |||||||
| BF‐GFKM | Precision | 0.667 | 0.700 | 0.744 | 0.843 | ||||||
| BF‐GAKM | 0.646 | 0.671 | 0.860 | 0.857 | |||||||
Note: The ranges of , which is the cutoff values of to determine a significant set and an insignificant set for BF‐GFKM across three hypothesis tests, are (6.730, 7.200), (1.489, 1.578), and (5.362, 5.704), while those for BF‐GAKM are (7.800, 8.583), (8.067, 9.490), and (6.777, 8.275).
The boxplots of four criteria for Cases 1–2 and Case 3 are displayed in Figure 3 and in Figure 4, respectively. In each Figure, we also display the four criteria of BF‐GFKM using the adj‐BF and traditional thresholds.
FIGURE 3.

The box plots of four criteria under simulation Cases 1–2. The orange, green, and blue box plots are obtained from BF‐GFKM and BF‐GAKM using various thresholds, including traditional BF thresholds and FT‐SAKM, respectively. The left top (a) is TPR; The right top (b) is FPR; The left bottom (c) is Accuracy. The right bottom (d) is Precision.
FIGURE 4.

The box plots of four criteria under simulation Case 3. The orange and sky blue box plots are obtained from BF‐GFKM and BF‐GAKM using various thresholds, including traditional BF thresholds. The left top (a) is TPR; The right top (b) is FPR; The left bottom (c) is Accuracy. The right bottom (d) is Precision.
Table 2 summarized the three hypothesis tests ( vs. , ) for Case 1 with . Among BF‐GFKM using various thresholds, BF‐GFKM using adj‐BF performs the best in terms of all criteria. The TPR values of BF‐GFKM using adj‐BF are for three hypothesis tests, while that of FT‐SAKM is . The FPR values of BF‐GFKM of three tests are , and , those of BF‐GAKM are , and 0%, and those of FT‐SAKM are , and 4%. The Accuracy values of BF‐GFKM are , and , and those of BF‐GAKM are , and , which are larger than those of FT‐SAKM, , and . The Precision values of BF‐GFKM are , those values of BF‐GAKM are , and those values of FT‐SAKM are .
The ranges of for BF‐GFKM across three hypothesis tests are (10.326, 32.217), (19.719, 26.923), and (5.634, 41.742), while those for BF‐GAKM are (3.144, 16.672), (9.326, 312.672), and (2.507, 6188.328). The BF‐GFKM and BF‐GAKM perform better than FT‐SAKM in terms of TPR, Accuracy, and Precision. The two approaches are comparable in terms of FPR since the FPR of the two approaches is the same for the first hypothesis testing ( vs. ).
On the other hand, Table 3 contains the multiple testing results for Case 1 with . We can observe that TPR, Accuracy, and Precision are increased, and FPR is decreased for BF‐GFKM and FT‐SAKM, as we expected, because of the larger sample size. However, the TPR and Accuracy of BF‐GAKM were slightly reduced, but FPR and Precision are better than those of a small sample size. The performance of BF‐GFKM and BF‐GAKM is better than that of FT‐SAKM in terms of the four criteria. The ranges of for BF‐GFKM across three hypothesis tests are (0.001, 120.373), (0.093, 15.559), and (0.006, 1656.376), while those for BF‐GAKM are (0.523, 5.581), (1.355, 2.855), and (, 4210.605).
Table 4 displayed hypothesis testing results for Case 2 with , whereas Table S1 summarized the results for Case 2 with . The results for Case 2 are similar to those for Case 1. BF‐GFKM and BF‐GAKM perform better than the FT‐SAKM in terms of the four criteria. The ranges of for BF‐GFKM across three hypothesis tests with small sample sizes are (8.173, 12.480), (1.408, 611.794), and (9.748, 75.762), while those for BF‐GAKM are (1.450, 2.166), (0.056, 1133.734), and (9.748, 75.762). The ranges of for BF‐GFKM across three hypothesis tests with large sample size are (, 6329.68), (0.015, ), and (0.021, ), while those for BF‐GAKM are (, 6329.68), (0.015, ), and (0.021, ). Among BF‐GFKM using various thresholds, BF‐GFKM using adj‐BF performs the best in terms of all criteria. Also, we can observe that BF‐GAKM using adj‐BF performs better than that using other traditional thresholds. The simulation results suggest that the performances of the three approaches for Case 2 are better than those for Case 1 because the function of Case 2 has only a quadratic form of z's, whereas the function of Case 1 has a more complex form of z's. The box plots of four criteria of BF‐GFKM using adj‐BF and traditional thresholds and other comparison methods for simulation Cases 1–2 are outlined in Figure 3. Overall, the performance of BF‐GFKM and BF‐GAKM using adj‐BF is better than that of FT‐SAKM.
The results on binary response variable for Case 3 with and with are summarized in Tables S2 and 5. Among BF‐GFKM using various thresholds, BF‐GFKM using adj‐BF performs the best in terms of all criteria. Also, we can observe that BF‐GAKM using adj‐BF performs better than that using other traditional thresholds. The TPR values of BF‐GFKM using adj‐BF are , and for three hypothesis tests, whereas those of BF‐GAKM are , and . The FPR values of BF‐GFKM of three tests are , and those of BF‐GAKM are . The Accuracy values of BF‐GFKM are , and , which are larger than those of BF‐GAKM: , and . The Precision values of BF‐GFKM are , and , and those values of BF‐GAKM are , and . The ranges of for BF‐GFKM across three hypothesis tests are (3.562, 3.909), (2.059, 2.303), and (6.729, 7.390), while those for BF‐GAKM are (6.211, 6.453), (5.007, 5.276), and (7.480, 7.815). The results on binary outcomes for larger sample sizes are similar to those for smaller ones. However, the performance of BF‐GFKM is worse than that of BF‐GAKM for versus in terms of FPR, Accuracy, and Precision. Also, BF‐GFKM performs worse than BF‐GAKM for the third hypothesis testing regarding the four criteria. The overall performance of BF‐GFKM on binary outcomes for a larger sample size is comparable to that of BF‐GAKM. The ranges of for BF‐GFKM across three hypothesis tests are (6.730, 7.200), (1.489, 1.578), and (5.362, 5.704), while those for BF‐GAKM are (7.800, 8.583), (8.067, 9.490), and (6.777, 8.275). The box plots of four criteria of BF‐GFKM using adj‐BF and traditional thresholds and other comparison methods for simulation Case 3 are also outlined in Figure 4. Again, we observe that BF‐GFKM performs better than BF‐GAKM in terms of the four criteria for Case 3. BF‐GFKM using adj BF performs better than that using other thresholds.
4.2. Independent Sets Without Shared Elements
This section presents a simulation study designed to investigate the performance of our approach under the condition of no correlation among sets. The core setup remains the same as in Section 4.1.1.
4.2.1. Setting
We consider Case 1 for testing three hypotheses ( vs. , ) with the setting . For each case, we ran 100 simulations and conducted 10 000 MCMC iterations.
We first generated the covariates , and set . We considered for , which contains two parts: shared common () and nonshared () parts. We also generated from with the correlation matrix , where , is an matrix with all‐ones components, and is the identity matrix. This means there is no correlation among sets. Each was generated from N(0,1) and then transformed to Uniform(−2,2). To generate a continuous response variable, we used the function of Case 1 in Section 4.1.1.
4.2.2. Simulation Result
Table S3 of Supporting Information presents the results of three hypothesis tests for Case 1 in a setting where there is no correlation between sets. Both BF‐GFKM and BF‐GAKM achieved a near‐perfect performance with the adj‐BF decision rule. The BF‐GFKM attained a TPR of 1, FPR of 0, Accuracy of 1, and Precision of 1 across all three hypotheses. The BF‐GAKM also achieved similar results to those of BF‐GFKM except for one hypothesis ( vs. ). The FT‐SAKM also performs much better in this setting. Its TPR improve to (up from in the correlated setting), while its FPR drops to for the first test ( vs. ) and for the other two tests. Consequently, its Accuracy and precision values also see a notable increase, reaching and , respectively.
The improved performance of the FT‐SAKM suggests that the presence of correlation between sets in the data generation process significantly hinders the frequentist approach's ability to accurately test the hypotheses.
4.3. Correlated Sets With Shared Elements
This section details a simulation study designed to investigate the performance of our approach under a shared element setting. The core setup remains the same as in Section 4.1.1.
4.3.1. Setting
Specifically, we consider Case 1 for testing three hypotheses ( vs. ), with the setting . The covariates were generated from , with .
The correlation structure is the same as in Section 4.1.1. The significant change in this setting is the introduction of a common element between the first two sets of variables, achieved by modifying the functional form for Case 1 as follows:
-
•Case 1: with , where , and .
-
·
-
·
-
·
-
·
The key modification is in the expression for , where the term has been replaced with . This change introduces a single common element from the first set of variables () into the functional form of the second set (). This simulation is designed to evaluate how each method performs when a true underlying relationship between sets exists through a shared element.
4.3.2. Simulation Result
Table S4 of Supporting Information presents the simulation results for Case 1 under a common element setting. As in previous simulations, we evaluate the performance of BF‐GFKM, BF‐GAKM, and FT‐SAKM methods based on their TPR, FPR, Accuracy, and Precision. The core distinction in this simulation is the introduction of a shared element between the first and second sets.
The results show that both Bayesian methods, BF‐GFKM and BF‐GAKM, maintain exceptional performance. For all three hypotheses, BF‐GAKM achieved a TPR of 1 and FPR of 0, resulting in a perfect Accuracy and Precision of 1. BF‐GFKM also achieved the same performance except for the second hypothesis test ( vs. ), but the performance of BF‐GFKM in the second hypothesis test also is close to TPR of 1 and FPR of 0.
The frequentist FT‐SAKM method also demonstrates strong performance, with a consistent TPR of 0.850 across all three hypotheses. While this is a good result, it is still lower than the near‐perfect TPR achieved by the Bayesian methods. The FPR for FT‐SAKM is notably lower in the third hypothesis ( vs. ) at 0.020, but higher for the first and second hypotheses (0.120 and 0.180, respectively). This variability suggests that the performance of the frequentist method is more sensitive to the specific functional form of the relationship, even in the presence of a common element.
5. Type II Diabetes Genetic‐Pathways Data
We applied our BF‐GFKM to type II diabetes genetic‐pathways data [1]. These data include 18 type II diabetes patients and 17 normal glucose levels. Several studies [8, 28] also reported significant multiple pathways either using a score test [8] or a hybrid test [28]. Their analysis is a single‐pathway analysis, ignoring dependence among pathways. Pang et al. [8] identified significant pathways associated with the glucose level, whereas Xu et al. [28] detected important pathways to distinguish between normal glucose‐level individuals and those with type II diabetes. In our study, we collected 21 pathways from these two studies.
A systematic selection framework was adopted to delineate pathways recurrently identified across both primary sources. Specifically, (i) all pathways reported as significant in both reference studies were retained; (ii) two pathways consistently classified as nonsignificant in both studies were incorporated to preserve representational balance; and (iii) six to eleven uniquely significant pathways identified in each study were additionally included to ensure comprehensive coverage. This procedure yielded a target set of 18–24 high‐priority pathways (corresponding to approximately 6–8 composite combinations). Selection of fewer combinations would curtail representation of validated candidates, whereas expansion beyond this range would necessitate inclusion of marginal pathways and introduce excessive computational and inferential burden, particularly under the limited sample size (). Guided by established metabolic and signaling interdependencies, the final configuration comprised seven composite combinations, partitioning the 21 pathways into seven functionally coherent clusters.
In our analysis, we considered two types of response variables: one for a continuous outcome, the log‐transformed glucose level, and the other for a binary outcome, representing the type II diabetes status. We fit BF‐GFKM using three pathways (). Let Z be the gene expression, where is 35, and , which varies from 16 to 140. Our goal is to identify significantly correlated pathways that affect the glucose level and pathways to distinguish between normal and type II diabetes patients. We decided on cutoff values of using the adj‐BF method for multiple testing.
Our analysis consists of two parts. The first is to conduct seven combinations of 21 pathways, summarized in Table 6. For each combination, we conducted the hypothesis tests. The second part is to conduct all possible three combinations out of the three pathways. Although the total number of orders for these three pathways is , the total number of tests is 3 () because some of the model estimations of the ordering of the three pathways are the same, as we explain in Section 3.2. We consider two combinations of three pathways shown in Table 8.
TABLE 6.
The seven combinations out of 21 pathways.
| Combination | Pathway | Name |
|---|---|---|
| 1 | 1 | Alanine and aspartate metabolism |
| 2 | ATP synthesis | |
| 3 | c22 U133 probes(user defined) | |
| 2 | 1 | c23 U133 probes(user defined) |
| 2 | JAK‐Stat Signaling Pathway | |
| 3 | Oxidative phosphorylation | |
| 3 | 1 | OXPHOS HG‐U133A probes |
| 2 | Parkinson's disease | |
| 3 | Ubiquinone biosynthesis | |
| 4 | 1 | gamma‐Hexachlorocyclohexane degradation |
| 2 | MAP00561 Glycerolipid metabolism | |
| 3 | Histidine metabolism | |
| 5 | 1 | Fructose and mannose metabolism |
| 2 | MAP00071 Fatty acid metabolism | |
| 3 | MAP00190 Oxidative phosphorylation | |
| 6 | 1 | MAP00380 Tryptophan metabolism |
| 2 | Starch and sucrose metabolism | |
| 3 | Wnt signaling pathway | |
| 7 | 1 | Ubiquitin mediated proteolysis |
| 2 | Apoptosis | |
| 3 | Complement and coagulation cascades |
TABLE 8.
The three orderings for two combinations of three pathways chosen from 5 pathways.
| Combination | Ordering | Pathway | Name |
|---|---|---|---|
| 1 | 1 | 1 | ATP synthesis |
| 2 | Ubiquitin mediated proteolysis | ||
| 3 | JAK‐Stat Signaling Pathway | ||
| 2 | 1 | Ubiquitin mediated proteolysis | |
| 2 | ATP synthesis | ||
| 3 | JAK‐Stat Signaling Pathway | ||
| 3 | 1 | ATP synthesis | |
| 2 | JAK‐Stat Signaling Pathway | ||
| 3 | Ubiquitin mediated proteolysis | ||
| 2 | 1 | 1 | Oxidative phosphorylation |
| 2 | gamma‐Hexachlorocyclohexane degradation | ||
| 3 | JAK‐Stat Signaling Pathway | ||
| 2 | 1 | Oxidative phosphorylation | |
| 2 | JAK‐Stat Signaling Pathway | ||
| 3 | gamma‐Hexachlorocyclohexane degradation | ||
| 3 | 1 | JAK‐Stat Signaling Pathway | |
| 2 | Oxidative phosphorylation | ||
| 3 | gamma‐Hexachlorocyclohexane degradation |
5.1. Identifying Significant Pathways Associated With Glucose Level
Using our adj‐BF method, we first decided on values, . We found that there were classes to decide of . The ranges of , , and of obtained from the randomly chosen three combinations out of 21 pathways are (3.638, 4.278), (7.092, 10.452), and (17.948, 23.315), respectively. The total number of combinations was 9 (7 combinations have no order, but the other two combinations have three orderings), as shown in Tables 6, 7, 8.
TABLE 7.
values of pathways in each combination: We calculated BF 100 times from 100 different MCMC samples.
| Id | Name of pathway | #genes | Combination | Continuous (BF) | Binary (BF) | |
|---|---|---|---|---|---|---|
| 4 | Alanine and aspartate metabolism | 18 | 1 | (5.733) | (10.797) | |
| 16 | ATP synthesis | 49 | 1 | (60.730) | 1.391 | |
| 43 | c22 U133 probes(user defined) | 95 | 1 | 1.789 | (4.658) | |
| 44 | c23 U133 probes(user defined) | 67 | 2 | 1.566 | (6.718) | |
| 110 | JAK‐Stat Signaling Pathway | 71 | 2 | (5.326) | (4.727) | |
| 229 | Oxidative phosphorylation | 133 | 2 | (77.441) | (3.726) | |
| 230 | OXPHOS HG‐U133A probes | 121 | 3 |
|
(6.956) | |
| 232 | Parkinson's disease | 40 | 3 | 3.121 | (9.518) | |
| 272 | Ubiquinone biosynthesis | 16 | 3 | 0.426 | 0.473 | |
| 88 | gamma‐Hexachlorocyclohexane degradation | 41 | 4 | (7.092) | (3.774) | |
| 177 | MAP00561 Glycerolipid metabolism | 43 | 4 | (23.316) | (7.752) | |
| 103 | Histidine metabolism | 37 | 4 | (39.060) | (3.585) | |
| 121 | Fructose and mannose metabolism | 23 | 5 | 3.638 | 2.781 | |
| 126 | MAP00071 Fatty acid metabolism | 65 | 5 | 2.157 | (8.269) | |
| 133 | MAP00190 Oxidative phosphorylation | 58 | 5 | (17.948) | (4.976) | |
| 154 | MAP00380 Tryptophan metabolism | 60 | 6 | (11.537) | (7.449) | |
| 259 | Starch and sucrose metabolism | 54 | 6 | (5.351) | 2.840 | |
| 278 | Wnt signaling pathway | 140 | 6 | (10.452) | (6.059) | |
| 273 | Ubiquitin mediated proteolysis | 46 | 7 | (6.070) | 2.643 | |
| 13 | Apoptosis | 92 | 7 | (29.370) | (7.904) | |
| 71 | Complement and coagulation cascades | 47 | 7 | 2.542 | (3.635) |
Note: The ranges of , , and of for continuous are (3.638, 4.278), (7.092, 10.452), and (17.948, 23.315), respectively. The ranges of those of BF for binary are (2.890, 3.097), (4.977, 6.058), and (9.004, 9.517), respectively. , , and mean the pathway is significant, strongly significant, or very strongly significant, respectively.
First, for each combination, we conducted BF‐GFKM to test the hypothesis ( vs. , ). The testing result is summarized in Tables 7, 8, 9. In total, the 16 pathways are insignificant, whereas 23 pathways are significant, strongly significant, or very strongly significant. Table 7 shows of 21 pathways with combinations 1–7. The largest was 77.443, in which the BF of “Oxidative phosphorylation,” namely pathway 229, and combination 2. Mootha et al. [1] found that the “Oxidative phosphorylation” pathway is coordinately decreased in human diabetic muscle. It has been hypothesized that PGC‐1α, which is expressed in skeletal muscle and powerfully induces mitochondrial biogenesis when expressed ectopically in skeletal and cardiac myocytes, activates the “Oxidative phosphorylation” pathway [29]. “ATP synthesis” and “OXPHOS HG‐U133A probes” are very strongly significant pathways because the “ATP synthesis” subset of the “Oxidative phosphorylation” and “OXPHOS HG‐U133A probes” are superset of “Oxidative phosphorylation” [1]. The “JAK‐Stat Signaling Pathway” is associated with glucose levels as its inhibition has been shown to prevent glucose‐induced growth in glomerular mesangial cells [30]. The “Ubiquitin mediated proteolysis” is associated with glucose levels without the inclusion of the “JAK‐Stat Signaling Pathway” and “ATP synthesis” in three pathways models [8]. The pathways 13, 88, 103, 133, 154, 177, 259, and 278 are identified and show that these pathways are associated with glucose levels. These pathways are identified to distinguish between normal and type II diabetes patients [28].
TABLE 9.
values of pathways in each combination: We calculated BF 100 times from 100 different MCMC samples.
| ID | NAME | #genes | Combination | Ordering | Continuous | Binary |
|---|---|---|---|---|---|---|
| 16 | ATP synthesis | 49 | 1 | 1 | (44.197) | (3.935) |
| 2 | (49.142) | (3.098) | ||||
| 3 | (43.051) | 2.587 | ||||
| 273 | Ubiquitin mediated proteolysis | 46 | 1 | 3.338 | 1.463 | |
| 2 | 2.723 | 2.890 | ||||
| 3 | 2.996 | 1.977 | ||||
| 110 | JAK‐Stat Signaling Pathway | 71 | 1 | 3.565 | (12.239) | |
| 2 | 3.237 | (11.456) | ||||
| 3 | (4.789) | (7.251) | ||||
| 2 | 1 | (4.278) | (11.791) | |||
| 2 | (4.453) | (6.890) | ||||
| 3 | 3.557 | (9.004) | ||||
| 88 | gamma‐Hexachlorocyclohexane degradation | 41 | 1 | 1.967 | 1.776 | |
| 2 | 1.706 | (3.436) | ||||
| 3 | 1.478 | (3.332) | ||||
| 229 | Oxidative phosphorylation | 133 | 1 | (172.314) | (4.918) | |
| 2 | (163.079) | (4.741) | ||||
| 3 | (134.874) | (3.229) |
Note: The ranges of , , and of for continuous are (3.638, 4.278), (7.092, 10.452), and (17.948, 23.315), respectively. The ranges of those of BF for binary are (2.890, 3.097), (4.977, 6.058), and (9.004, 9.517), respectively. , , and mean the pathway is significant, strongly significant, or very strongly significant, respectively.
Second, on the other hand, we investigated the ordering of three pathways that affect glucose levels. We selected five pathways and then considered three orderings from two combinations of five pathways. The results are summarized in Table 9. The “Ubiquitin mediated proteolysis” pathway is not a significant pathway with the inclusion of “ATP synthesis” and “JAK‐Stat Signaling Pathway”. The “Ubiquitin mediated proteolysis” pathway was not significant after the inclusion of “ATP synthesis” and “JAK‐Stat Signaling Pathway” under the additive model [8]. It appears that the pathways “Ubiquitin‐mediated proteolysis” and “JAK‐Stat Signaling Pathway” have significant crosstalk potential or are functionally related concerning glucose levels. These two pathways do not have any overlapping genes. However, the origin of this pathway from KEGG shows that “JAK‐Stat Signaling Pathway” pathway has linkage with “Ubiquitin‐mediated proteolysis” [8, 31].
“JAK‐Stat Signaling Pathway” with orderings 1–2 of combination 1 and ordering 3 of combination 2 is not a significant pathway. However, this pathway with ordering 3 of combination 1 and orderings 1–2 of combination 2 is significant. We can see that the testing results for glucose levels can be changed depending on the ordering of pathways.
The testing results of combinations 1, 2, and 4 and ordering 2 from combination 1 of the three pathways associated with the glucose level are in Figure 5. Figure 6 shows the testing results of orderings 1 and 3 for combinations 1–2 of three pathways associated with the glucose level. The solid blue circle indicates a significant pathway, whereas the dashed white circle indicates an insignificant pathway. The solid line denotes two pathways that are correlated with the fused structure, whereas the dashed line denotes that they are not.
FIGURE 5.

The hypothesis testing results of type II diabetes genetic pathway data. The solid blue circle indicates a significant pathway, while the dashed white circle indicates an insignificant pathway. The solid line denotes that two pathways are correlated with a fused structure, while the dashed line denotes that they are not. The left top (a) represents the result of combination 1 of three pathways associated with glucose level: pathways 4 and 16 correlate with fused dependence structure, but pathways 16 and 43 do not; pathways 4 and 16 are significant pathways to the glucose level. The right top (b) represents the result of combination 2 of three pathways associated with glucose level: pathways 110 and 229 are correlated with fused dependence structure, but pathways 110 and 44 are not; pathways 110 and 229 are significant pathways to glucose level related to diabetes. The bottom left (c) represents the result of combination 4 of three pathways associated with glucose level: pathways 88 and 177, as well as pathways 110 and 103, are correlated with a fused dependence structure. The right bottom (d) represents the result of ordering 2 for combination 1 of three pathways associated with glucose level: pathways 110 and 16, as well as pathways 110 and 273, are not correlated; pathway 110 is a significant pathway to the glucose level.
FIGURE 6.

The hypothesis testing results of type II diabetes genetic pathway data. The solid blue circle indicates a significant pathway, while the dashed white circle indicates an insignificant pathway. The solid line denotes that two pathways are correlated with a fused structure, while the dashed line denotes that they are not. The left top (a) represents the result of ordering 1 for combination 1 of three pathways associated with glucose level: pathways 16 and 273, as well as pathways 273 and 110, are not correlated; pathway 16 is a significant pathway to the glucose level. The right top (b) represents the result of ordering 3 for combination 1 of three pathways associated with glucose level: neither pathways 16 and 110 nor pathways 110 and 273 have correlation; pathways 16 and 273 are significant pathways to glucose level related to diabetes. The bottom left (c) represents the result of ordering 1 for combination 2 of three pathways associated with glucose level: pathways 229 and 88, as well as pathways 88 and 110, are not correlated; pathways 229 and 110 are significant. The right bottom (d) represents the result of ordering 3 for combination 2 of three pathways associated with glucose level: pathways 110 and 229, along with pathways 229 and 88, show no correlation; pathway 229 is a significant pathway to the glucose level.
5.2. Identifying Significant Pathways to Distinguish Normal and Type II Diabetes Patients
Again, we considered 21 pathways to identify significant pathways distinguishing between normal and type II diabetes patients. Using the adj‐BF method for multiple testing. We estimated classes to decide of , ,finding that the ranges of , , and of are (2.890, 3.097), (4.977, 6.058), and (9.004, 9.517), respectively.
The testing results are in Tables 7, 8, 9. In total, the 14 pathways are insignificant, whereas 25 pathways are significant, strongly significant, or very strongly significant.
First, for each combination, we conducted BF‐GFKM to test the hypothesis ( vs. , ). Table 7 shows the of 21 pathways with combinations 1–7. “Alanine and aspartate metabolism” has the largest value 10.797. “Alanine and aspartate metabolism” was a significant pathway in Xu et al. [28]. “Oxidative phosphorylation” and “OXPHOS HG‐U133A probes” were significant pathways, but “ATP synthesis” with combination 1 was not significant. “Parkinson's disease” and “MAP00071 Fatty acid metabolism” were not significant for glucose levels, but were significant to distinguish between normal and diabetes patients [28].
Second, we also investigated the ordering of three pathways that affect binary outcomes. Using five pathways, we conducted our BF‐GFKM to test the hypothesis ( vs. , ). These testing results are summarized in Table 9. The “Ubiquitin mediated proteolysis” pathway is not a significant pathway for the binary phenotype, as well as for continuous glucose levels, with the inclusion of “ATP synthesis” and “JAK‐Stat Signaling Pathway.” The “JAK‐Stat Signaling Pathway” pathway is significant with orderings 1–3 for combinations 1–2. “ATP synthesis” is not significant with ordering 3 of combination 1, whereas it is significant with orderings 1–2 of combination 1. The “gamma‐Hexachlorocyclohexane degradation” pathway is significant with orderings 2–3 of combination 2, whereas it is not significant with ordering 3 of combination 2.
5.3. Significant Pathways for Continuous and Binary Responses
The significant pathways for both continuous and binary responses are in Tables 10 and 11. Pathways 4, 13, 88, 103, 110, 133, 154, 177, 229, 230, and 278 are significant pathways associated with glucose levels and distinguish between normal and type II diabetes patients among 21 pathways with seven combinations. Pathway 16 with orderings 1–2 of combination 1, pathway 110 with ordering 3 of combination 1, and orderings 1–2 of combination 2, and pathway 229 with orderings 1–3 of combination 2 are significant pathways associated with the glucose level and distinguish between normal and type II diabetes patients among five pathways with three orderings of two combinations.
TABLE 10.
The significant pathways, both continuous response and binary response, among 21 pathways with 7 combinations.
| ID | NAME | #genes | Combination |
|---|---|---|---|
| 4 | Alanine and aspartate metabolism | 18 | 1 |
| 110 | JAK‐Stat Signaling Pathway | 71 | 2 |
| 229 | Oxidative phosphorylation | 133 | 2 |
| 230 | OXPHOS HG‐U133A probes | 121 | 3 |
| 88 | gamma‐Hexachlorocyclohexane degradation | 41 | 4 |
| 177 | MAP00561 Glycerolipid metabolism | 43 | 4 |
| 103 | Histidine metabolism | 37 | 4 |
| 133 | MAP00190 Oxidative phosphorylation | 58 | 5 |
| 154 | MAP00380 Tryptophan metabolism | 60 | 6 |
| 278 | Wnt signaling pathway | 140 | 6 |
| 13 | Apoptosis | 92 | 7 |
TABLE 11.
The significant pathways for both continuous response and binary response among 5 pathways with 3 orderings of 2 combinations.
| ID | NAME | #genes | Combination | Ordering |
|---|---|---|---|---|
| 16 | ATP synthesis | 49 | 1 | 1 |
| 2 | ||||
| 110 | JAK‐Stat Signaling Pathway | 71 | 3 | |
| 2 | 1 | |||
| 2 | ||||
| 229 | Oxidative phosphorylation | 133 | 1 | |
| 2 | ||||
| 3 |
The test results of the “JAK‐Stat Signaling Pathway” for binary response are not changed by orderings 1–3 for combinations 1–2, but the test results of “ATP synthesis” and “gamma‐Hexachlorocyclohexane degradation” pathway for binary response are changed depending on orderings. We can consider that the “JAK‐Stat Signaling Pathway” is not affected by the other pathways, but the “ATP synthesis” and “gamma‐Hexachlorocyclohexane degradation” pathways are affected by other pathways. Therefore, if researchers are interested in determining whether specific pathways, such as the “ATP synthesis” and “gamma‐Hexachlorocyclohexane degradation” pathway, are influenced by other pathways, unlike “JAK‐Stat Signaling Pathway”, our BF‐GFKM method can be utilized to provide valuable insights into these relationships.
6. Conclusion
In this article, we have developed a flexible Bayesian inference based on BF using generalized fused kernel machine regression to test significantly correlated high‐dimensional functions (pathways) with the response variable, which can be continuous or binary. We developed a data‐driven, flexible Bayesian inference for adjusting BF for multiple tests.
We apply our method to pathway‐based analysis. Because pathways depend on each other, it is essential to analyze multiple pathways together by incorporating correlated structures. We demonstrate the advantages of genetic pathway‐based analysis of type II diabetes. Although some pathways are identified as significant to glucose levels or diabetes, they need to be further validated biologically.
We further note that the fused‐lasso structure is not intended to encode biological ordering or mechanistic pathway relationships. Instead, it is employed as a parsimonious statistical coupling prior to stabilize inference across pathway‐specific latent models. While network‐based models have been extensively studied in high‐dimensional settings, including pathway analysis, the existing literature predominantly focuses on network estimation rather than formal hypothesis testing. Developing principled testing procedures for multilevel graph‐structured network models for pathway genetic analysis remains an important direction for future research. Our proposed Bayesian framework contributes to this direction by enabling hypothesis testing within a simple yet flexible dependence structure.
Since the dependency among pathways is unknown in our application, we designed the BF‐GFKM to incorporate correlation structures among three pathways at a time using the fused penalty, conducting all possible distinct tests—a manageable task when limited to three pathways simultaneously. Our approach balances computational feasibility with the ability of the fused structure to capture correlations among pathways. We also note that our GFKM can be fitted with a group penalty and an fused penalty under the assumption that only adjacent pathways are correlated. In this setting, the pathway is correlated with the pathway but not with the . By jointly modeling multiple pathways and accounting for this correlation structure, our method provides a more accurate detection of pathways significantly associated with the response. In contrast, traditional pathway‐based analyses typically adopt a myopic strategy that tests one pathway at a time and assumes independence across pathways, which can lead to inflated false positives and false negatives. However, as more pathways are incorporated, the dimension of increases, often causing numerical instability when computing its inverse.
In general, when pathways are considered and an ordering can be established from correlations among genes within pathways, the BF‐GFKM test needs to be performed only once. In the absence of such information, however, all distinct orderings must be examined. Because each ordering is equivalent to its reverse, the number of distinct tests is , which grows factorially with and quickly becomes computationally infeasible. The complexity arises from fitting the GFKM with an structure, which models dependence only between adjacent pathways. While this structure captures local sequential correlation, it may not represent broader patterns of interaction. An extension to an structure, allowing each pathway to depend on multiple neighbors, provides a more flexible framework for modeling plausible inter‐pathway relationships, such as clustered pathways contributing jointly to the response. Although larger increases computational demands, particularly for inverting due to multiple correlations among pathways, this extension offers promise for improving both interpretability and realism.
Further investigation is needed to develop strategies for pathway ordering that incorporate correlations among genes within pathways. One approach is to order pathways by maximum pairwise correlations, while another is to identify a core sequence with the strongest correlations and then extend the ordering sequentially outward from both ends. Comprehensive simulation studies will be required to assess the performance and robustness of these approaches.
Author Contributions
PhilGeun Jin conducted all numerical analysis and wrote the manuscript, YoungHo Yun reviewed the manuscript, and Inyoung Kim developed a method and wrote and reviewed the manuscript.
Funding
The authors have nothing to report.
Conflicts of Interest
The authors declare no conflicts of interest.
Supporting information
Data S1. Supporting Information.
Acknowledgments
We sincerely thank the referees and the editor for their valuable comments and constructive suggestions, which have greatly improved the manuscript.
Appendix A. Full Conditional Distributions
The full conditional distributions for all parameters except for have closed forms. So, we draw the MCMC sample using Metropolis‐Hastings (MH) for and Gibbs sampling for all parameters except for . The detailed forms of the conditional distributions for continuous and binary responses are summarized in below sections.
Full Conditional Distributions of Parameters for Continuous Response
Using our GFKM (1) and prior specification (2), a joint posterior distribution under continuous response is then
The full conditional distributions for parameters except for have closed forms because of Bayesian hierarchical structure and conjugate priors. The full conditional distribution of is
where . When we draw MCMC samples from the full distribution for using MH, we found a very slow mixing, which results in a convergence problem. To overcome this issue, we adopt a restricted maximum likelihood (REML) and consider a REML‐based posterior distribution. Let denote the vector for the parameters of marginal covariance of Y for continuous response, that is, . The marginal covariance of Y for continuous response is then . The REML‐based posterior distribution of is
We sample from the MH algorithm. The proposal distribution for is . We set , and .
We denote , where is a vector, is kronecker product, and is identity matrix.
The full conditional distributions for all parameters are derived as follows:
where is an matrix with all‐ones components, and InvG denotes the inverse Gaussian distribution.
Full Conditional Distributions of Parameters for Binary Response
Using our model described in Section 2.2, we can have the joint posterior distribution for a binary response as follows,
where if is true and otherwise, which is indicator function. The full conditional distributions for parameters except for have closed forms because of Bayesian hierarchical structure and conjugate priors. For sampling , we also use the REML‐based posterior distribution of . As we described in Appendix A, a REML‐based posterior distribution for binary response is then
The proposal distribution for is the same as Appendix A. The full conditional distributions for binary responses are derived as follows:
where the full conditional distributions of are truncated standard normal distributions. The probit‐link is applied to implement Gibbs sampling for binary responses since full conditional distributions for all parameters except for have a closed form. These full conditional distributions are similar to those for continuous response by replacing by and .
Data Availability Statement
The data that support the findings of this study are openly available in BF‐GFKM at https://github.com/pgj439/BF‐GFKM.
References
- 1. Mootha V. K., Lindgren C. M., Eriksson K. F., et al., “PGC‐1alpha‐Responsive Genes Involved in Oxidative Phosphorylation Are Coordinately Downregulated in Human Diabetes,” Nature Genetics 34 (2003): 267–273. [DOI] [PubMed] [Google Scholar]
- 2. Cheng L., Kim I., and Pang H., “Bayesian Semiparametric Model for Pathway‐Based Analysis With Zero‐Inflated Clinical Outcomes,” Journal of Agricultural, Biological, and Environmental Statistics 21 (2016): 641–662. [Google Scholar]
- 3. Fang Z., Kim I., and Jung J., “Semiparametric Kernel‐Based Regression for Evaluating Interaction Between Pathway Effect and Covariate,” Journal of Agricultural, Biological, and Environmental Statistics 23 (2018): 129–152. [Google Scholar]
- 4. Goeman J. J., Geer v. d S A., Kort d F., and Houwelingen H. C. V., “A Global Test for Groups of Genes: Testing Association With a Clinical Outcome,” Bioinformatics 20 (2004): 93–99. [DOI] [PubMed] [Google Scholar]
- 5. Kemp D. M., Nirmala N. R., and Szustakowski J. D., “Extending the Pathway Analysis Framework With a Test for Transcriptional Variance Implicates Novel Pathway Modulation During Myogenic Differentiation,” Bioinformatics 23 (2007): 1356–1362. [DOI] [PubMed] [Google Scholar]
- 6. Kim I., Pang H., and Zhao H., “Bayesian Semiparametric Regression Models for Evaluating Pathway Effects on Continuous and Binary Clinical Outcomes,” Statistics in Medicine 31 (2012): 1633–1651. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7. Kim I., Pang H., and Zhao H., “Statistical Properties on Semiparametric Regression for Evaluating Pathway Effects,” Journal of Statistical Planning and Inference 143 (2013): 745–763. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8. Pang H., Kim I., and Zhao H., “Random Effects Model for Multiple Pathway Analysis With Applications to Type II Diabetes Microarray Data,” Statistics in Biosciences 7 (2015): 167–186. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9. Carpenter C. M., Zhang W., Gillenwater L., et al., “PaIRKAT: A Pathway Integrated Regressionbased Kernel Association Test With Applications to Metabolomics and COPD Phenotypes,” PLoS Computational Biology 17 (2021): 1–22. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10. Hwangbo S., Lee S., Lee S., Hwang H., Kim I., and Park T., “Kernel‐Based Hierarchical Structural Component Models for Pathway Analysis,” Bioinformatics 38 (2022): 3078–3086. [DOI] [PubMed] [Google Scholar]
- 11. Wendel B., Heidenreich M., Budde M., et al., “Kalpra: A Kernel Approach for Longitudinal Pathway Regression Analysis Integrating Network Information With an Application to the Longitudinal PsyCourse Study,” Frontiers in Genetics 13 (2022): 13, 10.3389/fgene.2022.1015885. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12. Lin J. and Kim I., “Gaussian Process Selections in Semiparametric Multi‐Kernel Machine Regression for Multi‐Pathway Analysis,” Statistical Analysis and Data Mining 17 (2024): 1–17. [Google Scholar]
- 13. Liu D., Lin X., and GhoshLiu D., “Semiparametric Regression of Multi‐Dimensional Genetic Pathway Data: Least Square Kernel Machines and Linear Mixed Models,” Biometrics 63 (2007): 1079–1088. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14. Maity A. and Lin X., “Powerful Tests for Detecting a Gene Effect in the Presence of Possible Gene–Gene Interactions Using Garrote Kernel Machines,” Biometrics 67 (2011): 1271–1284. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15. Cai T., Tonini G., and Lin X., “Kernel Machine Approach to Testing the Significance of Multiple Genetic Markers for Risk Prediction,” Biometrics 67 (2011): 975–986. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16. Zhang L. and Kim I., “Semiparametric Bayesian Kernel Survival Model for Evaluating Pathway Effects,” Statistical Methods in Medical Research 28 (2019): 3301–3317. [DOI] [PubMed] [Google Scholar]
- 17. Stingo F. C., Chen Y. A., Tadesse M. G., and Vannucci M., “Incorporating Biological Information Into Linear Models: A Bayesian Approach to the Selection of Pathways and Genes,” Annals of Applied Statistics 5 (2011): 1–24. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18. Fang Z., Kim I., and Schaumont P., “Flexible Variable Selection for Recovering Sparsity in Nonadditive Nonparametric Models,” Biometrics 72 (2016): 1155–1163. [DOI] [PubMed] [Google Scholar]
- 19. Tibshirani R., Saunders M., Rosset S., Zhu J., and Knight K., “Sparsity and Smoothness via the Fused Lasso,” Journal of the Royal Statistical Society. Series B, Statistical Methodology 67, no. 1 (2005): 91–108. [Google Scholar]
- 20. Li C. and Li H., “Network‐Constrained Regularization and Variable Selection for Analysis of Genomic Data,” Bioinformatics 24, no. 9 (2008): 1175–1182. [DOI] [PubMed] [Google Scholar]
- 21. Kim I., Shan L., Lin J., Gao W., Kim B. J., and Mahmoud H., “Multiple and Multilevel Graphical Models,” Wiley Interdisciplinary Reviews: Computational Statistics 12 (2020): e1497. [Google Scholar]
- 22. Shan L., Cheng L., and Kim I., “Joint Estimation of Two‐Level Gaussian Graphical Models Across Multiple Classes,” Journal of Computational and Graphical Statistics 29 (2020): 562–579. [Google Scholar]
- 23. Kim B. J. and Kim I., “Joint Semiparametric Kernel Network Regression,” Statistics in Medicine 42 (2023): 5247–5265. [DOI] [PubMed] [Google Scholar]
- 24. Andrews D. F. and Mallows C. L., “Scale Mixtures of Normal Distributions,” Journal of the Royal Statistical Society. Series B, Statistical Methodology 36 (1974): 99–102. [Google Scholar]
- 25. Kass R. E. and Raftery A. E., “Bayes Factors,” Journal of the American Statistical Association 90 (1995): 773–795. [Google Scholar]
- 26. Weinberg M. D., “Computing the Bayes Factor From a Markov Chain Monte Carlo Simulation of the Posterior Distribution,” Bayesian Analysis 7, no. 3 (2012): 737–770. [Google Scholar]
- 27. Schweiger R., Weissbrod O., Rahmani E., et al., “RL‐SKAT: An Exact and Efficient Score Test for Heritability and Set Tests,” Genetics 207 (2017): 1275–1283. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28. Xu Y., Kim I., and Carroll R. J., “A Hybrid Omnibus Test for Generalized Semiparametric Single‐Index Models With High‐Dimensional Covariate Sets,” Biometrics 75 (2019): 757–767. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29. Lin J., Wu H., Tarr P. T., et al., “Transcriptional Co‐Activator PGC‐1 Alpha Drives the Formation of Slow‐Twitch Muscle Fibres,” Nature 418 (2003): 797–801. [DOI] [PubMed] [Google Scholar]
- 30. Wang X., Shaw S., Amiri F., Eaton D. C., and Marrero M. B., “Inhibition of the Jak/STAT Signaling Pathway Prevents the High Glucose‐Induced Increase in Tgf‐𝛽 and Fibronectin Synthesis in Mesangial Cells,” Diabetes 51 (2002): 3505–3509. [DOI] [PubMed] [Google Scholar]
- 31. Kanehisa M., Goto S., Kawashima S., Okuno Y., and Hattori M., “The KEGG Resource for Deciphering the Genome,” Nucleic Acids Research 32 (2004): 277–280. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data S1. Supporting Information.
Data Availability Statement
The data that support the findings of this study are openly available in BF‐GFKM at https://github.com/pgj439/BF‐GFKM.
