Abstract
Transcriptome-wide association study (TWAS) has successfully identified numerous complex disease susceptibility genes in the post-genome-wide association study (GWAS) era. Over the past 3 years, the focus of TWAS algorithms has shifted from merely identifying associations to understanding how single nucleotide polymorphisms (SNPs) regulate gene expression, with a growing emphasis on incorporating fine-mapping techniques. Additionally, the rapid increase in GWAS summary statistics, driven largely by the UK Biobank and other consortia, has made it essential to update our webTWAS resource. To address these challenges and meet the growing needs of researchers, we developed webTWAS 2.0, an updated platform for identifying susceptibility genes for human complex diseases using TWAS. Additionally, webTWAS 2.0 provides an online TWAS analysis tool that simplifies conducting TWAS analyses. The updated resource includes 7247 GWAS summary statistics covering 1588 complex human diseases from 192 publications. It also incorporates multiple TWAS methods, such as sTF-TWAS, 3′aTWAS and GIFT, along with an updated interactive visualization tool that allows users to easily explore significant associations across different methods. Other upgrades include a personalized online analysis tool for user-submitted GWAS data and a refined search function that makes it easier to identify relevant associations and meet diverse user needs more efficiently. webTWAS 2.0 is freely accessible at http://www.webtwas.net.
Graphical Abstract
Graphical Abstract.
Introduction
Transcriptome-wide association study (TWAS) was first introduced in 2015 (1), and is a powerful method for integrating genetic variation with gene expression to identify susceptibility genes for complex human diseases. TWAS uses a reference panel to train the weights between whole-genome sequences and gene expression, followed by independent datasets to predict the genetically regulated expression (GReX) component. The GReX is then associated with traits to identify susceptibility genes related to human complex diseases. TWAS addresses several limitations of genome-wide association study (GWAS), including difficulties in interpreting significant signals, as many are located in noncoding regions or result from linkage disequilibrium (LD). Additionally, TWAS significantly improves the power of association studies by more effectively aggregating genetic variants.
The traditional TWAS framework involves two key steps. The first is the training phase, where the relationship between genotype and transcriptome is modeled to determine GReX. The second step associates GReX with phenotypic traits in the target genotype data. In the initial development of TWAS, most efforts focused on improving the accuracy of predicted transcriptomes including PrediXcan (1), TWAS-FUSION (2), UTMOST (3), TIGAR (4) and kTWAS (5), among others (6–10). Over the past 3 years, the focus of TWAS algorithms has shifted toward investigating how SNPs regulate gene expression and identifying causal genes through methods such as fine mapping. This includes integrating prior biological knowledge, such as transcription factors binding to cis-regulatory elements (11), pathway networks (12), transcript isoforms (13), 3′ untranslated region alternative polyadenylation (14) and single-cell transcriptome data (15), among others (16–20). These improvements have significantly improved the power of TWAS to identify disease-associated genes. As a result, webTWAS needs to be updated to incorporate these cutting-edge methods and offer researchers a more powerful and accurate platform for gene–trait association analysis.
The aforementioned TWAS methods have been successfully applied to various complex human diseases, including mental disorders (1,3,21), cardiovascular diseases (22–24) and cancers (25–27), among others (20,28–30). These studies highlight the significant potential of TWAS in identifying disease-related susceptibility genes and unraveling the genetic mechanisms underlying complex diseases. Despite the growing interest in TWAS, only three related databases have been developed: TWAS-hub (http://twas-hub.org/), webTWAS (31) and TWAS Atlas (32) have been established to collect and restore TWAS-related datasets and gene–trait associations. TWAS-hub, the first TWAS database introduced in 2018, contains 75 951 gene–trait associations across 342 disease and nondisease traits. webTWAS, based on 1298 curated GWAS datasets, implements three TWAS tools—S-PrediXcan, TWAS-FUSION and UTMOST—and includes 235 064 gene–trait associations across 887 human diseases. TWAS Atlas manually reviewed 200 curated TWAS application publications on human traits, collecting 401 266 gene–trait associations along with online search and visualization tools. These databases offer valuable resources for studying complex human traits using TWAS. However, there are limitations in these TWAS-related databases. TWAS-hub provides analysis results from only one TWAS method, TWAS-FUSION, and has not been updated since 2019. TWAS Atlas collects gene–trait associations from 200 published papers but does not offer online TWAS analysis or results based on GWAS summary statistics using multiple state-of-the-art TWAS methods. webTWAS includes 1298 GWAS datasets and three TWAS methods. With the rapid increase in available GWAS summary statistics, the gradual release of large cohorts such as the UK Biobank (33) (UKBB), and the growing number of TWAS methods, there is an urgent need to update webTWAS. This update includes more GWAS summary statistics and TWAS methods to provide a comprehensive TWAS resource database and an online computational platform. Therefore, it is essential to update webTWAS with more GWAS summary statistics data and to recalculate the disease-susceptible genes using updated TWAS algorithms.
To address these needs, we have updated webTWAS from version 1.0 to 2.0, adding 7247 GWAS summary statistics across 1588 human complex diseases from 192 publications and implementing six representative TWAS methods for analysis. webTWAS 2.0 now restores 661 396 gene–trait associations, covering 1588 diseases, along with detailed disease-relevant tissue, gene and disease information. We addressed that diseases are tissue-specific and identified the most relevant tissue for each disease. Additionally, we developed an interactive visualization tool that allows users to easily explore gene–trait associations, along with an online TWAS analysis tool that provides free access to conduct TWAS analyses. All data in webTWAS 2.0 can be freely searched and downloaded, offering researchers a valuable resource for gene-disease associations. The complete information and resources in webTWAS 2.0 are available at http://www.webtwas.net/.
webTWAS version 2.0 design
Here, we present webTWAS version 2.0, a comprehensive database for disease susceptibility genes identified by multiple TWAS methods. Our platform offers gene–trait associations across a range of complex human diseases and provides an easy-to-use online TWAS analysis tool (Figure 1). Publicly available GWAS summary statistics were collected (Figure 1A), subjected to quality control (Figure 1B), and standardized to ensure clarity and avoid misinterpretation (Figure 1C). For TWAS analysis, six representative methods, including S-PrediXcan (27), TWAS-FUSION (2), UTMOST (3), sTF-TWAS (11), 3′aTWAS (14) and GIFT (17), were employed (Figure 1D). An important feature of webTWAS 2.0 is its ability to account for tissue specificity, enabling researchers to focus on the most relevant tissues for each disease for gene–trait association analysis. Additionally, a user-friendly interface was developed to support the main functions of webTWAS 2.0, including search, online analysis, visualization and data download (Figure 1E). Detailed gene–trait associations for any disease and gene can be easily accessed. The search interface was specifically designed to facilitate quick access to associations of interest. All data in webTWAS 2.0 can be freely downloaded from the database website.
Figure 1.
Overview of webTWAS 2.0. (A) GWAS summary statistics were collected from two main sources and subjected to preprocessing (B) to ensure data quality. (C) The datasets were then standardized through information mapping. (D) Six TWAS methods were employed for the analysis of the GWAS summary statistics. (E) The overall structure of the webTWAS 2.0 database.
Updates and new features
GWAS summary statistics update and pre-processing
The updated webTWAS 2.0 contains a total number of 7247 GWAS summary statistics for various diseases which were used to conduct multiple TWAS analyses. Following the collection criteria outlined in our previous work, CAUSALdb (34) and webTWAS (31), the public GWAS summary statistics were primarily sourced from two categories: UKBB and non-UKBB (Figure 1A). The UKBB GWAS summary statistics were obtained from Neale Lab UKBB v3 (http://www.nealelab.is/uk-biobank), Gene ATLAS (35) and GWAS ATLAS (36). These three sources differ in sample selection, quality control processes and types of association models. For non-UKBB GWAS summary statistics, data were collected from multiple public databases, including GWAS Catalog (37), LD Hub (38), GRASP (39), PhenoScanner (40), dbGaP (41), PGC (https://pgc.unc.edu), MAGIC (http://www.magicinvestigators.org), SSGAC (https://www.thessgac.org) and JENGER (http://jenger.riken.jp/en).
Quality control was conducted on these GWAS summary statistics. First, for each dataset, relevant information, including sample size and population-related details, was extracted from the original publication. Datasets lacking this information were excluded. For duplicate datasets collected from different sources, only the dataset with the most comprehensive information was retained. Specifically, GWAS summary statistics with unqualified columns, such as those missing standard rsIDs or beta values necessary for computation, were excluded. Missing data, such as Z-values, were imputed, and rsIDs and genomic coordinates were standardized to the GRCh37 version. For the trait information of each GWAS summary statistic, we manually mapped each reported trait extracted from the original publication to the medical subject headings (MeSH) term to ensure standardization and prevent misunderstanding. Regarding the population information, we used the five superpopulations defined by the 1000 Genomes Project (42) and, limited by the reference panel of TWAS analysis, retained only GWAS summary statistics of European ancestry (EUR) for analysis. Overall, webTWAS 2.0 contains a total number of 7247 GWAS summary statistics related to 27 complex human disease types and 192 publications (Figure 2A–C).
Figure 2.
Statistical overview of webTWAS 2.0. (A) Number of mapped MeSH terms categorized by disease type. (B) Number of gene–disease associations for each disease type. (C) Total number of GWAS summary statistics collected by year. (D) Distribution of gene–disease associations identified by each TWAS method. (E) Distribution of associated genes per disease. (F) Distribution of associated diseases per gene.
Newly integrated TWAS methods for analyzing GWAS summary statistics
With the rapid development of TWAS methods over the past decade, >20 methods have been introduced to identify trait-associated genes using various biological assumptions and different mathematical models. In webTWAS 2.0, we applied six popular and representative methods, including S-PrediXcan (27), TWAS-FUSION (2), UTMOST (3), sTF-TWAS (11), 3′aTWAS (14) and GIFT (17). To ensure the accuracy of TWAS analysis and prevent the misuse of unmatched tissue reference panels—which could affect results due to the tissue-specific feature of gene expression—we employed PASCAL (43) and deTS (44) algorithms to identify disease-specific tissues for each GWAS summary statistics (Figure 1D). First, PASCAL calculated disease-related gene scores, selecting those with P-value <0.05. Subsequently, deTS performed a chi-square association test on the disease-related gene set for each GWAS dataset to identify significantly related tissues. For TWAS analysis, we selected the top three tissues for each GWAS summary statistic to determine the appropriate reference panel for the TWAS method parameter settings.
S-PrediXcan computes PrediXcan results using GWAS summary statistics, utilizing elastic net-based models on GTEx v8 release data (45) for each GWAS summary statistics analysis. For TWAS-FUSION, gene weights from GTEx v8 multitissue expression data (http://gusevlab.org/projects/fusion/) are used, with the best model selected from BLUP, BSLMM, LASSO and Elastic Net, and top SNP within the default setting. For UTMOST, association tests are conducted on the top three tissues for each GWAS summary statistic, and gene–trait associations across these tissues are combined using the joint GBJ test with imputation models jointly trained on GTEx data provided by UTMOST. Additionally, two TWAS methods representing different biological assumptions, including transcription factors and 3′ untranslated region alternative polyadenylation were used to calculate gene–trait associations. For sTF-TWAS, weights were provided for four tissues (breast, lung, prostate and brain), and gene–trait associations were calculated for each GWAS summary statistics based on these tissues. For 3′aTWAS, 3′aQTLs from 49 tissues in the GTEx v8 cohort were used (46), with the top three tissues of each GWAS summary statistics selected for the association test. For GIFT, genome-wide eQTL summary statistics from GEUVADIS data provided by GIFT (17) and the two-stage version of GIFT were used to conduct the association tests. For each TWAS method, the default LD matrix provided by the respective method, all sourced from the 1000 Genomes Project (42), was used. The adjusted significance threshold was set at 0.05 divided by the total number of computed genes for each method, and default parameters were applied for all TWAS methods.
After performing TWAS analysis with these methods, a total of 661 396 significant associations were identified. The number of significant associations detected by each method is illustrated in a pie chart in Figure 2D. Specifically, the GIFT method identified the most associations (358 355), and the average number of associations per method was 110 232. For each disease, the average number of associated genes was 208.90. The detailed distribution of associated genes per disease is illustrated in Figure 2E. Conversely, for each gene, the average number of associated diseases was 19.65, with the detailed distribution of associated diseases per gene shown in Figure 2F.
Updated interactive visualization tool for assessing gene–trait associations
To facilitate user access to disease or gene associations of interest, an updated interactive visualization tool was added. On the detailed disease information page, a Manhattan plot displays all associations for each disease. When users hover over a node, detailed information about each association—including gene, gene location, method, reference tissue, P-value, Z-value and GWAS summary statistics ID—is shown (Figure 3A). Two customizable features are available to improve user experience. First, users can select different methods to show or hide associations identified by each, allowing for easy confirmation of consistency across methods. Second, users can select specific chromosomes to view detailed information about that region, which can be combined with the first feature for a clearer analysis of disease-related genes.
Figure 3.
Main pages in webTWAS 2.0. (A) Interactive visualization tool displaying detailed gene-disease associations for a selected disease. (B) Online analysis tool allowing users to perform customized TWAS analyses. (C) Trait page example for heart failure (webTWAS ID W04225). (D) Gene page example for the gene FUT11. (E) Search page offers three search options: traits, genes or publications. (F) Download page provides two types of downloadable files.
Updated online TWAS analysis tool for analyzing GWAS summary statistics
In addition to the precomputed associations and interactive visualization tool available for users, webTWAS 2.0 also offers an updated online TWAS analysis tool (Figure 3B). To lower the barrier to entry for TWAS analysis, the ‘TWAS Online’ page allows users to upload GWAS summary statistics files containing necessary columns such as rsID, effect allele, noneffect allele, P-value or Z-value. Users can set essential parameters for each method, including tissue reference panels, population, significant thresholds, etc. webTWAS 2.0 now supports four TWAS methods: S-PrediXcan, TWAS-FUSION, sTF-TWAS and 3′aTWAS, with the option to modify the default P-value cutoff as needed. To help users track the progress of their analysis jobs, each job is assigned a unique ID. Users can query job progress on the job search page, download results or follow the job’s URL provided via email if they have entered their email address during parameter setup.
Database construction and improved user interface
webTWAS 2.0 employs Spring Boot for the back-end architecture. On the front end, the interface is built with Vue.js, leveraging Element UI for a smooth user experience. MySQL serves as the database system, allowing for fast querying of GWAS summary statistics and TWAS-associated genes linked to diseases. The TWAS analysis server adopts an asynchronous approach to efficiently manage and schedule user-submitted tasks. All processes are logged and can be tracked through the webTWAS 2.0 dashboard. A detailed system layout is provided in Figure 1.
webTWAS 2.0 features a user-friendly interface and multiple data access options, allowing users to query the database in several steps:
On the ‘Home’ page, a central search box enables users to easily investigate gene–trait associations. Users can search by diseases of interest, with matching disease entries displayed in a drop-down box. All possible records are shown on the search results page, where each column supports sorting to help users more efficiently find the data of interest. The entire dataset can also be downloaded via a button at the top right of the page. For detailed information about specific associations, users can click on the ‘Reported Trait’ or ‘webTWAS ID’ column to access the detailed page for each disease entry.
On the ‘Disease’ page, the complete disease entry results of webTWAS 2.0 are displayed. A MeSH tree is available at the top left, allowing users to easily explore diseases of interest. By clicking on a specific disease, users are taken to a detailed gene–trait association page. The top table provides detailed information about the trait, including the ‘Reported Trait’ from the original publication, the mapped ‘MeSH Term’, ‘MeSH ID’, ‘Trait Type’, ‘Trait Description’ and ‘More Information’ hyperlinks to additional details about the disease. In the middle of the page, an interactive Manhattan plot displays all associations for the disease, allowing users to show or hide associations identified by different methods by clicking the corresponding method buttons. For user convenience, the Manhattan plot can be separated by chromosome to facilitate the examination of associated genes in different regions. Hovering over nodes in the Manhattan plot reveals detailed association information, including gene name, gene location, method, tissue reference, P-value, Z-value and webTWAS ID. The quantile–quantile (Q–Q) plot for visualizing P-value distributions for each method is presented below the Manhattan plot. This allows users to assess potential inflation and compare the performance of different methods. Both the Q–Q plot and the Manhattan plot can be downloaded for each method and GWAS summary statistic to facilitate user analysis. Furthermore, all gene associations for each method are available in the database, providing comprehensive access and enabling users to apply customized P-value thresholds as needed. The bottom table presents detailed information about each association, which can be freely downloaded by clicking the download button at the top right of the table. Users can also sort columns to browse detailed information about each association more easily (Figure 3C).
On the ‘Gene’ page, the complete gene entries from webTWAS 2.0 are displayed. The table provides detailed information about each gene, including ‘Gene Symbol’, ‘Ensembl ID’, ‘Location’, ‘Gene Type’, ‘Synonyms’ and ‘Number of Associated Traits’. Blue font entries are clickable, leading to a detailed gene information page, and each column supports sorting to easily find genes of interest. On the detailed gene page, the top table presents additional information about each gene, including a gene description and summary from the NCBI database, gene group and hyperlinks to the gene information pages on the ENSEMBL (47), GeneCards (48) and NCBI (49) databases. The bottom table displays associations related to the gene, with options for downloading the data and more detailed information about these associations (Figure 3D).
On the ‘Search’ page, users have three options to easily access associations of interest: search by trait, gene or publication. Each search box displays example prompt words to assist users. They can enter information such as trait name, gene name, Ensembl ID, gene location or PMID, and select matching entries from the drop-down box to more efficiently access the associations (Figure 3E).
On the ‘Downloads’ page, two download options are provided. Users can either enter a disease of interest, select the matching disease entry and click the download button next to the search box, or directly download all associations identified by each method. The table on the page displays detailed information about each download file, including file format, file size and download link (Figure 3F).
On the ‘TWAS Online’ page, an updated online TWAS analysis tool is provided for users to easily conduct TWAS analyses. The process is divided into two steps: upload the necessary GWAS summary statistics and configure method parameters. For user convenience, example data and demo results are available on the page to help users prepare their data and follow the job process. After uploading the GWAS data, click the ‘Load data’ button. Parameters will be auto-filled in the appropriate fields. Users then need to fill in essential parameters such as sample size, check other necessary parameters and optionally enter an email address to receive notifications about job progress. Afterward, click the ‘Next’ button to proceed to the configuration page. On the configuration page, users can choose the analysis methods, including S-PrediXcan, TWAS-FUSION, sTF-TWAS and 3′aTWAS, and set parameters for each method, such as significance thresholds, statistical models and tissue references. Specifically, MASHR (Multivariate Adaptive Shrinkage in R) (50) is well-suited for integrating data from multiple tissues, allowing for more accurate predictions by borrowing strength from shared SNP effects. This approach is recommended when users have access to multitissue data. Elastic-net is a regularization method that combines LASSO and Ridge regression, making it robust for single-tissue analyses with limited data. This distinction helps users choose based on data complexity and study design. Once all fields are completed, click ‘Submit’ to start the job. A page with a job ID will be generated, allowing users to track the job status. The final analysis results will also be displayed on this page. After the analysis is complete, results will be visualized in a Manhattan plot and also displayed in a table, and the results can be freely downloaded by the user. Alternatively, the ‘Job Search’ page can be used to check the status of a job by entering the job ID.
Conclusions, limitation and future directions
In the initial version of the webTWAS database, webTWAS 1.0, only 1298 GWAS summary statistics and 3 TWAS methods were included. However, with the development of TWAS methods and the release of large cohorts, the number of GWAS summary statistics and TWAS methods has increased significantly over the past 5 years. The rapid growth of TWAS research has demonstrated its power as a tool for decoding the genetic determinants of complex human diseases, highlighting the need for a more comprehensive TWAS database and an update to webTWAS 1.0. The updated webTWAS 2.0 now includes 7247 GWAS summary statistics of 1588 diseases from 192 publications. It now provides TWAS analysis results from six popular and representative methods, along with interactive visualizations for easy assessment of gene–trait associations. Additionally, a user-friendly online TWAS analysis tool has been introduced to facilitate these analyses. These resources are valuable for further exploration of genetic determinants of human complex diseases, discovery of disease targets and promotion of research. While multiple resources are available for GWAS, such as GWAS Catalog (37), GWASDB (51), GWAS Atlas (52), CAUSALdb (34) and GRASP (39), TWAS resources have been limited to TWAS-hub, TWAS Atlas and webTWAS. webTWAS 2.0 addresses this gap in comprehensive TWAS resources and aims to facilitate TWAS analysis. For performance, the average number of significant genes per tissue for S-PrediXcan in webTWAS 2.0 is 11.9, similar to TWAS Atlas at 10.8. For TWAS-FUSION, webTWAS 2.0 shows an average of 10.8 significant genes per tissue, higher than TWAS-hub at 5.5 but less than TWAS Atlas at 15.6. These results demonstrate that webTWAS 2.0 performs comparably to other TWAS platforms while providing a significantly larger dataset and incorporating a wider range of TWAS methods.
Although webTWAS 2.0 integrates six TWAS methods, several others are not yet included in the platform, particularly those that use kernel machines (5,9,10,53) for the second step of association analysis, which requires access to genotype data. In webTWAS 2.0, we have integrated the multitissue TWAS method UTMOST to leverage the advantages of using multiple tissues in the analysis, improving both power and accuracy. Associations for all tissues have been calculated using S-PrediXcan, providing comprehensive data access and allowing users to select their preferred tissues. In future versions of webTWAS, we plan to integrate additional multitissue methods, such as TisCoMM (8), to further improve the platform’s capability for robust gene–trait association studies using multiple tissue types. One shortcoming of traditional TWAS methods is their inability to account for uncertainty in the imputed gene expression, which can lead to reduced statistical power and less reliable results. To address this, future updates of webTWAS will integrate the CoMM series TWAS methods (8,54–56), which provide a more sophisticated framework by incorporating this uncertainty directly into the analysis. The CoMM methods, particularly CoMM-S2 (55), offer better precision and statistical power by jointly modeling both gene expression imputation and association with traits. Moving forward, it will be important to consider incorporating genotype data alongside GWAS summary statistics for more comprehensive TWAS analyses. Currently, webTWAS 2.0 only includes GWAS data for the EUR population due to limitations in available reference panels. To address this, we plan to regularly update the database with new GWAS data and TWAS methods. As larger cohorts become available, future updates will incorporate GWAS data from non-EUR populations and additional TWAS methods to ensure broader applicability and greater accuracy.
Acknowledgements
The computational resources generously provided by the High Performance Computing Center of Nanjing Medical University are greatly appreciated.
Author contributions: C.C. (Conceptualization, Funding acquisition, Project administration, Writing—original draft, Writing—review & editing); M.S. (Software, Visualization, Writing—original draft, Writing—review & editing); J.W. (Data curation, Writing—review & editing); Z.L. (Software, Visualization); H.C. (Writing—review & editing); T.Y. (Writing—review & editing); M.J.L. (Data curation, Writing—review & editing); Y.D. (Project administration, Writing—review & editing); and Q.Z. (Conceptualization, Funding acquisition, Project administration, Supervision, Writing—review & editing).
Contributor Information
Chen Cao, Key Laboratory for Bio-Electromagnetic Environment and Advanced Medical Theranostics, School of Biomedical Engineering and Informatics, Nanjing Medical University,101 Longmian Ave, Nanjing, Jiangsu 211166, China.
Mengting Shao, Key Laboratory for Bio-Electromagnetic Environment and Advanced Medical Theranostics, School of Biomedical Engineering and Informatics, Nanjing Medical University,101 Longmian Ave, Nanjing, Jiangsu 211166, China.
Jianhua Wang, Department of Systems Pharmacology and Translational Therapeutics, University of Pennsylvania—Perelman School of Medicine, 421 Curie Blvd, Philadelphia, PA 19104, USA.
Zhenghui Li, Key Laboratory for Bio-Electromagnetic Environment and Advanced Medical Theranostics, School of Biomedical Engineering and Informatics, Nanjing Medical University,101 Longmian Ave, Nanjing, Jiangsu 211166, China.
Haoran Chen, Key Laboratory for Bio-Electromagnetic Environment and Advanced Medical Theranostics, School of Biomedical Engineering and Informatics, Nanjing Medical University,101 Longmian Ave, Nanjing, Jiangsu 211166, China.
Tianyi You, Department of Pharmacology, School of Basic Medical Sciences, Tianjin Medical University, 22 Qixiangtai Road, Tianjin 300203, China.
Mulin Jun Li, Department of Pharmacology, School of Basic Medical Sciences, Tianjin Medical University, 22 Qixiangtai Road, Tianjin 300203, China.
Yijie Ding, Yangtze Delta Region Institute (Quzhou), University of Electronic Science and Technology of China, 1 Chengdian Road, Quzhou, Zhejiang 324003, China.
Quan Zou, Yangtze Delta Region Institute (Quzhou), University of Electronic Science and Technology of China, 1 Chengdian Road, Quzhou, Zhejiang 324003, China.
Data availability
The data underlying this article are available in webTWAS 2.0 (http://www.webtwas.net) and can be freely downloaded.
Funding
National Natural Science Foundation of China [Grant Nos. 62471240, 62102068, 62231013]. Funding for open access charge: National Natural Science Foundation of China.
Conflict of interest statement. None declared.
References
- 1. Gamazon E.R., Wheeler H.E., Shah K.P., Mozaffari S.V., Aquino-Michaels K., Carroll R.J., Eyler A.E., Denny J.C., Nicolae D.L., Cox N.J.et al.. A gene-based association method for mapping traits using reference transcriptome data. Nat. Genet. 2015; 47:1091–1098. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2. Gusev A., Ko A., Shi H., Bhatia G., Chung W., Penninx B.W., Jansen R., de Geus E.J., Boomsma D.I., Wright F.A.et al.. Integrative approaches for large-scale transcriptome-wide association studies. Nat. Genet. 2016; 48:245–252. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3. Hu Y., Li M., Lu Q., Weng H., Wang J., Zekavat S.M., Yu Z., Li B., Gu J., Muchnik S.et al.. A statistical framework for cross-tissue transcriptome-wide association analysis. Nat. Genet. 2019; 51:568–576. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4. Nagpal S., Meng X., Epstein M.P., Tsoi L.C., Patrick M., Gibson G., De Jager P.L., Bennett D.A., Wingo A.P., Wingo T.S.et al.. TIGAR: an improved bayesian tool for transcriptomic data imputation enhances gene mapping of complex traits. Am. J. Hum. Genet. 2019; 105:258–266. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5. Cao C., Kwok D., Edie S., Li Q., Ding B., Kossinna P., Campbell S., Wu J., Greenberg M., Long Q.. kTWAS: integrating kernel machine with transcriptome-wide association studies improves statistical power and reveals novel genes. Brief. Bioinform. 2021; 22:bbaa270. [DOI] [PubMed] [Google Scholar]
- 6. Parrish R.L., Gibson G.C., Epstein M.P., Yang J.. TIGAR-V2: efficient TWAS tool with nonparametric bayesian eQTL weights of 49 tissue types from GTEx V8. HGG Adv. 2022; 3:100068. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7. Barbeira A.N., Pividori M., Zheng J., Wheeler H.E., Nicolae D.L., Im H.K.. Integrating predicted transcriptome from multiple tissues improves association detection. PLoS Genet. 2019; 15:e1007889. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8. Shi X., Chai X., Yang Y., Cheng Q., Jiao Y., Chen H., Huang J., Yang C., Liu J.. A tissue-specific collaborative mixed model for jointly analyzing multiple tissues in transcriptome-wide association studies. Nucleic Acids Res. 2020; 48:e109. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9. Cao C., Kossinna P., Kwok D., Li Q., He J., Su L., Guo X., Zhang Q., Long Q.. Disentangling genetic feature selection and aggregation in transcriptome-wide association studies. Genetics. 2022; 220:iyab216. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10. Tang S., Buchman A.S., De Jager P.L., Bennett D.A., Epstein M.P., Yang J.. Novel variance-component TWAS method for studying complex human diseases with applications to Alzheimer’s dementia. PLoS Genet. 2021; 17:e1009482. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11. He J., Wen W., Beeghly A., Chen Z., Cao C., Shu X.O., Zheng W., Long Q., Guo X.. Integrating transcription factor occupancy with transcriptome-wide association analysis identifies susceptibility genes in human cancers. Nat. Commun. 2022; 13:7118. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12. Zhang L., Ju T., Jin X., Ji J., Han J., Zhou X., Yuan Z.. Network regression analysis for binary and ordinal categorical phenotypes in transcriptome-wide association studies. Genetics. 2022; 222:iyac153. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13. Bhattacharya A., Vo D.D., Jops C., Kim M., Wen C., Hervoso J.L., Pasaniuc B., Gandal M.J.. Isoform-level transcriptome-wide association uncovers genetic risk mechanisms for neuropsychiatric disorders in the human brain. Nat. Genet. 2023; 55:2117–2128. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14. Cui Y., Arnold F.J., Peng F., Wang D., Li J.S., Michels S., Wagner E.J., La Spada A.R., Li W. Alternative polyadenylation transcriptome-wide association study identifies APA-linked susceptibility genes in brain disorders. Nat. Commun. 2023; 14:583. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15. Song X., Ji J., Rothstein J.H., Alexeeff S.E., Sakoda L.C., Sistig A., Achacoso N., Jorgenson E., Whittemore A.S., Klein R.J.et al.. MiXcan: a framework for cell-type-aware transcriptome-wide association studies with an application to breast cancer. Nat. Commun. 2023; 14:377. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16. Zhao S., Crouse W., Qian S., Luo K., Stephens M., He X.. Adjusting for genetic confounders in transcriptome-wide association studies improves discovery of risk genes of complex traits. Nat. Genet. 2024; 56:336–347. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17. Liu L., Yan R., Guo P., Ji J., Gong W., Xue F., Yuan Z., Zhou X.. Conditional transcriptome-wide association study for fine-mapping candidate causal genes. Nat. Genet. 2024; 56:348–356. [DOI] [PubMed] [Google Scholar]
- 18. Liu L., Zeng P., Xue F., Yuan Z., Zhou X.. Multi-trait transcriptome-wide association studies with probabilistic mendelian randomization. Am. J. Hum. Genet. 2021; 108:240–256. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19. Dai Q., Zhou G., Zhao H., Võsa U., Franke L., Battle A., Teumer A., Lehtimäki T., Raitakari O.T., Esko T.et al.. OTTERS: a powerful TWAS framework leveraging summary-level reference data. Nat. Commun. 2023; 14:1271. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20. Zhang Z., Bae Y.E., Bradley J.R., Wu L., Wu C.. SUMMIT: an integrative approach for better transcriptomic data imputation improves causal gene identification. Nat. Commun. 2022; 13:6336. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21. Raj T., Li Y.I., Wong G., Humphrey J., Wang M., Ramdhani S., Wang Y.C., Ng B., Gupta I., Haroutunian V.et al.. Integrative transcriptome analyses of the aging brain implicate altered splicing in Alzheimer’s disease susceptibility. Nat. Genet. 2018; 50:1584–1592. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22. Al-Barghouthi B.M., Rosenow W.T., Du K.P., Heo J., Maynard R., Mesner L., Calabrese G., Nakasone A., Senwar B., Gerstenfeld L.et al.. Transcriptome-wide association study and eQTL colocalization identify potentially causal genes responsible for human bone mineral density GWAS associations. eLife. 2022; 11:e77285. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23. Thériault S., Gaudreault N., Lamontagne M., Rosa M., Boulanger M.C., Messika-Zeitoun D., Clavel M.A., Capoulade R., Dagenais F., Pibarot P.et al.. A transcriptome-wide association study identifies PALMD as a susceptibility gene for calcific aortic valve stenosis. Nat. Commun. 2018; 9:988. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24. Roselli C., Chaffin M.D., Weng L.C., Aeschbacher S., Ahlberg G., Albert C.M., Almgren P., Alonso A., Anderson C.D., Aragam K.G.et al.. Multi-ethnic genome-wide association study for atrial fibrillation. Nat. Genet. 2018; 50:1225–1233. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25. Wu L., Shi W., Long J., Guo X., Michailidou K., Beesley J., Bolla M.K., Shu X.O., Lu Y., Cai Q.et al.. A transcriptome-wide association study of 229,000 women identifies new candidate susceptibility genes for breast cancer. Nat. Genet. 2018; 50:968–978. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26. Gusev A., Lawrenson K., Lin X., Lyra P.C. Jr, Kar S., Vavra K.C., Segato F., Fonseca M.A.S., Lee J.M., Pejovic T.et al.. A transcriptome-wide association study of high-grade serous epithelial ovarian cancer identifies new susceptibility genes and splice variants. Nat. Genet. 2019; 51:815–823. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27. Barbeira A.N., Dickinson S.P., Bonazzola R., Zheng J., Wheeler H.E., Torres J.M., Torstenson E.S., Shah K.P., Garcia T., Edwards T.L.et al.. Exploring the phenotypic consequences of tissue specific gene expression variation inferred from GWAS summary statistics. Nat. Commun. 2018; 9:1825. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28. Gilchrist J.J., Makino S., Naranbhai V., Sharma P.K., Koturan S., Tong O., Taylor C.A., Watson R.A., de Los Aires A.V., Cooper R.et al.. Natural killer cells demonstrate distinct eQTL and transcriptome-wide disease associations, highlighting their role in autoimmunity. Nat. Commun. 2022; 13:4073. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29. Khunsriraksakul C., McGuire D., Sauteraud R., Chen F., Yang L., Wang L., Hughey J., Eckert S., Dylan Weissenkampen J., Shenoy G.et al.. Integrating 3D genomic and epigenomic data to enhance target gene discovery and drug repurposing in transcriptome-wide association studies. Nat. Commun. 2022; 13:3258. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30. Schmiedel B.J., Rocha J., Gonzalez-Colin C., Bhattacharyya S., Madrigal A., Ottensmeier C.H., Ay F., Chandra V., Vijayanand P.. COVID-19 genetic risk variants are associated with expression of multiple genes in diverse immune cell types. Nat. Commun. 2021; 12:6760. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31. Cao C., Wang J., Kwok D., Cui F., Zhang Z., Zhao D., Li M.J., Zou Q.. webTWAS: a resource for disease candidate susceptibility genes identified by transcriptome-wide association study. Nucleic Acids Res. 2022; 50:D1123–D1130. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32. Lu M., Zhang Y., Yang F., Mai J., Gao Q., Xu X., Kang H., Hou L., Shang Y., Qain Q.et al.. TWAS Atlas: a curated knowledgebase of transcriptome-wide association studies. Nucleic Acids Res. 2023; 51:D1179–D1187. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33. Bycroft C., Freeman C., Petkova D., Band G., Elliott L.T., Sharp K., Motyer A., Vukcevic D., Delaneau O., O’Connell J.et al.. The UK Biobank resource with deep phenotyping and genomic data. Nature. 2018; 562:203–209. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34. Wang J., Huang D., Zhou Y., Yao H., Liu H., Zhai S., Wu C., Zheng Z., Zhao K., Wang Z.et al.. CAUSALdb: a database for disease/trait causal variants identified using summary statistics of genome-wide association studies. Nucleic Acids Res. 2020; 48:D807–D816. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35. Canela-Xandri O., Rawlik K., Tenesa A.. An atlas of genetic associations in UK Biobank. Nat. Genet. 2018; 50:1593–1599. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36. Watanabe K., Stringer S., Frei O., Umićević Mirkov M., de Leeuw C., Polderman T.J.C., van der Sluis S., Andreassen O.A., Neale B.M., Posthuma D.. A global overview of pleiotropy and genetic architecture in complex traits. Nat. Genet. 2019; 51:1339–1348. [DOI] [PubMed] [Google Scholar]
- 37. Sollis E., Mosaku A., Abid A., Buniello A., Cerezo M., Gil L., Groza T., Güneş O., Hall P., Hayhurst J.et al.. The NHGRI-EBI GWAS Catalog: knowledgebase and deposition resource. Nucleic Acids Res. 2023; 51:D977–D985. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38. Zheng J., Erzurumluoglu A.M., Elsworth B.L., Kemp J.P., Howe L., Haycock P.C., Hemani G., Tansey K., Laurin C., Pourcain B.S.et al.. LD Hub: a centralized database and web interface to perform LD score regression that maximizes the potential of summary level GWAS data for SNP heritability and genetic correlation analysis. Bioinformatics. 2017; 33:272–279. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39. Leslie R., O’Donnell C.J., Johnson A.D.. GRASP: analysis of genotype-phenotype results from 1390 genome-wide association studies and corresponding open access database. Bioinformatics. 2014; 30:i185–i194. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40. Kamat M.A., Blackshaw J.A., Young R., Surendran P., Burgess S., Danesh J., Butterworth A.S., Staley J.R.. PhenoScanner V2: an expanded tool for searching human genotype-phenotype associations. Bioinformatics. 2019; 35:4851–4853. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41. Tryka K.A., Hao L., Sturcke A., Jin Y., Wang Z.Y., Ziyabari L., Lee M., Popova N., Sharopova N., Kimura M.et al.. NCBI’s Database of Genotypes and Phenotypes: dbGaP. Nucleic Acids Res. 2014; 42:D975–D979. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42. Auton A., Brooks L.D., Durbin R.M., Garrison E.P., Kang H.M., Korbel J.O., Marchini J.L., McCarthy S., McVean G.A., Abecasis G.R.. A global reference for human genetic variation. Nature. 2015; 526:68–74. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43. Lamparter D., Marbach D., Rueedi R., Kutalik Z., Bergmann S.. Fast and rigorous computation of gene and pathway scores from SNP-based summary statistics. PLoS Comput. Biol. 2016; 12:e1004714. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44. Pei G., Dai Y., Zhao Z., Jia P.. deTS: tissue-specific enrichment analysis to decode tissue specificity. Bioinformatics. 2019; 35:3842–3845. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45. PredictDB Team GTEx v8 models on eQTL and sQTL. 2021; PredictDB; https://predictdb.org/post/2021/07/21/gtex-v8-models-on-eqtl-and-sqtl/. [Google Scholar]
- 46. Li L., Huang K.L., Gao Y., Cui Y., Wang G., Elrod N.D., Li Y., Chen Y.E., Ji P., Peng F.et al.. An atlas of alternative polyadenylation quantitative trait loci contributing to complex trait and disease heritability. Nat. Genet. 2021; 53:994–1005. [DOI] [PubMed] [Google Scholar]
- 47. Harrison P.W., Amode M.R., Austine-Orimoloye O., Azov A.G., Barba M., Barnes I., Becker A., Bennett R., Berry A., Bhai J.et al.. Ensembl 2024. Nucleic Acids Res. 2024; 52:D891–D899. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48. Stelzer G., Rosen N., Plaschkes I., Zimmerman S., Twik M., Fishilevich S., Stein T.I., Nudel R., Lieder I., Mazor Y.et al.. The GeneCards Suite: from gene data mining to disease genome sequence analyses. Curr. Protoc. Bioinformatics. 2016; 54:1.30.1–1.30.33. [DOI] [PubMed] [Google Scholar]
- 49. Sayers E.W., Bolton E.E., Brister J.R., Canese K., Chan J., Comeau D.C., Connor R., Funk K., Kelly C., Kim S.et al.. Database resources of the national center for biotechnology information. Nucleic Acids Res. 2022; 50:D20–D26. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50. Urbut S.M., Wang G., Carbonetto P., Stephens M.. Flexible statistical methods for estimating and testing effects in genomic studies with multiple conditions. Nat. Genet. 2019; 51:187–195. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51. Li M.J., Wang P., Liu X., Lim E.L., Wang Z., Yeager M., Wong M.P., Sham P.C., Chanock S.J., Wang J.. GWASdb: a database for human genetic variants identified by genome-wide association studies. Nucleic Acids Res. 2012; 40:D1047–D1054. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 52. Tian D., Wang P., Tang B., Teng X., Li C., Liu X., Zou D., Song S., Zhang Z.. GWAS Atlas: a curated resource of genome-wide variant-trait associations in plants and animals. Nucleic Acids Res. 2020; 48:D927–D932. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53. Shao M., Tian M., Chen K., Jiang H., Zhang S., Li Z., Shen Y., Chen F., Shen B., Cao C.et al.. Leveraging random effects in cistrome-wide association studies for decoding the genetic determinants of prostate cancer. Adv. Sci. (Weinh.). 2024; 11:e2400815. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 54. Yang C., Wan X., Lin X., Chen M., Zhou X., Liu J.. CoMM: a collaborative mixed model to dissecting genetic contributions to complex traits by leveraging regulatory information. Bioinformatics. 2019; 35:1644–1652. [DOI] [PubMed] [Google Scholar]
- 55. Yang Y., Shi X., Jiao Y., Huang J., Chen M., Zhou X., Sun L., Lin X., Yang C., Liu J.. CoMM-S2: a collaborative mixed model using summary statistics in transcriptome-wide association studies. Bioinformatics. 2020; 36:2009–2016. [DOI] [PubMed] [Google Scholar]
- 56. Yang Y., Yeung K.F., Liu J.. CoMM-S(4): a collaborative mixed model using summary-level eQTL and GWAS datasets in transcriptome-wide association studies. Front. Genet. 2021; 12:704538. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
The data underlying this article are available in webTWAS 2.0 (http://www.webtwas.net) and can be freely downloaded.




