Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2026 Jun 23.
Published in final edited form as: Nat Genet. 2026 Jan 21;58(2):289–298. doi: 10.1038/s41588-025-02470-1

Multi-trait and multi-ancestry genetic analysis of comorbid lung diseases and traits improves genetic discovery and polygenic risk prediction

Yixuan He 1,2,3,4,*, Wenhan Lu 1,5, Yon Ho Jee 1,5, Mu-Yi Shih 6, Ying Wang 1,2,5, Kristin Tsuo 1,5,7, David C Qian 8, James A Diao 2,9,10, Hailiang Huang 1,2,5, Chirag J Patel 9, Jinyoung Byun 11,12,13, Bogdan Pasaniuc 14,15, Elizabeth G Atkinson 16,17, Christopher I Amos 11,12,13, Yen-Chen Anne Feng 6, Matthew Moll 18,19,20, Michael H Cho 18,19, Alicia R Martin 1,2,5,*
PMCID: PMC13284852  NIHMSID: NIHMS2167136  PMID: 41565855

Abstract

While respiratory diseases such as chronic obstructive pulmonary disease (COPD) and asthma share many risk factors, most studies investigate them in isolation and in predominantly European ancestry populations. Here, we conducted the most powerful multi-trait and multi-ancestry genetic analysis of respiratory diseases and auxiliary traits to date, identifying 25 novel loci associated with lung function in individuals of East Asian ancestry. Using these results, we developed PRSxtra (cross TRait and Ancestry), a multi-trait and multi-ancestry polygenic risk score (PRS) approach that leverages shared components of heritable risk via pleiotropic effects. PRSxtra significantly improved the prediction of asthma, COPD, and lung cancer compared to trait- and ancestry-matched PRS in a multi-ancestry cohort from the All of Us Research Program, especially in diverse populations. Our results present a new framework for multi-trait and multi-ancestry studies of respiratory diseases to improve genetic discovery and polygenic prediction.

Introduction

Respiratory diseases are leading causes of morbidity, mortality, and disability-adjusted life-years globally. Chronic obstructive pulmonary disease (COPD) is the third leading cause of death globally, asthma is the most common chronic disease of childhood, and lung cancer is the leading cause of cancer deaths worldwide13.

Existing models for predicting the risk of respiratory disease are limited and often do not consider environmental or genetic risk factors beyond smoking. The asthma predictive index (API) determines the likelihood of pediatric asthma mainly based on family history information4, and lung cancer and COPD risk are primarily assessed by age and smoking history.

While respiratory diseases such as COPD and lung cancer are strongly influenced by smoking, they are complex diseases shaped by many environmental and genetic risk factors. For instance, cumulative non-smoking factors can better predict and stratify COPD risk compared to smoking alone5, and genetic factors also play an important role. Family-based studies estimate heritability at 40% for COPD, 60% for asthma68, and around 18% for lung cancer9, indicating a substantial genetic component. Narrow sense heritability estimates between 5–40% for COPD and asthma10,11, and between 3–10% for lung cancer10,12.

Previous studies have shown that single-trait polygenic risk scores (PRS), which model cumulative genetic risk using genome-wide association study (GWAS) summary statistics, can identify and stratify individuals with risk of respiratory diseases13,14. While respiratory diseases share many comorbidities as well as genetic, clinical, and lifestyle risk factors, most PRS to date have been constructed for a single trait and ancestry group, ignoring dense genetic and phenotypic correlations between traits and linkage disequilibrium (LD) and allele frequency patterns between ancestries.

To address these limitations, we expanded genetic studies of spirometry, an important diagnostic and management tool for respiratory disease, and developed a novel multi-trait, multi-ancestry PRS framework integrating genetic correlations between respiratory diseases and auxiliary traits and LD patterns across diverse populations. This approach improves both the power for genetic discovery and PRS predictive accuracy, thereby providing a more comprehensive risk assessment tool that can be used in research contexts or potentially integrated into existing clinical models.

Previous efforts have used multi-trait approaches to improve genetic discovery and prediction15. For example, multi-trait analysis of GWAS (MTAG) has demonstrated improved power to detect signals for psychiatric disorders, cardiomyopathies, and tobacco and alcohol use, among others1618. Multi-trait PRS combining multiple cardiometabolic trait scores better predicts heart disease than the single trait PRS19,20. Similarly, lung function PRS from multiple spirometry measurements have been used in predicting asthma and COPD13,21. However, the majority of studies still report a single-trait PRS derived in primarily European ancestry cohorts. This lack of diversity in genomic studies has been shown to greatly reduce the generalizability of prediction models and exacerbate healthcare disparities22.

We hypothesize that the joint modeling of multi-trait and multi-ancestry information at the single-nucleotide polymorphism (SNP) level enhances genomic discovery and prediction in respiratory diseases with major global burden and disparities. Specifically, we jointly model genetic information of four ancestry groups, including African (AFR), Admixed American (AMR), East Asian (EAS), and European (EUR)23,24, and eight strongly correlated traits, including COPD, asthma, lung cancer, forced expiratory volume (FEV1), forced vital capacity (FVC), FEV1/FVC, smoking status, and smoking intensity.

To evaluate the utility of this approach for predicting respiratory diseases and interpreting genetic variant effects across traits and ancestries, we: (1) conducted the largest meta-analysis of lung function in EAS and a comprehensive multi-trait meta-analysis of respiratory disease and related traits to significantly improve genetic discovery; (2) compared shared and distinct genetic architectures and effects by modeling pleiotropy across traits; (3) developed and validated the PRSxtra (cross TRait and Ancestry) method in the All of Us Research Program, modeling genetic correlations between traits and ancestry-specific LD and allele frequency patterns between populations; and (4) quantified PRSxtra prediction accuracy and case stratification of asthma, COPD, and lung cancer risk across multiethnic populations, especially those traditionally underrepresented in genetic studies, and compared it to single-trait and single-ancestry PRS and clinical risk factors.

Results

We aim to improve the power of genetic discovery and prediction of respiratory diseases through multi-ancestry and multi-trait analyses of eight correlated traits—COPD, asthma, lung cancer, spirometry (FEV1, FVC, FEV1/FVC), smoking status, and smoking intensity—in AFR, AMR, EAS, and EUR ancestry populations. We conducted the largest meta-analysis GWAS of lung function in EAS to date and meta-analyzed additional GWAS from the Global Biobank Meta-analysis Initiative (GBMI), GWAS & Sequencing Consortium of Alcohol and Nicotine (GSCAN), and multi-population genome-wide meta-analyses of lung function and lung cancer. We then developed the most predictive PRS of asthma, COPD, and lung cancer to date across populations using the strategy and datasets outlined in Figure 1.

Figure 1 |. Study design overview.

Figure 1 |

Training PRS models consisted of three phases: Phase 1, multi-trait analysis (red); Phase 2, multi-ancestry analysis (blue); Phase 3, regularization to optimally predict traits (purple). We validated PRS in All of Us data for asthma, COPD, and lung cancer case status. The number of candidate scores is shown under each PRS. GBMI, Global Biobank Meta-analysis Initiative; LC/LF-GWMA, Lung Cancer/Lung Function multipopulation Genome-Wide Meta-Analysis; KCPS-II, Korean Cancer Prevention Study-II; TWB, Taiwan Biobank; GSCAN, GWAS & Sequencing Consortium of Alcohol and Nicotine use. Genetically defined ancestry group labels are based on population reference panels from the 1000 Genomes Project and Human Genome Diversity Project as follows: AFR, African; AMR, Admixed American; EAS, East Asian; EUR, European. META, meta-analysis across ancestries; PRSxt, PRS cross trait; PRSxa, PRS cross ancestry; PRSxtra, PRS cross trait and ancestry.

Genome-wide association study of lung function in East Asians identifies novel loci

Spirometry is a pulmonary test that evaluates lung function. As continuous measures, they are useful traits for studying the heritable basis of respiratory diseases. However, the largest genetic studies of spirometry to date have been Eurocentric. To identify novel loci associated with lung function in understudied populations, we performed the largest GWAS to date of EAS ancestry individuals of FEV1, FVC, and FEV1/FVC, combining GWAS summary statistics from the Korean Cancer Prevention Study-II (KCPS2) and Taiwanese Biobank (TWB) (Supplementary Table 1 and Methods)2527. We first conducted GWAS of all three continuous spirometry measures of lung function traits in both KCPS2 and TWB using a linear mixed model implemented in SAIGE (Methods). We then performed fixed-effects inverse-variance weighted meta-analysis for each lung function trait across the 130,000 total EAS individuals. After quality control filters (Methods), 8 million unique SNPs were included.

The EAS meta-analysis identified 13, 37, and 24 independent loci for FEV1, FVC, and FEV1/FVC, respectively (Fig. 2ac, Extended Data Figs. 13, and Supplementary Figs. 1 and 2). Of these, 25 have not been previously associated with each trait (Supplementary Tables 35), including three potentially novel loci for FEV1, 17 for FVC, and five for FEV1/FVC. Of these, three, 13, and three loci for FEV1, FVC, and FEV1/ FVC, respectively, were not previously reported with either FEV1, FVC, or FEV1/FVC. Across all new loci, six were associated with white blood cell count, five were previously associated with neutrophil count, and four were associated with educational attainment, systolic blood pressure, and type 2 diabetes. Potentially novel loci were defined such that the lead variant was at least 500 kb upstream and downstream of a previously discovered variant, excluding variants in linkage disequilibrium (LD) at r2 > 0.1, estimated using a 1000 Genomes Project reference panel (Methods). Previous reported associations for each locus compiled from the GWAS Catalog can be found in Supplementary Table 6.

Figure 2 |. Meta-analysis results of spirometry GWAS in East Asian (EAS) and multi-ancestry cohorts.

Figure 2 |

a-c, Frequency and effect size of risk alleles of 13, 37, and 24 index variants for FEV1 (a), FVC (b), and FEV1/FVC (c), respectively, based on an EAS-only meta-analysis (KCPS2 and TWB). Potentially novel associations are highlighted in red. Variants are annotated with the nearest gene. Effect sizes represent the change in FEV1 (Liters) and FVC (Liters) or FEV1/FVC per effect allele. d, The largest published GWAS of FEV1/FVC to date is depicted in the Manhattan plot in red (bottom, with its signals in orange dots). Integrating EAS results in a multi-ancestry meta-analysis identified new loci, depicted in the Manhattan plot in blue (top, with potentially novel loci in triangles). Unadjusted two-sided P values derived from METAL are on a −log10 scale. Novel loci with P < 10−10 are annotated with the nearest gene.

We then combined our EAS meta-analysis results with the largest multi-ancestry GWAS of lung function to date28, resulting in 505, 491, and 445 total loci discovered for FEV1, FVC, and FEV1/FVC, respectively, of which 70, 109, and 40 were not originally reported to be associated with each trait (Fig. 2d and Extended Data Figs. 4 and 5). Significantly associated loci are reported in Supplementary Tables 79. A very small minority of variants out of over 8 million previously associated with spirometry at P < 5 × 10−8 in the multi-ancestry GWAS did not reach genome-wide signficance after our meta-analysis with EAS cohorts (FEV1, nSNPs = 43; FVC, nSNPs = 32; FEV1/FVC, nSNPs = 19; Supplementary Table 10), suggesting strong consistency of signals across ancestries.

Pervasive pleiotropic effects across respiratory traits and diseases inform shared components of heritable risk

Measures of lung function are known clinical risk factors for respiratory diseases such as COPD, asthma, and lung cancer. Using the largest multi-ancestry GWAS summary statistics to date, we found that lung function (FEV1, FVC, FEV1/FVC), smoking behavior (smoking status and cigarettes/day), and respiratory diseases (asthma, COPD, and lung cancer) were all genetically significantly correlated with each other (P < 0.05), with the exception of FVC and lung cancer (Fig. 3a).

Figure 3 |. Shared and distinct heritable components inform the molecular basis of trait differences.

Figure 3 |

a, Genetic correlations between traits (top) and heritability estimates (bottom) analyzed in this study were calculated using linkage disequilibrium score regression (LDSC). Genetic correlation (rg) and unadjusted P-values were estimated by LDSC. Colors indicate magnitude of correlation, and asterisks indicate degree of significance. b-d, Comparison of effect sizes of variants from GWAS on EUR samples for asthma vs. cigarettes smoked per day (Cig/Day) (left), COPD vs. Cig/Day (middle) and lung cancer vs. Cig/Day (right). Effect sizes of variants on diseases (asthma, COPD and lung cancer) are on the y-axis, and effect sizes of variants on Cig/Day are on the x-axis. In a shared variant analysis, colored variants are those confidently identified (with posterior probability > 99%) by a Bayesian classifier (linemodels) to have smoking-predominant genetic effects (red) or disease-predominant genetic effects (blue). Effect sizes of variants are reoriented towards the direction that increases disease (when effect size on the disease phenotype is negative, flip the sign of effect size for both the disease and smoking status). Gray variants were not confidently assigned to either trait (posterior probability < 99%). The names of the nearest genes are labeled for up to 10 variants with the highest confidence for the disease phenotype per gene, and with the top 5 largest positive or negative effect sizes on the disease phenotypes. The error bands in b-d indicate the 95% probability regions of the fitted bivariate effect size distributions with each class.

To examine the distinct and shared variants among respiratory diseases and auxiliary traits, we first evaluated the consistency of lead SNPs between the traits. Many variants demonstrated pleiotropic effects across traits. In particular, we explored the shared and distinct genetic etiology among the three respiratory disease traits (asthma, COPD, and lung cancer), risk factors (smoking status, smoking intensity), and a measure of lung function (FEV1/FVC). For each trait pair, we selected variants significantly associated with either or both of the traits and computed the linear slopes of variant effect sizes using an expectation-maximization (EM) algorithm, which typically revealed more than one linear trend and divergent slopes (Methods, Fig. 3, and Extended Data Figs. 68). For example, between COPD and cigarettes smoked per day, 5,173 variants reached genome-wide significance across either trait, of which 582 were significant in both traits, and 15/634 and 690/3,957 smoking-specific and COPD-specific associations, respectively, reached posterior probability > 0.99. Similarly, between lung cancer and cigarettes smoked per day, 4,066 variants reached genome-wide significance across either trait, of which 684 were significant in both, and 486/3,191 and 88/191 smoking-specific and lung cancer-specific associations, respectively, reached posterior probability > 0.99. Furthermore, genetic variants associated with the amount and frequency of smoking have mostly distinct effects from asthma. This suggests that there is no uniform explanation of the genetic mechanisms of the two traits across those variants.

We then implemented a Bayesian algorithm, linemodels29, to classify variant effect sizes into confident associations with either trait, as smoking is a causal risk factor for both diseases. A model that fits three lines, including both traits, did not significantly improve the fit but is shown in Supplementary Figures 35. We used ancestry-specific (AFR, EAS, EUR) as well as meta-analyzed multi-ancestry GWAS results to identify variants having confident associations with a posterior probability above 0.99. We found 15 variants that were confidently associated with COPD and 88 with lung cancer, which are distinct from those associated with smoking intensity, as measured by cigarettes/day, in EUR. We identified 690 and 486 variants, respectively, confidently associated with only smoking intensity in each analysis. (Supplementary Tables 1113).

When we further investigated variants with predominant effects on lung cancer across EAS, EUR, and the multi-ancestry meta-analysis GWAS results, we identified six variants from TERT appearing in all three groups and 11 variants significant only in the meta-analysis results (Supplementary Figs. 6 and 7). These include rs11783093 and rs73229093 from GULOP, rs11778371 from CHRNA2, four variants from CLPTM1L, and five other variants within or near TERT.

PRSxtra jointly models multi-ancestry and multi-trait effects to predict diseases and exacerbations

To develop PRSxtra, we combined the results of our new GWAS meta-analysis of lung function in EAS with the largest and most diverse GWAS of COPD, asthma, lung cancer, and smoking (Fig. 1) to conduct ancestry-specific multi-trait GWAS (MTAG)15. We then used PRS-CSx to jointly model trait-specific MTAG results across ancestry groups to derive candidate scores30,31 to include in the final PRSxtra.

First, MTAG resulted in greater number of significant associations and loci discovered across traits (Supplementary Table 14), with the largest gains in EAS. For example, the number of loci identified to be associated with FEV1 and COPD in EAS increased from 20 to 28, and from 15 to 54, respectively. Across traits, more than 50% of the loci identified were not previously labeled. Of these new loci, the majority had consistent direction of effect as in the GWAS.

Next, we aggregated MTAG summary statistics to derive PRSxtra. For each trait, we used PRS-CSx to jointly model trait-specific MTAG summary statistics across ancestry groups to derive a total of 39 candidate scores for individuals in the All of Us Research Program (AoU), a longitudinal cohort study continuously enrolling adults in the US with genotype, self-reported, and linked health record data. The program places a strong emphasis on including diverse populations that have traditionally been underrepresented in biomedical research. The individuals from AoU are independent from the cohorts in which the GWAS was obtained. For each trait, we randomly split participants who passed quality control into 70% for training and 30% for validation (Fig. 1, Methods, and Supplementary Tables 1518). We then applied ridge regression to calculate linear combinations of the candidate scores (Supplementary Tables 1921 and Supplementary Figs. 810).

We evaluated the performance of PRSxtra in the held-out validation cohort for predicting asthma (ncases = 9,450, ncontrols = 64,070), COPD (ncases = 4,560, ncontrols = 66,605), and lung cancer (ncases = 578, ncontrols = 70,600) compared it to four benchmark PRS scores derived from: (1) single ancestry- and trait-matched summary statistics using PRS-CS32 (PRS); (2) PRS-CS scores of MTAG results within an ancestry group followed by regularization across traits (PRSxt); (3) PRS-CSx scores of GWAS results followed by regularization across ancestry groups (PRSxa); and (4) previous multi-trait methods (PRSmix+)33 (Fig. 1). The total validation cohort includes individuals of AFR, AMR, EAS, EUR, Middle Eastern (MID), and South Asian (SAS) genetic ancestries. In this multi-ancestry validation cohort, ancestry- and trait-matched PRS and PRSxtra were significantly correlated with each other (P < 0.0001), with r = 0.195, r = 0.150, and r = 0.120 for asthma, COPD, and lung cancer, respectively.

PRSxtra alone predicted asthma, COPD, and lung cancer more accurately than the trait- and ancestry-matched PRS alone (P < 0.0001) in the multi-ancestry validation cohort (Fig. 4 and Supplementary Tables 2224). For COPD, the AUC improved from 0.531 (95% CI = [0.523, 0.407]) to 0.6170 (95% CI = [0.599, 0.615]) with PRSxtra compared with a single trait- and ancestry-matched PRS. For lung cancer, AUC improved from 0.541 (95% CI = [0.518, 0.563]) to 0.612 (95% CI = [0.589, 0.635]). For asthma, AUC improved from 0.540 (95% CI = [0.535, 0.547]) to 0.584 (95% CI = [0.578, 0.590]) (Fig. 4). Across traits and ancestry groups, PRSxtra demonstrated marginally improved predictive performance compared to PRSmix+. Both scores outperformed PRSxt and PRSxa, and PRSxa tended to perform better than PRSxt. This difference was the largest for asthma, where PRSxt and PRSxa had AUCs of 0.554 (95% CI = [0.534, 0.574]) and 0.574 (95% CI = [0.568, 0.581]), respectively (Fig. 4). We derived a measure of %AUC to contextualize potential contributions of the cross-trait or cross-ancestry step to the prediction performance of PRSxtra (Supplementary Tables 2527 and Methods). PRSxa tended to have a greater contribution than PRSxt across traits and ancestries with higher %AUC.

Figure 4 |. PRSxtra significantly improves prediction for several respiratory diseases compared to PRS and clinical risk factors.

Figure 4 |

Diseases include COPD, lung cancer, and asthma in the full multi-ancestry held-out validation set and in ancestry-specific subgroups. Green dots indicate the performance of PRS, PRSxt, PRSxa, PRSmix+ or PRSxtra (in gradually darker shades) alone to predict diseases; blue dots represent the performance of the combined model with incremental variables added from sex and age, smoking status (for COPD and lung cancer) or family history (for asthma) (light blue) and PRSxtra (dark blue). The full validation set included individuals of AFR, AMR, EAS, EUR, MID, and SAS predicted genetic ancestry background. The area under the curve (AUC) point estimates and 95% CIs in each population are shown. Asterisks denote significant differences in prediction with DeLong's two-sided test P < 0.05. The number of case and control samples in each population for each trait can be found in Supplementary Tables 1618.

The AMR population showed the largest improvement in disease prediction. For COPD, AUC improved from 0.509 (95% CI = [0.479, 0.539]) to 0.633 (95% CI = [0.606, 0.660]), which could reflect increases in power from the additional EAS GWAS for spirometry with the relatively low genetic divergence between EAS and AMR groups. To this end, we observed a greater prediction improvement for COPD and asthma in AMR using PRSxa over PRSxt, with an even larger increase with PRSxtra. However, when we further assessed the robustness of PRSxtra by conducting leave-one-out analyses in the AMR cohort, we found that removing any single score did not significantly change the prediction ability of PRSxtra. For example, removing EUR lung cancer candidate PRS resulted in the largest decrease in AUC of only −0.0175 for predicting COPD, and removing EUR FVC candidate PRS resulted in the largest decrease in AUC of only −0.000421 for predicting asthma. We also conducted a principal component analysis of the AMR individuals, which identified roughly three axes of genetic variation (Supplementary Fig. 11).

PRSxtra improved overall stratification in the multi-ancestry validation cohort compared to PRS (Supplementary Fig. 12 and Supplementary Tables 2830). PRSxtra was also better at identifying individuals with higher risk of disease. The odds ratio (ORs) for asthma per standard deviation (SD) increase of PRS and PRSxtra were 1.13 (95%CI 1.11 to 1.15, P = 1.06 × 10−29) and 1.34 (95%CI 1.31 to 1.36, P = 3.81 × 10−166), respectively. The ORs for COPD per SD increase of PRS and PRSxtra were 1.10 (95%CI 1.07 to 1.14, P = 2.43 × 10−10) and 1.39 (95%CI 1.36 to 1.44, P = 7.25 × 10−121), respectively (Supplementary Tables 3133).

We compared the performance of PRS and PRSxtra alone to a joint model of the strongest clinical risk factors for each disease, i.e., family history of asthma, and smoking for COPD and lung cancer. Clinical risk factors provided the largest performance improvement beyond sex and age. In the multivariable model, PRSxtra remained significantly better than PRS for predicting asthma in the multi-ancestry validation cohort and for predicting asthma and COPD in the AMR subgroup. For lung cancer, risk scores added no predictive value beyond sex, age, and smoking status (Fig. 4). The predictive ability of all models across all populations can be found in Supplementary Tables 2224.

To investigate clinical subgroup performance, we tested PRSxtra associations with COPD and lung cancer in never and ever smokers, and with asthma in individuals with and without family history. Performance was similar in subgroups (Supplementary Tables 2830 and Supplementary Figs. 1315). PRSxtra was positively associated with COPD and asthma exacerbations, predicting COPD and asthma exacerbation significantly better than PRS, with AUC increasing from 0.545 to 0.604 (P < 0.0001) and from 0.531 to 0.571 (P < 0.0001), respectively (Supplementary Table 34).

Discussion

In this study, we conducted the most powerful multi-trait, multi-ancestry genetic analysis of respiratory diseases and comorbid traits to date. Our novel study framework contrasts with traditional genetic studies, which typically analyze a single trait within a single population, overlooking genetic correlations between traits and LD patterns across diverse populations.

Our innovative iterative modeling of genetic correlations across traits and ancestry groups, significantly increases the power to detect novel genetic associations and enhances disease prediction and risk stratification, especially in non-European ancestry populations. Applying MTAG resulted in a greater number of significant associations and loci discovered across traits, with the largest gains in AMR and EUR ancestry-specific analyses. We also conducted the largest meta-analysis of lung function to date in EAS to identify 25 novel loci associated with lung function. Of these, several are associated with immune function and inflammation, education attainment, blood pressure, and type 2 diabetes. These results support the well established role of inflammatory immune response in lung diseases like asthma34 and between lung dysfunction and disease risk and progression35,36.

Given that smoking is the leading risk factor for COPD and lung cancer, we further investigated distinct and shared variants between smoking intensity, COPD, and lung cancer. We identified variants that demonstrated vertical pleiotropy, influencing both smoking and disease. We also identified variants that distinctly influence disease but not smoking, including the CFTR Phe508del mutation rs113993960, a known pathogenic variant for cystic fibrosis37,38, which was confidently associated with COPD but not smoking; rs11571833, a rare variant of BRCA2-K3326X, was confidently associated with lung cancer but not smoking and has been shown to have large-effect genome-wide associations for squamous lung cancer (odds ratio = 2.47, P = 4.74 × 10−20)39. Furthermore, across EAS, EUR, and the multi-ancestry meta-analysis GWAS results, there were six variants from TERT appearing in all three groups and 11 variants in GULOP, CHRNA2, and CLPTM1L significant only in the meta-analysis results. TERT, CHRNA2, and CLPTM1L are known to affect respiratory function. TERT encodes a catalytic subunit of telomerase and sits at a highly pleiotropic locus. While the TERT variants highlighted in our study have been previously reported, they demonstrate how germline regulation of telomere biology can influence several, and sometimes opposing phenotypes. For example, longer genetically predicted telomere length increases lung cancer risk in EUR and EAS that is independent of smoking behaviors4042, whereas shorter telomeres predispose to COPD and are a strong risk factor for idiopathic pulmonary fibrosis43. CHRNA2, encoding a nicotinic acetylcholine receptor subunit, influences respiratory behavior and function by influencing nicotine addiction and smoking habits44,45., but has a smaller effect than the CHRNA3-CHRNA5 cluster, which has a stronger and more consistently observed association with smoking46,47. CLPTM1L encodes a mitochondrial membrane protein whose overexpression in cisplatin-sensitive cells causes apoptosis48. The CLPTM1L protein is more highly expressed in lung cancer and plays a pro-tumorigenic role critical for lung cancers driven by mutations in KRAS4850.

In our investigation of pleiotropic effects across pairs of traits, we applied linemodels to dissect the variant level pleiotropy. Using this method, we identified unique genetic features of lung cancer and COPD independent of smoking. The linemodels algorithm, tailored for pleiotropy dissection, uses a Bayesian framework to cluster variants based on linear effect size relationships across two outcomes and optimizes parameters using the EM algorithm and Gibbs sampling, accommodating correlated estimators due to sample overlaps. In contrast, techniques like Non-negative Matrix Factorization decompose data into factors to capture latent structures without modeling direct linear relationships, highlighting linemodels’s unique focus on linear effect size clustering.

We then developed a multi-trait and multi-ancestry score, PRSxtra, to model the genetic correlations between related traits and ancestry-specific LD and allele frequency patterns between populations. We derived and validated our score in the multi-ancestry All of Us cohort. Given that one of the strongest genetic risk factors for COPD and emphysema is alpha-1 antitrypsin deficiency due to rare variants in SERPINA1, we excluded individuals with homozygous polymorphisms of SERPINA1 from our study51,52. While PRSxtra and the trait- and ancestry-matched PRS are significantly correlated, PRSxtra demonstrates significantly better prediction and stratification for respiratory diseases and exacerbations. Part of the strength of PRSxtra is that many of the component traits are both genetically correlated and are also long established and mechanistically plausible risk factors of disease. For example, multiple studies have shown that COPD, reduced FEV1 and FEV1/FVC ratio exert a causal influence on lung cancer risk5355 through immune-mediated pathways beyond shared exposure and cigarette smoking55,56. By explicitly including the PRSs for these biologically grounded risk factors, we capture multi-factorial pathophysiologic processes that traditional single-trait PRSs miss. Furthermore, we observed that PRSxa performs better than PRSxt and that PRSxt had the smallest gain in AUC from a single-trait single-ancestry PRS, highlighting the importance of incorporating multiple ancestry groups into building an accurate PRS.

Although PRS have a high potential for clinical use, existing methods have faced challenges due to lack of generalizability across populations. Here, we show that PRSxtra significantly improves the prediction and stratification compared to existing PRS beyond common clinical risk factors (e.g., sex, age, family history and smoking status). In combination with clinical risk assessment, PRSxtra can be particularly useful in populations where the discovery population is smaller (e.g., AMR and AFR populations).

Previous studies have shown the utility of integrating large and diverse sources of data to construct PRS. For example, GPSmult, derived from the weighted sum of cardiometabolic trait components, predicts heart disease better than the single trait PRS19. Similarly, PRS derived from two spirometry measurements predict asthma and COPD13. We compared PRSxtra to a previously established multi-trait framework from Truong et al.33, which showed comparable prediction performance. Other recent methods like PRSmix+ and BridgePRS similarly model joint summary information from multiple traits and ancestry groups at the score level but do not account for genetic correlations across traits or model differences in LD across ancestries.

PRSxtra shares information across traits and ancestries at the SNP level, refining putatively causal loci and improving PRS accuracy. Our leave-one-out analysis demonstrates its robustness, showing minimal change in prediction when candidate scores were excluded. The enhanced predictive capacity of PRSxtra was particularly pronounced in the AMR population, which is not primarily explained by the candidate PRS from the new EAS spirometry GWAS. One potential explanation could be due to the phenotypic heterogeneity of asthma and COPD among Hispanic populations5759. For example, asthma, and COPD are more prevalent in individuals of Puerto Rican heritage than in other Hispanic populations. Additionally, there are significant differences in the social identities of AMR individuals, which can correlate with environmental exposures and contribute to increased phenotypic variation. Our principal component analysis of AMR individuals identified roughly three axes of genetic variation, highlighting the genetic diversity within this population. Further studies are needed to investigate subgroup differences.

Our study does have limitations. We relied on phenotype definitions based on ICD codes and self-reported data, which can be imprecise. For example, COPD defined by ICD-9 codes have been shown to misclassify patients compared to combining them with pharmacy data60. Therefore, the performance of PRSxtra may differ for COPD diagnosed by definition. While we aimed to mitigate this effect by combining spirometry with multiple ICD-9 and ICD-10 codes to define disease cases, it is unclear how adding spirometry would change disease definitions in AoU due to the lack of spirometry measures availability for all participants. A patient with respiratory symptoms and is a smoker may also be more likely to be labeled as COPD by physicians. We are limited by the study design of the existing cohorts and the measurements collected by the independent sites, such as that of FEV1 and FVC. While quality control was performed in these studies (e.g., filtering variants with low call rates and samples with excessive heterozygosity)25,26, comparison of association statistics suggests modest deflation in TWB for FEV1 and therefore also FEV1/FVC, as evidenced by lower LDSC intercepts and λGC values, whereas FVC shows greater consistency across cohorts. We compared our findings with previous reports from Shrine et al.28 and found that smaller cohorts in their study also had smaller values of LDSC intercepts. Therefore, we caution that some of the TWB FEV1 and FEV1/FVC signals may have conservative effect size estimates, potentially due to noisy phenotyping from the original study. We were also limited by the availability of ancestry-specific data beyond European populations. Our measures were also not explicitly post-bronchodilator lung function measures, although previous work demonstrated little impact of pre- versus post-bronchodilator definitions when predicting COPD61,62. Additionally, we used ridge regression, which assumes linear associations of candidate scores. It is unclear how non-linear methods of combining scores will compare. Our study sample is shaped by the All of Us Research Program’s cohort creation process, which relies on partnerships with universities, research centers, volunteers, and community engagement63. While AoU is diverse, there are demographic sampling biases that may not be representative of the general population64. For instance, our lung cancer study population had limited numbers of non-EUR cases, which may affect the generalizability of the results. Our results demonstrate significant improvements in a held-out cohort of diverse AoU participants but require validation in an independent cohort. In this study of the genetics of respiratory disease and traits, we included eight highly genetically correlated traits in four populations for pleiotropy analysis and PRS derivation. Given the role of inflammatory immune response in asthma, incorporating markers of inflammation, such as eosinophil count, could enable further discovery.

In summary, we conducted the largest multi-trait and multi-ancestry genetic analysis of respiratory diseases and auxiliary traits to date to discover numerous novel genetic signals. We propose PRSxtra as a method to model genetic correlation across traits and LD differences between ancestry groups, significantly improving disease prediction and stratification for asthma, COPD, and lung cancer. PRSxtra has the potential to reduce the disparities in risk stratification between populations for survival and outcome and to advance more equitable, generalizable prediction models for respiratory diseases.

Methods

Consent and ethical approval

For the All of Us Research Program, informed consents for all the participants are conducted in person or through an eConsent platform. The protocol was reviewed by the Institutional Review Board (IRB) of the All of Us Research Program. Data can be accessed through the All of US Research Workbench, a secure cloud-based analytic platform.

Meta-analysis of lung function across East Asian populations

We performed a fixed-effects meta-analysis with inverse variance weighting, as implemented in METAL v2001–03-25 software, for FEV1, FVC, and FEV1/FVC across two East Asian cohorts: Korean Cancer Prevention Study-II (KCPS2) and Taiwanese Biobank (TWB)25,27. The details of their study cohort have been described previously. In summary, the Korean Cancer Prevention Study-II Biobank (KCPS2) is a prospective cohort study based in Korea of 153,950 total participants with genotype data and phenotype measurements between 2004 and 2013. The Taiwanese Biobank is a prospective cohort study of the Taiwanese population with 149,894 total participants between the ages of 30–70 years old at recruitment (as of April 2021). FEV1 was defined as the total air blown between 0 and 1 second (measured in liters). FVC was defined as the vital capacity during forced expiration (measured in liters). In KCPS2, FEV1 and FVC were measured using pulmonary/metabolic systems Vmax 20, Carefusion, USA. For each trait, samples with measurements that were more than 6 standard deviations away from the sample average were excluded. We conducted GWAS of all three lung function traits in both KCPS2 and TWB using a linear mixed model implemented in SAIGE. We included sex, age, height, 10 principal components (PCs), and smoking as covariates.

Altogether, these cohorts had a total sample size of 129,685 individuals with spirometry measures that passed quality control. In conducting our meta-analysis, we excluded genetic variants with minor allele frequency < 0.01. We used FUMA v. 1.5.2 to annotate variants to the nearest gene in the meta-analysis. Genome-wide significance was defined using a threshold of P < 5 × 10−8. We defined independent lead SNPs in loci with an r2 threshold of 0.1 using the 1000 Genomes Project EAS as the reference panel. Genome positions are reported in build hg37 for index variants. To designate a locus as previously known or potentially novel, we extended 500 kb upstream and downstream of each previously discovered variant to define a previously known locus. We intersected these regions with our loci, and those with no overlap and had a low LD (r2 ≤ 0.1) with the index variant were considered novel. Previously discovered variants were compiled from Shrine et al.28.

We then combined the results of our EAS analysis with the largest GWAS to date for each spirometry trait from Shrine et al.28. For consistency, we re-defined independent lead SNPs of loci for both the previous GWAS and the multi-ancestry meta-analysis results with an r2 threshold of 0.1 using the 1000 Genomes Project reference panel and P-value threshold of 5 × 10−8. LD blocks within 500 kb were merged into a single locus.

Comparison of pleiotropic effect analysis

We used LDSC to estimate heritability and genetic correlation on a liability scale within and between phenotypes using summary statistics of EUR populations65,66. We then applied the linemodels package (https://github.com/mjpirinen/linemodels) to the GWAS summary statistics of three respiratory diseases (asthma, COPD, and lung cancer) and the three featured environmental and genetic risk factors (smoking status, smoking intensity (Cigarettes/Day), and FEV1/FVC) across three ancestry groups (AFR, EAS, and EUR) plus the corresponding meta-analysis results. We focused on comparisons for 12 pairs of traits selected from above. Three of these were between disease phenotypes (asthma and COPD, asthma and lung cancer, COPD and lung cancer), and nine pairs were between disease phenotype and smoking or lung function.

For each pairwise comparison, we considered variants present in the summary statistics of both traits and significantly associated with at least one of the traits being compared (Supplementary Table 35). We classified the variants into three classes (two when there is no variant associated with both traits) based on their association patterns: associated with trait 1 only, trait 2 only, and both. We then estimate the slopes of variants significantly associated with trait 1 and trait 2 only using an EM algorithm. Conditioning on these two classes, we ran the linemodels package on the GWAS effect sizes and standard errors of overlapped variants of the two traits, where we set the scale parameters determining the magnitude of effect sizes to 0.2, the correlation parameters determining the allowed deviation from the lines to 0.99 as default, and the slope parameter to the estimates from the previous EM step. The membership probabilities in the two classes were computed separately for each variant by assuming that the classes were equally probable a priori. This analysis was also repeated but for three classes (trait 1, trait 2, or both). We assumed no overlapping samples between the two GWASs being compared and set the correlation of their effect estimators to 0. Confident associations are defined as having a posterior probability above 0.99.

Multi-trait analysis

We conducted multi-trait genome-wide association studies as implemented in MTAG v. 1.0.7 for each ancestry group (AFR, AMR, EAS, EUR) by combining the ancestry-specific GWAS summary statistics for GBMI asthma, GBMI COPD, GWMA lung cancer, spirometry meta-analysis, and GSCAN smoking behaviors. MTAG performs a joint analysis of GWAS results from related traits to improve the number of genetic loci identified and the predictive power of polygenic scores. For each population, we used the ancestry-specific LD reference panel from the gnomAD reference panels v2.1.1. For each ancestry-specific MTAG, we included traits with χ2 > 1.02.

Study population

The All of Us Research Program is a longitudinal cohort study that has continuously enrolled US adults 18 years or older since May 2017. The program aims to engage in one million or more US participants and places a strong emphasis on including diverse populations that have traditionally been underrepresented in biomedical research. Details of the All of Us cohort have been previously described63. In summary, participants of the program opt to provide self-reported data, linked health record data, and biospecimen data to be made available for research uses. The program’s primary objective is to build a resource to help researchers understand individual differences in biological, clinical, social, and environmental determinants of health and disease to advance precision health care.

Informed consents for all the participants in the All of Us Research Program are conducted in person or through an eConsent platform. The protocol was reviewed by the Institutional Review Board (IRB) of the All of Us Research Program. Data can be accessed through the All of US Research Workbench, a secure cloud-based analytic platform. Whole genome sequencing, genotyping array variant data, variant annotations, computed ancestry, and quality reports are accessible through the Controlled Tier of the AoU. This project is registered in the All of Us program under the workspace name “PRSxtra AoU”. In our analysis, we included individuals with whole genome data in the v7 Data Release, self-reported sex, and date of birth along with additional disease-dependent filtering criteria. For COPD and lung cancer, we excluded individuals who did not self-report smoking status. For COPD, we additionally excluded individuals with homozygous polymorphism of SERPINA1 (rs6647, rs709932, rs28929474), which encodes for a serine protease inhibitor alpha 1 antitrypsin, as this is a known risk allele associated with COPD67. Participants were randomly split 70% for training and 30% for validation.

Phenotype ascertainment

We curated clinical phenotypes from All of Us using a combination of electronic health record data, and/or self-reported personal history data from the All of Us v7 Data Release. ICD codes for each phenotype and exacerbation are detailed in Supplementary Tables 3640. We define smoking status and family history based on self-reported data. Previous smokers are individuals who smoked more than 100 pack years but do not currently smoke. Individuals who have never smoked more than 100 cigarettes are considered never smokers. Family history included mother, father, and siblings with the same record of disease as the participants. No family history included those who did not explicitly self-report a family history.

PRSxtra construction

We constructed PRSxtra in a three-phase process. Phase 1 consisted of performing multi-trait meta-analysis across related traits for each ancestry population as previously described. In situations where ancestry-specific MTAG results were not available due to low χ2, we instead used the original GWAS results. In phase 2, we used PRS-CSx, which leverages linkage disequilibrium across discovery samples to jointly model the genetic effects across populations via a shared continuous shrinkage prior. We used the default parameters on PRS-CSx on AFR, AMR, EAS, and EUR ancestry-specific meta-analysis results. Only HapMap3 variants—a set of roughly 1.5 million variants compiled by the International HapMap Project which capture common patterns of variation in a variety of human populations—were included in calculating scores. In phase 3, we use ridge regression, as implemented by the “glmnet” R package68, to jointly model the 39 standardized PRS (with mean 0 and standard deviation 1) generated in phase 2 to construct an ancestry-specific PRSxtra for the three disease phenotypes: COPD, asthma, and lung cancer. In ridge regression, we used 10-fold cross-validation and minimum lambda value to estimate the weights of each PRS. PRSxtra was validated in the held-out multi-ancestry cohort from AoU.

Benchmark PRS construction

As a baseline comparison to PRSxtra, we derived three risk scores: (1) a single ancestry- and trait-matched PRS derived using PRS-CS, which uses a Bayesian regression framework to infer posterior effect sizes of SNPs, on the trait and ancestry matched GWAS summary statistics; (2) a cross-trait PRS derived using PRS-CS of MTAG results within an ancestry group followed by ridge regression regularization across traits (PRSxt); (3) a cross-ancestry PRS derived using PRS-CSx of GWAS results within one trait followed by ridge regression regularization across ancestry groups (PRSxa); and (4) PRS derived using the framework introduced by PRSmix+33, a multi-trait method to generate a risk score from a library of single-GWAS PRSs (PRSmix+). We implemented PRSmix+ from the candidate library of 31 single trait single ancestry PRS derived from PRS-CS. We used default parameters when running PRS-CS and PRSC-CSx.

Statistical analysis

For COPD, asthma, and lung cancer, we placed individuals into bins by their risk scoredeciles. In each decile, we calculated the prevalence of disease and disease exacerbation. We standardized all scores (to mean 0 and standard deviation 1) and calculated the risk of disease and exacerbation for each standard deviation of score using logistic regression models. We evaluated the performance of predicting diseases based on PRS, PRSxt, PRSxa, and PRSxtra alone, as well as in a joint multi-variable model with covariates. The baseline model included age and sex. We then subsequently added clinical risk factors (smoking or family history), and PRSxtra. We evaluated the predictive performance of each model using the area under the receiving operating curve. To contextualize the improvements in prediction achieved by PRSxt and PRxa, we calculated %AUC of PRSxtra by:

AUCPRS0.5AUCPRSxtta0.5

In the full population, we used the trait- and EUR-specific PRS as the baseline for comparison. All statistical analyses were two-sided and performed with the use of R software, version 3.5 (R Project for Statistical Computing).

Extended Data

Extended Figure 1:

Extended Figure 1:

Frequency and effect size of risk alleles of the 13 index variants associated with FEV1 with that reach genome-wide significance P<5×10−08 in meta-analyzed GWAS of East Asian ancestry population (P<5×10−08, derived from METAL). Dark blue shaded boxes represent variants that were present and significant (P<5×10−08) in the GWAS. Light blue shaded boxes represent variants that were present but not significant. Unshaded white boxes represent variants that were not present in the GWAS.

Extended Figure 2:

Extended Figure 2:

Frequency and effect size of risk alleles of the 37 index variants associated with FVC that reach genome-wide significancewith P<5×10−08 in meta-analyzed GWAS of East Asian ancestry population (P<5×10−08, derived from METAL). Dark blue shaded boxes represent variants that were present and significant (P<5×10−08) in the GWAS. Light blue shaded boxes represent variants that were present but not significant. Unshaded white boxes represent variants that were not present in the GWAS.

Extended Figure 3:

Extended Figure 3:

Frequency and effect size of risk alleles of the 24 index variants associated with FEV1/FVC that reach genome-wide significancewith P<5×10−08 in meta-analyzed GWAS of East Asian ancestry population (P<5×10−08, derived from METAL). Dark blue shaded boxes represent variants that were present and significant (P<5×10−08) in the GWAS. Light blue shaded boxes represent variants that were present but not significant. Unshaded white boxes represent variants that were not present in the GWAS.

Extended Figure 4:

Extended Figure 4:

The largest published GWAS of FEV1 to date is depicted in the Manhattan plot in red (bottom, with its signals in orange dots). Integrating EAS results in a multi-ancestry meta-analysis identified new signals, depicted in the Manhattan plot in blue (top, with potentially novel loci in triangles). Unadjusted two-sided P values derived from METAL are on a −log10 scale. Novel loci with P<10−10 are annotated with the nearest gene.

Extended Figure 5:

Extended Figure 5:

The largest published GWAS of FVC to date is depicted in the Manhattan plot in red (bottom, with its signals in orange dots). Integrating EAS results in a multi-ancestry meta-analysis identified new signals, depicted in the Manhattan plot in blue (top, with potentially novel loci in triangles). Unadjusted two-sided P values derived from METAL are on a −log10 scale. Novel loci with P<10−10 are annotated with the nearest gene.

Extended Figure 6:

Extended Figure 6:

Comparison of effect sizes of variants from GWAS for asthma vs. five other traits (columns) across all available ancestry groups (rows) in models fitted with two lines. Effect sizes of variants on asthma are on the x axis, and effect sizes of variants on the other traits are on the y axis. Each point represents a variant significantly associated (P< 5×10−8) with at least one of the corresponding pair of traits. In a shared variants analysis, variants predominantly (with posterior probability >99%) associated with asthma are colored blue, and variants predominantly associated with the other trait are colored red. Gray variants were not confidently assigned to either traits (posterior probability <99%). The colored shaded ellipse range indicates the 95% probability regions of the fitted bivariate effect size distributions with each class. Empty space means either the two traits do not have enough overlapped variants or GWAS results are not applicable for the corresponding ancestry group.

Extended Figure 7:

Extended Figure 7:

Comparison of effect sizes of variants from GWAS for COPD vs. four other traits (columns) across all available ancestry groups (rows) in models fitted with two lines. Effect sizes of variants on COPD are on the x axis and effect sizes of variants on the other traits are on the y axis. Each point represents a variant significantly associated (P< 5×10−8) with at least one of the corresponding pair of traits. In a shared variants analysis, variants predominantly (with posterior probability >99%) associated with COPD are colored blue, and variants predominantly associated with the other trait are colored red. Gray variants were not confidently assigned to either traits (posterior probability <99%). The colored shaded ellipse range indicates the 95% probability regions of the fitted bivariate effect size distributions with each class. Empty space means either the two traits do not have enough overlapped variants or GWAS results are not applicable for the corresponding ancestry group.

Extended Figure 8:

Extended Figure 8:

Comparison of effect sizes of variants from GWAS for Lung cancer vs. three other traits (columns) across all available ancestry groups (rows) in models fitted with two lines. Effect sizes of variants on lung cancer are on the x axis and effect sizes of variants on the other traits are on the y axis. Each point represents a variant significantly associated (P< 5×10−8) with at least one of the corresponding pair of traits. In a shared variants analysis, variants predominantly (with posterior probability >99%) associated with lung cancer are colored blue, and variants predominantly associated with the other trait are colored red. Gray variants were not confidently assigned to either traits (posterior probability <99%). The colored shaded ellipse range indicates the 95% probability regions of the fitted bivariate effect size distributions with each class. Empty space means either the two traits do not have enough overlapped variants or GWAS results are not applicable for the corresponding ancestry group.

Supplementary Material

Supplementary Figures
Supplementary Tables

ACKNOWLEDGEMENTS

This study was supported by the National Human Genome Research Institute (T32HG010464 to Y.H., K99HG013969 to Y.W., U01HG011719 to A.R.M.), the National Institute of Environmental Health Sciences (R01ES032470 and R01DK137993 to C.J.P.), the National Cancer Institute (U19CA203654 and R01CA243483 to C.I.A. and J.B.), the National Heart, Lung, and Blood Institute (R01HL179112 to A.R.M. and W.L., R01HL168199, R01HL162813, R01HL153248, and R01HL135142 to M.H.C.), and the National Institute of Mental Health (K99/R00MH117229 to A.R.M.). We are grateful for the All of Us participants for their contributions. We also thank the National Institutes of Health’s All of Us Research Program for making available the participant data examined in this study.

Footnotes

COMPETING INTERESTS

M.H.C. has received grant support from GSK, consulting fees from Apogee and BMS, and speaking fees from Illumina. M.M. has received consulting fees from TheaHealth, 2ndMD, Axon Advisors, Verona Pharma, and Sanofi. A.R.M. has received speaker fees from Novartis. All other authors declare no competing interests.

DATA AVAILABILITY

The individual-level genotype and phenotype data of All of Us are available on the Researcher Workbench. Researchers can register to access at: https://www.researchallofus.org/. The GWAS summary statistics for the East Asian and multi-ancestry meta-analyses of lung function are available at the GWAS Catalog (https://www.ebi.ac.uk/gwas) under the accession codes GCST90705067, GCST90705068, GCST90705069, GCST90705070, GCST90705071, and GCST90705072. Data sources for ancestry-specific summary statistics for each trait used in this study are available in Supplementary Table 1. Weights of PRSxtra for each trait are available in Supplementary Tables 1921.

CODE AVAILABILITY

Analyses were conducted using publicly available software: MTAG v.2018 (https://github.com/JonJala/mtag), METAL v.2011–03-25 (https://genome.sph.umich.edu/wiki/METAL), PLINK v.2.0 (https://www.cog-genomics.org/plink/2.0/). PRS-CS v1.1.0 (https://github.com/getian107/PRScs). PRS-CSx v1.1.0 (https://github.com/getian107/PRScsx). Scripts for data analyses are available at https://github.com/yixuanh/lung-mutitrait-multiancestry and at https://zenodo.org/records/17452013.

References

  • 1.Chen S et al. The global economic burden of chronic obstructive pulmonary disease for 204 countries and territories in 2020–50: a health-augmented macroeconomic modelling study. Lancet Glob. Health 11, e1183–e1193 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Sung H et al. Global Cancer Statistics 2020: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA. Cancer J. Clin. 71, 209–249 (2021). [DOI] [PubMed] [Google Scholar]
  • 3.Vos T et al. Global burden of 369 diseases and injuries in 204 countries and territories, 1990–2019: a systematic analysis for the Global Burden of Disease Study 2019. Lancet 396, 1204–1222 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Castro-Rodriguez JA The Asthma Predictive Index: early diagnosis of asthma. Curr. Opin. Allergy Clin. Immunol. 11, 157–161 (2011). [DOI] [PubMed] [Google Scholar]
  • 5.He Y et al. Prediction and stratification of longitudinal risk for chronic obstructive pulmonary disease across smoking behaviors. Nat. Commun. 14, 8297 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Duffy DL, Martin NG, Battistutta D, Hopper JL & Mathews JD Genetics of asthma and hay fever in Australian twins. Am. Rev. Respir. Dis. 142, 1351–1358 (1990). [DOI] [PubMed] [Google Scholar]
  • 7.Ingebrigtsen T et al. Genetic influences on chronic obstructive pulmonary disease – a twin study. Respir. Med. 104, 1890–1895 (2010). [DOI] [PubMed] [Google Scholar]
  • 8.Silverman EK Genetics of COPD. Annu. Rev. Physiol. 82, 413–431 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Lichtenstein P et al. Environmental and heritable factors in the causation of cancer — analyses of cohorts of twins from Sweden, Denmark, and Finland. N. Engl. J. Med. 343, 78–85 (2000). [DOI] [PubMed] [Google Scholar]
  • 10.Karczewski KJ et al. Pan-UK Biobank genome-wide association analyses enhance discovery and resolution of ancestry-enriched effects. Nat. Genet. 57, 2408–2417 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Zhou JJ et al. Heritability of chronic obstructive pulmonary disease and related phenotypes in smokers. Am. J. Respir. Crit. Care Med. 188, 941–947 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Gorman BR et al. Multi-ancestry GWAS meta-analyses of lung cancer reveal susceptibility loci and elucidate smoking-independent genetic risk. Nat. Commun. 15, 8629 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Moll M et al. Chronic obstructive pulmonary disease and related phenotypes: polygenic risk scores in population-based and case-control cohorts. Lancet Respir. Med. 8, 696–708 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Tsuo K et al. Multi-ancestry meta-analysis of asthma identifies novel associations and highlights the value of increased power and diversity. Cell Genomics 2, 100212 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Turley P et al. Multi-trait analysis of genome-wide association summary statistics using MTAG. Nat. Genet. 50, 229–237 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Grove J et al. Identification of common genetic risk variants for autism spectrum disorder. Nat. Genet. 51, 431–444 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Tadros R et al. Shared genetic pathways contribute to risk of hypertrophic and dilated cardiomyopathies with opposite directions of effect. Nat. Genet. 53, 128–134 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Liu M et al. Association studies of up to 1.2 million individuals yield new insights into the genetic etiology of tobacco and alcohol use. Nat. Genet. 51, 237–244 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Patel AP et al. A multi-ancestry polygenic risk score improves risk prediction for coronary artery disease. Nat. Med. 29, 1793–1803 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Inouye M et al. Genomic risk prediction of coronary artery disease in 480,000 adults. J. Am. Coll. Cardiol. 72, 1883–1893 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Moll M et al. Polygenic risk scores identify heterogeneity in asthma and chronic obstructive pulmonary disease. J. Allergy Clin. Immunol. 152, 1423–1432 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Martin AR et al. Clinical use of current polygenic risk scores may exacerbate health disparities. Nat. Genet. 51, 584–591 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.The 1000 Genomes Project Consortium et al. A global reference for human genetic variation. Nature 526, 68–74 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Bergström A et al. Insights into human genetic variation and population history from 929 diverse genomes. Science 367, eaay5012 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Feng Y-CA et al. Taiwan Biobank: a rich biomedical research database of the Taiwanese population. Cell Genomics 2, 100197 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Jee YH et al. Genome-wide association studies in a large Korean cohort identify quantitative trait loci for 36 traits and illuminate their genetic architectures. Nat. Commun. 16, 4935 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Jee YH et al. Cohort Profile: The Korean Cancer Prevention Study-II (KCPS-II) Biobank. Int. J. Epidemiol. 47, 385–386f (2018). [DOI] [PubMed] [Google Scholar]
  • 28.Shrine N et al. Multi-ancestry genome-wide association analyses improve resolution of genes and pathways influencing lung function and chronic obstructive pulmonary disease risk. Nat. Genet. 55, 410–422 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Pirinen M linemodels: clustering effects based on linear relationships. Bioinformatics 39, btad115 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Ge T et al. Development and validation of a trans-ancestry polygenic risk score for type 2 diabetes in diverse populations. Genome Med. 14, 70 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Ruan Y et al. Improving polygenic prediction in ancestrally diverse populations. Nat. Genet. 54, 573–580 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Ge T, Chen C-Y, Ni Y, Feng Y-CA & Smoller JW Polygenic prediction via Bayesian regression and continuous shrinkage priors. Nat. Commun. 10, 1–10 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Truong B et al. Integrative polygenic risk score improves the prediction accuracy of complex traits and diseases. Cell Genomics 4, 100523 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Hammad H & Lambrecht BN The basic immunology of asthma. Cell 184, 1469–1485 (2021). [DOI] [PubMed] [Google Scholar]
  • 35.Ramalho SHR & Shah AM Lung function and cardiovascular disease: a link. Trends Cardiovasc. Med. 31, 93–98 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.El-Azeem IAA, Hamdy G, Amin M & Rashad A Pulmonary function changes in diabetic lung. Egypt. J. Chest Dis. Tuberc. 62, 513–517 (2013). [Google Scholar]
  • 37.Çolak Y, Nordestgaard BG & Afzal S Morbidity and mortality in carriers of the cystic fibrosis mutation CFTR Phe508del in the general population. Eur. Respir. J. 56, 2000558 (2020). [DOI] [PubMed] [Google Scholar]
  • 38.Pereira SV-N, Ribeiro JD, Ribeiro AF, Bertuzzo CS & Marson FAL Novel, rare and common pathogenic variants in the CFTR gene screened by high-throughput sequencing technology and predicted by in silico tools. Sci. Rep. 9, 6234 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Wang Y et al. Rare variants of large effect in BRCA2 and CHEK2 affect risk of lung cancer. Nat. Genet. 46, 736–741 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Cortez Cardoso Penha R et al. Common genetic variations in telomere length genes and lung cancer: a Mendelian randomisation study and its novel application in lung tumour transcriptome. eLife 12, e83118 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Kachuri L et al. Mendelian Randomization and mediation analysis of leukocyte telomere length and risk of lung and head and neck cancers. Int. J. Epidemiol. 48, 751–766 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Shi J et al. Genome-wide association study of lung adenocarcinoma in East Asia and comparison with a European population. Nat. Commun. 14, 3043 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Duckworth A et al. Telomere length and risk of idiopathic pulmonary fibrosis and chronic obstructive pulmonary disease: a mendelian randomisation study. Lancet Respir. Med. 9, 285–294 (2021). [DOI] [PubMed] [Google Scholar]
  • 44.Xu K et al. Genome-wide association study of smoking trajectory and meta-analysis of smoking status in 842,000 individuals. Nat. Commun. 11, 5302 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Wang S et al. Significant associations of CHRNA2 and CHRNA6 with nicotine dependence in European American and African American populations. Hum. Genet. 133, 575–586 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Thorgeirsson TE et al. A variant associated with nicotine dependence, lung cancer and peripheral arterial disease. Nature 452, 638–642 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Saccone NL et al. Multiple distinct risk loci for nicotine dependence identified by dense coverage of the complete family of nicotinic receptor subunit (CHRN) genes. Am. J. Med. Genet. Part B Neuropsychiatr. Genet. Off. Publ. Int. Soc. Psychiatr. Genet. 150B, 453–466 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Ni Z et al. CLPTM1L is overexpressed in lung cancer and associated with apoptosis. PLoS One 7, e52598 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Chen XF et al. Multiple variants of TERT and CLPTM1L constitute risk factors for lung adenocarcinoma. Genet. Mol. Res. 11, 370–378 (2012). [DOI] [PubMed] [Google Scholar]
  • 50.James MA, Vikis HG, Tate E, Rymaszewski AL & You M CRR9/CLPTM1L regulates cell survival signaling and is required for Ras transformation and lung tumorigenesis. Cancer Res. 74, 1116–1127 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Ortega VE et al. The effects of rare SERPINA1 variants on lung function and emphysema in SPIROMICS. Am. J. Respir. Crit. Care Med. 201, 540–554 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Stoller JK & Aboussouan LS α1-antitrypsin deficiency. Lancet 365, 2225–2236 (2005). [DOI] [PubMed] [Google Scholar]
  • 53.Brenner DR, McLaughlin JR & Hung RJ Previous lung diseases and lung cancer risk: a systematic review and meta-analysis. PloS One 6, e17479 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Denholm R et al. Is previous respiratory disease a risk factor for lung cancer? Am. J. Respir. Crit. Care Med. 190, 549–559 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Kachuri L et al. Immune-mediated genetic pathways resulting in pulmonary function impairment increase lung cancer susceptibility. Nat. Commun. 11, 27 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Zhang D et al. Exploring the relationship between Treg-mediated risk in COPD and lung cancer through Mendelian randomization analysis and scRNA-seq data integration. BMC Cancer 24, 453 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57.Barr RG et al. Pulmonary disease and age at immigration among Hispanics. Results from the Hispanic Community Health Study/Study of Latinos. Am. J. Respir. Crit. Care Med. 193, 386–395 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58.Pino-Yanes M et al. Genetic ancestry influences asthma susceptibility and lung function among Latinos. J. Allergy Clin. Immunol. 135, 228–235 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59.Kachuri L et al. Gene expression in African Americans, Puerto Ricans and Mexican Americans reveals ancestry-specific patterns of genetic architecture. Nat. Genet. 55, 952–963 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60.Cooke CR et al. The validity of using ICD-9 codes and pharmacy records to identify patients with chronic obstructive pulmonary disease. BMC Health Serv. Res. 11, 37 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 61.Buhr RG et al. Reversible airflow obstruction predicts future chronic obstructive pulmonary disease development in the SPIROMICS cohort: an observational cohort study. Am. J. Respir. Crit. Care Med. 206, 554–562 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62.Hobbs BD et al. Genetic loci associated with chronic obstructive pulmonary disease overlap with loci for lung function and pulmonary fibrosis. Nat. Genet. 49, 426–432 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 63.The All of Us Research Program Investigators. The “All of Us” Research Program. N. Engl. J. Med. 381, 668–676 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64.He Y & Martin AR We need more-diverse biobanks to improve behavioural genetics. Nat. Hum. Behav. 8, 197–200 (2023). [DOI] [PubMed] [Google Scholar]

Methods-only References

  • 65.Bulik-Sullivan BK et al. LD Score regression distinguishes confounding from polygenicity in genome-wide association studies. Nat. Genet. 47, 291–295 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 66.Bulik-Sullivan B et al. An atlas of genetic correlations across human diseases and traits. Nat. Genet. 47, 1236–1241 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 67.Li X et al. Genome-wide association study of lung function and clinical implication in heavy smokers. BMC Med. Genet. 19, 134 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 68.Friedman J, Hastie T & Tibshirani R Regularization paths for generalized linear models via coordinate descent. J. Stat. Softw. 33, (2010). [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Figures
Supplementary Tables

Data Availability Statement

The individual-level genotype and phenotype data of All of Us are available on the Researcher Workbench. Researchers can register to access at: https://www.researchallofus.org/. The GWAS summary statistics for the East Asian and multi-ancestry meta-analyses of lung function are available at the GWAS Catalog (https://www.ebi.ac.uk/gwas) under the accession codes GCST90705067, GCST90705068, GCST90705069, GCST90705070, GCST90705071, and GCST90705072. Data sources for ancestry-specific summary statistics for each trait used in this study are available in Supplementary Table 1. Weights of PRSxtra for each trait are available in Supplementary Tables 1921.

RESOURCES