Skip to main content
Communications Biology logoLink to Communications Biology
. 2026 Sep 28;9:1226. doi: 10.1038/s42003-026-10875-9

Recurrent deletions and regulatory disruption of the Y chromosome in cancer

Trini Nguyen 1, Aditi Kuchi 1, Hongyu Liu 2,3, Nicholas Tatonetti 1,2,3,✉, Dan Theodorescu 4,5,✉
PMCID: PMC13619554  PMID: 42806068

Abstract

Somatic loss of the Y chromosome (LOY) is the most frequent acquired genomic alteration in aging males and occurs across multiple cancer types. While common, its functional role in tumorigenesis is only beginning to emerge. Most studies treat LOY as a binary event, ignoring partial deletions that may selectively remove gene-rich euchromatic regions. To address this, we used high-coverage whole-genome sequencing, large-scale transcriptomics, and eQTL to map male cancer cell lines, eliminating confounding non-malignant cells. We reveal a region-specific LOY landscape with recurrent euchromatic deletions affecting protein-coding genes and non-coding elements. Losses were quantified using a new Y EroSion (YES) Score, which is associated with transcriptional differences and clinical outcome patterns suggestive of functional relevance. Retained Y-linked loci remain transcriptionally active, enriched for proliferation, immune signaling, and stress adaptation functions. These findings position LOY as a structured genomic event with regulatory and clinical implications, motivating validation in primary tumor cohorts.

Subject terms: Cancer genomics, Cancer genomics


Using high-coverage whole-genome sequencing and transcriptomics of male cancer cell lines, the authors map the landscape of somatic Y chromosome loss at single-gene resolution. Y chromosome erosion is structured, region-specific and associated with transcriptional dysregulation and adverse clinical outcomes across multiple cancer types.

Introduction

Somatic loss of the Y chromosome (LOY) in peripheral blood mononuclear cells (PBMCs) is the most frequent acquired chromosomal alteration in aging males and is associated with cancer mortality1–4, the mechanism of which has recently begun to be defined5,6. While initially observed in hematopoietic cells, LOY has now been documented in solid tumors such as bladder, lung, prostate, and colorectal cancers1–7, where its prevalence increases with age and tobacco exposure8,9. LOY in both PBMCs and cancer cells is associated has been implicated in tumor progression, poor prognosis, and immune dysregulation6,10,11.

Despite its small size, the Y chromosome encodes regulators of chromatin remodeling, transcription, and immune function beyond their roles in sex determination and spermatogenesis6,12–14, and their loss may impair cellular homeostasis, facilitate immune evasion, or promote oncogenesis6,15,16. We have recently shown that loss of specific genes such as UTY and KDM5D can significantly impact tumor growth and response to checkpoint inhibitor immunotherapy6. Beyond protein-coding genes, the Y chromosome harbors lincRNAs and pseudogenes that have been largely overlooked in prior LOY studies, despite growing evidence that non-coding RNAs regulate transcription through miRNA sponging, chromatin remodeling, and competitive endogenous RNA networks17,18. Nevertheless, most studies reduce LOY to a binary state-present or absent-overlooking partial deletions that may selectively remove gene-rich euchromatic regions while retaining others.

Efforts to resolve LOY at higher resolution have been hampered by the repetitive and palindromic nature of the Y chromosome, which complicates alignment and variant calling12,19. Recent advances in high-coverage whole-genome sequencing (WGS) and large-scale transcriptomic profiling now enable more nuanced investigations, including detection of mosaicism, gene-specific losses, and downstream regulatory effects20,21.

Here, we integrate high-coverage WGS and RNA-seq data from the Cancer Cell Line Encyclopedia (CCLE) to characterize the genomic and transcriptional landscape of LOY in 160 male cancer cell lines-a controlled discovery framework that eliminates normal cell contamination inherent to bulk tumor tissue. Using bin-level read densities across chromosome Y, we identify recurrently deleted euchromatic regions and develop the Y EroSion (YES) Score, a continuous metric of Y-linked genomic loss. We validate these findings against two external datasets: a pan-cancer clinical gene panel and single-nucleus LOY profiles from bladder tumors10. We further characterize recurrently lost Y-linked lincRNAs and pseudogenes, perform eQTL analysis linking Y-linked variants to genome-wide transcriptional changes, and map LOY burden across TCGA tumor types. Together, these analyses position LOY as a structured genomic event with coordinated coding and non-coding losses that may disrupt tumor-associated transcriptional networks.

Results

Whole genome sequencing identifies specific regions of Y chromosome loss in cancer cell lines

We analyzed high-coverage WGS data from 160 male human cancer cell lines in CCLE22. Normalized read count distributions revealed that chromosome Y consistently displayed markedly reduced coverage relative to autosomes and chromosome X (Fig. 1A). Chromosomes X, Y, 6, 21, and 22 served as internal controls; chromosomes 21 and 22 were selected for size similarity to chromosome Y, and chromosome 6 as a mid-sized autosome reference. Although this normalization does not explicitly correct for ploidy variation, the consistent leftward shift of chromosome Y relative to all other chromosomes provides robust qualitative evidence of widespread LOY across the dataset2. Y to chromosome 22 read count ratios confirmed that the majority of samples fell below 1.0, consistent with widespread Y chromosome loss (Fig. S1); a small subset at or above 1.0 may reflect diploid retention or rare Y chromosome gain and represents a potential source of heterogeneity within the non-LOY group.

Fig. 1. Evidence of widespread and variable LOY across male cell lines.

Fig. 1

A Distribution of normalized read counts across chromosomes X, Y, 6, 21, and 22 in CCLE WGS data. Read counts were normalized by chromosome length. B Read density along chromosome Y. The histogram shows mapped read positions across the euchromatic region, aggregated across all samples. Regions of high or low coverage reflect retention or loss, respectively, with a schematic of chromosome Y below for genomic context. C Sparse read coverage across chromosome Y. Reads were binned into 1 kb intervals; each dot represents mapped reads in a bin for a sample, with samples ordered by the number of bins lacking coverage (“sparse regions”). The Y EroSion (YES) Score quantifies this pattern, with high YES indicating extensive sparse regions and 0 indicating no loss. D Relationship between hazard ratios (HR) and association significance with the YES score for ~ 30 differentially expressed genes. Linear regression lines are shown for genes upregulated with increasing YES (blue; Spearman’s two-sided ρ = −0.102, p = 0.5091, n = 14 genes) and downregulated genes (red; Spearman’s ρ = 0.56, p = 0.02, n = 16 genes).

To localize affected regions, we generated read density profiles across chromosome Y in 1 kb bins, revealing reproducible focal losses interspersed with areas of retained reads (Fig. 1B). Retained regions clustered around the centromere, consistent with prior observations9. Four lines of evidence support the biological credibility of these retention patterns: multiple samples showed at least one mapped read in every bin, excluding fixed mappability constraints; retained regions corresponded to areas of detectable gene expression in matched RNA-seq data (Fig. S2); retention patterns were consistent across hg19 and GRCh38 (Fig. S3); and only two translocation events involving chromosome Y were identified across all samples (Fig. S4), GC bias analysis confirmed stable normalized coverage across the 20–60% GC range (Fig. S5), indicating sparse coverage patterns reflect biological loss rather than technical artifact. We additionally aggregated read depth into 100 kb windows to quantify large-scale depletion, enabling both fine-scale regional analysis and robust sample-level assessment (Fig. 1B, C). Figure 1B shows cumulative retention and loss across all 160 samples, while Fig. 1C shows sample-level coverage ordered by YES score, capturing the full spectrum of individual LOY patterns. Samples were ranked by the number of these sparse bins, defining the YES Score-a continuous descriptor of Y chromosome erosion heterogeneity intended to stratify samples for transcriptional and regulatory analyses and not as a validated prognostic biomarker (Fig. 1C).

Differential gene expression analysis comparing high versus low YES score samples identified 30 LOY-associated differentially expressed (DESeq2, adjusted p < 0.05, ∣log2fold change∣ > 1; Fig. S5). Downregulated genes showed a positive association between hazard ratio (HR) and statistical significance in external Kaplan–Meier analyses, whereas upregulated genes showed a weaker or inverse relationship (Fig. 1D); these hazard ratios do not account for treatment, age, or other clinical covariates. Cross-assembly validation confirmed that 18 recurrently lost genes were consistently identified in both hg19 and GRCh38, including USP9Y, KDM5D, and EIF1AY, overlapping with the 24-gene set derived from the full CCLE cohort (Table S1 and Fig. S3) dataset using GRCh38; genes appearing on both lists-including USP9Y, KDM5D, and EIF1AY-represent those most robustly identified as recurrently lost regardless of reference genome or sample subset. Technical artifact was excluded by uniformly high genome-wide alignment rates (Fig. S6) and stable Y-to-22 and 21-to-22 read ratios (Figs. S1 and S7). No systematic bias in LOY across tissue types was observed (Fig. S8), indicating that YES variability reflects biological rather than tissue composition effects.

In 26 matched tumor-normal WGS pairs, most tumors showed lower YES scores, though this pattern was not uniform across all pairs (Fig. S9). In cases where normal samples displayed elevated YES scores, paired tumors also showed Y chromosome loss, consistent with LOY representing a possible early premalignant event, though this interpretation is limited by the small sample size (n = 26 pairs) and inter-individual variability (Figs. S10, S11)23.

Finally, WGS read density correlated significantly with gene expression levels (Spearman’s ρ = 0.36, p < 0.0001; Fig. S2)15, indicating that sparse coverage is associated with reduced transcriptional output. Together, these data reveal a heterogeneous, region-specific landscape of Y chromosome loss in cancer cell lines, characterized by recurrent focal deletions rather than uniform chromosome loss.

Recurrently lost Y-linked genes are shared across datasets and associated with poor prognosis

To identify Y-linked genes consistently affected by LOY, we evaluated bin-level read density across the euchromatic Y chromosome using a CDF-derived threshold of 2000 reads to classify low-coverage bins (Fig. S12)1,23. This threshold marked a natural inflection point separating high- and low-coverage regions and was consistent across samples. Genes overlapping sparse regions yielded 24 recurrently lost protein-coding genes from the full 160-sample CCLE dataset. Of these, 15 overlapped with at least one external source: single-nucleus sequencing of bladder tumors6 and the pan-cancer TEMPUS dataset (Figs. 2A and S13). Genes identified in both the 24-gene CCLE set and the 18-gene cross-assembly set-including USP9Y, KDM5D, and EIF1AY-represent the highest-confidence recurrently lost loci regardless of sample subset or reference genome.

Fig. 2. Recurrently lost Y-linked genes identified in CCLE show substantial overlap with in vivo and TEMPUS datasets and display consistent transcriptional loss across cancer cell lines.

Fig. 2

A Overlap of Y-linked genes with low read density in CCLE and two external datasets. The Venn diagram shows intersections between 24 protein-coding Y-linked genes recurrently lost in CCLE WGS data, genes frequently lost in the in vivo study by ref. 6, and genes with low coverage in the TEMPUS panel. Fifteen genes overlap with at least one external dataset. A summary of gene-level overlap is provided in Table S2. B Heatmap of Y-linked gene expression across CCLE male cancer cell lines. The 15 overlapping genes are shown, with samples grouped by LOY status (LOY vs non-LOY). Colors represent z-scores of log-transformed raw counts45. C Distribution of Y-linked gene loss across TCGA cancer types. The y-axis shows the average number of affected genes (of the 23 most frequently lost) per male cell line. Each box represents the interquartile range (IQR), with the central line indicating the median; whiskers extend to 1.5 × the IQR, and individual points beyond the whiskers represent outliers. Cancer types are ordered by median gene loss frequency.

Expression of Y-linked genes in recurrently deleted regions was markedly reduced in LOY samples relative to non-LOY samples in a gene- and sample-specific manner, consistent with cis dosage effects of focal genomic deletion rather than global transcriptional suppression (Fig. 2B). USP9Y, KDM5D, EIF1AY, and UTY were consistently downregulated or undetectable in cell lines where genomic coverage indicated locus loss.

To assess prognostic relevance, we plotted per-gene retention frequency against average HR from KMplot Kaplan–Meier analyses. A statistically significant inverse correlation was observed (Pearson’s r = −0.42, p = 0.038), indicating that frequently lost genes tend toward higher hazard ratios (Fig. S14)1,4. These associations are indirect and do not account for treatment, age, or other clinical covariates.

We then quantified Y-linked gene loss burden across TCGA tumor types by counting how many of the 24 recurrently lost genes were affected per sample, normalized by cohort size. Lymphoma and mesothelioma showed the highest burden (Fig. 2C); a parallel analysis restricted to the 15-gene cross-platform overlap set yielded broadly similar patterns (Fig. S15). Tumor ploidy did not confound these estimates (Spearman’s ρ = −0.09, p = 0.37; Fig. S16), supporting focal biological loss rather than genome-wide copy number instability as the driver. Cancers previously associated with worse LOY prognosis4 appeared at both extremes of the YES distribution, underscoring that LOY burden alone is insufficient to predict clinical outcome without accounting for tumor type.

Recurrently lost non-coding Y-linked elements exhibit broad genomic associations

Beyond protein-coding genes, we examined Y-linked lincRNAs and pseudogenes overlapping sparse WGS coverage regions (Fig. 3A). These elements have been systematically excluded from prior LOY analyses despite their substantial representation in the Y chromosome euchromatic content and established roles in trans-regulatory gene expression through miRNA sponging, chromatin remodeling, and ceRNA mechanisms17,18. Several elements-including TTTY13, TTTY14, RBMY2EP, and AC010086.1-were consistently located within recurrently deleted regions24.

Fig. 3. Recurrently lost Y-linked non-coding elements and their putative trans-regulatory targets across the genome.

Fig. 3

A Non-coding Y-linked elements recurrently lost across CCLE samples. Long intergenic non-coding RNAs (lincRNAs, red) and pseudogenes (blue) are shown along the euchromatic region of chromosome Y. Genomic coordinates correspond to regions of sparse read coverage identified from WGS data. B Genomic distribution of autosomal genes associated with Y-linked non-coding elements. Predicted targets of recurrently lost Y-linked lincRNAs and pseudogenes were identified using the ENCORI database. C Expression of recurrently lost Y-linked lincRNAs and pseudogenes across LOY and non-LOY CCLE samples. Boxplots show log2-tranformed raw counts (log2(count + 1)) for each non-coding element identified in (A), with individual sample points shown as jittered dots. Samples were classified as LOY or non-LOY based on whether their YES score exceeded the median YES score across all samples. Statistical comparisons were performed by Wilcoxon rank-sum two-sided test; significance is indicated as ns not significant, * p < 0.05, ** p < 0.01, *** p < 0.001, **** p < 0.0001.

Using the ENCORI database25, which predicts RNA-RNA and RNA-protein interactions from CLIP-seq data, we identified putative autosomal targets of recurrently lost Y-linked non-coding elements. These computationally predicted targets spanned nearly all autosomes, with the exception of chromosomes 10, 11, 13, and 21 (Fig. 3B); this distributional gap may reflect annotation or coverage limitations in ENCORI rather than genuine biological specificity. Consistent with genomic deletion driving transcriptional silencing, TTTY14 and TTTY13 showed statistically significant expression reduction in LOY samples by Wilcoxon rank-sum test (Fig. 3C), though not all elements reached significance, possibly reflecting low baseline expression or inter-sample variability. Assessment of the highest-confidence ENCORI-predicted autosomal targets of TTTY14, LINC00279, and TTTY11 revealed no statistically significant differential expression after multiple testing correction, indicating that predicted element-target relationships do not produce detectable transcriptional changes in the current dataset.

Together, these findings identify recurrently lost Y-linked non-coding elements with reduced transcriptional output in LOY samples and broadly distributed putative autosomal targets. Their functional trans-regulatory consequences remain to be established experimentally.

Y-linked eQTLs in sparse regions influence transcriptional programs across the genome

To investigate transcriptional consequences of Y chromosome loss, we performed eQTL analysis using Matrix eQTL26, restricted to variants within sparse Y-linked regions to assess how their loss might affect autosomal gene expression in trans. Y-linked eQTLs clustered within discrete, commonly deleted areas of chromosome Y (Fig. 4A), suggesting that recurrent deletions remove critical regulatory loci. Target genes were distributed across all chromosomes, though chromosomes 13, 21, 22, and Y harbored the fewest (Fig. 4B), echoing the trans-regulatory breadth observed in our non-coding RNA analysis. GO enrichment analysis on the eQTL target genes identified overrepresented biological processes, including cell motility, wound healing, extracellular matrix organization, and stress response (Fig. 4C; fold enrichment range 20–60), highlighting pathways frequently co-opted during tumor progression27. These enrichments may help prioritize candidates for future functional investigation, since GO analysis alone does not establish mechanistic links between Y-linked regulatory loss and pathway disruption.

Fig. 4. Distribution, functional enrichment, and differential expression of genes targeted by Y-linked eQTLs in regions affected by mosaic loss of chromosome Y.

Fig. 4

A Distribution of eQTL loci across recurrently lost CCLE regions of the Y chromosome. Histogram of Y chromosome positions (1 kb bins) showing eQTLs only within sparse euchromatic regions. B Chromosomal distribution of genes associated with Y-linked eQTLs. C GO enrichment of Y-linked eQTL target genes. Bar plot shows significantly enriched biological processes among Y-lined eQTL target genes, ranked by fold enrichment. D Volcano plot of differential expression between LOY and non-LOY samples. Non-significant genes are grey, downregulated genes in LOY are blue, upregulated genes are red (two-sided adjusted p < 0.05 and ∣log2 fold change∣ > 1). Y-linked eQTL target genes in LOY-affected regions are highlighted in orange, regardless of significance.

Intersecting eQTL results with differential expression analysis identified 268 autosomal genes meeting both criteria (Y-linked sparse region eQTL association at p < 0.05 and differential expression at adjusted p < 0.05). Of these, 168 were downregulated and 100 were upregulated in LOY samples, indicating bidirectional rather than uniformly repressive effects. These eQTL-identified targets represent a subset of total LOY-associated differentially expressed genes; additional mechanisms, including indirect transcriptional cascades, altered oncogenic master regulator activity, and post-transcriptional regulation, likely contributed to the broader expression differences observed. A volcano plot highlighting the eQTL targets in sparse Y regions shows that 12.65% met thresholds for strong differential expression (adjusted p < 0.05 and ∣log2 fold change∣ > 1; Fig. 4D).

Together, retained Y-linked regions are not functionally inert but are associated with gene regulatory networks related to immune modulation, proliferation, and stress adaptation14, with potential selective and clinical relevance in the tumor context.

Y-linked regions retained in LOY samples exhibit distinct regulatory functions

Despite widespread Y chromosome erosion, a subset of protein-coding and non-coding genes was consistently retained across LOY-positive samples (Fig. 5A). GO enrichment analysis of retained genes revealed significant overrepresentation of terms related to nucleosome assembly, chromatin organization, cell differentiation, and reproductive development (Fig. 5B), consistent with the hypothesis that selective pressure maintains elements relevant to chromatin regulation and cellular identity6. eQTL analysis of single-nucleotide polymorphisms (SNPs) in retained Y regions identified autosomal target genes broadly distributed across the genome, with a chromosomal pattern similar to that observed for deleted-region eQTLs (Fig. 5C). GO analysis of these targets identified functional categories including immune surveillance, inflammatory signaling, cell differentiation, tissue organization, stress response, and programmed cell death28. Of retained-region eQTL targets, 331 were upregulated and 207 were downregulated in non-LOY samples (Fig. S17). This bidirectional patterns mirrors that observed for lost-region eQTL targets and likely reflects the complexity of trans-regulatory effects and indirect downstream consequences rather than distinct regulatory mechanisms; these results are therefore exploratory and require functional validation.

Fig. 5. Regulatory activity and functional relevance of Y chromosome regions retained in LOY-positive CCLE samples.

Fig. 5

A Y-linked genes consistently preserved in LOY-positive cell lines. Retained protein-coding genes are red, RNA genes are blue, and pseudogenes are yellow. B GO enrichment of retained genes. C Chromosomal distribution of autosomal targets of Y-linked eQTLs from retained regions.

Together, these findings demonstrate that recurrent loss of Y-linked regulatory elements is associated with broad transcriptional consequences across the genome, particularly in pathways related to stress adaptation and cellular motility14.

Discussion

This study moves beyond binary LOY classification to reveal a structured, region-specific landscape of Y chromosome erosion in cancer, with multilayered transcriptional consequences. The YES score, defined here as a quantitative descriptor of Y chromosome erosion heterogeneity, enables continuous stratification of samples for hypothesis generating analyses rather than serving as a validated prognostic biomarker. Convergent support from the bladder tumor dataset, the TEMPUS cohort, TCGA burden analysis, and paired tumor-normal comparison corroborate a subset of these findings in primary tumor contexts, but extrapolation to human tumor biology will require validation in primary patient cohorts with matched WGS and transcriptomic data.

The recurrent, non-random nature of the focal deletions we describe is conceptually distinct from stochastic age-related mosaic LOY in blood. Cancer-associated Y chromosome erosion is characterized by selective loss of specific euchromatic gene-rich regions, somatic origin as evidenced by discordance between tumor and matched normal YES scores (Figs. S9–S11), and absence in healthy males from the 1000 Genomes Project. These features are inconsistent with passive chromosomal instability and instead suggest active selective pressure operating on specific Y-linked loci during tumorigenesis. Structural rearrangements are unlikely to account for these patterns, as only two translocation events involving chromosome Y were identified across all CCLE samples (Fig. S4). The physical configuration of a Y chromosome bearing multiple focal deletions-whether shortened, internally gapped, or forming extrachromosomal DNA complexes29-remains an open question. Recurrently lost regions encompass both protein-coding genes and non-coding elements, including lincRNAs and pseudogenes implicated in transcriptional regulation, chromatin remodeling, and immune modulation12,13,24, several of which overlap with genes identified in bladder tumor tissue, the TEMPUS dataset, and experimental Y chromosome depletion in murine cancer cells6.

The regulatory consequences of LOY extend across nearly all autosomes. eQTL analysis links recurrently deleted Y-linked loci to transcriptional alterations enriched in pathways related to detoxification, tissue remodeling, immune signaling, and stress response2,6,10. Expression changes among eQTL targets are bidirectional rather than uniformly repressive, indicating that a model of global transcriptional repression is an oversimplification6,15. eQTL-identified targets account for only a subset of total LOY-associated expression differences, indicating that additional regulatory mechanisms contribute to the broader transcriptional landscape. These may include indirect transcriptional cascades, post-transcriptional regulation, and altered master regulator activity; however, TF enrichment analysis of differentially expressed genes did not yield statistically significant results after multiple testing correction, and their specific contributions remain to be determined. Retained Y-linked regions are likewise not transcriptionally inert, associating through eQTLs with genes involved in immune surveillance, cell differentiation, and apoptosis28. The mechanistic basis of both lost- and retained-region eQTL associations, whether mediated through chromatin accessibility, three-dimensional genome organization, or RNA-DNA interactions-remains to be established through future work, including ATAC-seq profiling and chromatin conformation capture in isogenic LOY and non-LOY models.

The distributional gap in non-coding RNA targets and eQTL associations on chromosomes 10, 11, 13, and 21 may reflect technical limitations, including limited CLIP-seq coverage, under-annotation, acrocentric chromosome organization, or low gene expression in cancer cell lines, rather than genuine biological specificity, and should be treated as exploratory observations.

Our findings extend prior work4 linking LOY to poor clinical outcomes in diffuse large B-cell lymphoma, bladder urothelial carcinoma, sarcoma, and prostate adenocarcinoma, with these cancer types also showing the highest Y-linked gene loss burden in our TCGA analysis. The discordance between our conservative bin-based approach-which may classify high-level mosaic LOY as partial rather than complete-and prior bulk copy number estimates suggests that region-specific attrition rather than full aneuploidy is the predominant form of Y-linked disruption in solid tumors. This also reconciles FISH-based LOY detection with sequencing approaches; when key genes, including DAZ1, USP9Y, KDM5D, and UTY, are lost, fluorescent signal from alpha-satellite repeats in Yp11.1-q11.1 may fall below detection thresholds23. The 24-gene CCLE-derived set applied to TCGA requires independent validation in primary tumor WGS cohorts, where tumor purity and stromal admixture may influence detection.

Several limitations warrant consideration. All WGS-based analyses presented here are derived from this controlled cell line discovery framework and require validation in primary tumor cohorts before generalization to in vivo tumor biology. Survival associations from KMplot aggregate heterogeneous populations without covariate adjustment. Passage number and replicative history are not systematically recorded in CCLE and cannot be fully excluded as contributors to chromosomal instability in individual cell lines, though the recurrent, region-specificity, and cross-dataset consistency of our findings argue against this. The use of chromosomes 21 and 22 as sequencing depth controls was determined a priori and is not circular with the observation that these chromosomes harbor fewer eQTL. A small number of CCLE samples showed Y-to-chromosome-22 consistent with Y chromosome gain; these were retained in all analyses but may attenuate observed LOY-associated differences. Finally, while our analyses focus on the euchromatic Y chromosome, the T2T assembly30 may provide additional resolution for ampliconic and heterchromatic regions in future studies.

Together, these findings position LOY as a structured genomic event that reshapes the transcriptome through coordinated loss and retention of regulatory elements, challenging its classification as a passive marker of genomic instability. Future studies should prioritize functional validation through chromatin accessibility profiling, chromatin conformation capture, and CRISPR-based restoration of specific Y-linked loci in isogenic models, alongside prospective validation of the YES score and Y-linked gene loss burden in clinically annotated primary tumor cohorts.

Methods

Data sources

We used multiple publicly available and proprietary datasets to investigate recurrent LOY across cancer types.

Cancer cell line encyclopedia (CCLE)

We obtained WGS, RNA-seq, and alignment metrics for male tumor-derived cell lines from the CCLE22, which contained cells from a variety of cancer types (Table S2). RNA-seq data were downloaded as raw gene-level counts for differential expression and gene set enrichment analysis. Cell line metadata, including sex annotation and lineage, were retrieved from the DepMap portal (https://depmap.org). A complete summary of tumor types represented in the male CCLE cell line dataset analyzed here, including sample counts and LOY status distribution per tumor type, is provided in Table S3. BAM alignment files and associated read statistics were used to quantify Y chromosome coverage and assess LOY. Cell line quality control in the CCLE dataset is performed by the Broad Institute as part of standard data generation procedures and includes short tandem repeat (STR) profiling for cell line authentication and mycoplasma testing to confirm mycoplasma-negative status prior to sequencing. These quality control measures are documented in the CCLE data release notes and have been confirmed for the cell lines used in this analysis. Additionally, all WGS samples were confirmed to have high genome-wide alignment rates (fraction aligned > 0.8 across all samples; Fig. S18), providing further quality assurance for the sequencing data analyzed here.

ENCORI database

To identify predicted target genes of Y-linked lincRNAs and pseudogenes, we queried the ENCORI database (https://starbase.sysu.edu.cn/)25, which compiles experimentally supported and computationally predicted RNA-RNA and RNA-protein interactions from CLIP-seq and related assays.

TEMPUS pan-cancer immunotherapy cohort dataset

We used a transcriptomic and clinical dataset of 1019 male patients from TEMPUS who received anti-PD-1 or anti-PD-L1 immune checkpoint inhibitors, including cases of bladder, lung, kidney, colorectal, head and neck, and gastroesophageal cancers. RNA sequencing was performed using a proprietary exome-capture-based assay that incorporates custom spike-in probes to generate targeted gene expression profiles from tumor samples. The assay (RS.v2) uses IDTxGene Exome Research Panel v2.0, targeting 19,433 genes and produces quantified expression data for 20,061 genes. RNA-seq data were generated from bulk tumor tissue and quantified using transcripts per million. To assess Y chromosome loss at the transcriptomic level, we applied single-sample Gene Set Enrichment Analysis via the GSVA R package31 using a curated gene set of Y-linked protein-coding genes lacking X chromosome homologs, which resulted in 22 genes total (Fig. S13B). Normalized Enrichment Scores for this LOY gene set were computed per sample. Patients were stratified into LOY-high and LOY-low groups using thresholds determined by the surv_cutpoint function from the survminer R package32, as well as alternative mean and median-based cutoffs. Because TEMPUS LOY assessment relied on RNA-seq-based enrichment scoring rather than WGS read density, the detection of Y-linked gene loss in this dataset is subject to different sensitivity and specificity characteristics than the WGS-based approach applied to CCLE. Genes identified as lost in TEMPUS may reflect transcriptional silencing rather than genomic deletion, and this distinction should be considered when interpreting cross-dataset overlaps. Treatment duration was defined as the interval between therapy initiation and last available follow-up.

Cross-dataset gene loss comparisons were performed using platform-appropriate thresholds for each dataset rather than a unified normalized metric, given the substantial differences in sequencing modality, coverage and biological context across CCLE, the in vivo bladder tumor dataset and TEMPUS. In CCLE, gene loss was defined by WGS read density falling below 2000 reads in 1 kb bins across the euchromatic Y chromosome region. In the in vivo bladder dataset. Y chromosome status was inferred from single-nucleus RNA-seq data using previously established methods5. In TEMPUS, LOY was assessed via GSVA enrichment scores applied to a curated set of Y-linked protein-coding genes lacking X chromosome homologs. Because these thresholds and modalities are not directly comparable, the overlap analysis in Fig. 2A should be interpreted as convergent evidence across independent platforms rather than a normalized quantitative comparison.

Single-cell bladder tumor LOY dataset

We also analyzed a previously published single-cell RNA-seq dataset profiling LOY in human bladder tumors5. This dataset includes single-cell transcriptomes with inferred Y chromosome status, enabling cell-type-specific analysis of LOY effects. Because Y chromosome status in this dataset was inferred from transcriptomic data rather than directly measured from DNA-level read coverage, comparisons with CCLE WGS-derived loss calls reflect convergent evidence across modalities rather than equivalent measurements of the same biological phenomenon.

Colon cancer tumor-normal dataset

To assess whether recurrent Y-linked deletions could reflect inherited polymorphisms, we analyzed a paired colon cancer dataset containing matched tumor and adjacent normal tissue samples from male patients33. WGS data were used to identify Y chromosome read positions, and read coverage patterns were compared between tumor and normal pairs to detect focal deletions. The lack of concordance between tumor and normal samples within the same individual supports a somatic origin for LOY.

1000 genomes project healthy controls

As a control for baseline Y chromosome integrity, we analyzed WGS data from randomly selected healthy male samples from the 1000 Genomes Project34. Y-linked read positions were mapped and visualized to assess structural integrity. No evidence of Y chromosome loss was detected in any healthy individuals, supporting the conclusion that LOY is associated with tumorigenesis rather than representing inherited or population-level variation.

Read mapping and processing

We used WGS BAM files to extract reads aligned to chromosome Y35. Because only the euchromatic portion (0–30 Mb) was available, we restricted analyses to this region. We recorded read start positions and used them to compute read density profiles across chromosome Y. To align samples to GRCh38, all raw paired-end FASTQ files were passed the QC control applied by FastQC, and aligned to GRCh38 using STAR version 2.7.11b. Before alignment, the GRCh38 genome index was generated using the corresponding GTF annotation file from GENCODE v49. Reads were mapped using parameters –twopassMode Basic –outSAMtype BAM SortedByCoordinate. The resulting sorted BAM files were indexed with samtools and assessed for alignment quality using log.final.out metrics, including total uniquely mapped reads and mismatch rates.

We divided the chromosome into 1 kb bins and aggregated read counts within each bin to visualize consistent patterns of read loss and retention across samples. The 1 kb bin size was selected to balance two competing considerations: bins that are too small produce unstable read counts due to low coverage per bin, while bins that are too large obscure focal deletions by averaging across regions with heterogeneous coverage. At the sequencing depths present in the CCLE WGS data, 1 kb bins provide sufficient reads per bin in retained regions to distinguish true coverage loss from stochastic sampling variation, while remaining small enough to resolve focal deletions affecting individual genes. We defined sparse regions as bins with no mapped reads and quantified these per sample to assess LOY extent.

Coverage-based filtering of BAM Files

To reduce noise from low-coverage regions on the Y chromosome, we applied coverage-based filtering to BAM files using samtools and BEDTools36. For each BAM file, we first computed per-base read depth along chromosome Y and retained only positions with a coverage of at least 30 reads. The minimum coverage threshold of 30 reads was chosen based on standard practices for WGS variant calling and copy number analysis, at which depth stochastic coverage variation is sufficiently low to reliably distinguish covered from uncovered positions. This threshold is conservative relative to the overall sequencing depth of the CCLE WGS data and ensures that retained positions reflect genuine coverage rather than sporadic read mapping. Sensitivity analyses using thresholds of 10 and 50 reads yielded consistent sparse region identification, with the primary effect of lower thresholds being inclusion of additional low-confidence positions at the boundaries of deletion regions. Consecutive high-coverage bases were then merged into contiguous intervals, and only reads overlapping these intervals were extracted while excluding secondary, supplementary, and unmapped reads.

This approach ensured that downstream analyses focused on reliably covered regions of the Y chromosome, reducing technical noise and improving detection of recurrent Y-linked gene loss, providing a robust basis for quantifying Y chromosome erosion across samples.

Chromosome-level normalization and read count ratios

We extracted read counts for chromosomes X, Y, 6, 21, and 22, and normalized them by chromosome length to account for differences in genomic size. The choice of control chromosomes was made a priori based on the following rationale: chromosome 22 was selected as the primary autosomal control due to its small size, well-characterized gene content, and consistently high mappability, facilitating direct comparison with chromosome Y after length normalization; chromosome 21 was included as a second small autosome control to assess coverage patterns at a similar size range; and chromosome 6 was included as a representative mid-sized autosome to confirm that length normalization performs appropriately across chromosomes of different sizes. We note that37 demonstrated that chromosome 21 loss is common in solid tumors, which could in principle, affect its utility as a control. Examination of chromosome 21 to chromosome 22 read count ratios across our dataset (Fig. S7) revealed tight clustering around 1.0 in the majority of samples, indicating that chromosome 21 coverage is stable relative to chromosome 22 in the CCLE cell lines analyzed here and does not represent a widespread confound. We also confirm explicitly that the selection of these control chromosomes was made prior to and independently of the eQTL target distribution analysis shown in Fig. 4B, and is therefore not circular with the observation that chromosomes 21 and 22 show fewer eQTL targets in that analysis. We used chromosome 22 as an autosomal control because of its size and gene density. For each sample, we calculated read count ratios-Y:22 and 21:22-to distinguish biological mosaic LOY from technical artifacts1,38.

Gene expression analysis

We analyzed RNA-seq expression data from CCLE to quantify transcript levels of Y-linked genes19. After log-transforming raw counts and converting them to z-scores, we compared expression across samples. We categorized samples into LOY and non-LOY groups based YES score, using the top 100 most eroded samples as the LOY group and the remaining samples as non-LOY. This threshold was chosen to provide balanced group sizes while capturing samples with the most extensive Y chromosome erosion. We note that the dichotomization of samples into LOY and non-LOY groups based on YES score thresholds is an analytical convenience rather than a reflection of a true biological binary. The YES score is a continuous descriptor of erosion heterogeneity, and the choice of threshold for group classification introduces analytical flexibility that should be considered when interpreting downstream differential expression and survival analyses. We used reduced or absent expression of Y-linked genes in LOY samples to support the presence of Y chromosome loss39. In parallel, we performed differential expression analysis using linear regression, with the YES score as a continuous predictor.

Regulatory target prediction and genomic distribution analysis

We used the ENCORI database25 to identify predicted target genes of recurrently lost Y-linked lincRNAs and pseudogenes. Predictions were based on CLIP-seq-supported RNA-RNA and RNA-protein interactions. We cross-referenced predicted targets with CCLE expression data to evaluate relevance. To assess trans-regulatory potential, we visualized the chromosomal distribution of target genes. We obtained Y chromosome gene annotations, including lincRNAs and pseudogenes, from GENCODE v37.

Survival analysis

To evaluate the prognostic significance of recurrent Y-linked gene loss, we retrieved HRs from Kaplan–Meier analyses using the KMplot web tool40. We examined the relationship between gene retention frequency in CCLE and associated HRs to identify potential tumor-suppressive roles for lost genes3.

Statistical analyses and visualization

We conducted all data processing and statistical analyses in R (version 4.2.0) using the tidyverse suite41. We generated visualizations-including read density plots, gene expression heatmaps, Venn diagrams, and survival curves-using ggplot242. We selected data filtering thresholds based on empirical distributions and guidance from prior studies6.

eQTL analysis

To investigate trans-regulatory effects of somatic Y chromosome loss, we performed eQTL analysis using matched genotype and gene expression data from male CCLE cell lines.

Variant calling from whole-genome sequencing data

We performed variant calling on high-coverage WGS BAM files using bcftools and the GRCh37 (hg19) reference genome. We processed only coordinate-sorted BAM files and called variants on a per-sample basis. We then compressed and indexed VCF files for integration. We retained all SNPs for initial analysis.

Genotype matrix construction

We extracted genotypes from VCF files and converted genotype calls (e.g., 0/0, 0/1, 1/1) into numeric format (0, 1, 2). We matched sample IDs parsed from filenames to CCLE cell line names using a mapping file. After harmonizing the genotypes across samples, we created a unified matrix with SNPs as rows and cell lines as columns. We calculated minor allele frequency (MAF) and excluded SNPs with MAF ≤0.05 or zero variance. We converted the final genotype matrix to SlicedData format for use with Matrix eQTL.

Gene expression preprocessing

We retrieved RPKM gene expression data for the same cell lines from the CCLE repository. We applied a log2(RPKM + 1) transformation to stabilize variance and matched gene identifiers to genomic coordinates using Ensembl BioMart43. We excluded genes with zero variance and converted the final matrix to SlicedData format.

Principal component analysis and covariate generation

To control for hidden confounders, we ran principal component analysis (PCA) separately on expression and genotype matrices. We selected the top 32 PCs from the expression data and the top 20 from the genotype data. We combined and transposed these PCs to generate the covariate matrix for Matrix eQTL.

eQTL mapping

We performed eQTL mapping using the Matrix eQTL main function in linear model mode from the MatrixEQTL R package26. We focused on trans-eQTLs originating from Y chromosome SNPs, restricting analysis to variants in recurrently deleted regions identified from read depth analysis. We tested associations between these Y-linked SNPs and autosomal gene expression. We set a significance threshold of p < 0.05 and adjusted for multiple testing using the Benjamini–Hochberg procedure. All reported eQTLs passed the MAF filter and were located outside regions retained in LOY-positive samples.

Statistics and reproducibility

All statistical analyses were performed in R (version 4.2.0) using the tidyverse suite41. No new experimental data were generated in this study; all analyses were conducted on publicly available or previously published datasets, and reproducibility is therefore determined by the computational pipelines and thresholds described below rather than by experimental replication.

WGS-based read density analyses were performed across 160 male cancer cell lines from CCLE. Sparse coverage bins were defined using a threshold of 2000 reads per 1 kb bin, derived from the inflection point of the cumulative distribution function of genome-wide read density values (Fig. S12). Sensitivity of sparse region identification to coverage thresholds was assessed using values of 10, 50, and 30 reads per base for coverage-based BAM filtering, yielding consistent results across thresholds. The YES score was computed per sample as the count of sparse 1 kb bins across the euchromatic Y chromosome region and treated as a continuous variable throughout. For group-level analyses, LOY and non-LOY groups were defined by YES score, with the top 100 most eroded samples designated as LOY and the remainder as non-LOY. This threshold was chosen to provide balanced group sizes while capturing samples with the most extensive Y chromosome erosion; it represents an analytical convenience rather than a biological binary, and downstream results should be interpreted accordingly.

Differential gene expression between LOY and non-LOY groups was performed using DESeq2, with significance defined as adjusted p < 0.05 and ∣log2 fold change∣ > 1. Multiple testing correction used the Benjamini–Hochberg procedure throughout. In parallel, differential expression was assessed using linear regression with the YES score as a continuous predictor. eQTL mapping was performed using MatrixEQTL26 in linear model mode, with trans-eQTLs tested between Y-linked SNPs in sparse regions and autosomal gene expression. SNPs with MAF ≤0.05 or zero variance were excluded. Hidden confounders were controlled using the top 32 principal components from expression data and top 20 from genotype data as covariates. All reported eQTLs passed the MAF filter and Benjamini–Hochberg correction at p < 0.05. Survival associations were retrieved from the KMplot database and reflect external Kaplan–Meier analyses aggregated across heterogeneous patient populations without adjustment for treatment, age, or comorbidities; these should be interpreted as hypothesis-generating rather than as direct clinical validation. The inverse correlation between gene retention frequency and hazard ratio was assessed by Pearson correlation. The correlation between WGS read density and gene expression was assessed by Spearman correlation. Ploidy confounding was assessed by Spearman correlation between ABSOLUTE ploidy estimates and YES score in TCGA samples. Non-coding element expression differences between LOY and non-LOY samples were assessed by Wilcoxon rank-sum test.

Robustness of WGS-derived findings was assessed by repeating alignment and read density analyses on a representative sample subset using both GRCh38 and hg19 reference genomes, yielding consistent sparse region identification across assemblies (Fig. S3 and Table S1). No systematic bias in Y chromosome erosion across tissue types was observed (Fig. S8). All CCLE samples were confirmed to have high genome-wide alignment rates (fraction aligned > 0.8; Fig. S6). GC bias analysis confirmed stable normalized coverage across the 20–60% GC range (Fig. S5). Cell line quality control, including STR profiling and mycoplasma testing, was performed by the Broad Institute as part of standard CCLE data generation procedures.

Reporting summary

Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article.

Supplementary information

Reporting Summary (8.1MB, pdf)

Acknowledgements

We acknowledge TEMPUS for providing access to high-throughput sequencing datasets and as-sociated clinical annotations used in this study. We also thank Paul Wang and Yizhou Wang (Cedars-Sinai Medical Center) and Floris Barthel and Rhyker Ranallo-Benavidez (Translational Genomics Research Institute) for their valuable insights.

Author contributions

T.N., N.T. and D.T. conceived the study. T.N., A.K. and H.L. performed data curation. T.N. and A.K. performed formal analysis, investigation, and visualization. T.N. drafted the manuscript. T.N. and D.T. revised the manuscript. N.T. and D.T. provided domain expertise and supervision.

Peer review

Peer review information

Communications Biology thanks the anonymous reviewers for their contribution to the peer review of this work. Primary Handling Editor: Huiying Zhao, Rosie Bunton-Stasyshyn, Aylin Bircan.

Funding

N.T. was funded by the National Institutes of Health (NIH) under grant number R35 GM131905. D.T. was funded by the NIH under grant number R35CA294022.

Data availability

Publicly available datasets analyzed in this study include Cancer Cell Line Encyclopedia (CCLE) data obtained from the DepMap portal, colon cancer whole-genome sequencing (WGS) data from ref. 33, and healthy male samples randomly selected from the 1000 Genomes Project34. Survival analysis data were obtained from the Kaplan–Meier Plotter. Data generated by TEMPUS were used under a data use agreement and are not publicly available because of restrictions associated with that agreement. Processed datasets supporting the findings of this study are publicly and permanently available through Zenodo at44.

Code availability

Custom analysis scripts used to generate and analyze the results reported in this study are publicly and permanently available through Zenodo at44. The archived Zenodo record corresponds to the specific version of the code used for this study and provides a persistent DOI for citation and reproducibility. Software versions and study-specific variables or parameters required to reproduce the analyses are documented in the “Methods” and accompanying repository documentation.

Competing interests

The authors declare no competing interests.

Consent for publication

All authors read and approved the final manuscript.

Footnotes

Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Contributor Information

Nicholas Tatonetti, Email: Nicholas.Tatonetti@csmc.edu.

Dan Theodorescu, Email: theodorescu@arizona.edu.

Supplementary information

The online version contains supplementary material available at https://doi.org/10.1038/s42003-026-10875-9.

References

  • 1.Abdel-Hafiz, H., Hoelzen, L. & Theodorescu, D. Beyond sex determination: the Y chromosome in male cancers. Nat. Rev. Cancer26, 586–603 (2026). [DOI] [PubMed]
  • 2.Forsberg, J. et al. Mosaic loss of chromosome Y in peripheral blood is associated with shorter survival and higher risk of cancer. Nat. Genet.46, 624–628 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Wright, D. J. et al. Genetic variants associated with mosaic Y chromosome loss highlight cell cycle genes and overlap with cancer susceptibility. Nat. Genet.49, 674–679 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Qi, M. et al. Loss of chromosome Y in primary tumors. Cell186, 3125–3136 (2023). [DOI] [PubMed] [Google Scholar]
  • 5.Chen, X. et al. Concurrent loss of the Y chromosome in cancer and T cells impacts outcome. Nature642, 1041–1050 (2025). [DOI] [PMC free article] [PubMed]
  • 6.Abdel-Hafiz, H. et al. Y chromosome loss in cancer drives growth by evasion of adaptive immunity. Nature619, 624–631 (2023). [DOI] [PMC free article] [PubMed]
  • 7.Forsberg, J. et al. Age-related somatic structural changes in the Y chromosome in blood cells. J. Clin. Investig.122, 1340–1347 (2012). [Google Scholar]
  • 8.Danielsson, M. et al. Longitudinal changes in the frequency of mosaic chromosome Y loss in peripheral blood cells of aging men varies profoundly between individuals. Eur. J. Hum. Genet.28, 349–357 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Dumanski, J. P. et al. Smoking is associated with mosaic loss of chromosome Y. Science347, 81–83 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Dumanski, J. P. et al. Loss of Y chromosome in blood is associated with alzheimer disease. Nat. Commun.12, 1–11 (2021). [Google Scholar]
  • 11.Sano, S. et al. Hematopoietic loss of Y chromosome leads to cardiac fibrosis and heart failure mortality. Science377, 292–297 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Bellott, D. W. et al. Mammalian Y chromosomes retain widely expressed dosage-sensitive regulators. Nature508, 494–499 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Wilson, M. The Y chromosome and its impact on health and disease. Hum. Mol. Genet.30, 296–300 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Charlesworth, B. The organization and evolution of the human Y chromosome. Genome Biol.4, 226 (2003). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Dumanski, J. et al. Immune cells lacking Y chromosome show dysregulation of autosomal gene expression. Cell. Mol. Life Sci.78, 4019–4033 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Mattisson, J. et al. Loss of chromosome Y in regulatory T cells. BMC Genomics25, 243. (2024). [DOI] [PMC free article] [PubMed]
  • 17.Cesana, M. et al. A long noncoding RNA controls muscle differentiation by functioning as a competing endogeous RNA. Cell147, 358–369 (2011). [DOI] [PMC free article] [PubMed]
  • 18.Wang, K. & Chang, H. Molecular mechanisms of long noncoding RNAs. Mol. Cell43, 904–914 (2012). [DOI] [PMC free article] [PubMed]
  • 19.Hallast, P. et al. Assembly of 43 human Y chromosomes reveals extensive complexity and variation. Nature621, 355–364 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Zhou, W. et al. Mosaic loss of chromosome Y is associated with common variation near TCL1A. Nat. Genet.48, 563–568 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Vermeulen, M., Pearse, R., Young-Pearse, T. & Mostafavi, S. Mosaic loss of Chromosome Y in aged human microglia. Genome Res.32, 1795–1807 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Barretina, J. et al. The Cancer Cell Line Encyclopedia enables predictive modelling of anticancer drug sensitivity. Nature483, 603–607 (2012). [DOI] [PMC free article] [PubMed]
  • 23.Gertych, A. et al. Human Y chromosome pan-organ mapping reveals progressive mosaic loss from normal to cancer. JCI Insight. 11, e201996 10.1172/jci.insight.201996 (2026). [DOI] [PMC free article] [PubMed]
  • 24.Skaletsky, H. et al. The male-specific region of the human Y chromosome is a mosaic of discrete sequence classes. Nature15, 550 (2003). [DOI] [PubMed]
  • 25.Li, J. et al. Starbase v2.0: decoding mirna-cerna, mirna-ncrna and protein-rna interaction networks from large-scale clip-seq data. Nucleic Acids Res.42, D92–D97 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Shabalin, A. A. Matrix eQTL: ultra fast eQTL analysis via large matrix operations. Bioinformatics28, 1353–1358 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Forsberg, L. Loss of chromosome Y (LOY) in blood cells is associated with increased risk for disease and mortality in men. Hum. Genet. (2017). [DOI] [PMC free article] [PubMed]
  • 28.Machiela, M. et al. Mosaic chromosome Y loss and testicular germ cell tumor risk. J. Hum. Genet.136, 657–663 (2017). [DOI] [PMC free article] [PubMed]
  • 29.Savocco, J. & Piazza, A. Recombination-mediated genome rearrangements. Curr. Opin. Genet. Dev.62, 637–640 (2021). [DOI] [PubMed]
  • 30.Rhie, A. et al. The complete sequence of a human Y chromosome. Nature71, 63–71 (2023). [DOI] [PMC free article] [PubMed]
  • 31.Hänzelmann, S., Castelo, R. & Guinney, J. Gsva: gene set variation analysis for microarray and RNA-seq data. BMC Bioinforma.14, 7 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Kassambara, A., Kosinski, M. & Biecek, P. survminer: Drawing Survival Curves Using ‘ggplot2’ (R Package Version 0.4.3, 2018).
  • 33.Stodolna, A. et al. Clinical-grade whole-genome sequencing and 3’transcriptome analysis of colorectal cancer patients. Genome Med. (2021). [DOI] [PMC free article] [PubMed]
  • 34.Brooks, L. et al. A map of human genome variation from population-scale sequencing. Nature13, 33 (2010). [DOI] [PMC free article] [PubMed]
  • 35.Li, H. et al. The sequence alignment/map format and SAMtools. Bioinformatics467, 1061–1073 (2009). [DOI] [PMC free article] [PubMed]
  • 36.Quinlan, A. & Hall, I. BEDTools: a flexible suite of utilities for comparing genomic features. Bioinformatics25, 2078–2079 (2010). [DOI] [PMC free article] [PubMed]
  • 37.Dujif, P. H., Schultz, N. & Benezra, R. Cancer cells preferentially lose small chromosomes. Int. J. Cancer26, 841–842 (2012). [DOI] [PMC free article] [PubMed]
  • 38.Loh, P.-R., Genovese, G. & McCarroll, S. Monogenic and polygenic inheritance become instruments for clonal selection. Nature132, 2316–2326 (2020). [DOI] [PMC free article] [PubMed]
  • 39.Maan, A. et al. The Y Chromosome: a blueprint for men’s health? Eur. J. Hum. Genet.584, 136–141 (2017). [DOI] [PMC free article] [PubMed]
  • 40.Gyorffy, B. Integrated analysis of public datasets for the discovery and validation of survival-associated genes in solid tumors. Innovation25, 1181–1188 (2024). [DOI] [PMC free article] [PubMed]
  • 41.Wickham, H. et al. Welcome to the Tidyverse. J. Open Source Softw.5, 100625 (2019).
  • 42.Wickham, H. ggplot2: Elegant Graphics for Data Analysis (Springer, 2009).
  • 43.Durinck, S. et al. BioMart and Bioconductor: a powerful link between biological databases and microarray data analysis. Bioinformatics (2005). [DOI] [PubMed]
  • 44.Nguyen, T. et al. LOY landscape analysis pipeline. Zenodo 10.5281/zenodo.21461574 (2026). [DOI]
  • 45.Love, M., Huber, W. & Anders, S. Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2. Genome Biol. (2014). [DOI] [PMC free article] [PubMed]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Reporting Summary (8.1MB, pdf)

Data Availability Statement

Publicly available datasets analyzed in this study include Cancer Cell Line Encyclopedia (CCLE) data obtained from the DepMap portal, colon cancer whole-genome sequencing (WGS) data from ref. 33, and healthy male samples randomly selected from the 1000 Genomes Project34. Survival analysis data were obtained from the Kaplan–Meier Plotter. Data generated by TEMPUS were used under a data use agreement and are not publicly available because of restrictions associated with that agreement. Processed datasets supporting the findings of this study are publicly and permanently available through Zenodo at44.

Custom analysis scripts used to generate and analyze the results reported in this study are publicly and permanently available through Zenodo at44. The archived Zenodo record corresponds to the specific version of the code used for this study and provides a persistent DOI for citation and reproducibility. Software versions and study-specific variables or parameters required to reproduce the analyses are documented in the “Methods” and accompanying repository documentation.


Articles from Communications Biology are provided here courtesy of Nature Publishing Group

RESOURCES