SUMMARY
T-lineage acute lymphoblastic leukemia (T-ALL) is a high-risk tumor1 that has eluded comprehensive genomic characterization, in part due to the high frequency of non-coding genomic alterations resulting in oncogene deregulation2,3. Here we report integrated analysis of genome and transcriptome sequencing of tumor and remission samples from over 1300 uniformly treated children with T-ALL, coupled with epigenomic and single cell analysis of malignant and normal T cell precursors. This identified 15 subtypes with distinct genomic drivers, gene expression, developmental state and outcome. Analysis of chromatin topology elucidated multiple mechanisms of enhancer deregulation that involve enhancers and genes in a subtype-specific fashion, demonstrating widespread involvement of the noncoding genome. We show that the immunophenotypically-described, high-risk entity of early T-cell precursor ALL is superseded by a broader category of “ETP-like” leukemia with variable immunophenotype and diverse genomic alterations of a core set of genes encoding regulators of hematopoietic stem cell development. Using multivariable outcome models, we show that genetic subtypes, driver and concomitant genetic alterations independently predict treatment failure and survival. These findings provide a roadmap for the classification, risk stratification and mechanistic understanding of this disease.
INTRODUCTION
The prognosis for patients with relapsed or refractory T-cell acute lymphoblastic leukemia (T-ALL) is dismal1. A critical need is the identification of patients at risk for recurrence allowing intervention with novel therapies. Genomic characterization of B cell progenitor ALL (B-ALL) and integration of genomic data into treatment approaches has been transformative4–6. In contrast, prior attempts to identify genetic aberrations that are reproducibly prognostic independent of treatment response in T-ALL have failed for multiple reasons. First, several small studies have suggested a significant percentage of biologically relevant alterations in T-ALL occur in non-coding regions of the genome2, yet few cases have been subjected to whole genome sequencing (WGS)3, limiting identification of key genomic drivers. Secondly, no study has been large enough to power the identification of prognostic genetic alterations independent of treatment response. Thirdly, the only widely used treatment stratification factor except for response to therapy is early T-cell precursor (ETP) ALL, a subset of high-risk leukemia defined by immunophenotype, rather than by biological or genomic features7,8. Notably, the impact of ETP-ALL on outcome remains poorly defined7,9. Finally, most T-ALL genomic profiling efforts excluded patients with refractory disease3, due to the lack of remission bone marrow (BM) or blood samples to serve as matched non-tumor controls. Collectively, these issues have prevented accurate patient risk stratification and a comprehensive understanding of the biologic basis of T-ALL.
To overcome these limitations, we embarked on comprehensive WGS, whole exome sequencing (WES) and whole transcriptome sequencing (WTS) of tumor and germline samples, on more than 1300 T-ALL cases treated on the Children’s Oncology Group (COG) AALL0434 trial (NCT00408005; Supplementary Table 1–2)10. This multidimensional approach aimed to uncover both coding and non-coding alterations, provide a comprehensive genomic classification of the T-ALL, and ascertain predictors of relapse and treatment failure (Extended Data Fig.1a).
RESULTS
Fifteen genomic subtypes of T-ALL
We used the uniform manifold approximation and projection (UMAP) technique11 and Leiden algorithm12 to project and cluster WTS gene expression data in two dimensions (Fig.1a, Extended Data Fig.1b–c, Supplementary Table 3–6), and examined associations of clusters with sequence and structural DNA variants (SV). Integrated analysis delineated 15 subtypes with distinct drivers and patterns of oncogene expression (Extended Data Fig.1b–e). Several of these subtypes were previously recognized3,5, including those with deregulation of the TAL1/TAL2/LMO1/LMO2/LYL1 transcriptional regulators, or deregulation of TLX1, TLX3, NKX2-1 or HOXA9. Putative drivers were identified in 95.1% (1245 of 1309) of cases, 59% (777 of 1309) of which were in non-coding regions of the genome (Fig.1b–c, Supplementary Table 3). WGS was required to identify drivers for 29% (378) cases, particularly for detecting inversions, translocations, and enhancer single nucleotide variants or small insertions/deletions (SNV/indels) in non-coding regions. Subtype-defining drivers were not identified in 4.9% of cases, 39% of which had low tumor purity clustered in a group enriched for myeloid signatures (“TME-enriched group”, Extended Data Fig.1f).
Figure 1. Classifying drivers and gene expression define 15 T-ALL subtypes.

a, UMAP scatterplot depicting T-ALL subtypes and ETP status: Each subtype indicated by color and immunophenotype by shape; b, Bar plot illustrating percentages of classifying driver alterations and proportion of coding/intergenic alteration per driver with “Unknown” cases Shown in blue; “Rare” category includes all drivers with <0.8% frequency; c, Bar plot of driver alteration type percentages: differentiating coding and non-coding/intergenic alterations; d, Gene set score of subtype gene expression signatures in normal BM/thymic scRNA data: the dot plot depict signature detection percentage (dot size) and mean gene set score by normal cell type (color).
We analyzed the differentiation stage of each T-ALL subtype by projecting gene expression signatures onto single cell RNA (scRNA) and assay for transposase-accessible chromatin (ATAC) sequencing data of normal thymi (N=3) and BM samples (N=5; Fig.1d, Extended Data Fig.1g). The 15 T-ALL subtypes spanned a continuum of normal T-cell maturation, ranging from immature HSC or progenitor cells (HSPC), common lymphoid progenitors (CLP), lympho-myeloid primed progenitors (LMPP) or ETP for the BCL11B and ETP-like subtypes; pro/pre-T for the KMT2A, TLX3 and MLLT10 subtypes; cycling double positive (DP) for HOXA9 TCR, TLX1, NKX2-1 and TAL1 DP-like subtypes; TCRA expressing single positive αβ/mature αβ for TAL1 αβ-like; and γδ/effector T for LMO2 γδ-like that was also associated with TCRγδ rearrangements (Extended Data Fig.1h). Enrichment for myeloid signatures were also observed, including dendritic cells (DC) for the SPI1 subtype, granulocyte/monocyte progenitors (GMP) for NKX2-5, and megakaryocytic/erythroid precursor (MEP) for the STAG2/LMO2 subtype.
We identified additional genomic alterations for previously recognized subtypes (Fig.2). For example, deregulation of TLX3 in the TLX3 subtype is commonly due to hijacking of the BCL11B (ThymoD) enhancer13, but we identified multiple additional enhancers that were rearranged (R) and deregulated TLX3, including the T-cell receptor (TCR) β locus, CDK6, the NOTCH1-driven MYC enhancer (N-Me)14,15 and SATB1 enhancers. TLX1 activation commonly arises from rearrangement to the TCRβ or δ loci, but we identified a diverse range of intergenic losses, translocations, and inversions that resulted in TLX1 deregulation. Similarly, NKX2-1 activation was achieved by TCRδ loci rearrangements, along with recurrent chromosome 14 chromothripsis, intergenic losses, rare enhancer hijacking events, and enhancer gains. The NKX2-5 subtype harbored rearrangements of NKX2-5 to TCRβ/δ loci or BCL11B enhancer hijacking16. A subtype of cases with HOXA9 deregulation was characterized by HOXA9 hijacking the TCRβ locus enhancers17. SPI1 fusions with TCF7/STMN118, and YWHAE were hallmarks of the SPI1 subtype. The BCL11B-activated subtype harbored rearrangements juxtaposing hematopoietic stem cell (HSC) enhancers to BCL11B, previously identified in lineage ambiguous leukemia and ETP-ALL19,20. STAG2/LMO2 T-ALL most commonly harbor a LMO2::STAG221 rearrangement that activates LMO2 and inactivates the cohesin gene STAG2.
Figure 2: Oncoprint of classifying drivers.

Drivers are grouped by subtype and ordered by mutual exclusivity, with dark red row labels denoting non-coding drivers. Above the oncoprint, oncogene and differentiation markers gene expression Z-scores, TCR rearrangements and ETP status are shown.
Several subtypes were divided into subgroups based on genomic and clinical features. Cases with deregulation of the TAL1/TAL2/LMO1/LMO2/LYL1 core transcriptional circuitry could be divided into two subtypes previously termed TAL1-RA and -RB17. Considering our tumor and normal progenitor genomic data, we term these TAL1 αβ-like and TAL1 double positive (DP)-like subtypes. The TAL1 DP-like subtype exhibits higher expression of RAG1/2, CD4 and CD8, whereas the TAL1 αβ-like subtype is characterized by higher expression of the TCR alpha constant (TRAC) gene and TCRα/β rearrangements (Fig.2, Extended Data Fig.1h). Known drivers in these subtypes are STIL::TAL1, rearrangements of TAL1, LMO1, LMO2 and LYL1 to TCRδ/β enhancers, and non-coding sequence mutations creating a TAL1 neoenhancer2. Here we identified multiple additional mechanisms of oncogene deregulation, including copy number variation (CNV) gains and SNV/indel mutations generating neoenhancers for TAL1, LMO2, and LYL1, intergenic inversions and deletions that result in enhancer hijacking-mediated deregulation of TAL1 and LMO2.
Two additional subtypes were defined: The ETP-like subtype was enriched for cases of ETP immunophenotype and diverse genomic driver alterations, and the LMO2 gamma delta (γδ)-like subtype that has diverse alterations, including LMO2 activation from BCL11B enhancer hijacking, FOXP1 rearrangement, rearrangement of MYC to TCRδ, and enhancer SNV/indels activating LMO2 (Extended Data Fig.1h–j, Supplementary Table 3,7–8).
We observed associations between subtypes and clinical features. Patients within STAG2/LMO2, NKX2-5 and SPI1 T-ALL were younger at diagnosis and the BCL11B and HOXA9 TCR subtypes included a higher proportion of older patients (Extended Data Fig.1k–l). Sex distribution varied across subtypes, with a higher female proportion in the STAG2/LMO2, ETP-like, HOXA9 TCR, LMO2 γδ-like and NKX2-5 subtypes, and a higher male proportion in the TAL1 subtypes (Extended Data Fig.1m). Collectively, our findings reveal that T-ALL subtypes are primarily delineated by non-coding driver alterations and the stage of cell differentiation.
Detection of genetic alterations
Using DnDScv, Gistic222,23 and GRIN224, we identified 163 significantly mutated genes (q value<0.05) and 46 broad CNVs (Extended Data Fig.2a, Supplementary Table 9–25). Many of these recurrent alterations showed subtype specificity (Extended Data Fig.2b). Twenty-one genes had recurrent alterations outside of coding regions that were only detected by GRIN2. Alterations of CDKN2A (71% of cases) and NOTCH1 (69%) were most common. PHF6, PTEN, LEF1, MYB, MYC, and RUNX1 each had at least 9 mechanisms of deregulation (Extended Data Fig.2a). We identified recurrent sequence mutations for 16 previously unreported coding genes, including putative loss of function stop/frameshift mutations for CUL1 (1.1%), PSIP1 (0.8%), NOL4L (0.8%), KMT2E (0.8%), KAT6A25 (0.8%), DDX39B (0.7%) and MYL1 (0.6%) and mutation hotspots for CD99 (0.8%), RBMX (1.1%), LCK (0.8%) and E2F1 (0.5%).
Oncogene-activating enhancer alterations
Enhancer-mediated oncogene activation was observed in 70.5% of cases, and WGS was required to detect these in 43.9% of the cohort. These included translocations, inversions, chromothripsis and deletions juxtaposing 12 enhancers to 62 oncogenes observed in 50.8% of cases, and sequence alterations generating neoenhancers for 8 genes in 34.8% of cases (Fig.3a).
Figure 3: Diverse mechanisms of oncogene and enhancer driver alterations in T-ALL.

a, Bar plot of non-coding inter-/intragenic alterations. Right: recurrently targeted genes (frequency in brackets, >0.5% shown) with monoallelic expression −log10 q value (one-tailed Fisher’s exact test); b, TAL1 enhancer alterations, CD34+ HSPC and DP T-cell H3K27ac and CTCF coverage; c, Diagnostic (D) and germline (G) WGS, H3K27ac, Nucleosome-free and ATAC-seq nucleosome cut fragments (ATAC-free, ATAC-nuc) coverage for the sample exhibiting TAL1 enhancer gain; d, H3K27ac HiChIP, WGS and ATAC-seq in PAVHBC with RAG2::LMO2 intergenic loss. Arcs represent interaction strength at 5kb resolution. Left: Boxplot of LMO2 expression and P value for altered and wild type (wt) samples (two-tailed Student t-test), PAVHBC in cyan. The box includes the median, hinges mark the 25th and 75th percentiles, and whiskers extend 1.5 times the interquartile range; Right: heatmaps showing double positive (DP) and CD34+ HSPCs H3K27ac HiChIP interactions with color scale bar showing interaction strength at 5kb resolution, with H3K27ac coverage below. Lesion breakpoints are shown as black lines; hijacked enhancers color coded. The intensity of each arc/heatmap represents contacts at 5kb resolution between a pair of loci. Maximum intensity is indicated in the color scale; e, Same as d for SOX4::HOXA13 enhancer hijacking, by chr7-chr6 and chr7-chr17 translocations; f, Splice junction reads and coverage for NOTCH1 exon 28-29 intronic SNV are shown for mutated patient sample (red) and wt controls; g, AlphaFold2 models showing the structural difference between the wt (grey) and NOTCH1 intronic SNV mutant (blue). The mutation results in a 43-residue insertion (magenta), as compared to wt connector (green), between the heterodimerization (HD) and the transmembrane (TM) domains; h, Relative luciferase activity for NOTCH1 intronic SNV, HD/PEST domain mutations and controls are shown (N=3 each), with black line denoting the median and two-tailed Student t-test P values shown on top. Assay was repeated three times with consistent results.
We used ATAC-seq and histone 3 lysine acetylation (H3K27ac) HiChIP in 19 T-ALL samples encompassing 19 different alterations in 6 subtypes to examine chromatin accessibility and interactions putative enhancers with oncogenes. Fifty genes, 23 of which were significantly mutated (q value<0.05) hijacked TCR enhancers; events deregulating HOXA13, HOXA9 and LMO3 were verified by HiChIP (Extended Data Fig.3a–c, Supplementary Table 20). The BCL11B enhancer was hijacked by 16 genes (6 recurrent; q value<0.05); LMO2, HOXA13, NKX2-1, HOXB13 were validated by HiChIP (Extended Data Fig.3d–g). Several other enhancers were hijacked by specific genes, such as TLX1 hijacking of the LINC00592 enhancer via intergenic loss (Extended Data Fig.3h); NKX2-1 hijacking of the NFKBIA enhancer (Extended Data Fig.3i); LMO2 hijacking of MIR2117HG enhancer (Extended Data Fig.3j), TAL1 inversion to the DHX9 enhancer (Extended Data Fig.3k); intergenic 11p deletions resulting in LMO2 deregulation driven by hijacking of RAG2, CAPRIN1 and CELF1 enhancers (Extended Data Fig.3l–m) and ARID1B enhancer hijacking by BCL11B20. By contrast, some enhancers were hijacked by many different oncogenes: IGH (5 oncogenes), MIR181A1HG (4, including MIR181A1HG::HOXA13 and MIR181A1HG::LMO2, Extended Data Fig.3n–o), SATB1, N-Me14 and CDK6 (3 each; Fig.3a).
Non-coding alterations also deregulated 15 genes that were not initiating/classifying drivers in 14.6% of cases. These included MYC deregulation by the N-Me enhancer14,15 (8.3%) and deletion of the RUNX1 enhancer (1.2%; Extended Data Fig.4a–b). A recurrent deletion (1.2%) within FTO removed a putative regulator of the downstream IRX3 and IRX5 loci, resulting in their deregulation (Extended Data Fig.4c). Recurrent intergenic deletion between ZNF219 and the HNRNPC promoter (1.2%) was linked to increased monoallelic expression of ZNF219, in tandem with deletions at the TCRδ locus positioned 520 kB downstream (Extended Data Fig.4d–e).
Mutational generation of recurrent neoenhancers was frequent in TAL1 DP/αβ-like ALL (24.7% of cases; TAL1, 4 regions; LMO2, 4 regions; LMO1, 1 region) and also observed in NKX2-1 (N=2) and LYL1 (N=5). Unlike other mechanisms that activate oncogenes, certain LMO2 deregulation mechanisms do not lead to increased expression (Extended Data Fig.4f); small gains in intron 1 (median size 77 bp) and previously reported SNV/indels26 were associated with LMO2 neomorphic promoter generation and non-canonical LMO2 isoform expression (Extended Data Fig.4g–h). Another mutation hotspot was found 1.8kb upstream of LMO2. ATAC-seq and HiChIP revealed that a 6bp deletion at this enhancer resulted in increased H3K27 acetylation compared to other LMO2 alterations, open chromatin, and heightened expression, consistent with the generation of a neoenhancer (Extended Data Fig.4g–h). In addition, we detected recurrent TAL1 enhancer gains (median size 133bp) 28kb downstream of TAL1, enhancer SNV/indels 20kb downstream of TAL1, SNV/indels in the first intron of TAL1 and the enhancer Indel upstream of TAL12 all of which were active enhancers in CD34+ and not DP cells (Fig.3b). We validated TAL1 enhancer gains using ATAC-seq, HiChIP and isoform sequencing (isoseq; Fig.3c, Extended Data Fig.4i–j).
Several enhancer hijacking events were associated with developmental stage. Intergenic inversion of TAL1 to the CD1 locus resulted in hijacking of CD1E enhancers that are active in normal DP cells and thymic CD34+CD1a+ cells, but not BM CD34+ cells (Extended Data Fig.5a–b). Similarly, Intergenic deletions between LMO2 and RAG2 were observed in the TAL1 DP-like subtype, and the RAG2 enhancer that drives LMO2 deregulation was highly active in normal DP and not in CD34+ cells (Fig.3d, Extended Data Fig.5c). Conversely, a HOXA13-deregulated case in the ETP-like subtype hijacked SOX4 enhancers located within the CASC15 locus that are preferentially active in normal BM/thymic CD34+ cells as compared to DP cells, and frequently hijacked MIR181A1 enhancer active in γ/δ or CD34+ cells (Fig.3e, Extended Data Fig.5d–e). Breakpoint analysis of SVs linked to enhancer hijacking revealed that TCR and various other enhancer hijacking SVs, are likely facilitated by RAG activity. Conversely, many enhancer hijacking events in the ETP-like and BCL11B-activated subtypes occurred independently of RAG activity, illustrating dependency between subtype differentiation stage and RAG-mediated SVs (Extended Data Fig.5f, Supplementary Table 26–27). Together, these data emphasize the key role of enhancer-mediated activation of oncogenes and co-occurring alterations in a subtype and developmental stage-associated manner.
Intragenic and intronic SVs/SNVs
We observed non-coding and intragenic events in 50 genes that we demonstrated, or are predicted, to have diverse mechanisms of gene activation and perturbation. For example, NOTCH1 frequently harbors coding sequence mutations that activate NOTCH1 signaling. Here, we also identified intronic SNVs in 1.6% of cases. Using RT-PCR and long read WTS (isoseq), we showed that these SNVs generate a non-canonical splice acceptor in intron 28 that results in proximal extension of the coding sequence of exon 28 and predicted27 expansion of the disordered region between the transmembrane and HD domains (Fig.3f–g, Extended Data Fig.6a–b). This was reminiscent of NOTCH1 tandem duplications28 that disrupt signaling or alter cleavage of the extracellular domain (Extended Data Fig.6c). Accordingly, the intronic SNVs drove increased NOTCH1 activation compared to common coding mutations (Fig.3h). Recurrent NOTCH1 exon 3-27 and 16-27 deletions disrupt NOTCH1 extracellular domains or splicing likely resulting in the constitutive activation of the NOTCH1 intracellular domain (Extended Data Fig.6d–e).
Intragenic SVs could simultaneously alter the structure, isoform expression and activity of multiple genes simultaneously. For example, deletion of the IL7R transcription start site (TSS, 0.6% of cases) resulted in loss of the IL7R allele harboring the mutation, and also IL7R enhancer hijacking by PRLR that is likely to augment JAK/STAT signaling29–31 (Extended Data Fig.6f). Furthermore, in 2% of samples, CCND3 and TAF8 exhibited simultaneous TSS losses, with the deleted site affecting only one allele of the long isoform of CCND3 and one TAF8 allele (Extended Data Fig. 6g, Extended Data Fig. 7, see Supplementary Results).
TLX3, NKX2-1, TAL1/LMO2 genomics
We performed comparative analysis of driver mechanisms and co-lesions to understand the genetic composition of TLX3, NKX2-1 and TAL1/LMO2 subtypes facilitating a refined classification into genetic subgroups.
TLX3 was divided into two subgroups; TLX3-immature exhibited a higher incidence of WT1 (20%) alterations, 16q22.1/CTCF losses (30%), DNM2 (20%), and enrichment for kinase pathway alterations, including NUP214::ABL1 fusions (20%) and JAK/STAT alterations (60%) (Extended Data Fig.8a–g). In contrast, more differentiated TLX3 DP-like exhibited distinct features like 14q gains (20%), LEF1 (10%), and MYB alterations (20%) (Supplementary Table 28). These groups also differed in TCR rearrangements and ETP status confirmed by normal cell scRNA signature analysis (Extended Data Fig.8h).
NKX2-1 TCR exhibited TCRD hijacking events (90%) and RPL10 mutations (80%), while NKX2-1 Other was characterized by chromothripsis leading to BCL11B enhancer hijacking (40%), NFKBIA enhancer hijacking (20%), and TCRB::MYB rearrangements (20%) (Extended Data Fig.8i–m, Supplementary Table 29).
TAL1/LMO2 subtypes were divided into genetic subgroups based on drivers and co-lesions (See Supplementary Results, Extended data Fig.9a–d). TAL1 DP-like was classified into subgroups characterized by: RPL10 mutations (100%), frequent DDX3X (23%) and MYB alterations (48%); JAK/STAT alteration with frequent IL7R (53%) and STAT5B mutations (33%); LEF1 SV/Del (74%) or LYL1 SV (38%) genetic subgroup; and a diverse “Other” subgroup featuring a higher frequency of FBXW7 (38%), CCND3 (19%), TAL2 alterations (7%), and TCRD::MYC (3%) (Extended Data Fig.9e–f, Supplementary Table 30–31). TAL1 αβ-like was subdivided into NOTCH1 wild type with frequent PI3K pathway alterations (72%, PTEN alterations 68%); a group marked by 6q loss (100%); and an “Other” category with NOTCH1 (100%) and PI3K pathway alterations (70%) but lacking 6q loss. TAL1 subtypes and LMO2 γδ-like had distinct flow cytometry-based immunophenotypes and TCR rearrangements (Extended Data Fig.9f–l, Supplementary Table 8). Although driver genes were often shared between these subtypes, differences in co-lesions, hijacked enhancers and oncogene activation mechanisms suggests that the maturational arrest state is major determinant of the phenotype.
Characterization of ETP-like ALL
One of the few diagnostic features used to stratify risk in T-ALL is ETP immunophenotype (cytoplasmic CD3+, CD1a-, CD8-, CD5 dim/-, with expression of stem cell or myeloid antigens)7. “Near-ETP” T-ALL cases have similar immunophenotype except positivity for CD5. Prior genomic studies of ETP ALL have identified recurrent alterations of genes encoding regulators of hematopoietic development, kinase signaling and chromatin modification3,8, but have failed to identify unifying genomic alterations that distinguish such cases. Several observations from this study strongly indicate that this immunophenotypic classification is suboptimal. First, four subtypes were enriched for cases of ETP/near-ETP immunophenotype. Also, 94% of the cases belonging to the BCL11B-activated subtype were of ETP immunophenotype19,20 and harbored 13.6% ETP cases in the cohort. Most strikingly, a distinct group of cases included 70.9% of ETP and 41% of near-ETP cases in the cohort, but conversely, comparable proportions of each immunophenotypic group: 38.2% were ETP, 33.8% Near-ETP and 27.9% non-ETP. This “ETP-like” subtype had multiple recurrent driver alterations of genes with known or putative roles in HSC development: activating rearrangements of HOXA13 (18.7%) to TCR, BCL11B, MIR181A1HG, SATB1, CDK6 enhancers, cases with HOXA9/10/11 deregulation driven by rearrangements of MLLT10 (18.3%), KMT2A (11%), NUP214 (5.1%), NUP98 (3.4%); loss-of-function mutations of MED12 (14.4%); ZFP36L2 (8.9%) rearrangements, and alterations of ETV6 (7.2%; Fig.4a). KMT2A, MLLT10, and NUP98/214 rearrangements are also drivers of non-ETP ALL enriched subtypes. KMT2A fusions within the ETP-like subtype harbored mostly KMT2A::AFDN and other fusion partners, whereas the non-ETP KMT2A subtype exclusively had KMT2A::MLLT1 fusions (Extended Data Fig.10a). Moreover, each ETP-like driver had distinct patterns of concomitant alterations: 2q alterations in the ETV6 subgroup; RUNX1, JAK, SUZ12, ASXL1 mutations in the ZFP36L2 subgroup (Extended Data Fig.10b, Supplementary Table 32); ETV6, TP53, SATB1, SH2B3 alterations in the HOXA13 subgroup; PSIP1 mutations in cases with MLLT10 fusions; KAT6A and MBNL1 mutations in the KMT2A subgroup, and gains of chromosomes 8, 10 and 19, loss of 5q and mutation of IKZF1 and GATA3 in the MED12 subgroup. ETP-like cases with MLLT10/KMT2A/NUP98/NUP214 driver alterations had a higher frequency of CNVs and alterations of ETV6, GATA3, IKZF1, RUNX1 and RAS signaling, and fewer NOTCH pathway alterations than non-ETP-like cases with these drivers. Collectively, these results indicate that cell of origin, fusion partner and genomic co-alterations cooperate to determine gene expression patterns and T-ALL subtype.
Figure 4: Genomic classification of ETP/ETP-like ALL.

a, UMAP plot displaying the distribution of ETP-like genetic subgroups along with HOXA9, MLLT10, and KMT2A groups, separated from the ETP-like subtype by the dotted curved line. Right: HOXA9 and HOXA13 gene expression Z-scores, with dotted curved lines separating the samples with high HOXA9 or HOXA13 expression; b, Heatmap of select differentially expressed genes between MED12wt and MED12KO LOUCY cells, with FDR shown on the left; c, H3K27ac HiChIP raw interactions and H3K27ac and CTCF coverage in normal CD34+ HSPCs showing different TADs for HOXA9 and HOXA13 with CTCF boundaries marked in yellow and topologically associated domains annotated by black dotted triangles; d, UMAP plot depicting the distribution of ETP status and TCR rearrangement status among the ETP-like samples; e, Dot plot visualizing the gene set detection percentage (dot size) and mean gene set enrichment by normal cell type (color) for ETP-like genetic subtype and MLLT10, KMT2A, HOXA9 subtypes in BM/thymic scRNA data. Signatures were generated by dividing ETP-like by ETP status, TCR rearrangement status, or genetic subtypes; f, Heatmap displaying the expression of immunophenotypic cell surface markers as positive cell percentages for ETP-like cases, and KMT2A, MLLT10, and HOXA9 subtypes. ETP status and TCR rearrangement status are indicated on top of the heatmap; g, Volcano plot representing differential flow cytometry markers between ETP-like cases compared to the rest of the samples (ETP-like upregulated markers log2 fold change positive and downregulated negative). Significantly altered markers (from two-tailed Wilcoxon rank-sum test) are labeled with their respective protein names and log2 fold change (0.25) and P value cutoffs (0.1) are denoted by dotted lines; h, Similar to g, differential markers between Non-ETP vs ETP cases within ETP-like subtype are shown; i, Similar to g, differential markers between ETP-like Non-ETP cases compared to all other Non-ETP cases.
The MED12 alterations observed in ETP-like cases were observed across the coding region of MED12, suggesting loss of function (Extended Data Fig.10c). To test this, we inactivated MED12 using genome editing in the LOUCY (SET::NUP214) cell line (Extended Data Fig.10d) with immunophenotypic similarity to ETP ALL, and observed upregulation of histone deacetylase pathway gene expression (Extended Data Fig.10e). Intersection of these data with the gene expression profile of MED12 altered ETP-like cases showed common reduced expression of the T-cell differentiation markers CD5 and CD28, and increased expression of the stem cell markers LMO2 and HHEX (Fig.4b, Extended Data Fig.10f–g, Supplementary Table 33–35), indicating loss of MED12 function directly contributes to the immaturity characteristic of ETP-like ALL.
Integrated genomic analysis also elucidated mechanisms driving differential deregulation of specific HOXA genes in ETP-like ALL. Specifically, enhancer hijacking alterations driving HOXA9, but not HOXA13 deregulation such as TCRB::HOXA9 showed that rearrangement breakpoints were always located between two CTCF peaks that demarcate a topologically associating domain (TAD) boundary between the HOXA9 and HOXA13 loci in CD34+ HSPC cells (Fig.4c, Extended Data Fig.10h). By contrast, all breakpoints of rearrangements deregulating HOXA13 were confined to the HOXA13 TAD, thus constraining activation of HOXA9.
Although 28% of cases in the ETP-like subtype did not fulfil the immunophenotypic criteria for ETP/near-ETP ALL, they exhibited immunophenotypic trends (lower expression of T-cell and myeloid/stem cell markers) and commonalities including absence of TCR rearrangements, similar maturational stage, genomic drivers, and outcome (Fig.4d–i, Extended Data Fig.10h–j, see Supplementary Results). Finally, near-ETP cases were more dispersed across subtypes, and enriched in the ETP-like, TLX3 Immature and TAL1 αβ-like subtypes (Extended Data Fig.1i–k). Thus, ETP-like ALL is a subtype of ALL with distinct, heterogeneous drivers, variable diagnostic immunophenotype and a likely HSPC origin. Our data indicates that genomic classification should replace immunophenotypic classification.
Outcome analysis of genomic risk factors
We examined genomic features associated with clinical outcome (Supplementary Table 36–45). Positivity for residual disease (RD; MRD≥0.01%9) was particularly common in the ETP-like and LMO2 γδ-like subtypes (Fig.5a, Extended Data Fig.11a). We examined associations between subtypes, genetic drivers, co-lesions, broad CNV changes, and altered pathways with RD (Extended Data Fig.11b). ETP-like drivers and co-lesions (SH2B3, ETV6, NRAS, WT1) were associated with higher RD risk, and TAL1 subtype-related features (LEF1, USP7, PI3K, CCND3) with lower RD risk. JAK-STAT and Ras signaling pathway alterations were associated with higher RD risk, whereas NOTCH, ribosome, and PI3K pathway with lower RD risk.
Figure 5: T-ALL genomic risk factors and clinical outcomes.

a, Residual disease category proportions per subtype; b, Forest plot displaying hazard ratios as box (HR) and error bars indicate 95% confidence intervals (CI) comparing cases with subtype or genetic subtype present (N) vs. absent, considering EFS and MRD as outcomes. Significant associations are highlighted in red for EFS (P<0.1, two-tailed, Firth-penalized Cox regression that also adjusted for MRD (MRD.adj)) and for MRD (P<0.1, two-tailed, Firth-penalized logistic regression). P values were not adjusted for multiple comparisons; c, Similar to b, Forest plot of select variants significantly associated with outcomes (EFS, DFS, OS, two-tailed Firth-penalized Cox regression); d, Cumulative incidence plot of SPI1 fusions and second cancer, with contingency table and malignant type shown below and P value from two-tailed Gray’s Test; e, Clonal evolution of T-ALL to Langerhans cell histiocytosis (LCH) for case TALL023, with T-ALL driver alterations shown for each clone; f, Volcano plot depicting differential flow cytometry markers between SPI1 subtype cases and other samples, with significantly altered markers (from two-tailed Wilcoxon rank-sum test) labeled by protein names (SPI1 upregulated markers log2 fold change positive and downregulated negative, and log2 fold change (0.25) and P value (0.1) cutoffs are denoted by dotted lines); g, Dot plot showing immunophenotype markers from f in bone marrow and thymic dendritic cells.
In analyses of event-free (EFS), disease-free (DFS) and overall survival (OS), the SPI1 and LMO2 γδ-like subtypes had dismal outcomes, and the NKX2-5, as well as ETP-like KMT2A, MLLT10, HOXA13 genetic subtypes had adverse outcome (Fig.5b). The non-ETP-like KMT2A subtype had higher MRD but a very favorable prognosis, unlike KMT2A cases within the ETP-like subtype. Similarly, non-ETP-like MLLT10 cases had a better prognosis compared to ETP-like MLLT10 cases. Analogous patterns emerged for TLX3, where TLX3 DP-like had favorable prognosis and TLX3 Immature exhibited worse outcomes. The ZFP36L2 subgroup had higher rates of RD, yet a favorable outcome, highlighting that early poor disease response alone should not be the sole factor driving decisions such as HSC transplant. Notably, heterogeneity was evident within TAL1 genetic subgroups, as TAL1 DP-like subgroups (‘LEF1/LYL1’ and ‘Other’) were associated with inferior EFS and DFS, and TAL1 αβ-like subgroups (‘NOTCH wt’ and ‘Other’) were associated with inferior OS, whereas the ‘RPL10’ subgroup had an excellent outcome (Extended data Fig.11c).
We observed that specific types of genetic alterations of individual genes modulated outcome (Fig.5c, Extended Data Fig.12a). Most NOTCH1 variants had favorable prognosis (Extended Data Fig.12b), but intronic SNV and intragenic losses were associated with inferior OS and EFS (Fig.5c). Cases with TCR::MYC rearrangements had inferior DFS, whereas cases with MYC enhancer gains had favorable DFS. Cases with PTEN deletions had markedly worse outcomes compared to other PI3K pathway alterations. Within the TAL1 subtype, LYL1 TCR and LMO2 intergenic losses leading to RAG/CAPRIN1 enhancer hijacking were associated with worse outcomes, in contrast to 6q deletion and RPL10 mutations (Fig.5c).
Through competing risk (CR) models, we identified risk factors such as LMO2 intergenic loss, TCR::MYC, PTEN deletions and NOTCH1 intragenic deletions associated with relapse (Extended Data Fig.12c–g). TAL1 upstream indel were associated with higher relapse risk compared to other TAL1-deregulating mechanisms and differential relapse risks across TAL1 subtypes (Extended Data Fig.12h–i). In a recent study25, TAL1 upstream indels and TCR::LMO1 were associated with induction failure; however, we found no association in our cohort (Extended Data Fig.12j).
Second malignancies were uncommon but enriched in the SPI1 subtype, in which 4 of 11 cases developed fatal histiocytosis or myeloid sarcoma within a year of diagnosis (Fig.5d). Tissue from the secondary malignancies in this cohort was not available, but we identified identify two other children with SPI1-rearranged T-ALL who developed secondary histiocytic lesions and had available tissue. Case TALL-HS-1 had TCF7::SPI1 present in the diagnostic T-ALL sample and in the subsequent histiocytic sarcoma (T-ALL-LCH-2; Supplementary Table 46). A second case, TALL023, had STMN1::SPI1 positive T-ALL and subsequent Langerhans cell histiocytosis (LCH)18,32,33. WGS and WTS of both samples confirmed that the initiating clone harbored STMN1::SPI1, CDKN2A deletion and KRAS mutation in both the diagnostic T-ALL and LCH (Fig.5e, Extended Data Fig.12k, Supplementary Table 47–48). The LCH sample had high expression of markers of LC2 type Langerhans cells that are driven by SPI134, including CD207 and HLA-D, and low expression of T-cell markers (Extended Data Fig.12l–m). The T-ALL samples expressed genes also exhibited by DCs (HLA-D, CD1a, CD38, CD45, CD7, CD5, and sCD3) and the SPI1 T-ALL signature showed enrichment of the signature of normal DC (Fig.1d, Fig.5f–g). These findings suggest that SPI1 fusion-driven malignancies arise in a progenitor with T and DC potential, and frequently transform into high-risk DC-related second malignancy.
Multivariable risk stratification models
We explored the utility of multiple multivariable modelling approaches incorporating clinical variables, treatment response and genomic features to predict outcome and risk stratify patients. Random Survival Forest (RSF) and Penalized Cox Regression (pCox) had the highest accuracy when each model was fitted using numeric MRD, clinical variables (sex, WBC count at diagnosis >2x105 cells/μl and central nervous system (CNS) status), and subtype/variant level genomic features for the pCox model, and genetic subtype for the RSF model (Extended Data Fig.13a, Supplementary Table 49, see Supplementary Results).
The pCox model enables accurate risk stratification using a complete set of the most prognostic genomic and clinical features. The model incorporated three clinical features, five subtypes and 18 genomic alterations to stratify patients into four equally sized groups with five-year EFS ranging from 65 to 97% (Fig.6a,b, Supplementary Table 50). SPI1 subtype, ETP-like subtypes, PTEN deletions/loss, PIK3CD SNV/indels, and LMO2 intergenic loss were associated with worse outcomes, while KMT2A subtype, 6q loss, RPL10 and NOTCH1 SNV/indels were associated with favorable outcomes both in univariable and pCox models, highlighting their value as independent prognostic biomarkers (Fig.5b–c,6a).
Figure 6. Multivariable outcome models.

a, The Kaplan-Meier plots illustrate the risk score (divided into quartiles) derived from a penalized Cox regression multivariable model fitted using features from b; b, The oncoprint displays the genomic or clinical features selected by the penalized Cox regression model (N=1299, excluding samples with missing data). Genomic data are stratified according to risk score quartiles, while day 29 MRD (percentage of leukemic cells after induction) is shown as bar plots. Binary genomic features are categorized by data type, with corresponding model coefficients presented on the left ordered by association (>0 adverse, <0 favorable). EFS events are depicted at the bottom of the figure to illustrate the association between risk score quantiles and EFS events; c, This illustration depicts the fitted survival tree (N=1287, excluding samples with missing data), initially categorized into four groups based on subtype or genetic subtype (superscript), and subsequently subdivided according to day 29 MRD (≥0.1% for MRD positive, <0.1% negative). Bottom: Kaplan-Meier curves show the survival relationship within each of these divisions.
The RSF model was simplified into four node survival tree (ST) model that enables clinical implementation using focused selection of features. The ST model was able to risk stratify patients into eight groups with 5-year EFS ranging from 45-98% (Fig.6c, Extended Data Fig.13b, Supplementary Results). Several features, including the ETP-like drivers KMT2A, MLLT10, NUP98 and rare drivers, and the SPI1, LMO2 γδ-like and NKX2-5 subgroups had poor outcome (5-year EFS <60%) regardless of MRD response. These patients should be considered for HSC transplant or novel immunotherapies as outcomes are poor despite intensive multi-agent chemotherapy. In contrast, several other features, including ETP-like with ZFP36L2 alterations, TLX3 DP-like, TAL1 DP-like RPL10, NKX2-1, TLX1, KMT2A, HOXA9 TCR had very favorable outcomes (5-year EFS >98%) if day 29 MRD was <0.1%. This large group of patients (N=260; ~20% of cohort) may benefit from a reduction in intensity of chemotherapy.
Both the pCox model and ST proved effective in accurately predicting outcomes for individuals within both the ETP-like and groups classified by immunophenotype (Fig.6c, Extended Data Fig.13c–j). In contrast, immunophenotype was not prognostic in the ETP-like subtype (Extended Data Fig.13g) and BCL11B subtype had favorable outcome, even though immunophenotype was ETP (Fig.5b). These findings underscore the necessity of employing genomic-based multivariable prognostic classification.
DISCUSSION
We classified T-ALL into 15 subtypes with distinct gene expression and genomic drivers, including previously undefined ETP-like and LMO2 γδ-like subtypes. We also refined the classification of known subtypes such as TAL1 and TLX3, and have shown that driver lesions, concomitant genomic alterations, and cell of origin interact to determine genomic subtype, clinical and biologic features. Notably, approximately 60% of identified leukemia drivers involved non-coding regions, requiring WGS for identification in 28% of cases. This highlights a substantial shift in our understanding of T-ALL biology, with a predominant genomic landscape driven by non-coding alterations. Notably, enhancer hijacking emerged as a prevalent method for oncogene activation, impacting 50% of cases and 61 oncogenes. Thus, stage-specific deregulation of lineage-appropriate transcription factors or hijacking of T lineage enhancers, such as TCR (28%) and BCL11B (17%) for both lineage-appropriate or inappropriate TFs, is a hallmark of most T-ALL subtypes. By contrast, chromosomal rearrangements leading to the generation of chimeric fusion oncoproteins (48% B-ALL, 17% T-ALL), and deleterious alterations of lineage-appropriate TFs (8%) are more common in B-ALL5. In contrast to T-ALL, B-ALL drivers include lineage inappropriate TFs with neomorphic activity, such as DUX4, MEF2D, and ZNF384.
In total, we identified recurrent pathogenic or likely pathogenic genetic alterations in 163 genes and uncovered numerous recurrently altered genes and variants for known genes not previously documented in T-ALL. Many potentially targetable genes with biologic pathway relevance were altered and enriched in various subtypes (Extended Data Fig.13k). For instance, TLX3 Immature T-ALL was associated with various co-lesions deregulating kinase signaling, including FLT3 ITD and other JAK/STAT alterations and NUP214::ABL1 fusions. TAL1 αβ-like and NKX2-5 T-ALL were enriched for PI3K pathway alterations, while TAL1 DP-like JAK subgroup and HOXA9 TCR were strongly associated with JAK/STAT alterations. These subgroups could be targeted through FDA-approved tyrosine kinase inhibitors such as gilteritinib (FLT3i), ruxolitinib (JAK/STATi), dasatinib (ABL1i), and idelalisib (PI3Ki).
We developed two complementary multivariable models incorporating comprehensive or limited genomic features to assess the prognostic significance of genomic subtypes, altered genes, and dysregulated pathways. Future studies must examine the utility of these models in other cohorts, but these provide a framework for diagnostic implementation. T-ALL is typically classified according to immunophenotype; however, immunophenotypic classification has not been informative in risk stratification and our results show the importance of genomic characterization in accurate risk stratification in T-ALL. We showed that the ETP-like subtype exhibits poor early MRD response and inferior EFS, and is notably enriched with genetic drivers such as KMT2A-R and MLLT10-R. ETP-like and non-ETP-like KMT2A-R or MLLT10-R patients revealed a stark divergence in outcomes, underscoring the importance of separating these subgroups. We identified genetic alterations such as the BCL11B and ZFP36L2 ETP-like subtypes that associated with poor MRD response but good outcome. Conversely, other alterations such as SPI1-R subtype and PI3K alterations displayed favorable MRD response but poorer survival, highlighting that MRD alone is insufficient to risk stratify patients. Finally, we observed a significant link between the type of gene alterations and outcomes in T-ALL. Although mutations in the NOTCH pathway are typically associated with a favorable prognosis, our study revealed that patients with intronic SNVs or intragenic deletions in NOTCH1 exhibited inferior survival, and notably, these mutations showed no subtype specificity. Similar patterns emerged for other genes, including MYC and PTEN. In summary, through the largest comprehensive sequencing effort in T-ALL performed to date, our data elucidate insights into T-ALL disease biology and underscore the need for comprehensive risk stratification that considers incorporates genomic subtypes, variant type, as well as coding and non-coding alterations.
METHODS
Patient cohort, Inclusion and Ethics Statement
AALL0434 Clinical Trial and Samples Used for Genomic Analyses
AALL0434 (NCT04408005) was a Children’s Oncology Group (COG) phase 3 international clinical trial for patients with newly diagnosed T-cell acute lymphoblastic leukemia (T-ALL) and lymphoblastic lymphoma (T-LL) aged 1-30 years. Subjects were enrolled from January 22 2007 until July 25 2014 at 214 centers in the United States, Canada, Australia, New Zealand, and Switzerland. All subjects with T-ALL were required to enroll on a companion classification study that was used for sample banking and risk stratification, AALL03B1 (NCT00482352) from January 22 2007 until August 08 2010 or AALL08B1 (NCT01142427) from August 9, 2010, until July 25 2014. AALL0434, AALL03B1, and AALL08B1 were approved by the Pediatric Central Institutional Review Board (IRB), local IRBs at all participating centers, and NCI Cancer Evaluation and Therapeutic Program (CTEP). Written informed consent/assent was obtained from all study participants and/or their legally authorized representative in accordance with the Declaration of Helsinki. Genomic studies performed for this work were approved by COG, CTEP, and the local IRBs at the Children’s Hospital of Philadelphia and St Jude Children’s Research Hospital. Samples were decoded and assigned a unique study identifier (USI). Samples were banked at the COG biorepository at Nationwide Children’s Hospital in Chicago, IL.
Details on the chemotherapy treatment backbone, inclusion and exclusion criteria for the clinical trial, and results of the AALL0434 clinical trial have been previously published10,37. A total of 1562 eligible and evaluable subjects with T-ALL were enrolled on AALL0434. Of these, 1409 subjects consented to correlative research and had samples banked for genomic analyses. Diagnostic specimens were used for extracting somatic/tumor DNA and RNA and day 29/end of induction/remission samples were used for matched normal control DNA. Details on isolation of normal control DNA for subjects who did not attain remission at day 29/end of induction are detailed in Supplementary Methods.
Complete sequencing was performed for 1309 subjects, defined as WGS, WES, and WTS of tumor and WGS of matched normal sample. An additional 53 subjects had successful WTS sequencing without either WGS/WES of tumor or normal matched control and were included in transcriptome only analyses. A comparison of studied genomics cohort and whole cohort was performed to demonstrate that the sequenced cohort was representative of the overall trial cohort (Supplemental Table 2). This comparison included important clinical and demographic features, including self-reported sex, race, and ethnicity as well as outcomes between the eligible and evaluable T-ALL patients. T-ALL has an expected 3:1 male to female predominance and is more common in persons who self-identify as Black or African. Expected rates enrollment based on sex, race, and ethnicity matched actual rates of enrollment as detailed in the clinical trial publications10,37. Biospecimen reporting for improved study quality (BRISQ) criteria were followed 38. Details on biospecimen type, anatomical site, disease status of patients, clinical characteristics of patients, clinical diagnosis of patients, pathology diagnosis, collection mechanism, and type of stabilization, type of long-term preservation, storage temperature, storage duration, shipping temperature, and composition assessment are included in Methods, Supplemental Methods, and Supplemental Tables 1 and 2.
Subjects with newly diagnosed T-ALL were eligible based on either >25% leukemic blasts on bone marrow aspirate or by a complete blood count (CBC) documenting the presence of at least 1,000/μl circulating peripheral blasts. A bone marrow examination was required unless there was a medical contraindication to having the test. Bone marrow and/or peripheral blood samples were collected and banked at diagnosis and bone marrow and/or peripheral blood were also banked at the end of induction (after one month of chemotherapy). Bone marrow samples collected at diagnosis were prioritized for use for tumor genomic studies; however, peripheral blood samples were used if bone marrow samples were not available.
Central determination of immunophenotype and minimal residual disease (MRD) by flow cytometry
Immunophenotype and MRD of patient samples were performed centrally on AALL0434 by 8-9 color flow cytometry. Comprehensive immunophenotype including determination of ETP, near-ETP and non-ETP status was performed on diagnostic bone marrow or peripheral blood samples, as previously published7.
ETP ALL was defined by T-lymphoblasts that were CD8 negative and CD1a negative (<5% positive), weakly expressed CD5 (either <75% positive or median intensity more than 1 log less than mature T cells) with expression of one or more myeloid or stem cell markers (>25% positive) including CD13, CD33, CD34, CD117 or HLA-DR9. Near-ETP was defined by meeting the ETP-immunophenotype but having stronger CD5 expression. The remaining cases were defined as non-ETP. The panel of antibodies is included in Supplemental Table 7. Of note, the panel changed in 2008. ETP status was determined on 82.3% of subjects treated on AALL0434 (1256 of 1526) and 87.2% of subjects (1141 of 1309) who had complete sequencing performed. ETP ALL was described as an entity in 2009 and therefore subjects enrolled prior to its description were uncharacterized for ETP status by immunophenotype7,9. MRD was assessed by flow cytometry at Day 29/end of induction in all patients who remained on protocol therapy at that time point.
DNA/RNA isolation
Matched normal cell DNA isolation for all samples except those that underwent flow-sorting (N=24) was performed at Nationwide Children’s Hospital, using the Qiagen QiaAMP DNA kit (Mini kit or Maxi kit depending on number of cells in the sample). Matched normal control DNA for those that underwent flow sorting was performed at CHOP, using the Qiagen QiaAMP DNA kit (Mini kit or Micro kit depending on number of cells in the sample). DNA and RNA from Ficoll-enriched viably preserved tumor samples was extracted at the Fred Hutchinson Cancer Center using the Qiagen AllPrep Extraction Kit.
Sequencing
Ribosomal RNA (rRNA) reduction RNA-seq library preparation and sequencing:
The concentration and integrity of the total RNA was estimated by Ribogreen assay (Invitrogen), and Fragment Analyzer (Agilent), respectively. Approximately 500ng of total RNA from each sample was taken into library prep using the Illumina Stranded Total RNA Prep with Ribo-Zero Plus kit (Illumina) as per manufacturer’s recommended protocol. Final Library concentration was measured by Picogreen Assay (Invitrogen), and the library size was estimated by utilizing a DNA High Sense chip on a LabChip Gx (PerkinElmer). Accurate quantification of the final libraries for sequencing applications was determined using the qPCR-based KAPA Biosystems Library Quantification kit (Roche). 2x100 PE Sequencing was performed on an Illumina NovaSeq 6000 instrument (Illumina).
Whole Exome library preparation:
Approximately 500ng DNA from each sample was sheared on a Covaris focused-ultrasonicator (Covaris Inc, USA) with a target yield of 200bp fragment size. Following this the fragmented DNA was taken into standard library preparation protocol using KAPA HyperPrep Kits (Roche, USA) with slight modifications. Post-ligated material was individually barcoded with unique in-house primers and amplified PCR using KAPA HiFi HotStart Ready Mix (Roche, USA). The concentration of the libraries was then measured by Picogreen assay (Thermo, USA), and the average fragment size of the libraries was estimated by utilizing a DNA High Sense chip on a LabChip GX Touch Nucleic acid analyzer (PerkinElmer, USA), respectively. KAPA qPCR assay (Roche, USA) was then performed to assess the nanomolar amounts of ligated libraries.
Post Library prep, approximately 600ng library per sample was hybridized using the xGen Exome Hyb Panel v2 (Integrated DNA Technologies, USA) as per manufacturer’s recommendation. This probe set consists of 415,115 probes that span a 34 Mb target region (19,433 genes) of the human genome and 39 Mb of probe space. Post hybridized libraries were amplified through PCR and the concentration of the libraries was then measured by Picogreen assay and the average fragment size of the libraries was estimated by utilizing a DNA High Sense chip on a LabChip GX Touch Nucleic acid analyzer. KAPA qPCR assay was then performed to assess the nanomolar amounts of ligated libraries. Final libraries were pooled and then sequenced as 2x100bp Paired-end sequencing on the NovaSeq 6000 instrument using an S4 200 cycle flow cell.
Whole Genome library preparation and sequencing:
Approximately, 500ng DNA from each sample was sheared on a Covaris focused-ultrasonicator (Woburn, MA, USA) with a target yield of 500bp fragment size. Following this the fragmented DNA was taken into standard library preparation protocol using KAPA HyperPrep kit as per manufacturer’s recommendation. The concentrations of the libraries were assessed by Picogreen, and the average fragment size of the libraries was estimated by utilizing LabChip® GXII Touch (Caliper), respectively. Accurate quantification for sequencing applications was determined using the qPCR-based KAPA Biosystems Library Quantification kit (Roche). Paired End (PE) (150bp) sequencing was performed on an Illumina NovaSeq 6000 (Illumina).
Other sequencing:
Amplicon, HiChIP, ATACseq, and long read RNA sequencing are described in the supplementary methods (Supplementary Table 51).
Data analysis
See Supplementary Methods for data analysis and variant calling approaches (Supplementary Table 52).
Subtyping
UMAP
A series of filtering steps were applied to the features as follows: (1) Genes were required to exhibit expression of over 1 transcript per million (TPM) ≥ 5 samples. (2) Only genes located on chromosomes 1 to 22 and X were included. (3) Genes were filtered based on gene biotype, with the requirement that they be classified as protein_coding, lncRNA, TR_C_gene, or TR_V_gene. (4) the 300 most variable genes were selected solely from diagnostic samples with a blast percentage exceeding 70%, determined by the standard deviation. Subsequently, UMAP11 (Uniform Manifold Approximation and Projection) dimensionality reduction analysis was conducted on the entire cohort (N=1362) using the ‘uwot’ R package (v0.1.14), with parameters set to n. neighbors=15 and min.dist=0.1, with seed.use=1. The robustness of the resulting projection was evaluated by performing the UMAP projection across varying variable feature counts, ranging from 100 to 2000 genes.
Clustering
We applied community detection-based clustering using the same set of 300 most variable genes (Supplementary Table 5) employed in the UMAP analysis. The Leiden12 algorithm, implemented within the ‘igraph’ R package (v.1.3.5), was employed for this purpose. Initially, SNN (Shared Nearest Neighbor) graphs were constructed using the ‘scran’ v.1.26.2 ‘buildSNNGraph‘ function, with a parameter of k=7 and ‘set.seed(1)’ for reproducibility. We conducted clustering at two different resolutions, namely 0.1 and 0.5, to delineate both primary subtypes and finer subgroups. Additionally, we extended our clustering approach to include frequently altered oncogene expressions, specifically TAL1, TAL2, LMO1, LMO2, NKX2-1, NKX2-5, TLX1, TLX3, HOXA9, HOXA13 and LYL1. For this, we constructed the SNN graph with k=20 and set the resolution to 0.8.
Defining classifying drivers
We detected Leiden cluster-specific genetic lesions using a one-tailed Fisher’s exact test with an FDR threshold of <0.05 (Supplementary table 53). Subsequently, mutual exclusivity and co-occurrence analyses were performed for these lesions, employing the discover R package (v.0.9.4) with default parameters and a q value cutoff of 0.05 (Supplementary Tables 54–55). We filtered out co-occurring alterations that were likely co-lesions, except for TAL1/LMO1/LMO2/LYL1 alterations that were allowed to co-occur with each other but not with other subtype-defining driver alterations. Finally, a manual curation process included recurrent subtype-specific variants, as defined above, together with rare variants on a case-by-case basis. This process resulted in the compilation of the list of putative classifying drivers (Supplementary Table 3).
Defining subtypes using integrative analysis
The subtypes were defined by Leiden clusters at a resolution of 0.1, also having high concordance with the oncogene-based expression analysis (Extended Data Fig.1b). However, certain smaller subgroups, such as SPI1, LMO2 γδ-like, NKX2-5, and STAG2/LMO2, did not exhibit distinct boundaries within the low-resolution clusters. To address this, a combined approach involving clustering at a resolution of 0.5 and driver classification was used to delineate these subtypes. This approach was validated by the identification of highly distinct gene expression patterns for each subtype. In addition, we observed that the MLLT10, HOXA9, and KMT2A subtypes also included rare instances of NUP98 and NUP214 cases. Importantly, these cases did not cluster with the ETP-like subtype NUP98/NUP214 cases, warranting their separate analysis. Meanwhile, the TME-enriched subtype, lacking a unifying driver, exhibited a connection with low blast percentages (Extended data Fig.1f). Further analysis of gene expression signatures revealed an elevated presence of monocytes and other components of the tumor microenvironment (TME), accompanied by an absence of T-cell gene expression. This evidence suggests that low blast percentages pose challenges for driver discovery and is likely to affect transcriptional profiling of these samples due to TME signals.
The Leiden algorithm at a resolution of 0.5 was employed to identify subclusters within the main subtypes. Specifically, the ETP-like, TAL1, TLX3, NKX2-1 groups were subjected to subgrouping based on genetic alterations in alignment with these subclusters. The ETP-like subtype exhibited different drivers per subcluster, whereas the TAL1, NKX2-1 and TLX3 groups showed notable disparities in gene expression subclusters and underlying driver variants or co-lesions, providing a basis for their refined classification.
Bulk and scRNA expression analysis
T-ALL gene expression, surface protein expression analysis and assembly of normal thymus and bone marrow cell scRNA reference set are described in the Supplementary Methods.
Assembly of final genetic alteration dataset
The final dataset employed for subsequent statistical analysis encompassed harmonized data from 1309 patients with WGS and RNA. This comprehensive dataset integrated significant alterations (See GRIN2, GISTIC2, DnDScv and data harmonization in Supplementary Methods), along with classifying driver alterations for each patient sample. This was done to also keep rare variants that were putative drivers. Additionally, manually reviewed broad copy number variations (CNVs) were included for each sample in addition to GISTIC2/GRIN2 identified focal events. Each variant underwent classification as either Coding or Non-coding, with three levels of annotations, ranging from simple to more granular annotations (Supplementary Table 9–12). The subsequent step involved condensing the data at the gene, variant, or pathway level as binary features. Pathway curation and PubMed PMIDs are shown in Supplementary Table 13.
Outcome Analysis
Overall survival (OS) was defined as the time from study enrollment or postinduction randomization to death or date of last contact. Event-free survival (EFS) was defined as time from study enrollment to first event (induction failure, induction death, relapse, second malignant neoplasm, or remission death) or date of last contact. Disease free survival was defined as the time from postinduction randomization to first event (relapse, second malignant neoplasm, or remission death) or date of last contact10. Minimal residual disease (MRD) was treated as a numeric proportion or as binary variable with class “Negative” if MRD<0.1% and “Positive” if MRD≥0.1%. We also created an ordinal MRD variable that further divided the “Positive” group to “Low Positive” if MRD<1% and “High Positive” if MRD≥1%.
Univariable Screening and Competing Risks
Multivariable Analysis
We explored multiple statistical methodologies to build prognostic models with each of the survival endpoints and multiple candidate predictors. To minimize overfitting, we split the data 100 times into 70% training and 30% test datasets. Then for each method, we built a predictive model using the training dataset and calculated the concordance of this model in the test dataset39. The 100 replicates allow us to capture the uncertainty associated with the model concordance, which was summarized by mean and 95% interval from the replicates. The prognostic model methods considered were Penalized Cox Models, Random Survival Forests, Survival Trees (See Supplementary methods).
In Vitro Experiments
Engineering MED12 knockout cells
LOUCY-MED12KO (MED12 knockout) cells were generated using CRISPR/Cas9 technology, employing Cas9-gRNA ribonucleoprotein (RNP) delivery. A mixture consisting of 3 μl of Cas9 protein (20 μM) and sgRNAs (60 μM) was incubated for 15 minutes at room temperature and subsequently combined with LOUCY cells at a concentration of 1x105 cells/ml in Buffer T (Invitrogen, #MPK1025K). Electroporation was executed using the Neon Transfection System (Thermo Fisher Scientific) under the conditions of 1600V, 10ms, and three pulses. Following a 72-hour incubation, the electroporated cells underwent single-cell sorting for precise clonal isolation. The genetic modifications were assessed through Sanger Sequencing, and the loss of MED12 protein expression was confirmed by western blotting. Furthermore, WGS was conducted to validate the gene edits and to scrutinize potential off-target effects.
Western blotting
Engineered LOUCY-MED12KO cells were lysed in radioimmunoprecipitation assay buffer supplemented with protease and phosphatase inhibitors (Thermo Fisher Scientific, #1861281). 20 μg of protein of the cell lysate was electrophoresed through 3-8% NuPage Tris-acetate gels (Life Technologies) at 110 V for 120 minutes. Blots were probed with anti-MED12 (Cell Signaling technology, #14360, 1:1000 dilution) and anti-Lamin B (Abcam, #133741, 1:2000 dilution) antibodies. For imaging and quantitation, Odyssey DLx (LI-COR) and Image Studio (LI-COR) were used.
Sequencing validation of NOTCH1 exon 27-28 intronic SNV
We analyzed two patient samples from St Jude Children’s Research Hospital (Memphis, USA) Biorepository. DNA/RNA extraction was performed by Quick-DNA/RNA Microprep™ Plus Kit (#d7005, Zymo Research). Primers were designed upstream and downstream the SNV and used to amplify DNA by PCR using VeriFi™ Hot Start Mix Red (#PB10.47, PCR Biosystems). Purification of the PCR products was obtained by Wizard® SV Gel and PCR Clean-Up System (#A9282, Promega) and sequences were verified by Sanger sequencing using 3730 DNA analyzer (Applied Biosystems) and BigDye™ Terminator v3.1 Cycle Sequencing Kit (Applied Biosystems). To validate the extended transcript due to the intronic SNV, reverse transcription of RNA of the two patient samples and two cell lines not harboring the intronic SNV, PEER (#ACC6, DSMZ) and LOUCY (#CRL-2629, ATCC), was performed using SuperScript™ III First-Strand Synthesis System (#18080051, Invitrogen). Amplification of the region of interest was performed by PCR using VeriFi™ Hot Start Mix Red (#PB10.47, PCR Biosystems) and primers designed upstream and downstream the retained intronic region (Supplementary table 56). Purification of the PCR products was obtained by Wizard® SV Gel and PCR Clean-Up System (#A9282, Promega) after separation by molecular size through 0.8% agarose gel electrophoresis. Sequences were verified by Sanger sequencing as described above. The results were analyzed by CLC Genomic Workbench v. 22.0 (Qiagen).
Luciferase reporter assays
pcDNA6.2 NOTCH1 expression constructs were transiently transfected in HEK293T (CRL-11268 293T/17 #3531893, ATCC) cells using FuGENE HD Transfection Reagent (#E2311, Promega) together with RBPJ/CSL luciferase reporter construct (pGa981-6)40 and Renilla luciferase expressing vector (pRL-CMV) (#E2261, Promega). Luciferase assays were performed using Dual Luciferase Assay System (#E1910, Promega) 24 hours after transfection on Biotek Synergy HTX Multimode Reader (Agilent) according to manufacturer’s protocols. Assays were performed in triplicate and repeated at least three times with consistent results.
Data Availability
Primary sample whole-genome, exome and transcriptome data are available under database of Genotypes and Phenotypes (dbGaP) accession number phs002276.v2.p1 (phs000218, phs000464 for T-ALL TARGET samples) and the Kids First data portal (https://portal.kidsfirstdrc.org/dashboard). Clinical data processed genomic data and statistical analysis results can be found in the Supplementary Tables excel file. HiChIP, Isoseq, and ATACseq data are available in European Genome Phenome (EGA) data portal, accession number EGAS50000000016, https://ega-archive.org/studies/EGAS50000000016. Processed gene expression data can be accessed from synapse, accession number syn54032669, https://doi.org/10.7303/syn54032669. Somatic alterations, recurrent mutations and HiChIP/ATACseq tracks can also be explored interactively using ProteinPaint41 and GenomePaint42 on St. Jude Cloud at https://viz.stjude.cloud/mullighan-lab/collection/the-genomic-basis-of-childhood-t-lineage-acute-lymphoblastic-leukemia~29. Recurrent mutation plots and oncoprints are provided as supplementary data.
Code Availability
This study did not involve the development of software. Code to reproduce key parts of the analysis can be accessed from GitHub: https://github.com/ppolonen/genomic_basis_TALL.
Supplementary Information is available for this paper.
Extended Data
Extended Data Figure 1. T-ALL cohort composition and subtyping.

a, Schematic depiction of the cohort composition, analysis framework, and study outcomes; b, UMAP scatter plots illustrating Leiden algorithm clusters at 0.1 resolution, annotated by subtype; c, Leiden algorithm clusters at 0.5 resolution, further delineating distinct groups; d. UMAP plot showing types of TCR loci involved in structural variants (SV) that lead to oncogene activation by enhancer hijacking; e, Leiden algorithm clusters showing oncogene expression patterns; f, Bar plot presenting blast percentage distribution proportions, categorized by subtypes; g, Dot plot of marker genes from Fig.2 in normal cell scRNA data, depicting average (mean) expression (dot color) and detection percentage (dot size); h, UMAP plot showing the identification of clonal TCR rearrangements in cancer cells; i, UMAP plot of ETP status; j, Bar plot of ETP status proportions for cases with known immunophenotype, categorized by subtypes, see panel i for legend (two-tailed Chi-square test P=3.25e-104, see N for each subtype from l, whole cohort N=1309).; k, Bar plot showing age distribution proportions, categorized by subtypes and ordered by median age; l, Box plot, as defined in Fig.3d legend, presenting age distributions per subtypes, with P value from one-way ANOVA. (two-sided pairwise t-test Holm-adjusted P for significant pairs using TAL1 DP-like as comparison group: BCL11B P=0.0053, ETP-like P=0.00016, HOXA TCR P=0.019, NKX2-1 P=0.0092, NKX2-5 P=0.0014, SPI1 P=0.049, STAG2/LMO2 P=0.00023). Median age in the cohort is shown as red dotted line; m, Bar plot depicting gender distribution proportions, categorized by subtypes (two-tailed Chi-square test P=0.000000417, see N for each subtype from l, whole cohort N=1309).
Extended Data Figure 2. Alterations in T-ALL and subtype-associated genetic changes.

a, Stacked bar plot depicting proportions of alteration types for significantly altered genes (q<0.05 and altered in >1% of cases). Left: frequency of alteration types. Right: alteration frequencies; b, Dot plot illustrating altered genes, broad CNVs, and pathways significantly associated with subtypes, displaying odds ratio as dot color and −log10 FDR as dot size. Odds ratios and −log10 FDR were truncated at 20. Genes/lesions are ordered by significance within each subtype.
Extended Data Figure 3. Genomic characterization of enhancer hijacking events.

a, Complex chromosome 7 rearrangements: Depiction of H3K27ac HiChIP, H3K27ac coverage, diagnostic sample, and germline WGS coverage in a T-ALL patient, showing interactions between HOXA13 and TCRβ loci after chromosome 7 rearrangements. Arcs represent interactions between genomic loci, with color coding denoting interaction strength at 500 kb resolution. Breakpoints are annotated by orange lines. Left: Boxplot, as defined in Fig.3d legend, of HOXA13 gene expression and P value for TCR::HOXA13 mutated (mut) vs HOXA13 unaltered wild type (wt) samples (two-tailed Student t-test, P values were not adjusted for multiple comparison), with PAVHBC sample labeled in cyan. The intensity of each arc represents contacts at 5kb resolution between a pair of loci. Maximum intensity is indicated in the color scale; b, TCR::HOXA9 Inversion: Similar representation as a; c, TCR:: LMO3 translocation: similar representation as a; d, BCL11B::LMO2 translocation: Similar representation as a; e, HOXA13 BCL11B translocation: similar representation as a; f, NKX2-1 Chromosome 14 chromothripsis and BCL11B locus rearrangement: similar representation as a; g, BCL11B::HOXB13 translocation: similar representation as a; h, LINC00592::TLX1 enhancer hijacking by 10Mb intergenic loss. WGS and RNA coverage for the representative cases, normal double positive and CD34+ HSPCs H3K27ac and CTCF coverage is shown, with putative enhancers highlighted in blue. Top: Heatmap showing normal double positive and CD34+ HSPCs H3K27ac HiChIP raw interactions as heatmaps, with topologically associated domains annotated by black dotted triangles. Left: boxplot as in a; i, NFKBIA::NKX2-1 enhancer hijacking: similar representation as a; j, MIR2117HG::LMO2 translocation: Similar representation as a; k, DHX9::TAL1 inversion: Similar representation as a; l, CELF1::LMO2 intergenic loss: Similar representation as a; m, CAPRIN1::LMO2 enhancer hijacking is shown as in h, with additional highlight of CTCF boundary by blue rectangle; n, MIR181A1HG::HOXA13 Translocation: similar representation as a; o, MIR181A1HG::LMO2 translocation: similar representation as a; Abbreviations: diagnostic and germline/remission sample WGS (WGS-D, WGS-G). Nucleosome-free and nucleosome cut fragments (ATAC-free, ATAC-nuc).
Extended Data Figure 4. Genomic characterization of enhancer and promoter alterations.

a, H3K27ac HiChIP, H3K27ac coverage, cancer (WGS-D) and remission/germline (WGS-G) coverages in T-ALL patient, showing interactions between MYC and the amplified N-me enhancer. Color scale bar showing interaction strength at 5kb resolution. The intensity of each arc represents contacts at 5kb resolution between a pair of loci. Maximum intensity is indicated in the color scale; b, WGS coverage tracks showing recurrent LINC00649 losses in T-ALL patients. H3K27ac HiChIP, H3K27ac and CTCF coverage shown as in a. LINC00649 is a putative regulatory region for RUNX1 in CD34+ cells; c, Illustration of representative cases featuring WGS coverage tracks in FTO/IRX loci, along with two WGS germline control samples. Color scale bar showing interaction strength at 5kb resolution. Boxplots, as defined in Fig.3d legend, showing IRX3 and IRX5 gene expression comparing FTO/IRX loci deletions and wild type (wt) (two-tailed Student t-test, not adjusted for multiple comparisons). Allelic expression denoted by red dots. Top: H3K27ac HiChIP, H3K27ac and CTCF coverage in CD34+ cells, showing interactions between FTO enhancers and IRX3 and IRX5. Arcs represent interactions between genomic loci, with color coding denoting strength of interaction; d, Similar representation as c for ZNF219 loci with annotation of TCRδ loci; e, WGS, H3K27ac HiChIP, H3K27ac and CTCF coverage in CD34+ cells, showing interactions between the HNRNPC promoter and enhancers and ZNF219 loci. Below, WGS and RNA coverage showing copy number loss at ZNF219 loci and revealing RNA coverage match with breakpoint location (red) and reads clip with HNRNPC promoter (not shown); f, Gene expression boxplot by LMO2 variant type with samples with monoallelic expression annotated as red dots; g, WGS, H3K27ac HiChIP, ATAC seq, WTS coverage tracks are shown for T-ALL patient PASPLG with 6bp enhancer deletion (light blue line) and PAWGYX 85bp intronic gain (red line), revealing high levels of H3K27ac marking neoenhancer formation for PASPLG and neomorphic promoter generation for PAWGYX; h, RNA coverage tracks for samples with LMO2 intronic gains (red), intronic SNV/indel (blue) and intergenic SNV/indels (light blue), showing generation of alternative transcript start site for intronic gains and intronic SNV/indel, but not intergenic SNV/indels. Top: junction reads for representative case PARMMV are shown; i, WGS coverage tracks for TAL1 enhancer gains (cyan line) are shown in 9 patient samples and one germline control (PATSIL_G) sample; j, WGS coverage for T-ALL patient and remission/germline samples are shown with four RNA isoform sequencing reads visualized in IGV, capturing the 89 bp gain of the TAL1 enhancer. Orange rectangle: highlight of the increased WGS coverage matching 89 bp gain/insert sequence size. Right: boxplot as in c of TAL1 gene expression comparing TAL1 enhancer gains with TAL1 wt samples with PASXMF sample labeled in cyan.
Extended Data Figure 5. Enhancer hijacking is associated with differentiation stage.

a, H3K27ac HiChIP, H3K27ac, WGS and ATAC-seq coverages in T-ALL sample PASLPH showing CD1E::TAL1 intergenic inversion. Arcs depict interaction between genomic loci, with colors indicating interaction strength. The intensity of each arc represents contacts at 5kb resolution between a pair of loci. Maximum intensity is indicated in the color scale. Left: boxplot, as defined in Fig.3d legend, of TAL1 gene expression for CD1E::TAL1 and TAL1 wild type (wt) samples, with the PASLPH sample labeled in cyan. Right: heatmap showing H3K27ac HiChIP interactions for normal double positive (DP) and CD34+ HSPCs as heatmaps, with H3K27ac coverage shown below. Lesion breakpoints are shown as black lines; hijacked enhancers color coded; topologically associated domains annotated by black dotted triangles; b, H3K27ac coverage tracks of CD1E::TAL1 enhancer hijacking (marked by cyan line) is shown in different thymic T-cells and in CD34+ cells. Top: heatmap showing H3K27ac HiChIP raw interactions for normal double positive and CD34+ HSPCs; c, same as b for RAG2::LMO2 enhancer hijacking; d, same as b for SOX4/CASC15::HOXA13 enhancer hijacking; e, same as b for MIR181A1::HOXA13 enhancer hijacking; f, Expression of enhancer hijacking-driven leukemia gene expression signatures in normal BM/thymic scRNA data: Dot plots depict mean gene set score of each gene expression signature by normal cell type (color) and mean gene set detection percentage (dot size). Left: subtype annotation for each alteration. Top: RAG1, RAG2 gene expression is shown, with mean expression (dot color) and detection percentage (dot size). Right: RAG1/RAG2 gene expression log2 fold change (FC) of for each alteration (altered vs. wild type) is shown as a heatmap. Right to the heatmap, horizontal bar plots show the percentage of each enhancer hijacking structural variant (SV) event predicted to be RAG mediated. Bonferroni adjusted −log10 P value is shown (one-tailed Fisher’s exact test, comparing RAG mediated breakpoints compared to non-RAG mediated breakpoints for each lesion category).
Extended Data Figure 6. Intragenic and intergenic non-coding alterations.

a, NOTCH1 amino acid position, mRNA bp position, exon number, and previously described35 protein domains are shown. Heterodimerization domain (HD) and transmembrane (TM), TM-connector domains are color coded and locations of the proteolytic cleavage sites (S2 and S3) that lead to intracellular domain release from the membrane are shown as black lines. Below, exons 27-28 are shown, with NOTCH1 intronic SNV location marked by a red line, with mutation position and types (forward strand C-A or C-G), resulting 3’ splice site sequences (TAG and CAG mutated splice site, GAG wild type (wt)) are shown. The mutated splice site results in 129 bp or 43 amino acid insertion (green rectangle) at exon 28 and the insert sequence is shown below; b, NOTCH1 intergenic SNV validation by RT-PCR. Two patient samples, SJTALL031662 and SJTALL031904, have longer 538bp transcript due to the intronic SNV, compared to NOTCH1 wt 409bp transcript for PEER and LOUCY cell lines, primers at Exon 26 (9:136504848-9:136504827), Exon 28 (9:136502426-9:136502405). This experiment was not repeated, but variant was confirmed using Sanger sequencing; c, NOTCH1 exon 28 33 bp duplication sites are shown using cancer patient sample (D) and matched germline/remission sample (G) WGS and LR isoform sequencing. Duplicated site sequence match with increased DNA coverage, highlighted in blue; d, NOTCH1 locus showing WGS coverage tracks and intragenic deletions between exon 16-27 and 3-27 in 11 representative cases and matched non-tumor sample controls; e, NOTCH1 splicing plots for T-ALL patients with exon 16-27 and 3-27 deletions, compared to wt control, are visualized. Atypical splicing reads are highlighted in red; f, Top: Coverage tracks for IL7R diagnostic and matched non-tumor WGS and WTS coverage, along with SNPs. Below: Isoform sequencing reads, indicating the loss of one allele due to TSS loss. Right: IL7R and PRLR loci deletion is shown for representative cases, accompanied by IL7R and PRLR gene expression and P value (two-tailed, Student t-test); g, Similar representation as f for CCND3: showing the loss of the long isoform (red) of CCND3 and TAF8 allele, while both alleles of short isoforms (blue) are expressed; WGS-D, whole genome sequencing of diagnosis sample, WGS-G, matched non-tumor sample. Nucleosome-free (ATAC-free) and nucleosome cut fragments (ATAC-nuc).
Extended Data Figure 7. Effect of CCND3/TAF8 TSS deletion and SNV/indels on protein structure.

a, the distribution of detected isoforms, showing that ENST00000372991 (short isoform of CCND3) is predominantly expressed; b, boxplots, as defined in Fig.3d legend, depict the gene expression comparisons of CCND3, TAF8, short isoform (ENST00000372991), and long isoforms (ENST00000372988, ENST00000415497), with P value (two-tailed Student t-test) shown for comparison between CCND3/TAF8 TSS loss (N=27) and CCND3 wild type (wt, N=1193). Additional control groups were samples with CCND3 SNV/indel (N=79), and CCND3 wt TAL1 DP-like (N=253), and TAL1 αβ-like (N=206) subtypes (these subtypes are enriched for CCND3/TAF8 TSS loss). CCND3/TAF8 TSS loss is associated with decreased long isoform expression while total CCND3 expression remains constant due to high short isoform expression; c, Exon composition, CDD/PFAM protein domains, and amino acid residue positions of CCND3 SNV/indel for different isoforms; d, AlphaFold2 model of the long (blue) and short (yellow) isoforms of CCND3 superimposed onto the crystal structure of the CDK4/CCND3 (PDB code 3G-33); e, Detailed views of the contacts between α1 and α2 of CCND3 showing that the long isoform of CCND3 lacks the first two α-helices of the short isoform protein and α1 and α2 engage in essential hydrophobic contacts within the core of CCND3. FoldX predicted that the long isoform of CCDN3’s protein is unstable, with an estimated destabilization of 100 kcal/mol, likely leading to an unfolded protein; f, SNV/indel at the C-terminus of CCND3 is unlikely to affect the interaction with CDK4. Mutations at the C-terminus may affect CCND3 stability by interfering with CCND3 protein homeostasis; there is a conserved TPTDV motif at the C-terminus of CCND3 that recruits the E3 ligase AMBRA1. The TPTDV is maintained in both the K268R and 264_274del mutants, but not in the R217fs variant.
Extended Data Figure 8. Subtyping and genomic analysis of TLX3 and NKX2-1 subgroups.

a, Oncoprint showing alterations that are significantly different in frequency between TLX3 DP-like and Immature subgroups. Left: −Log10 FDR (two-tailed Fisher’s exact test); TCR and ETP status are shown above lesions; b, UMAP plots representing TLX3 DP-like and Immature subgroups; c, UMAP plot displaying TCR rearrangement status for TLX3 cases; d, UMAP plot illustrating ETP status for TLX3 cases; e, Oncoprint of NUP214::ABL1 fusion and JAK/STAT or RAS pathway genes. Gene expression of differentially expressed (FDR<0.01) surface markers and kinase genes are shown above; f, Dot plot depicting significantly altered genes, broad CNVs, and pathways associated with TLX3 DP-like and Immature subgroups. Dot color represents odds ratio, and dot size reflects −log10 FDR values. Odds ratio values are values are truncated at a maximum of 10. Genes are ordered by significance within each subtype; g, UMAP plots showing selected gene and lesion alterations (green) for TLX3 samples; h, Dot plot visualizing expression of TLX3 DP-like and Immature gene expression signatures in normal BM/thymic single cell RNA data. Detection percentage is represented by dot size, while mean signature enrichment by normal cell type is indicated by color. The TLX3 subgroup is further divided based on ETP status and TCR rearrangement status; i, Similar to a, showing significantly different lesions between NKX2-1 TCR and Other subgroups; j, UMAP plot illustrating labels for NKX2-1 TCR and Other subgroups; k, WGS coverage tracks in chromosome 14 showing chromothripsis and annotation for the NKX2-1 locus; l, similar as g, Dot plot depicting significantly altered co-lesion genes, broad CNVs, and pathways associated with NKX2-1 TCR and Other subgroups; m, UMAP plots displaying selected gene and lesion alterations (green) for NKX2-1 samples.
Extended Data Figure 9. Subtyping and genomic analysis of TAL1/LMO2 genomic subgroups.

a, Oncoprint representing driver alteration prevalence between TAL1 DP-like and αβ-like subtypes. Left: −log10 FDR significance levels are shown (Two-tailed Fisher’s exact test); b, Oncoprint of co-lesion prevalence as in a; c, Heatmap showing mutually exclusive and co-occurring lesions within the TAL1 DP-like and αβ-like subtypes. Color represents log2 odds ratio. Significant pairs are indicated by stars (one-sided Fisher’s exact test, *** = FDR<0.001, ** = FDR<0.01, * = FDR<0.05); d, UMAP plots representing the alterations associated with the TAL1 groups; e, Oncoprint of genetic subgroup-associated alterations categorized by genes, lesions, and pathways. The subdivisions are based on subtypes and further divided by genetic subgroup. Oncogene gene expression and expression of MYC, MYB and MYCN are represented as Z scores; f, UMAP plots illustrating the distribution of TAL1 genetic subgroups and LMO2 γδ-like and STAG2/LMO2 subgroups. TCR rearrangement status and ETP status are indicated on the right; g, Volcano plot depicting differential flow cytometry markers between TAL1 αβ-like and DP-like subtypes (TAL1 αβ-like upregulated markers log2 fold change positive and downregulated negative); Significantly altered markers (two-tailed Wilcoxon rank-sum test) are labeled by protein name, with P value (0.1) and log2 fold change cutoffs (0.25) are shown as black dotted lines; h, As in g, volcano plot comparing LMO2 γδ-like samples to the rest of the samples, showing differential flow cytometry markers. i, As in g, Volcano plot showing differential flow cytometry markers between LMO2 γδ-like and TAL1 αβ-like subtypes; j, Similar to g, volcano plot comparing flow cytometry markers between LMO2 γδ-like and TAL1 DP-like subtypes; k, Similar to g, differential flow cytometry markers in STAG2/LMO2 group; l, Similar to g, differential flow cytometry markers between STAG2/LMO2 and TAL1 DP-like subtypes.
Extended Data Figure 10: ETP/ETP-like genomics, immunophenotype and clinical associations.

a, UMAP plot revealing distinct MLLT10 and KMT2A fusion partners; b, Dot plot showing significant gene, CNV (gain/amplification/loss/deletion), and pathway alterations linked with ETP-like genetic subtypes, with dot color indicating odds ratio and dot size reflecting −log10 FDR. Odds ratio values are truncated at a maximum of 10, with genes/lesions ordered by significance within each subtype; c, Location of MED12 SNV/indel amino acid position in T-ALL; d, Western blot of MED12 in knockout in PER117 and LOUCY cell line. These uncropped images were obtained from LI-COR odyssey by defining the area to be scanned. Assay was repeated three times with consistent results; e, Dot plot displaying GSEA enrichment result for significant pathways in MED12KO vs. wild type (wt) and ETP-like MED12 subgroup, where dot size is −log10 FDR and dot color is Normalized Enrichment Score (NES); f, Volcano plot illustrating significantly differentially expressed genes (two-tailed Wald test) between LOUCY MED12 knockout (KO) and wt samples (LOUCY MED12 KO upregulated markers log2 fold change positive and downregulated negative). The analysis is restricted to genes that exhibited differential expression in ETP-like MED12 subgroup; g, Dot plot visualizing the gene set detection percentage (dot size) and mean gene set enrichment by normal cell type (color) for MED12 knockout and MED12 subtype upregulated (N=60) and downregulated (N=104) genes in normal BM/thymic scRNA data. Below is a dot plot of select significantly differentially expressed genes (as in Fig. 4b) in normal cell scRNA, depicting mean expression (dot color) and detection percentage (dot size); h, Heatmap of HOXA gene expression, samples ordered by distance to HOXA9 TSS, and driver/alteration type annotations on top; i, Oncoprint of ETP-like group altered drivers and pathways, divided by ETP status, with day 29 MRD bar plot, clinical characteristics, morphological response, relapse, and TCR rearrangement status annotations. FDR between ETP categories and genomics indicated on the left; j, Volcano plot depicting differential expression of cell surface markers detected by flow cytometry between ETP-like cases without TCR rearrangement (no TCR-R) versus those with TCR γδ rearrangement (ETP-like TCRγδ) upregulated markers log2 fold change positive and downregulated negative, with P value (0.1) and log2 fold change (0.25) cutoffs shown as black dotted lines), with significantly altered (from two-tailed Wilcoxon rank-sum test) markers labeled by protein names; k, Similar to j, this plot compares flow cytometry markers between ETP-like MLLT10 cases and other MLLT10 cases.
Extended Data Figure 11. T-ALL genomic risk factors and clinical outcomes considering MRD, DFS and OS.

a, Proportions of induction failure and residual disease cases by genetic subgroup; b, Oncoprint of variants significantly associated with MRD. Day 29 MRD displayed as a bar plot on top with other clinical annotations. Genomic data was divided to residual disease (day 29 MRD≥0.01%), and negative groups (day 29 MRD<0.01%). Left: number of present (N) vs. absent lesions, −log10 q value from Firth penalized logistic regression, log odds ratios (OR) and then forest plot with center box presenting log OR and error bars indicating 95% confidence intervals (CI); c, Forest plot showing hazard ratios (HR) and 95% CI comparing cases with subtype or genetic subtype present (N) vs. absent, considering DFS and OS as outcomes. Significant associations (two-tailed, Firth-penalized Cox models that also adjusted for MRD (MRD.adj), P<0.1) are highlighted in red. These P values were not adjusted for multiple comparisons.
Extended Data Figure 12. Genomic risk factors and cumulative incidence of relapse in T-ALL.

a, Left: number of present vs. absent (N) lesions, −log10 P value, hazard ratios (HR). Right: forest plot with center box presenting HR and error bars indicating 95% confidence intervals (CI) for variants that were present vs. absent considering EFS, DFS, and OS as outcomes. Significant associations (two-tailed, Firth-penalized Cox models that also adjusted for MRD, P<0.1) are highlighted in red. These P values were not adjusted for multiple comparisons; b, The Kaplan-Meier curve of NOTCH pathway mutated vs. not mutated cases, stratified by MRD percentage (%), with P value from two-tailed Log-rank test; c, Cumulative incidence of relapse plot, revealing distinct risk between PI3K (NOTCH wild type (wt)) and NOTCH (PI3K wt) pathway alterations, with P value from two-tailed Holm-adjusted Gray’s test; d, as in c, Comparison of PTEN deletions to other PTEN alterations in terms of cumulative relapse incidence; e, as in c Comparison of NOTCH1 intragenic loss to other NOTCH1 alterations regarding cumulative relapse incidence; f, as in c, Comparison of TCR:: MYC alterations to other MYC alterations in relation to cumulative relapse incidence; g, as in c, Comparison of LMO2 intergenic loss to other LMO2 alterations with respect to cumulative relapse incidence; h, as in c, Comparison of TAL1 upstream enhancer Indel to other TAL1 alterations in terms of cumulative relapse incidence; i, Comparison of TAL1 genetic subtypes in relation to cumulative relapse incidence with P value from two-sided global Gray’s test; j, Contingency table comparing TAL1 upstream enhancer Indel and LMO1 TCR to induction failure, demonstrating no association (P=0.91 and P=0.92, one-tailed Fisher’s exact test); k, Scatter plot of variant allele frequency (VAF) of all alterations in TALL0223 case, with T-ALL shown on the y axis and Langerhans cell histiocytosis (LCH) in the x axis with annotations for T-ALL driver mutations and clusters denoting three distinct clones; l, analysis of differentially expressed genes (two-tailed Wald test) between the T-ALL SPI1 subtype (N=11) and the TALL023 Langerhans cell histiocytosis (LCH) case (N=1) reveals the downregulation of T-cell genes and the upregulation of the LCH marker CD207; m, A heatmap showcases select SPI1 subtype markers (HLA-D, CD7, CD5, CD1A, CD3E, CD2, PTPRC/CD38), LCH markers (CD207)34, dendritic cell progenitor marker (IRF8)34, T-cell markers (ZAP70, LCK, GATA3, RAG1, CD4, CD8A, TRDC, TRGC1, TRBC1, TRAC), and myeloid markers (CD33, CD14) for the SPI1 subtype and the TALL023 T-ALL and LCH cases. Right: same markers are shown in a normal cell scRNA, showing mean expression (dot color) and detection percentage (dot size).
Extended Data Figure 13. Multivariable outcome model evaluation.

a, The mean validation concordance and associated 95% confidence intervals (CI) are shown, derived from 100 random splits (70% training, N=915, 30% validation, N=394) of the data for three distinct survival models. These models were fitted using different data types, with binary Day 29 MRD (≥0.1%) employed as the baseline, shown as red line. The baseline concordance results are denoted by grey rectangles, where binary MRD, numeric MRD, and a combination of clinical features (MRD, sex, WBC, CNS) and NOTCH1/FBXW7/RAS/PTEN classifier36 are shown. The blue rectangle represents models fitted using single data types, including numeric MRD, Sex, WBC, and CNS as predictors. The pink rectangles denote combinations of genomic and clinical features. Model coefficient numbers are shown on top of the figure as bar plots for penalized Cox regression and survival Trees. Model concordance is shown on top of the best performing model for each algorithm and highest concordance annotated as blue line. Select subtype denotes data where the main subtype is further divided into genetic subtypes when available; b, The stacked bar plot portrays a four-node survival tree that was generated from 1000 bootstrap resamples. The proportion of bootstrap samples in which each subtype was classified into various risk groups is shown; c, The Kaplan-Meier curve illustrates the risk score (divided into four quartiles) derived from a penalized Cox regression multivariable model, limited to ETP cases, showing models ability to risk-stratify patients within ETP group. Log-rank test P value is shown; d, Same as c, for Near-ETP cases; e, Same as c, for ETP and Near-ETP cases; f, Same as c, for Non-ETP cases; g, The Kaplan-Meier curve dividing ETP-like group by ETP status, revealing no difference in outcomes; h, Same as c, for ETP-like subtype ETP cases; i, Same as c, for ETP-like subtype ETP and Near-ETP cases; j, Same as c, for ETP-like subtype Non-ETP cases; k, Dot plot illustrating altered pathways and pathway genes significantly associated with subtypes, displaying odds ratio as dot color and −log10 FDR as dot size. Odds ratio and −log10 FDR were truncated at maximum of 20.
Supplementary Material
Acknowledgements
We would like to thank the Gabriella Miller Kids First Pediatric Data Research Program and Data Resource Center, including Marcia Fournier, PhD, James Coulombe, PhD, Emily Boja, PhD, Jamie Guidry Auvil, PhD, Valerie Cotton, BSc, and David Higgen, PhD; the Biopathology Center at Nationwide Children’s Hospital, including Alexis Cameron, BS, Yvonne Moyer, MBA, and Tyler Jones, BS; Children’s Oncology Group (COG) operations including Sarah Vargas, PhD, Mary Beth Sullivan, MPH, Michael Thomas, BA, and Chelsee Sauni, MPH; COG leadership including Douglas Hawkins, MD, Lia Gore, MD, and Peter Adamson, MD; the Cancer Therapy Evaluation Program (CTEP) including Malcolm Smith, MD, PhD; Hudson Alpha Genomics Project Management, including Salina Kuhafa-Hall, MSc; and, the Flow Cytometry Core at the Children’s Hospital of Philadelphia, including Florin Tunic, MD, PhD, and Jennifer Murray, BS; Phil Raess, M.D., Ph.D. Oregon Health & Science University for facilitating SPI1 case sequencing; Computational Biology Training in Hematology (CBTH) program mentors. The AALL0434 clinical trial was supported by Novartis.
Competing interests statement
D.T.T. received research funding from BEAM Therapeutics, NeoImmune Tech and serves on advisory boards for BEAM Therapeutics, Janssen, Servier, Sobi, and Jazz. D.T.T. has multiple patents pending on CAR-T. C.G.M. serves on the scientific advisory board and honoraria for Illumina, and received research funding from Pfizer, equity from Amgen and royalties from Cyrus. E.R Institutional research funding from Pfizer and serving on a Data and Safety Monitoring Board for Bristol Myers Squibb. I.I. Travel and accommodation expenses reimbursed by Mission Bio. I.I and P.P Consultancy fee by Arima Genomics. K.M.B research funding from Syndax.
Funding
National Institutes of Health Gabriella Miller Kids First Pediatric Research X01HD100702 (D.T.T., C.G.M., P.P., M.L.L., S.P.H., S.W., E.A.R., B.L.W., M.D., S.P.B., K.P.D., J.J.Y.), R03CA256550 (D.T.T., C.G.M., P.P., M.L.L., S.P.H., S.W., E.A.R., B.L.W., M.D., S.P.B., K.P.D., J.J.Y.), R01CA193776 (D.T.T., B.W., K.T., C.G.M., S.P.H., J.J.Y., R.S., M.D.), U10CA180886 (D.T.T., M.L.L.), R01CA264837 (D.T.T., J.J.Y., C.G.M., K.T., B.L.W., R.S.), U10CA180899 (M.D.), U24CA114766 (D.T.T., M.L.L.), U24CA196173 (D.T.T.), R01GM115634 (R.W.K.), U54CA243124 (R.W.K., C.G.M.), P30CA021765 (C.G.M., G.W.), R35CA197695 (C.G.M.), T32CA236748 (C.G.M), F32CA254140 (L.E.M), K12CA076931 (C.D). the Aiden Everett Davies Innovation Fund (D.T.T), Alex’s Lemonade Stand Foundation (D.T.T., K.T., S.P.H.), American Lebanese and Syrian Associated Charities of St. Jude Children’s Research Hospital, Canadian Institutes of Health Research Young Investigator Award (C.D), the Harrison Willing Memorial Research Fund (D.T.T), The Henry Schueler 41&9 Foundation (P.G., C.G.M.), Hyundai Hope of Wheels (D.T.T., K.T., R.S.), the Invisible Prince Foundation (D.T.T), Leukemia and Lymphoma Society (D.T.T.), Pennsylvania Department of Health (D.T.T.), St. Baldricks Foundation (D.T.T), the St. Jude Children’s Research Hospital Chromatin Collaborative, and the St. Jude Children’s Hospital Hematological Malignancies Program Garwood Fellowship (S.K.). This research was supported in part by the National Cancer Institute grants. The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health.
REFERENCES
- 1.Summers RJ & Teachey DT SOHO State of the Art Updates and Next Questions | Novel Approaches to Pediatric T-cell ALL and T-Lymphoblastic Lymphoma. Clin Lymphoma Myeloma Leuk 22, 718–725 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Mansour MR et al. Oncogene regulation. An oncogenic super-enhancer formed through somatic mutation of a noncoding intergenic element. Science 346, 1373–1377 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Liu Y et al. The genomic landscape of pediatric and young adult T-lineage acute lymphoblastic leukemia. Nat Genet 49, 1211–1218 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Roberts KG et al. Targetable kinase-activating lesions in Ph-like acute lymphoblastic leukemia. N Engl J Med 371, 1005–1015 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Brady SW et al. The genomic landscape of pediatric acute lymphoblastic leukemia. Nat Genet 54, 1376–1389 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Holmfeldt L et al. The genomic landscape of hypodiploid acute lymphoblastic leukemia. Nat Genet 45, 242–252 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Coustan-Smith E et al. Early T-cell precursor leukaemia: a subtype of very high-risk acute lymphoblastic leukaemia. Lancet Oncol 10, 147–156 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Zhang J et al. The genetic basis of early T-cell precursor acute lymphoblastic leukaemia. Nature 481, 157–163 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Wood BL et al. Prognostic significance of ETP phenotype and minimal residual disease in T-ALL: a Children’s Oncology Group study. Blood 142, 2069–2078 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Dunsmore KP et al. Children’s Oncology Group AALL0434: A Phase III Randomized Clinical Trial Testing Nelarabine in Newly Diagnosed T-Cell Acute Lymphoblastic Leukemia. J Clin Oncol 38, 3282–3293 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Becht E et al. Dimensionality reduction for visualizing single-cell data using UMAP. Nat Biotechnol 37, 38–44 (2019). [DOI] [PubMed] [Google Scholar]
- 12.Traag VA, Waltman L & Van Eck NJ From Louvain to Leiden: guaranteeing well-connected communities. Scientific Reports 9 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Isoda T et al. Non-coding transcription instructs chromatin folding and compartmentalization to dictate enhancer-promoter communication and T cell fate. Cell 171, 103–119.e118 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Herranz D et al. A NOTCH1-driven MYC enhancer promotes T cell development, transformation and acute lymphoblastic leukemia. Nat Med 20, 1130–1137 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Yashiro-Ohtani Y et al. Long-range enhancer activity determines Myc sensitivity to Notch inhibitors in T cell leukemia. PNAS 111, E4946–E4953 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Nagel S, Kaufmann M, Drexler HG & MacLeod RA The cardiac homeobox gene NKX2-5 is deregulated by juxtaposition with BCL11B in pediatric T-ALL cell lines via a novel t(5;14)(q35.1;q32.2). Cancer Res 63, 5329–5334 (2003). [PubMed] [Google Scholar]
- 17.Soulier J et al. HOXA genes are included in genetic and biologic networks defining human acute T-cell leukemia (T-ALL). Blood 106, 274–286 (2005). [DOI] [PubMed] [Google Scholar]
- 18.Seki M et al. Recurrent SPI1 (PU.1) fusions in high-risk pediatric T cell acute lymphoblastic leukemia. Nat Genet 49, 1274–1281 (2017). [DOI] [PubMed] [Google Scholar]
- 19.Di Giacomo D et al. 14q32 rearrangements deregulating BCL11B mark a distinct subgroup of T-lymphoid and myeloid immature acute leukemia. Blood 138, 773–784 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Montefiori LE et al. Enhancer Hijacking Drives Oncogenic BCL11B Expression in Lineage-Ambiguous Stem Cell Leukemia. Cancer Discov 11, 2846–2867 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Chen S et al. Novel non-TCR chromosome translocations t(3;11)(q25;p13) and t(X;11)(q25;p13) activating LMO2 by juxtaposition with MBNL1 and STAG2. Leukemia 25, 1632–1635 (2011). [DOI] [PubMed] [Google Scholar]
- 22.Martincorena I et al. Universal Patterns of Selection in Cancer and Somatic Tissues. Cell 171, 1029–1041.e1021 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Mermel CH et al. GISTIC2.0 facilitates sensitive and confident localization of the targets of focal somatic copy-number alteration in human cancers. Genome Biol 12, R41 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Cao X, Elsayed AH & Pounds SB Statistical Methods Inspired by Challenges in Pediatric Cancer Multi-omics. Methods Mol Biol 2629, 349–373 (2023). [DOI] [PubMed] [Google Scholar]
- 25.O’Connor D et al. The Clinicogenomic Landscape of Induction Failure in Childhood and Young Adult T-Cell Acute Lymphoblastic Leukemia. J Clin Oncol 41, 3545–3556 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Rahman S et al. Activation of the LMO2 oncogene through a somatically acquired neomorphic promoter in T-cell acute lymphoblastic leukemia. Blood 129, 3221–3226 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Jumper J et al. Highly accurate protein structure prediction with AlphaFold. Nature 596, 583–589 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Sulis ML et al. NOTCH1 extracellular juxtamembrane expansion mutations in T-ALL. Blood 112, 733–740 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Liu Y et al. Discovery of regulatory noncoding variants in individual cancer genomes by using cis-X. Nat Genet 52, 811–818 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Ali S & Ali S Prolactin receptor regulates Stat5 tyrosine phosphorylation and nuclear translocation by two separate pathways. J Biol Chem 273, 7709–7716 (1998). [DOI] [PubMed] [Google Scholar]
- 31.Goffin V Prolactin receptor targeting in breast and prostate cancers: New insights into an old challenge. Pharmacol Ther 179, 111–126 (2017). [DOI] [PubMed] [Google Scholar]
- 32.Kato M et al. Genomic analysis of clonal origin of Langerhans cell histiocytosis following acute lymphoblastic leukaemia. Br J Haematol 175, 169–172 (2016). [DOI] [PubMed] [Google Scholar]
- 33.Kimura S et al. Biologic and clinical features of childhood gamma delta T-ALL: identification of STAG2/LMO2 γδ T-ALL as an extremely high risk leukemia in the very young. Cancer Discov, in press (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Liu X et al. Distinct human Langerhans cell subsets orchestrate reciprocal functions and require different developmental regulation. Immunity 54, 2305–2320.e2311 (2021). [DOI] [PubMed] [Google Scholar]
- 35.Gordon WR, Arnett KL & Blacklow SC The molecular logic of Notch signaling - a structural and biochemical perspective. J Cell Sci 121, 3109–3119 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Trinquand A et al. Toward a NOTCH1/FBXW7/RAS/PTEN–based oncogenetic risk classification of adult T-cell acute lymphoblastic leukemia: a group for research in adult acute lymphoblastic leukemia study. J Clin Oncol 31, 4333–4342 (2013). [DOI] [PubMed] [Google Scholar]
- 37.Winter SS et al. Improved survival for children and young adults with T-lineage acute lymphoblastic leukemia: results from the Children’s Oncology Group AALL0434 Methotrexate randomization. J Clin Oncol 36, 2926–2934 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Moore HM et al. Biospecimen reporting for improved study quality (BRISQ). Cancer Cytopathol 119, 92–101 (2011). [DOI] [PubMed] [Google Scholar]
- 39.Harrell FE Jr., Lee KL & Mark DB Multivariable prognostic models: issues in developing models, evaluating assumptions and adequacy, and measuring and reducing errors. Stat Med 15, 361–387 (1996). [DOI] [PubMed] [Google Scholar]
- 40.Oswald F et al. p300 acts as a transcriptional coactivator for mammalian Notch-1. Mol Cell Biol 21, 7761–7774 (2001). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Zhou X et al. Exploring genomic alteration in pediatric cancer using ProteinPaint. Nat Genet 48, 4–6 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Zhou X et al. Exploration of coding and non-coding variants in cancer using GenomePaint. Cancer Cell 39, 83–95.e84 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
Primary sample whole-genome, exome and transcriptome data are available under database of Genotypes and Phenotypes (dbGaP) accession number phs002276.v2.p1 (phs000218, phs000464 for T-ALL TARGET samples) and the Kids First data portal (https://portal.kidsfirstdrc.org/dashboard). Clinical data processed genomic data and statistical analysis results can be found in the Supplementary Tables excel file. HiChIP, Isoseq, and ATACseq data are available in European Genome Phenome (EGA) data portal, accession number EGAS50000000016, https://ega-archive.org/studies/EGAS50000000016. Processed gene expression data can be accessed from synapse, accession number syn54032669, https://doi.org/10.7303/syn54032669. Somatic alterations, recurrent mutations and HiChIP/ATACseq tracks can also be explored interactively using ProteinPaint41 and GenomePaint42 on St. Jude Cloud at https://viz.stjude.cloud/mullighan-lab/collection/the-genomic-basis-of-childhood-t-lineage-acute-lymphoblastic-leukemia~29. Recurrent mutation plots and oncoprints are provided as supplementary data.
