Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2025 Dec 31.
Published in final edited form as: Trends Genet. 2025 Dec 4;42(1):4–6. doi: 10.1016/j.tig.2025.11.007

AlphaGenome: a Swiss-army knife for exploring non-coding DNA

Judit García-González 1,2, Krzysztof Gogolewski 1,3
PMCID: PMC12750497  NIHMSID: NIHMS2125991  PMID: 41350205

Abstract

AlphaGenome, recently announced in a preprint, is Google DeepMind’s powerful “Swiss army knife” for predicting molecular effects from non-coding DNA. Remarkably, it does so with base-pair resolution while maintaining long-range context. Here, we discuss AlphaGenome’s promise and limitations in prioritizing and functionally interpreting non-coding variants underlying human traits and diseases.

Keywords: AlphaGenome, Non-coding DNA, Artificial Intelligence, Deep-learning models


Deciphering the three billion letters of the human genome has long been biology’s ultimate puzzle. With the advent of the genome-wide association study (GWAS) era, this task has gained renewed urgency, as thousands of variants associated with complex traits, such as blood pressure and lipid levels, and diseases, such as diabetes or schizophrenia, have been identified within non-coding regions of the genome but remain functionally uncharacterized. The functional characterization of these variants has thus become a central, rate-limiting step in translating GWAS findings into biological insight.

Comprising nearly 98% of the genome, non-coding regions can influence cellular function through diverse mechanisms such as altering transcription factor binding or chromatin accessibility, reshaping 3D genome interactions, or modulating splicing. Understanding how genetic variation affects these molecular processes, and ultimately traits and diseases, remains a daunting task [1]. Computational methods are stepping up to that challenge. In particular, deep learning-based sequence-to-function models, trained on existing experimental data, aim to predict the molecular consequences (e.g., splicing, gene expression, chromatin state) of DNA variation directly from the sequence. However, these models face two key limitations: First, while some models achieve base pair resolution, they overlook distal genomic interactions. Other models such as Enformer or Borzoi can capture such distal effects, but this is achieved at the cost of resolution. Second, most models are trained on a single or few modalities, requiring multiple pipelines to obtain a comprehensive view of the regulatory landscape of a locus.

AlphaGenome overcomes both limitations. Within a single model, it processes up to one megabase of DNA and can capture long-range effects, while maintaining inference at the single base resolution. Furthermore, it predicts 11 molecular modalities, including gene expression, splicing, chromatin state and contact maps, with state-of-the-art performance [2]. This versatility, akin to a “Swiss army knife” for predicting the molecular effects of DNA, arises from a deep learning architecture that combines convolutional layers to detect fine-scale detail, and transformers to capture long-range sequence context [2]. Previous models, such as Enformer or Borzoi, set the foundations for this architecture, but lack single-base resolution outputs and cannot process such long input sequences, limiting their capacity to model and predict long-range interactions. Available via API for non-commercial use (https://deepmind.google.com/science/alphagenome/), AlphaGenome enables researchers to generate comprehensive predictions of the molecular consequences of a particular DNA sequence variant.

Since AlphaGenome can predict how genetic variants affect molecular processes underlying human traits and diseases, an appealing extension is to explore its use for variant prioritization. Variant prioritization aims to identify the specific genetic variants that are most likely to be the true biological cause of a disease or trait, driving the association of a particular locus in a GWAS, or causing rare genetic diseases. It is a critical step toward the functional follow-up and understanding of disease biology.

Training and benchmarking deep learning sequence-to-function models for variant prioritization requires reliable positive controls (variants already known to cause a trait or disease). Since these are limited in number, variants associated with distinct complex diseases and traits must be aggregated, introducing heterogeneity that may decrease performance. Defining negative controls (variants unlikely to affect the trait) is equally complex. For this, AlphaGenome authors compared two strategies: one using variants with a low probability of causality, and another using a “matched” set of variants chosen to have similar baseline features (such as allele frequency and linkage disequilibrium) to the positive controls. To evaluate AlphaGenome, its predicted molecular effects were used as features for training two classifiers distinguishing positive from negative variants: a random forest, which builds multiple decision trees during training, and a zero-shot model, which makes inferences for previously unseen data using pre-existing knowledge. The results show that performance is generally better with the random forest than with the zero-shot model. More importantly, performance is highly sensitive to the choice of negative controls, showing how difficult it is to reliably solve the variant prioritization problem. Overall, this analysis, included as a Supplementary Note, highlights that robust variant prioritization remains an open challenge.

Complex traits and diseases emerge from gene-gene interactions, developmental dynamics, as well as cell-type specific regulatory programs [3]. AlphaGenome captures only part of this complexity, as it is primarily trained on large bulk-tissue functional genomics datasets such as GTEx, ENCODE, and FANTOM. Consequently, predictions for underrepresented tissues or rare cell types remain limited and would require model retraining. Performance also decreases for distal regulatory elements, which are frequently identified in GWAS. Even if AlphaGenome made perfect predictions across all tissues and cell types and was not limited for distal effects, fundamental questions would remain: which genes mediate these variant effects on phenotypes? which other genes or pathways are involved, and how do they interact? are the effects tissue-specific, developmental-stage dependent, or environmentally modulated?

Answering these questions will require groundbreaking multidisciplinary advances. In the meantime, AlphaGenome can still be a valuable tool for complex trait genomics. For instance, its comprehensive molecular predictions can inform and enhance statistical fine-mapping approaches. This has already proven to be a promising approach for casual variant prioritization [4]. AlphaGenome can also indirectly aid variant prioritization by predicting the regulatory effects of GWAS-associated variants in cellular models of disease. By systematically predicting the molecular consequences of all possible genetic variants within a GWAS-associated locus [2], AlphaGenome could guide labor-intensive post-GWAS experiments such as CRISPR-based functional validation [5]. AlphaGenome’s multimodal framework, which allows scoring variant effects across diverse molecular layers, can also be particularly valuable for dissecting complex loci, as demonstrated in the preprint, where it recapitulated the oncogenic effects of variants in the TAL1 locus [2]. The mechanistic insights informed by AlphaGenome’s predictions could complement gene function annotations and variant effect prediction scores like CADD [6], to improve variant and gene prioritization or assess their pathogenicity. While many challenges remain in understanding how genetic variation influences molecular processes and ultimately traits and diseases, AlphaGenome represents an important step toward bridging this gap. It encourages researchers to move beyond a mere annotation of variant function, and enables them the generation of testable hypotheses about the mechanisms orchestrating gene regulation.

AlphaGenome extends DeepMind’s growing toolkit for biology. It joins models such as AlphaFold, which predicts 3D protein structures and their interaction with other molecules [7,8]; AlphaMissense [9], which estimates the pathogenicity of missense mutations; and AlphaProteo [10], which designs de novo proteins. Together, these models advance our understanding of molecular processes across different levels of the central dogma of molecular biology: AlphaGenome and AlphaMissesense at the DNA and RNA level, tackling non-coding and coding variation, and AlphaFold and AlphaProteo at the protein level, tackling protein folding and interactions.

Despite these groundbreaking advances, a long road remains to understand how these molecular processes influence most human traits and diseases (Figure 1). For rare diseases, where genetic variants often have larger effects, the improved predictions of non-coding variant effects and other molecular processes could be clinically actionable sooner. However, for complex traits and diseases, further improvements, integrating single-cell, spatial, and temporal data will be essential to fully realize AlphaGenome’s potential to help researchers better understand their genetic underpinnings. Enhancing the model interpretability (e.g., by using intrinsic attribution methods that assign importance scores to individual variants) could also be valuable for researchers, since it may help identify the specific nucleotides or sequence motifs driving the model’s predictions. These improvements would require the public release of AlphaGenome’s code and model weights, along with user agreements that enable researchers re-training and fine-tuning of the model. We believe such open and collaborative developments will determine whether AlphaGenome becomes a truly accessible tool for a comprehensive non-coding variant prediction, for variant prioritization, and ultimately, for advancing our understanding of human traits and diseases.

Figure 1.

Figure 1.

AlphaGenome expands DeepMind’s biology toolkit, including AlphaFold, AlphaFold 3, AlphaMissense, and AlphaProteo, advancing molecular insights from DNA to protein. However, fully understanding how these molecular processes influence human traits and disease remains largely unknown.

Acknowledgments

This work was supported by the National Human Genome Research Institute (grant 1K99HG013547–01 to J.G.G.) and by the National Science Centre, Poland (SONATA grant 2023/51/D/NZ2/02892 to K.G.).

We thank Li Shen, Clive Hoggart, Max Drabkin, Laura Sloofman and Paulina Szymczak for their valuable feedback on the manuscript.

Footnotes

Publisher's Disclaimer: This is a PDF file of an unedited manuscript that has been accepted for publication. As a service to our customers we are providing this early version of the manuscript. The manuscript will undergo copyediting, typesetting, and review of the resulting proof before it is published in its final form. Please note that during the production process errors may be discovered which could affect the content, and all legal disclaimers that apply to the journal pertain.

Competing interests

The authors declare no competing interests.

References

  • 1.Monti R and Ohler U (2023) Toward Identification of Functional Sequences and Variants in Noncoding DNA. Annual Review of Biomedical Data Science 6, 191–210 [Google Scholar]
  • 2.Avsec Ž et al. (2025) AlphaGenome: advancing regulatory variant effect prediction with a unified DNA sequence model. bioRxiv, 2025.06.25.661532 [Google Scholar]
  • 3.Lappalainen T et al. (2024) Genetic and molecular architecture of complex traits. Cell 187, 1059–1075 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Srivastava D et al. (2025) Borzoi-informed fine mapping improves causal variant prioritization in complex trait GWAS. bioRxiv, 2025.07.09.663936 [Google Scholar]
  • 5.Morris JA et al. (2023) Discovery of target genes and pathways at GWAS loci by pooled single-cell CRISPR screens. Science 380, eadh7699. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Schubach M et al. (2024) CADD v1.7: using protein language models, regulatory CNNs and other nucleotide-level scores to improve genome-wide variant predictions. Nucleic Acids Res 52, D1143–D1154 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Jumper J et al. (2021) Highly accurate protein structure prediction with AlphaFold. Nature 596, 583–589 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Abramson J et al. (2024) Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature 630, 493–500 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Cheng J et al. (2023) Accurate proteome-wide missense variant effect prediction with AlphaMissense. Science 381, eadg7492. [DOI] [PubMed] [Google Scholar]
  • 10.Zambaldi V et al. (2024) De novo design of high-affinity protein binders with AlphaProteo. arXiv [Google Scholar]

RESOURCES