Skip to main content
BMC Bioinformatics logoLink to BMC Bioinformatics
. 2026 Jun 12;27:188. doi: 10.1186/s12859-026-06491-3

Protein language models are accidental taxonomists

Logan Hallee 1,2, Tamar Peleg 3, Nikolaos Rafailidis 1, Jason P Gleghorn 1,2,4,✉
PMCID: PMC13487993  PMID: 42286461

Abstract

Protein-protein interactions (PPIs) are fundamental to nearly all biological processes, yet their experimental characterization remains costly and time-consuming. While computational methods, particularly those using protein language models (pLMs), offer higher-throughput solutions, they often report unexpectedly high performance on multi-species datasets. Here, we introduce the accidental taxonomist hypothesis, proposing that neural networks can exploit the phylogenetic distances across labels in protein datasets rather than genuine interaction features. We show that in standard multi-species PPI datasets, positive pairs typically share a taxonomic origin, while randomly sampled negatives do not. We then demonstrate that pLM embeddings can be used to accurately distinguish whether two proteins share a taxonomic origin, allowing models to “cheat” by learning phylogeny instead of genuine PPI features. By employing a strategic sampling strategy that restricts negative examples to protein pairs from the same species, we reveal a marked drop in model performance, confirming our hypothesis. Compellingly, these strategically trained models still outperform single-species models, suggesting that multi-species data can improve performance if carefully curated. These findings suggest that accidental taxonomist behavior is a particularly influential confounder for PPI, and it is also broadly applicable to any supervised-learning protein dataset.

Supplementary Information

The online version contains supplementary material available at 10.1186/s12859-026-06491-3.

Keywords: Protein language modeling, Protein-protein interactions, Taxonomy, Phylogenetics, Negative sampling, Confounders, Reward hacking

Introduction

Proteins are ubiquitous macromolecules that drive biochemical reactions, compose biological structures, and enable physiological signaling, all from organized but straightforward physicochemical forces [1–11]. They often accomplish these vital tasks through intricate cross-talk with each other, known as Protein-Protein Interactions (PPIs). In its most basic form, a PPI involves two proteins exhibiting physical contact that mediates chemical or conformational change, especially with non-generic function [12–14]. At scale, paired PPIs can form complex networks and pathways, and are central to life sciences due to their implications in studying fundamental biology, disease, and potential therapeutics. However, traditional characterization of PPIs through in vitro and in vivo methods is expensive and time-consuming. Established higher-throughput methods like Yeast 2-Hybrid (Y2H) can identify thousands of PPIs once libraries are prepared at relatively low costs. More recent methods, such as Cell-Free Two-Hybrid (CF2H), enable the coupling of cell-free protein expression and precise detection capabilities that can correlate to binding affinity [15, 16]. Despite these advances, the scale of all possible human pairwise PPIs (over 200 million possible combinations from approximately 20,000 proteins, only a small fraction of which represent true interactions) necessitates a major improvement in throughput. A comprehensive, computable, and accurate representation of the interactome could serve as the base blueprint of cellular function. To this end, many groups are researching computational methods to predict PPIs from protein identifiers [17, 18]. Some methods use protein or gene IDs to search and extrapolate from established databases; others leverage protein structure solutions or predictions to infer PPIs, and many recent methods leverage Protein Language Models (pLMs) to infer PPIs from primary sequences [14, 19–22]. These platforms, particularly pLMs, offer an attractive solution with the possibility for massively high-throughput screening, amino acid- or atom-level analysis of PPI mechanisms, and inference of functionality [18].

pLMs have shown various levels of claimed success in predicting PPI. However, as with any machine learning model, there are many possible biases in dataset composition that can lead to inflated model performance. PPIs are an interesting machine learning and data compilation problem; while numerous PPI databases curated from the literature contain hundreds of thousands of real PPIs (and millions of likely ones), PPIs are extremely dependent on the biochemical context of the experiment. Factors such as post-transcriptional and post-translational modifications (PTMs), chaperone-mediated folding, protein cofactors, mutagenesis, pH, and many others ensure that recorded PPIs between two entities are rarely guaranteed to be universal true positives. Because of this, we reason that PPI machine learning models trained on general PPI data actually learn the relative plausibility of PPI, without accounting for biochemical context. In other words, we reason that pLM-based methods may infer that there exists a combination of biochemical factors, like PTMs and allosteric changes, that allow for a PPI without guaranteeing that the combination of characteristics ever occurs in reality. Therefore, positive PPI data and prediction come with a high level of ambiguity. Additionally, it is impossible to prove empirically that something cannot happen, so there are no “real” negative PPI examples to contrast against PPI positives. The Negatome project has improved the availability of extremely unlikely PPIs through manual analysis; however, even then, all of the Negatome−2.0 data splits include only 6,000 negative PPI protein pairs when combined [23]. This extreme data imbalance ensures that any binary prediction machine learning model trained on the unmodified available data will almost always predict “positive” [14, 24]. Previous popular negative curation methods leveraged information about protein subcellular compartment localization to estimate negatives, under the logic that proteins that localize to different cellular regions are unlikely to interact [25, 26]. While this logic is reasonable, proteins that do not interact in vivo due to their location may be able to interact in vitro, which is still important information in the coming age of protein design [14, 27]. In our earlier work around SYNTERACT, we posited that pLMs are accidental localizers: pLM features can be used to predict the subcellular compartment localization of proteins from their primary sequence with high accuracy [28–30]. Thus, pLM-based PPI models trained on real positives and subcellular negatives “cheat” and learn to predict the difference in localization, rather than PPI (or in addition to PPI) [14, 31].

More recently, papers by Bernett et al. established a “gold-standard” dataset built with different sequence-similarity-based biases in mind, attempting to control for pLMs memorizing domains or motifs [20]. Their balanced binary PPI dataset was constructed using positives from the Human Integrated Protein-Protein Interaction Reference (HIPPIE), and negatives generated through randomly selecting protein pairs to match the number of positives [20]. This random selection method for PPI negatives is currently the most popular approach, which leverages the assumption that randomly generated or paired amino acid sequences should not induce any meaningful PPI when expressed together, a good assumption when considering the highly specific nature of many PPIs. It also, at least partially, alleviates the accidental localizer issue by ensuring a mix of subcellular localization trends in the negative examples. Therefore, a model cannot purely rely on the localization confounder. Evaluation splits of the Bernett dataset were strategically partitioned so that there were no node degree biases, no sequence overlap (referred to as C3 by Park and Marcotte [32]), and a maximum sequence similarity of 40% between any of the train, validation, and test sequences. Works from Ko et al. and Bernett et al. have benchmarked many established pLM- or neural network-based PPI methods on this gold-standard dataset and have shown that performance is roughly limited to 0.7 AUC and 65% accuracy with current modeling capabilities [21, 22, 33, 34].

Yet, we make a simple observation that on multi-species datasets, unlike Bernett’s gold-standard human-protein-only PPI dataset, pLM-based methods showcase much higher performance metrics on withheld data. Even our SYNTERACT model, trained on a modified version of BioGRID, showcased over 90% accuracy on our withheld test set [14, 35]. In Ko et al.’s paper, we also observe high AUC scores ranging from 0.8 to over 0.9 with multi-species performance (STRING-db v11, C3, 40% CD-HIT) [22]. We believe this trend holds for much of the literature [25, 34, 36–41]. Recent work has also shown that pre-training data leakage from pLMs can inflate PPI scores, though the effect is modest (approximately 3% AUROC) and distinct from the confounders we explore here [42]. Of course, to a certain degree, higher multi-species performance is expected, as more species means more data and sequence diversity, and machine learning models should improve accordingly. However, the performance difference is staggering.

In this paper, we highlight observations and experiments that explore PPI performance discrepancies and put forward a hypothesis for the suspiciously high multi-species PPI performance reported in the literature. Specifically, we expand upon work analyzing the encoded phylogenetic relationships in pLM embeddings by demonstrating that neural networks trained on pLM embeddings can predict the taxonomic origin of protein primary sequences, even at extreme levels of discrimination [43–45]. We further postulate and support that it should be even easier than predicting the precise taxonomic classification of one protein to distinguish whether two proteins share a taxonomic origin. If it is possible for pLM-based methods to model the phylogenetic distance of two protein sequences, it implies that PPI models trained on multi-species data could be inadvertently learning taxonomic or phylogenetic features, and not PPIs. We refer to this hypothesis as the “accidental taxonomist,” where multi-species PPI models exploit the phylogenetic distance of negative examples to artificially inflate their performance. We further support our hypothesis by training a variety of pLM-based models in two conditions, where (1) the negatives are sampled randomly from BioGRID and (2) the negatives are sampled randomly from within the same species of BioGRID data. We show on rigorously split evaluation sets, similar to Bernett’s gold-standard (C3, 40% CD-HIT), that option (1) will “cheat” and option (2) cannot; implying a straightforward solution to this problem through strategic sampling.

Unfortunately, while we uncover the accidental taxonomist for PPI prediction, it is also potentially a key confounder in any supervised protein dataset, where certain classes or sections of labels may belong to a phylogenetic region, thereby compromising the underlying performance of the model despite reasonable evaluation set metrics. We intend for researchers to carefully consider phylogenetic distances and taxonomic biases to establish standardized controls that can be implemented to improve the performance and utility of pLM-based supervised learning.

Results

Multi-species PPI datasets are subject to inherent phylogenetic biases

BioGRID is a curated database of experimentally validated protein interactions containing only positive PPI edges [35]. A common approach for generating negative examples is to randomly pair proteins that are not known to interact. To characterize the taxonomic composition of such randomly generated negatives, we independently shuffled the protein A and protein B columns of BioGRID 100 times, effectively simulating random negative pairing. The resulting randomly paired examples shared the same organism of origin only about 31% of the time on average (Fig. 1). In contrast, nearly all real (positive) PPIs in BioGRID involve two proteins from the same species, since interactions are typically characterized within a single organism. Therefore, assuming a deep learning model could perfectly predict whether two proteins share a taxonomic origin and using the BioGRID ratios as reference, the perfect taxonomic classifier would identify positive examples 96.5% correctly (both proteins are intra-species) and negative examples with 68.2% accuracy (by predicting “negative” whenever it detects an inter-species pair). On a balanced dataset of 100,000 examples, a simple simulation reveals that this taxonomic shortcut alone would yield 82% accuracy, an F1 score of 0.85, an MCC of 0.67, and an AUROC of 0.82. More metrics and associated plots can be found in Supplemental Fig. S1.

Fig. 1.

Fig. 1

Taxonomic composition of positive PPI pairs versus randomly generated negatives in BioGRID. A Proportion of intra-species and inter-species pairs among positives (96.5% intra-species) versus shuffled negatives (31.8% ± 0.02% intra-species, mean ± s.d. over 100 shuffles). B Radial bar chart comparing organism pair frequencies (log scale) between positives (blue) and shuffled negatives (orange). Positives are dominated by intra-species pairs (Hs-Hs, Sc-Sc, Ec-Ec), while negatives are dominated by inter-species pairs (Hs-Sc, Ec-Hs). Species abbreviations: Hs = H. sapiens, Sc = S. cerevisiae, Ec = E. coli, Sp = S. pombe, At = A. thaliana, Mm = M. musculus, Dm = D. melanogaster, Ce = C. elegans, SCV2 = SARS-CoV-2

pLM features can be used to predict precise taxonomic origin

We benchmarked several popular pLMs on large protein datasets with labels based on taxonomic discrimination, ranging from domain to species. Datasets were embedded and fed to MLP probes, highlighting the intrinsic correlation between the embeddings and the tasks. We also trained on a similar balanced binary classification dataset where pairs of sequences were either from the same species or not. Performance on these compiled datasets is showcased in Fig. 2.

Fig. 2.

Fig. 2

Performance heatmap of all models and datasets evaluated comparing pLM embedding correlations to taxonomic discrimination. Weighted F1 scores are presented for each model-dataset combination, chosen because of the large class imbalance. The different datasets constructed by taxonomy are arranged along the rows, with individual model performance within each column

The random vector column in Fig. 2 shows truly random performance across the test datasets, which varies with the number of classes in each dataset. For multi-class tasks, random vectors yield near-zero weighted F1 because the probe collapses to predicting a single dominant class; for the balanced binary taxonomyInline graphic task, the weighted F1 of 0.33 reflects the same single-class prediction behavior (predicting one class on a balanced binary dataset yields precision of 0.5, recall of 0.5 for the predicted class and 0.0 for the other, resulting in a weighted F1 of 0.33). A transformer with random weight initialization and one-hot encoding demonstrated basic performance gains from the addition of basic sequence features alone, also known as a homology controls [30]. All other probes were fed embeddings from a pLM that had undergone substantial semi-supervised denoising, which clearly enhanced the feature vectors for taxonomic classification, as evidenced by test F1 scores relative to random and homology controls.

Highlighted in Fig. 2, most pLM embeddings could be used to almost perfectly predict sequences with vast phylogenetic distances at the domain level in taxonomyInline graphic (corresponding to Archaea, Bacteria, and Eukarya), with the best performer at 0.97 F1 (ProtT5). At the kingdom level, taxonomyInline graphic, this performance dropped to 0.69 F1 (ESMCInline graphic), 0.49 F1 (ProtT5) at phylum, 0.36 at order (ESMCInline graphic), 0.28 at family (ESMCInline graphic), and 0.21 at genus (ESMCInline graphic). At the most discriminative level, taxonomyInline graphic (among 618 species classes), the best-performing pLM embeddings achieved significantly better-than-random performance, with an F1 score of 0.18, compared with 0.00.

While taxonomyInline graphic and taxonomyInline graphic are different task formulations (multi-class classification of a single protein vs. binary classification of a protein pair), they probe the same underlying capability: the degree to which pLM embeddings encode phylogenetic information. The key insight is that even when precise species classification is unreliable (0.18 F1), the binary same-vs-different discrimination is highly accurate. The model probes performed much better on taxonomyInline graphic, identifying if two input sequences share a species or not, compared to precise species classification in taxonomyInline graphic. Most pLMs exceeded 0.8 F1 on this task, with ESMCInline graphic embeddings yielding the best F1 score of 0.87. This is directly relevant to PPI prediction, where the model receives two protein sequences and must produce a binary output, mirroring the taxonomyInline graphic formulation. Because pLMs achieve near-perfect domain-level discrimination (0.97 F1) even in the harder single-protein multi-class setting, paired binary tasks at coarser taxonomic levels would trivially approach perfect performance, making species-level discrimination the most conservative and informative test of the accidental taxonomist hypothesis.

Strategic sampling prevents “cheating” on PPI sets

Having demonstrated that pLM embeddings encode sufficient phylogenetic information to reliably distinguish whether two proteins share a species of origin, we next tested whether this capability actually inflates PPI prediction performance. Specifically, we compared two negative sampling strategies on a multi-species PPI dataset derived from BioGRID to determine whether models exploit phylogenetic distance as a shortcut for PPI classification.

We attempted to rectify the accidental taxonomist behavior by never feeding a “negative” PPI to a model from a different species. Instead, we randomly selected sequence pairs from within the same species to construct negatives; we refer to this approach as Strategic Sampling (SS). This simple change from the Normal Sampling (NS) approach, in which negatives are sampled randomly from the entire multi-species dataset, yielded substantially different performance metrics on the validation and test sets (Fig. 3). The validation sets follow the same sampling as the training set (NS validation for NS models, SS validation for SS models), but both models are evaluated on an SS test set to reveal whether the NS model has learned to cheat. The NS models averaged 0.71 MCC on the NS validation set, compared with 0.39 MCC for the SS models on the SS validation set, indicating that the NS models rely heavily on phylogenetic distance measurements. This is confirmed on the shared SS test set, where the NS model achieved only a 0.23 MCC, compared to 0.34 MCC for the SS models, a significantly lower performance from the NS model, as it is unable to effectively exploit phylogenetic distances on intra-species negatives. Both differences are statistically significant, with a t-test p < 0.0001 for each (n = 5). Notably, the 0.23 MCC score from NS is much better than random performance (equivalent to 62% accuracy, 0.61 F1 score, and 0.63 AUROC), meaning that the NS models do not solely rely on phylogenetic distances and appear to still learn aspects of PPI. Excitingly, the best test SS MCC of 0.37 is approximately 23% higher than the current best MCC score on Bernett’s single-species dataset (0.30 MCC). While these datasets differ in their underlying data sources (BioGRID vs. HIPPIE), approximately 45% of Bernett’s positive pairs overlap with BioGRID, and both employ the same evaluation rigor (C3 splitting, 40% sequence identity clustering). This comparison is therefore suggestive, though not definitive, that incorporating multi-species data with proper phylogenetic controls is an effective way to enhance PPI prediction performance and generalizability.

Fig. 3.

Fig. 3

Comparison of SS (blue) and NS (orange) approaches with Violin plots of MCC scores on the validation and test sets over five PPI training runs. The five NS and SS score distribution differences are highly statistically significant, with NS greatly outperforming SS on the validation set and underperforming on the test set

Training dynamics reveal distinct reward hacking

We analyzed the training dynamics of the NS and SS training runs by tracking the training loss, average predicted probability of PPI for positive examples, and average predicted probability of PPI for negative examples. After calculating a time-weighted exponential moving average and the 99.99% confidence interval (Fig. 4), these data reveal that the training trajectories of NS and SS are completely disjoint.

Fig. 4.

Fig. 4

Comparison of training dynamics between SS (blue) and NS (orange) across five training runs. Left: Cross-entropy training loss over training steps, with the random baseline (ln 2 ≈ 0.693) shown as a dotted line. Middle: Average predicted probability of PPI for positive examples, showing NS models drive predictions to 0.94 while SS models plateau at 0.75. Right: Average predicted probability for negative examples, where both approaches converge to approximately 0.5. Shaded regions represent 99.99% confidence intervals; faint black lines show individual seed trajectories. The divergent positive prediction trajectories reveal reward hacking behavior in NS models

In particular, the average positive PPI probability approaches 0.94 for NS compared to 0.75 for SS, presenting clear reward hacking behavior [46], as the loss was driven extremely low by increasing the associated logit when the NS models measure a zero phylogenetic distance. Interestingly, the average negative PPI prediction probabilities are remarkably similar, approximately 0.5, implying that negatives are much harder for PPI models to classify for NS and SS models.

Inter-species interactions are invisible regardless of sampling

A natural question arising from the accidental taxonomist hypothesis is what happens to inter-species PPI prediction. To address this, we trained SS and NS models on BioGRID Multi Validated (BioGRID-MV) (version 5.0.253) [35], a subsection of BioGRID where all examples have been experimentally validated by at least two methods from at least two sources. BioGRID-MV also contains a significant amount of inter-species positives, 35,970 ( 7%) of the 531k total pairs. We trained new NS and SS models and evaluated both models on a shared test set containing inter-species positives and intra-species positives with SS negatives.

The test set contained 2,815 examples (2,762 positives, 53 negatives), including 1,091 inter-species positives dominated by human-SARS-CoV-2 pairs (982 of 1,091). The extreme positive-to-negative ratio reflects that inter-species PPIs in BioGRID are predominantly viral-host interactions, for which few intra-species (SS) negatives can be generated from the limited set of viral proteins. On the full test set, the SS model achieved an AUROC of 0.80 (95% CI: 0.73-−0.86) compared to 0.62 (95% CI: 0.56-−0.67) for the NS model (DeLong test: Z = 3.60, Inline graphic; Fig. 5).

Fig. 5.

Fig. 5

ROC curves comparing SS (blue) and NS (orange) models on the full BIOGRID-MV inter-species experiment test set (Inline graphic). Shaded regions indicate 95% bootstrap confidence intervals. The SS model achieves significantly higher discriminative ability (AUROC = 0.80) than the NS model (AUROC = 0.62; DeLong Inline graphic), consistent with the accidental taxonomist effect inflating NS validation performance while degrading generalization

The inter-species subgroup reveals the most striking result: the SS model classified zero of 1,091 inter-species positives as positive interactions (0% recall), and the NS model classified only two (0.2% recall). Both models assigned inter-species positives a mean predicted probability near 0.46, barely distinguishable from a coin flip, indicating that inter-species pairs are effectively invisible to both sampling strategies. On intra-species positives, the SS model substantially outperformed the NS model (AUROC 0.81 vs. 0.49; DeLong Z = 11.37, Inline graphic; Supplemental Fig. S4), consistent with our earlier findings.

Discussion

Computational PPI prediction is a vital problem due to its widespread applicability across life science domains. With performance plateauing around 65% accuracy and 0.3 MCC on rigorously constructed single-species datasets [20, 21], the strikingly higher performance reported on multi-species datasets demands explanation. As we demonstrated, the taxonomic composition of randomly generated negatives in multi-species datasets such as BioGRID creates a shortcut that closely aligns with reported multi-species results from the literature [22, 25, 34, 37–41]. It is worth noting that BioGRID contains over 44% human data, with other large percentages coming from model organisms. Much larger and more diverse repositories, like STRING-db, would exacerbate this issue further, with an even smaller fraction of protein pairs coming from the same species if the entire dataset were shuffled.

To explore the capabilities of pLM-based methods to distinguish the taxonomic origin of primary sequences, we expanded upon existing work [43, 44] on pLM embeddings and encoded phylogenetic distances to characterize how well various neural network probes can classify taxonomic origin across a wide range of discrimination, from broad domain to species. Unsurprisingly, the best-performing model on average, ESMCInline graphic [47, 48], achieved almost perfect performance in predicting sequences with vast phylogenetic distances at the domain level in taxonomyInline graphic (corresponding to Archaea, Bacteria, and Eukarya). We observed a gradual decline in overall performance as the phylogenetic distance of the labels decreased. Interestingly, at the most discriminative label set in taxonomyInline graphic, among 618 species classes, ESMCInline graphic maintained a much better-than-random performance of 0.18 F1 score (vs. 0.00 random), and sometimes the exact species of origin could be determined from only a single primary sequence. Considering the training dataset utilized sequences that were maximally 40% similar to the validation and test sets, we were surprised to achieve better-than-random performance at the species level and encourage future work to explore which specific combination of learned features from semi-supervised denoising drives this performance. Nevertheless, despite being better than random, a 0.18 F1 score is far from reliable performance.

However, consistent with our hypothesis, we demonstrated that when fed two sequences in taxonomyInline graphic, aligned with the standard PPI-prediction workflow, neural networks could effectively predict whether each sequence came from a different species. In fact, most pLMs tested exceeded 0.8 F1 score on this balanced binary prediction task, with ESMCInline graphic achieving 0.87 F1. Even the homology controls exceed 0.7 F1. Thus, although neural networks cannot predict the exact species from pLM embeddings, they can reliably detect whether two sequences have a nonzero phylogenetic distance. We refer to this finding as the accidental taxonomist, where pLM-based methods exploit phylogenetic distance distributions across classes to inflate their performance. We reasoned that this should be particularly applicable to PPI datasets, where typical datasets have pairs labeled “1” with no phylogenetic distance and most pairs labeled “0” with some distance, and confirmed this by comparing the performance of NS- and SS-trained models. In particular, the reward hacking associated with accidental taxonomist behavior is clearly evident in the evaluation metrics in Fig. 3 and in the training loss and logit distributions shown in Fig. 4.

We acknowledge that SS likely introduces slightly higher label noise than NS, as two random proteins from the same species are more likely to interact (but remain uncatalogued) than two proteins from different species. However, the magnitude of the performance collapse, from  0.7 MCC to  0.2 MCC on the SS test set, is too drastic to be explained by label noise alone. If label noise were the primary driver, we would expect the SS-trained models to suffer similarly, yet they maintain consistent performance across validation and test sets. This asymmetry strongly supports the accidental taxonomist hypothesis. We also note that, although we employed a custom transformer-based architecture rather than simpler alternatives to ensure our comparison was representative of state-of-the-art methods, the accidental taxonomist is fundamentally a data issue rather than a model issue; we expect similar NS/SS discrepancies with any reasonable architecture.

Importantly, despite comparable dataset construction rigor, our multi-species SS PPI models all exceed the 0.3 MCC plateau for single-species PPI, suggesting that adding sequence diversity may improve PPI performance and generalizability when properly controlled. For fairness, we note that our multi-species dataset is substantially larger than typical single-species datasets, and part of this improvement could be attributable to scale rather than phylogenetic diversity alone. Future work with size-matched controls would be needed to fully disentangle these effects, but for our purposes it is clear that the training loss curves between NS and SS have exceedingly different trends early into training, well within the size of typical training datasets and runs. More reported evaluation metrics from our runs are shown in Supplemental Figs. S2 and S3.

From this work, our previous identification of accidental localizers, and the outstanding contributions of Bernett, Ko, Rost, and many others, exploring the flaws and biases behind PPI prediction, we can safely assume that very high PPI performance metrics (for example above 0.95 AUC, 0.8 MCC, 0.9 F1, 90% accuracy, etc.) are likely due to some form of data leakage or underlying confounders, in addition to (or instead of) genuine learning of PPI. High-throughput computational PPI remains an unsolved problem, and dataset construction and curation challenges remain major obstacles.

Our findings also have direct implications for host-pathogen PPI prediction. Our inter-species experiment on BioGRID-MV confirmed that both NS and SS models exhibit near-zero recall on inter-species interactions (Fig. 5): neither model correctly identified more than 0.2% of inter-species positives, which were dominated by human-SARS-CoV-2 pairs (supported by [42]). This represents a significant blind spot, as viral entry mechanisms and bacterial effector-host protein interactions are inter-species PPIs of great biomedical interest. Importantly, even the SS model, which cannot exploit phylogenetic distance, fails on inter-species positives, suggesting that dedicated inter-species training data and evaluation protocols may be necessary for host-pathogen applications.

Unfortunately, our analysis underscores a more fundamental flaw in pLM dataset compilation beyond PPI prediction. For example, one common application of pLM embeddings is predicting Enzyme Commission (EC) numbers, a categorization scheme for different enzymatic reactions. If a documented EC number occurs only in a specific taxonomic family, a model trained on such data may inadvertently learn to classify by taxonomy rather than by enzymatic function. This same principle applies to any supervised protein-learning task, including regression, where a specific label range could belong exclusively to a taxonomic range.

It is conceptually intuitive to imagine an attention-based neural network exploiting confounders during multiple-input tasks, such as PPI, as these models directly compare two inputs via dot product [49]. However, even simpler neural networks and single-protein input datasets are not immune to the accidental taxonomist effect. Neural networks are capable of storing various training distributions into their weights via gradient descent [50–52]. For instance, a simple feed-forward network could organize one layer of weights to correspond to the average taxonomic features for each class, causing inputs with specific matching taxonomic features to produce higher values after this layer.

Additionally, we suspect that genomic language models, leveraging DNA or codon vocabularies to study proteins, will be even more susceptible to the accidental taxonomist phenomenon than amino acid–based pLMs. Codon usage bias is a powerful feature for predicting taxonomic origin on its own [45, 53], adding another layer of exploitable signal on top of the amino acid-level features already present in pLM embeddings.

Finally, there may be other “accidental _______” phenomena where unintended physical, chemical, biological, or even social trends, surrounding which proteins researchers choose to study, distract machine learning models from learning the underlying categorizations and principles of interest. We hope to see future work exploring how neural networks “cheat” on protein datasets and how improved dataset design and controls can mitigate these confounders.

Methods

Prototyping phylogenetic bias impact

The February 2025 release of BioGRID (4.4.243) was downloaded (BIOGRID-ALL−4.4.243.tab3.zip), and protein identifiers were mapped to UniProt accessions, with Swiss-Prot prioritized. IDs were mapped to sequences using the UniProt ID mapper tool https://www.uniprot.org/id-mapping and the primary sequence and organism of origin were recorded for each protein pair [35, 54]. The six columns in our dataset include each UniProt ID (A, B), each sequence (SeqA, SeqB), and each organism (OrgA, OrgB). To characterize the taxonomic composition of randomly generated negatives, we independently shuffled the protein A and protein B columns (and their associated organism labels) 100 times, which simulates the random protein pairing commonly used for negative generation. For each shuffle, we recorded the fraction of randomly paired proteins that shared the same organism of origin, as shown in Fig. 1. This analysis quantifies what fraction of random pairings would be intra-species vs. inter-species, revealing the taxonomic imbalance that models can exploit if a dataset used uncontrolled random sequence pairing to generate PPI negatives. We simulated the classification metrics for a perfect OrgA = OrgB vs. OrgA ≠ OrgB classifier by building random predictions and labels. The 100,000 data points consisted of 50% labeled one and 50% labeled zero, where 31% of the zeros were false negatives per the results of the BioGRID shuffling. The actual predicted probabilities were randomly generated from a uniform distribution between 0.51 and 1.00 for true positives, 0.00 to 0.49 for true negatives, and 0.51 to 1.00 for false negatives, allowing for the estimation of the ROC AUC and PR AUC. Accuracy, precision, recall, F1, and MCC were also calculated (Supplemental Fig. S1).

Probing taxonomic features

We constructed two types of taxonomic probing tasks. The first type (taxonomyInline graphic through taxonomyInline graphic) is a single-protein multi-class classification task: given one protein sequence, predict its taxonomic label. The second type (taxonomyInline graphic) is a paired-protein binary classification task: given two protein sequences, predict whether they originate from the same species or not (i.e., whether protein A comes from species X and protein B comes from species Y, where X ≠ Y). The paired formulation directly mirrors the input structure of PPI prediction.

For the single-protein tasks, datasets at different taxonomic discrimination levels were constructed from Swiss-Prot, downloaded on 7/22/2025 (release 2025_03) [54]. We kept the taxonomic lineage column to build labels and retained sequences of lengths greater than 20 and less than 2048. Sequences were deduplicated and clustered at 40% sequence identity with a word size of two using CD-HIT [55], and clusters were randomly assigned to validation and test splits until their count exceeded 10,000 entries, respectively. For each taxonomic discrimination, we only kept labels that had at least 100 examples total, resulting in three domain, 17 kingdom, 68 phylum, 126 class, 248 order, 371 family, 545 genus, and 618 species classes. The total training set sizes consisted of 444k domain, 450k kingdom, 458k phylum, 451k class, 446k order, 434k family, 419k genus, and 265k species examples. We enforced that there was no sequence or cluster overlap between any splits to ensure evaluation rigor. The datasets are referred to as taxonomyInline graphic, taxonomyInline graphic, taxonomyInline graphic, taxonomyInline graphic, taxonomyInline graphic, taxonomyInline graphic, taxonomyInline graphic, and taxonomyInline graphic throughout the work for clarity of the discrimination level.

The taxonomyInline graphic dataset was compiled by taking taxonomyInline graphic and randomly constructing protein pairs from within each data split. 50,000 random sequences from the same species were sampled for the training split, 500 for validation, and 500 for testing. Then, the dataset was balanced by randomly selecting sequences from different species. The final dataset, comprising 100,000 training, 1000 validation, and 1000 test sequences, presented a binary prediction problem: determining if two sequences had any phylogenetic distance between them, where protein pairs were labeled as zero if they were from the same species and one otherwise.

Models and datasets were probed using the Protify package, an open source project aimed at standardized, effective, and high-throughput analyses with chemical language models [56, 57]. We embedded all datasets using mean and variance pooling on the last hidden state of various pLMs, resulting in a consistently sized vector for each protein describing the distribution of the models’ last hidden state. For single-protein tasks (taxonomyInline graphic through taxonomyInline graphic), each example consisted of one pooled vector with o equal to the number of taxonomic classes. For the paired-protein task (taxonomyInline graphic), each example consisted of two pooled vectors stacked together, with o = 2 for binary classification. We chose many popular pLMs for analysis, including ESM2 eight million parameter version (ESM2Inline graphic), ESM2Inline graphic, ESM2Inline graphic, ESM2Inline graphic, ESM2 three billion parameter version (ESM2Inline graphic), ESMCInline graphic, ESMCInline graphic, DSMInline graphic, ProtBERT (big fantastic dataset version), ProtT5-encInline graphic (ProtT5), ANKH-BaseInline graphic, and ANKH-LargeInline graphic [47, 57–60]. For the T5-based models, ANKH and ProtT5, only the encoders were used for embedding generation. Three controls were also used: random vectors, which embed each protein with a randomly generated vector (Inline graphic); random transformer, which embeds proteins through a re-initialized equivalent of ESM2Inline graphic; and OneHot protein, which replaces the pLM embeddings with one-hot encoded strings [57].

Within Protify, MLP probes were trained with the architecture:

graphic file with name d33e1027.gif

where Inline graphic denotes Layer Normalization, mapping input Inline graphic (with d being the embedding dimensionality of the pLM) to output Inline graphic where o was the number of classes in the dataset. Inline graphic, Inline graphic, Inline graphic, Inline graphic, b corresponded to the associated bias term in the linear layer, σ was ReLU [61], h was 8,192, p was calculated with

graphic file with name d33e1083.gif

and a dropout of 20% was applied before every MatMul, except the first one [62]. We optimized cross-entropy loss with a batch size of 64, learning rate of Inline graphic, AdamW optimizer, weight decay of 0.01, and evaluation on the validation set until a patience of five was exceeded, comparing the validation loss. The best weights were kept and evaluated on the test set [56].

Strategic sampling experiment

The protein sequences in BioGRID 4.4.243 were deduplicated and clustered at 40% sequence identity with a word size of two using CD-HIT [55]. Clusters were added to a list until their paired sequences exceeded 5,000 BioGRID examples for validation and testing, respectively. We ensured that every data split was disjoint in terms of sequence and cluster ID, following the C3 strategy from Park and Marcotte [32]. Negatives were generated within each data split by either sampling random proteins from the clusters (Normal Sampling / NS) or sampling random proteins from the clusters that are from the same organism (Strategic Sampling / SS). Each candidate negative pair was cross-referenced against the full set of known positive interactions; if a candidate pair was found among the positives, it was discarded and resampled to avoid false negatives. The final datasets for NS and SS both consisted of 4,523,432 training examples, 10,070 validation examples, and 10,034 test examples, with a perfectly balanced label distribution; the organism distributions between the positive and negative examples were approximately the same, and the datasets used a total of 76,045 unique protein sequences.

We trained a hierarchical transformer neural network (Fig. 6) to map pLM embeddings to binary logits. The embedding model of choice was ESMCInline graphic [47, 48], which we froze and extracted the last hidden state for each sequence, truncating it to a length of 512 to minimize computational expense. This followed recent findings that larger maximum lengths have a minimal effect on PPI performance [42]. Because these are not intended as production models and are used for comparison between NS and SS, this truncation was acceptable.

Fig. 6.

Fig. 6

Architecture of the PPI prediction model. Two protein sequences are embedded by a frozen ESMCInline graphic protein language model, producing per-residue embeddings. Each embedding passes through a separate encoder track (non-Siamese) consisting of a linear projection (1152 → 512), a transformer block with SwiGLU feed-forward networks, and a cross-attention pooler that compresses variable-length sequences into 32 fixed tokens Inline graphic. The pooled representations are concatenated (Inline graphic) and processed by four transformer blocks with progressive dimensionality reduction (Inline graphic) via interleaved projection layers (LayerNorm → Linear → GELU → Dropout). The final Inline graphic representation is mean-pooled and passed to a linear classification head. Input orientation (A, B) is randomly swapped with 50% probability during training

For each sequence A and B, there was a separate encoder track (not shared-weights / siamese network) consisting of a linear layer projection from ESMCInline graphic hidden size of 1,152 to 512, one transformer block with hidden size 512, and then token-parameter cross-attention, which pooled every protein sequence into consistently sized matrices Inline graphic, Inline graphic [63, 64]. During training, the input orientation (A, B) was randomly swapped with 50% probability to prevent order-dependent biases; the separate encoder tracks exist to give each protein its own representational subspace, which we believe is particularly important for homodimer modeling. Averaging logits from both orientations of the independent tracks at inference is a technique we use in local experiments to marginally improve performance but did not significantly affect results in our experiments. A and B were then concatenated into Inline graphic, which served as the input to the rest of the model. We suspect it is better to concatenate the protein representations instead of combining elementwise, which is popular in the literature [36, 39], enabling the final layers the ability to leverage independent information from both entities. The remainder of the model consisted of four transformer blocks, interleaved with linear layers, which reduced the dimensionality from Inline graphic to Inline graphic, projecting down by a factor of two at every transformer block. The final Inline graphic matrix was mean-pooled into a Inline graphic vector and fed to a final linear layer to build the logits Inline graphic. An expansion ratio of Inline graphic, rounded to the nearest multiple of 256, was used for all transformer block MLPs and all linear layers (aside from attention layers) used a dropout of 0.1. Binary cross-entropy was employed alongside the AdamW optimizer [65], a batch size of 128, learning rate of Inline graphic, evaluation on the validation set every 5000 steps, and all models were trained for one epoch on a single GH200 GPU; the weights with the best validation MCC were kept and test scores were reported.

In case each model curated slightly different decision boundaries, and different thresholds were optimal to split logits into positives and negatives, we split the precision-recall curve into 100 discrete points and tried each threshold. The threshold used for the metrics calculations was chosen based on which one produced the maximal F1 score. Technically, macro F1 was used for threshold selection and weighted averages were reported for F1, precision, and recall (Supplemental Fig. S3); however, because the datasets were balanced, this did not impact the final result.

Inter-species positives experiment

To evaluate model performance on inter-species PPIs, we used BIOGRID-MV version 5.0.253, a curated subset of BioGRID enriched for multi-species and viral-host interactions [35]. Protein sequences were clustered at 40% identity with 80% coverage using MMseqs2 [66]. An inter-species-aware C3 split was constructed by greedily assigning clusters to validation and test sets, prioritizing clusters containing inter-species positive pairs (where protein A and protein B originate from different species) until each split contained at least 500 total pairs and 100 inter-species positives. The remaining clusters formed the training set. Protein and cluster disjointness was enforced across all splits. Negatives were generated using the same procedure as the main experiment: SS negatives from within the same species and NS negatives from random pairing. Both models shared an identical SS test set to enable direct comparison. The same model architecture and hyperparameters described above were used. ROC curves were compared using the DeLong test for paired AUROCs, with 95% confidence intervals computed via the DeLong method and bootstrap confidence bands (2,000 replicates) visualized on the ROC plots.

Supplementary Material

Acknowledgements

The authors thank Katherine M. Nelson, Ph.D., for reviewing and commenting on drafts of the manuscript.

Author contributions

Conceived and designed the experiments (LH, JPG), performed the experiments (LH, TP, NR), analyzed the data (LH, TP, NR), contributed materials/analysis tools (LH, JPG), wrote the paper (LH, JPG), edited and reviewed the manuscript (LH, TP, NR, JPG).

Funding

This work was partly supported by the University of Delaware Graduate College through the Unidel Distinguished Graduate Scholar Award (LH), the National Science Foundation through NAIRR pilot 240064 (JPG), and the National Institutes of Health through NIGMS T32GM142603 (LH, NR), R01HL178817 (JPG), R01HL133163 (JPG), and R01HL145147 (JPG).

Data availability

Datasets are available at https://github.com/Gleghorn-Lab/PLMConfounders. Protify and its associated datasets are accessed at https://github.com/Gleghorn-Lab/Protify and on Hugging Face at https://huggingface.co/collections/GleghornLab/protify. The taxonomic classification datasets are available on Hugging Face: Domain (https://huggingface.co/datasets/GleghornLab/taxonomy_domain_0.4_clusters), Kingdom (https://huggingface.co/datasets/GleghornLab/taxonomy_kingdom_0.4_clusters), Phylum (https://huggingface.co/datasets/GleghornLab/taxonomy_phylum_0.4_clusters), Class (https://huggingface.co/datasets/GleghornLab/taxonomy_class_0.4_clusters), Order (https://huggingface.co/datasets/GleghornLab/taxonomy_order_0.4_clusters), Family (https://huggingface.co/datasets/GleghornLab/taxonomy_family_0.4_clusters), Genus (https://huggingface.co/datasets/GleghornLab/taxonomy_genus_0.4_clusters), Species (https://huggingface.co/datasets/GleghornLab/taxonomy_species_0.4_clusters), and Different (https://huggingface.co/datasets/GleghornLab/diff_phylo). The sequences and taxonomic origin were originally gathered from UniProt release 2025_03. PPI training, validation, and test datasets are compiled automatically at training time; the exact examples used in the study can be found at https://github.com/Gleghorn-Lab/PLMConfounders/tree/main/processed_datasets, sourced from BioGRID version 4.4.243. The BioGRID IDs, mapped to UniProt IDs and sequences, are also hosted on Hugging Face: https://huggingface.co/datasets/Synthyra/BIOGRID. Selected code is available at https://github.com/Gleghorn-Lab/PLMConfounders. Protify and its code can be accessed here: https://github.com/Gleghorn-Lab/Protify.

Code availability

Selected code is available at https://github.com/Gleghorn-Lab/PLMConfounders. Protify and its code can be accessed here: https://github.com/Gleghorn-Lab/Protify.

Declarations

Ethics approval and consent to participate

NA

Consent for publication

Not applicable.

Competing interests

LH and JPG are co-founders of and have an equity stake in Synthyra, PBLLC.

Footnotes

Publisher's Note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

References

  • 1.LaPelusa A, Kaushik R. Physiology, Proteins. StatPearls Publishing, Treasure Island (FL) 2025. http://www.ncbi.nlm.nih.gov/books/NBK555990/. [PubMed]
  • 2.Roberts RJ, Hallee L, Lam CK. The potential of hsp90 in targeting pathological pathways in cardiac diseases. J Pers Med. 2021;11(12):1373. 10.3390/jpm11121373. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Sen S, Hallee L, Lam CK. The potential of gamma secretase as a therapeutic target for cardiac diseases. J Pers Med. 2021;11(1212):1294. 10.3390/jpm11121294. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Millar-Haskell CS, Dang AM, Gleghorn JP. Coupling synthetic biology and programmable materials to construct complex tissue ecosystems. MRS Commun. 2019;9(2):421–32. 10.1557/mrc.2019.69. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Khan S, Sharifi M, Gleghorn JP, Babadaei MMN, Bloukh SH, Edis Z, et al. Artificial engineering of the protein corona at bio-nano interfaces for improved cancer-targeted nanotherapy. J Control Release. 2022;348:127–47. 10.1016/j.jconrel.2022.05.055. [DOI] [PubMed] [Google Scholar]
  • 6.Shen Y, Gleghorn JP. Class iii phosphatidylinositol-3 kinase/vacuolar protein sorting 34 in cardiovascular health and disease. J Cardiovasc Transl Res. 2025;18(2):392–407. 10.1007/s12265-024-10581-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Shen Y, Gonyea CR, Hallee L, Gleghorn JP. Akt-dependent vacuolar protein sorting 34 activation in pulmonary arterial vascular is a shared molecular mechanism in adult pulmonary hypertension and congenital diaphragmatic hernia-associated pulmonary vascular remodeling. Am J Respir Crit Care Med 211(Abstracts), 2025;2180–2180 10.1164/ajrccm.2025.211.Abstracts.A2180.
  • 8.Shirazi J, Donzanti MJ, Nelson KM, Zurakowski R, Fromen CA, Gleghorn JP. Significant unresolved questions and opportunities for bioengineering in understanding and treating covid-19 disease progression. Cell Mol Bioeng. 2020;13(4):259–84. 10.1007/s12195-020-00637-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Nelson KM, Minahan DJ, Edwards VL, Glomski IJ, Delgado Diaz DJ, Thomas K, Walker FC, Bavoil PM, Derré I, Criss AK, Ravel J, Gleghorn JP. A microphysiologic model of the cervical epithelium recapitulates microbial, immunologic, and pathogenic properties of sexually transmitted infections. bioRxiv: The Preprint Server for Biology, 2025;0721665989 10.1101/2025.07.21.665989.
  • 10" GrobidStyleApplied="Changed.Felsenstein J, Nelson KM, Shirazi J, Gleghorn JP. Therapeutic nanoparticle safety in pregnancy: Bridging knowledge gaps with environmental insights and a translational roadmap. J Control Release. 2025;385:114026. 10.1016/j.jconrel.2025.114026. [DOI] [PMC free article] [PubMed]
  • 11.Gilbert RM, Gleghorn JP. Connecting clinical, environmental, and genetic factors point to an essential role for vitamin a signaling in the pathogenesis of congenital diaphragmatic hernia. Am J Physiol Lung Cell Mol Physiol. 2023;324(4):456–67. 10.1152/ajplung.00349.2022. (Publisher: American Physiological Society. Accessed 2023-06-26). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12" GrobidStyleApplied="Changed.Soleymani F, Paquet E, Viktor H, Michalowski W, Spinello D. Protein-protein interaction prediction with deep learning: a comprehensive review. Comput Struct Biotechnol J. 2022;20:5316–41. 10.1016/j.csbj.2022.08.070. [DOI] [PMC free article] [PubMed]
  • 13.Hamp T, Rost B. Evolutionary profiles improve protein–protein interaction prediction from sequence. Bioinformatics. 2015;31(12):1945–50. 10.1093/bioinformatics/btv077. [DOI] [PubMed] [Google Scholar]
  • 14.Hallee L, Gleghorn JP. Protein-protein interaction prediction is achievable with large language models. bioRxiv. 2023. p. 2023–0607544109. 10.1101/2023.06.07.544109.
  • 15.Brückner A, Polge C, Lentze N, Auerbach D, Schlattner U. Yeast two-hybrid, a powerful tool for systems biology. Int J Mol Sci. 2009;10(6):2763–88. 10.3390/ijms10062763. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Capin J, Mayonove P, DeVisch A, Becher A, Ngo G, Courbet A, Ragotte RJ, Gonsaud MC, Espeut J, Bonnet J. Cf2h: a cell-free two-hybrid platform for rapid protein binder screening. bioRxiv, 2025;2025–0716665152 10.1101/2025.07.16.665152 [DOI] [PMC free article] [PubMed]
  • 17.Han J, Cen J, Wu L, Li Z, Kong X, Jiao R, et al. A survey of geometric graph neural networks: data structures, models and applications. Front Comp Sci. 2025;19(11):1911375. 10.1007/s11704-025-41426-w. [Google Scholar]
  • 18.Zhang J, Durham J, Cong Q. Revolutionizing protein-protein interaction prediction with deep learning. Curr Opin Struct Biol. 2024;85:102775. 10.1016/j.sbi.2024.102775. [DOI] [PubMed] [Google Scholar]
  • 19.Szklarczyk D, Kirsch R, Koutrouli M, Nastou K, Mehryary F, Hachilif R, et al. The string database in 2023: protein-protein association networks and functional enrichment analyses for any sequenced genome of interest. Nucleic Acids Res. 2023;51(D1):638–46. 10.1093/nar/gkac1000. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Bernett J, Blumenthal DB, List M. Cracking the black box of deep sequence-based protein–protein interaction prediction. Brief Bioinform. 2024;25(2):076. 10.1093/bib/bbae076. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Reim T, Hartebrodt A, Blumenthal DB, Bernett J, List M. Deep learning models for unbiased sequence-based ppi prediction plateau at an accuracy of 0.65. Bioinformatics. 2025;590–8. 10.1093/bioinformatics/btaf192. [DOI] [PMC free article] [PubMed]
  • 22.Ko YS, Parkinson J, Liu C, Wang W. Tuna: an uncertainty-aware transformer model for sequence-based protein–protein interaction prediction. Brief Bioinform. 2024;25(5):359. 10.1093/bib/bbae359. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Blohm P, Frishman G, Smialowski P, Goebels F, Wachinger B, Ruepp A, Frishman D. Negatome 2.0: a database of non-interacting proteins derived by literature mining, manual annotation and protein structure analysis. Nucleic Acids Res 2014;42(D1):396–400 10.1093/nar/gkt1079. [DOI] [PMC free article] [PubMed]
  • 24.Hamp T, Rost B. More challenges for machine-learning protein interactions. Bioinformatics. 2015;31(10):1521–5. 10.1093/bioinformatics/btu857. [DOI] [PubMed] [Google Scholar]
  • 25.Li X, Han P, Wang G, Chen W, Wang S, Song T. Sdnn-ppi: self-attention with deep neural network effect on protein-protein interaction prediction. BMC Genomics. 2022;23(11):1–14. 10.1186/s12864-022-08687-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Ben-Hur A, Noble WS. Choosing negative examples for the prediction of protein-protein interactions. BMC Bioinformatics. 2006;7(S1):2. 10.1186/1471-2105-7-S1-S2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Cotet T-S, Krawczuk I, Stocco F, Ferruz N, Gitter A, Kurumida Y, et al. Crowdsourced protein design: Lessons from the adaptyv egfr binder competition. bioRxiv. 2025. p. 2025–0417648362. 10.1101/2025.04.17.648362.
  • 28.Stärk H, Dallago C, Heinzinger M, Rost B. Light attention predicts protein location from the language of life. Bioinform Adv. 2021;1(1):035. 10.1093/bioadv/vbab035. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Thumuluri V, Almagro Armenteros JJ, Johansen A, Nielsen H, Winther O. Deeploc 2.0: multi-label subcellular localization prediction using protein language models. Nucleic Acids Res. 2022;50(W1):228–34. 10.1093/nar/gkac278. [DOI] [PMC free article] [PubMed]
  • 30.Hallee L, Rafailidis N, Horger C, Hong D, Gleghorn JP. Annotation vocabulary (might be) all you need. bioRxiv. 2024. p. 2024–0730605924. 10.1101/2024.07.30.605924.
  • 31.Dubourg-Felonneau G, Wesego DM, Akiva E, Varadan R. Pinui: A dataset of protein–protein interactions for machine learning. bioRxiv, 2023;2023–1212571298 10.1101/2023.12.12.571298.
  • 32.Park Y, Marcotte EM. Flaws in evaluation schemes for pair-input computational predictions. Nat Methods. 2012;9(12):1134–6. 10.1038/nmeth.2259. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Tartici A, Nayar G, Altman RB. Pool parti: A pagerank-based pooling method for robust protein sequence representation in deep learning. bioRxiv. 2024. p. 2024–1004616701. 10.1101/2024.10.04.616701.
  • 34.Piochi LF, Tang D, Malmström J, Karami Y, Khakzad H. ppiris: deep learning for proteome-wide prediction of bacterial protein-protein interactions. bioRxiv. 2025. p. 2025–0922677885. 10.1101/2025.09.22.677885.
  • 35.Oughtred R, Rust J, Chang C, Breitkreutz B-J, Stark C, Willems A, et al. The biogrid database: A comprehensive biomedical resource of curated protein, genetic, and chemical interactions. Protein Sci A Publication of the Protein Society. 2021;30(1):187–200. 10.1002/pro.3978. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Sledzieski S, Singh R, Cowen L, Berger B. D-script translates genome to phenome with sequence-based, structure-aware, genome-scale predictions of protein-protein interactions. Cell Syst. 2021;12(10):969–9826. 10.1016/j.cels.2021.08.010. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Zheng X, Du H, Xu F, Li J, Liu Z, Wang W, Chen T, Ouyang W, Li SZ, Lu Y, Dong N, Zhang Y. Pring: Rethinking protein-protein interaction prediction from pairs to graphs. arXiv 2025. 10.48550/arXiv.2507.05101.arXiv:2507.05101 [cs] [DOI]
  • 38.Zhao S, Cui Z, Zhang G, Gong Y, Su L. Mgppi: multiscale graph neural networks for explainable protein–protein interaction prediction. Front Genet. 2024;15:1440448 10.3389/fgene.2024.1440448. [DOI] [PMC free article] [PubMed]
  • 39.Szymborski J, Emad A. Intrepppid—an orthologue-informed quintuplet network for cross-species prediction of protein–protein interaction. Brief Bioinform. 2024;25(5):405. 10.1093/bib/bbae405. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Liu S, Liang B, Fang Y, Jiang Z, Xu R. Hierarchical multi-label contrastive learning for protein-protein interaction prediction across organisms. arXiv 2025. 10.48550/arXiv.2507.02724. arXiv:2507.02724 [cs]. [DOI] [PMC free article] [PubMed]
  • 41.Dang TH, Vu TA. xcapt5: protein–protein interaction prediction using deep and wide multi-kernel pooling convolutional neural networks with protein language model. BMC Bioinformatics. 2024;25(1):106. 10.1186/s12859-024-05725-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Szymborski J, Emad A. A flaw in using pretrained protein language models in protein–protein interaction inference models 8(2):197–208 10.1038/s42256-025-01176-7 Accessed 2026–04-03.
  • 43.Lupo U, Sgarbossa D, Bitbol A-F. Protein language models trained on multiple sequence alignments learn phylogenetic relationships. Nat Commun. 2022;13(1):6298. 10.1038/s41467-022-34032-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Tule S, Foley G, Bodén M. Do protein language models learn phylogeny? Brief Bioinform. 2025;26(1):047. 10.1093/bib/bbaf047. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Hallee L, Khomtchouk BB. Machine learning classifiers predict key genomic and evolutionary traits across the kingdoms of life. Sci Rep. 2023;13(1):2088. 10.1038/s41598-023-28965-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Skalse J, Howe NHR, Krasheninnikov D, Krueger D. Defining and characterizing reward hacking 2025. 10.48550/arXiv.2209.13085. arXiv:2209.13085 [cs]
  • 47.ESM Team: Esm cambrian: Revealing the mysteries of proteins with unsupervised learning. Evolutionary Scale Blog 2024. Evolutionary Scale, blog post.
  • 48.Hallee L, Bichara D, Gleghorn JP. Esm++: Efficient and hugging face compatible versions of the esm cambrian models. Hugging Face 2024. 10.57967/hf/3726.
  • 49.Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser L, Polosukhin I. Attention is all you need. 2017. 10.48550/arXiv.1706.03762 Number: arXiv:1706.03762. Accessed 2022–06-23.
  • 50.Lu X, Li X, Cheng Q, Ding K, Huang X, Qiu X. Scaling laws for fact memorization of large language models. arXiv 2024. 10.48550/arXiv.2406.15720. arXiv:2406.15720 [cs].
  • 51.Hopfield JJ. Neural networks and physical systems with emergent collective computational abilities. Proc Natl Acad Sci. 1982;79(8):2554–8. 10.1073/pnas.79.8.2554. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Hallee L, Kapur R, Patel A, Gleghorn JP, Khomtchouk BB. Contrastive learning and mixture of experts enables precise vector embeddings in biological databases. Sci Rep. 2025;15(1):14953. 10.1038/s41598-025-98185-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Hallee L, Rafailidis N, Gleghorn JP. cdsbert - extending protein language models with codon awareness. bioRxiv 2023. 10.1101/2023.09.15.558027.
  • 54.Consortium, T.U. Uniprot: the universal protein knowledgebase in 2025. Nucleic Acids Res. 2025;53(D1):609–17. 10.1093/nar/gkae1010. [DOI] [PMC free article] [PubMed]
  • 55.Li W, Godzik A. Cd-hit: a fast program for clustering and comparing large sets of protein or nucleotide sequences. Bioinformatics (Oxford England). 2006;22(13):1658–9. 10.1093/bioinformatics/btl158. [DOI] [PubMed] [Google Scholar]
  • 56" GrobidStyleApplied="Changed.Rafailidis N, Hallee L, Peleg T, Horger C, Haviland L, Gleghorn JP. Protify: an accessible platform for large-scale protein language model benchmarking. (arXiv preprint).
  • 57.Hallee L, Rafailidis N, Bichara DB, Gleghorn JP. Diffusion sequence models for enhanced protein representation and generation. arXiv 2025. 10.48550/arXiv.2506.08293. arXiv:2506.08293 [q-bio].
  • 58.Lin Z, Akin H, Rao R, Hie B, Zhu Z, Lu W, et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science. 2023;379(6637):1123–30. 10.1126/science.ade2574. [DOI] [PubMed] [Google Scholar]
  • 59.Elnaggar A, Heinzinger M, Dallago C, Rehawi G, Wang Y, Jones L, et al. Prottrans: Toward understanding the language of life through self-supervised learning. IEEE Trans Pattern Anal Mach Intell. 2022;44(10):7112–27. 10.1109/TPAMI.2021.3095381. [DOI] [PubMed] [Google Scholar]
  • 60.Elnaggar A, Essam H, Salah-Eldin W, Moustafa W, Elkerdawy M, Rochereau C, Rost B. Ankh: Optimized protein language model unlocks general-purpose modelling. 2023. 10.48550/arXiv.2301.06568. arXiv:2301.06568 [cs, q-bio].
  • 61.Agarap AF. Deep learning using rectified linear units (relu). arXiv 2019. 10.48550/arXiv.1803.08375. arXiv:1803.08375 [cs].
  • 62.Hinton GE, Srivastava N, Krizhevsky A, Sutskever I, Salakhutdinov RR. Improving neural networks by preventing co-adaptation of feature detectors. 2012. 10.48550/arXiv.1207.0580. arXiv:1207.0580 [cs].
  • 63.Wang H, Fan Y, Naeem MF, Xian Y, Lenssen JE, Wang L, Tombari F, Schiele B. Tokenformer: Rethinking transformer scaling with tokenized model parameters. arXiv 2025. 10.48550/arXiv.2410.23168. arXiv:2410.23168 [cs].
  • 64.Zhou X, Han C, Zhang Y, Su J, Zhuang K, Jiang S, Yuan Z, Zheng W, Dai F, Zhou Y, Tao Y, Wu D, Yuan F. Decoding the molecular language of proteins with evolla. bioRxiv, 2025;2025–0105630192 . 10.1101/2025.01.05.630192.
  • 65.Loshchilov I, Hutter F. Decoupled weight decay regularization. arXiv (2019) 10.48550/arXiv.1711.05101. arXiv:1711.05101 [cs].
  • 66.Steinegger M, Söding J. Mmseqs2 enables sensitive protein sequence searching for the analysis of massive data sets. Nat Biotechnol. 2017;35(11):1026–8. 10.1038/nbt.3988. [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Data Availability Statement

Datasets are available at https://github.com/Gleghorn-Lab/PLMConfounders. Protify and its associated datasets are accessed at https://github.com/Gleghorn-Lab/Protify and on Hugging Face at https://huggingface.co/collections/GleghornLab/protify. The taxonomic classification datasets are available on Hugging Face: Domain (https://huggingface.co/datasets/GleghornLab/taxonomy_domain_0.4_clusters), Kingdom (https://huggingface.co/datasets/GleghornLab/taxonomy_kingdom_0.4_clusters), Phylum (https://huggingface.co/datasets/GleghornLab/taxonomy_phylum_0.4_clusters), Class (https://huggingface.co/datasets/GleghornLab/taxonomy_class_0.4_clusters), Order (https://huggingface.co/datasets/GleghornLab/taxonomy_order_0.4_clusters), Family (https://huggingface.co/datasets/GleghornLab/taxonomy_family_0.4_clusters), Genus (https://huggingface.co/datasets/GleghornLab/taxonomy_genus_0.4_clusters), Species (https://huggingface.co/datasets/GleghornLab/taxonomy_species_0.4_clusters), and Different (https://huggingface.co/datasets/GleghornLab/diff_phylo). The sequences and taxonomic origin were originally gathered from UniProt release 2025_03. PPI training, validation, and test datasets are compiled automatically at training time; the exact examples used in the study can be found at https://github.com/Gleghorn-Lab/PLMConfounders/tree/main/processed_datasets, sourced from BioGRID version 4.4.243. The BioGRID IDs, mapped to UniProt IDs and sequences, are also hosted on Hugging Face: https://huggingface.co/datasets/Synthyra/BIOGRID. Selected code is available at https://github.com/Gleghorn-Lab/PLMConfounders. Protify and its code can be accessed here: https://github.com/Gleghorn-Lab/Protify.

Selected code is available at https://github.com/Gleghorn-Lab/PLMConfounders. Protify and its code can be accessed here: https://github.com/Gleghorn-Lab/Protify.


Articles from BMC Bioinformatics are provided here courtesy of BMC

RESOURCES