Skip to main content
Briefings in Bioinformatics logoLink to Briefings in Bioinformatics
. 2026 Jul 31;27(4):bbag408. doi: 10.1093/bib/bbag408

SSAS-GO: structure-sequence adaptive synergy network for protein function prediction

Dong Wang 1,2,3,✉, Hailong Wang 4, Tao Jiang 5,✉, Bin Lu 6,7,8, Fujun Xiang 9, Qiang Wang 10,11,12
PMCID: PMC13435231  PMID: 42537002

Abstract

Protein function prediction is essential and fundamental for drug discovery and disease treatment. In recent years, deep learning methods have achieved notable improvements by exploiting either protein sequence or structural features. Specifically, Convolutional Neural Networks often fail to apprehend global protein topologies due to restricted receptive fields, while Graph Convolutional Networks excel at processing graph-structured data, a singular network paradigm fundamentally lacks the capacity to fully integrate diverse, multimodal features. Furthermore, combining modalities via static feature aggregation frequently limits model efficacy and causes modality interference. To resolve these challenges, we propose the Structure-Sequence Adaptive Synergy network (SSAS-GO), which employs a Multi-Scale Motif Block to extract localized sequence semantic anchors, alongside a parallel Dual-Stream Graph Encoder to capture spatial topologies. Subsequently, these representations are fed into a Task-Adaptive Cross-Modal gating mechanism. This core module dynamically recalibrates the weights of sequence and structural features. On the PDBch test set, SSAS-GO leverages native topologies to achieve state-of-the-art Area Under the Precision-Recall Curve (AUPR) scores of 0.463 for Biological Process and 0.559 for Cellular Component. Remarkably, on the AFch test set, the model achieves a substantial 17.7% relative AUPR improvement for Molecular Function tasks.

Keywords: protein function prediction, gene ontology, graph convolutional networks, convolutional neural networks, multimodal fusion

Introduction

Proteins are the direct executors of cellular activities, underpinning virtually all biological systems through catalysis, molecular transport, signal transduction, and immune defense [1]. A precise understanding of protein function is therefore foundational for interpreting genotype–phenotype relationships, elucidating disease mechanisms, and supporting drug discovery and personalized medicine. To enable consistent and interoperable functional descriptions across studies and species, the Gene Ontology [2] Consortium developed the GO database, a structured vocabulary that organizes functional knowledge into three domains: Molecular Function (MF), Biological Process (BP), and Cellular Component (CC).

While accurate functional annotation is essential, experimental identification of protein functions is time consuming and costly [3]. Classical wet-lab approaches, including X-ray crystallography, NMR spectroscopy, and cryogenic electron microscopy [4, 5], provide the indispensable ground truth for functional annotation. However, their intrinsically low-throughput and resource-intensive nature precludes systematic application across expanding sequence databases [6]. To bridge this throughput discrepancy, computational inference frameworks have become paramount, offering scalable algorithmic proxies for proteome-wide functional screening.

Based on the information modalities utilized, current methods for protein function prediction can be categorized into four primary paradigms: alignment-based, sequence-based, protein–protein interaction (PPI) network-based and structure-based approaches. Alignment-based methods, such as BLAST [7], PSI-BLAST [8] and FunFams [9], operate on the premise that homologous sequences share similar functions. However, functional similarity often persists even among proteins with low sequence identity. To address this, sequence-based deep learning methods have emerged to capture complex motifs and semantic dependencies without relying on explicit alignment. Early approaches like DeepGOPlus [10] utilized Convolutional Neural Networks (CNNs) to capture sequence motifs. More recently, methods like DeepGO-SE [11] have leveraged Transformer-based Pre-trained Protein Language Models (PLMs), such as ESM-2 [12] and ProtTrans [13], to generate high-dimensional embeddings, significantly enhancing the representation of sequence semantics that essential for accurate function prediction.

Beyond sequence analysis, researchers have utilized PPI networks. In this domain, Classic methods, such as DeepGO [14] and NetGO [15] have evolved to integrate network topology with sequence features, leveraging graph embeddings to enhance prediction accuracy. Nevertheless, high-throughput PPI data are often fraught with noise and suffers from incomplete coverage [16, 17]. To circumvent the sparsity and uncertainty of PPI networks, protein function prediction has increasingly shifted toward 3D structure-based approaches utilizing Graph Neural Networks (GNNs) [18, 19], such as DeepFRI [20], GAT-GO [21], and HEAL [22]. While high-throughput predicted structures from databases like AlphaFold2 [23] provide high-resolution topological templates for localized active sites, they inevitably introduce a specific conformational bias: these predictions are predominantly isolated monomers stripped of their native multimeric contexts. Recent structural interactome analyses demonstrate that interfacial regions in isolated monomers often collapse into conditionally folded states or artificial coils [24, 25]. Consequently, for BP annotations, which intrinsically govern systems-level quaternary interactions, GNNs that aggregate these detached, spurious correlations, thereby injecting inductive noise into the prediction pipeline.

Recognizing that sequence and structure offer highly complementary insights, recent research has shifted toward multi-modal fusion methods. Pioneer frameworks like TAWFN [26] and MAEF-GO [27] have laid a solid foundation by integrating these modalities. MAEF-GO, for instance, employs cross-attention mechanisms to interactively integrate sequence features with structural embeddings. However, these architectures suffer from a critical architectural flaw: context-agnostic feature entanglement. By compelling sequence and structural representations to densely interact via forced cross-attention or rigid concatenation, these models fail to dynamically decouple modalities when one source is dominated by physical artifacts. Crucially, this fundamental biophysical asymmetry dictates that robust functional inference cannot rely on homogeneous multimodal fusion. On one hand, computationally predicted structures predominantly yield isolated monomers devoid of native multimeric inter-protein contexts, employing their global topologies for interactome-dependent BP annotations inevitably assimilates spurious inductive noise [24]. On the other hand, while protein language models capture deep co-evolutionary patterns, their sequence representations exhibit documented statistical biases when processing orphan proteins or targets with shallow evolutionary depth [28].

Consequently, resolving this bi-directional modality interference demands a paradigm shift from static feature aggregation to dynamic, context-aware modality routing. The optimal architecture must possess the capacity to act as a mathematical filter that is capable of quantitatively assessing modality reliability and selectively decoupling the corrupted pathway based on the specific ontological task. To systematically resolve this issue, we propose the Task-Adaptive Cross-Modal (TACM) gating mechanism. By calculating task-specific routing weights, TACM effectively uncouples corrupted pathways. For topology-sensitive BP annotations with available native structures, it adaptively suppresses sequence noise to prioritize valid spatial interactomes. Conversely, when processing computationally predicted monomers, which predominantly consist of orphan proteins with shallow evolutionary depths, TACM executes a fail-safe mechanism. Furthermore, to complement this global routing at the local semantic level, SSAS-GO incorporates a sequence-based Multi-Scale Motif Block (MSMB) [29]. By extracting highly conserved, multi-length semantic anchors directly from pre-trained language models, MSMB ensures that localized functional pockets remain explicit [30]. Experimental evaluations validate this dual strategy: on the high-confidence PDBch test set, SSAS-GO effectively harnesses native topologies to achieve a superior AUPR of 0.463 on BP tasks. Crucially, on the AFch test set, the synergy of TACM’s noise suppression and MSMB’s local semantic anchoring secures a substantial 17.7% relative AUPR improvement for MF tasks, providing a robust computational solution for noise-resistant multimodal protein annotation.

Materials and methods

Datasets

To facilitate a fair comparison, we utilized the standard benchmark dataset originally constructed by HEAL [22]. The data comprises two distinct subsets derived from SIFTS [31] and UniProtKB [6]: the PDB dataset (PDBch), containing representative experimental protein chains filtered at a 95% sequence identity, and the AF dataset (AFch), consisting of AlphaFold2-predicted structures designed to evaluate the model’s generalization capabilities. To ensure strict data independence for hard generalization tasks, the AFch was clustered using a stringent 25% sequence identity threshold. Detailed data preprocessing steps, subset partitioning ratios, and comprehensive dataset statistics are provided in Supplementary Material (S1).

Model overview

The structure of SSAS-GO, as shown in Fig. 1, comprising three core components: Sequence Branch, Structure Branch, Task-Adaptive Cross-Modal gating mechanism. Initially, the input features were processed through parallel Sequence and Structure Branches within the Dual-Branch architecture. Subsequently, a Task-Adaptive Cross-Modal gating mechanism was employed to dynamically fuse the cross-modal representations then after passing through a classifier, logits were obtained to predict protein functions.

Figure 1.

(A) The Dual-Branch Multimodal Architecture processes protein sequence and structural inputs in parallel. (B) Within the Sequence Branch, the Multi-Scale Motif Block (incorporating a simplified SE-Block) extracts localized semantic anchors from ESM-1b embeddings. (C) Concurrently, within the Structure Branch, a Dual-Stream Graph Encoder (parallel GCN/GAT) combined with a Structural Attention Block [49, 50] captures global macroscopic spatial topologies. (D) Positioned at the terminus of the dual branches, the Task-Adaptive Cross-Modal (TACM) gating mechanism dynamically evaluates and integrates the extracted representations. This core module computes adaptive routing weights to selectively filter out modality-specific noise prior to the final functional classification.

Overall architecture of SSAS-GO. (A) The Dual-Branch Multimodal Architecture processes protein sequence and structural inputs in parallel. (B) Within the Sequence Branch, the Multi-Scale Motif Block (incorporating a simplified SE-Block) extracts localized semantic anchors from ESM-1b embeddings. (C) Concurrently, within the Structure Branch, a Dual-Stream Graph Encoder (parallel GCN/GAT) combined with a Structural Attention Block [49, 50] captures global macroscopic spatial topologies. (D) Positioned at the terminus of the dual branches, the Task-Adaptive Cross-Modal (TACM) gating mechanism dynamically evaluates and integrates the extracted representations. This core module computes adaptive routing weights to selectively filter out modality-specific noise prior to the final functional classification.

Input data

Sequence features

We employed the ESM-1b [32] model to extract residue-level sequence features. ESM-1b is a Transformer-based language model pre-trained on the massive UniRef50 [33] protein sequence database. For sequences exceeding the ESM-1b maximum length limit, we truncated them to the first 1000 residues. We selected this model for its ability to capture deep evolutionary patterns and physicochemical properties via unsupervised learning, generating embedding vectors rich in high-dimensional semantic information. Specifically, for a protein sequence of length Inline graphic, we extracted the hidden states from the final layer of ESM-1b as the node feature matrix Inline graphic, where the feature dimension Inline graphic is 1280.

Contact maps

We modeled protein 3D structures as contact maps which we represented as a 2D matrix to explicitly represent all pairs acid residues in the protein structure. Specifically, we calculated the distance between the Inline graphic-carbon atoms (Inline graphic) of all amino acid pairs. Following previous works such as DeepFRI [20] we set the distance threshold to 10 Å and established an undirected edge between two residues Inline graphic and Inline graphic if their distance Inline graphic was less than 10 Å. In this rule, the structural contacts were compiled into an edge index matrix, Inline graphic, where Inline graphic represent the total count of edges within the protein graph.

Sequence branch

Biologically, protein functional semantics are inherently multi-scale. Short sequential motifs captured by smaller kernels often define highly localized catalytic active sites or tight binding pockets [33]. Conversely, extended regions or secondary structure elements necessitate larger receptive fields, while discontinuous residues that co-evolve or function together in 3D space require dilated convolutions to capture their non-adjacent sequence dependencies. To mathematically parameterize these heterogeneous biological properties, we designed a sequence-specific MSMB which employed parallel 1D convolutions with varying kernel sizes and dilation rates to explicitly map multi-length sequence semantics into high-dimensional anchors. Let Inline graphic denote the initial ESM-1b embeddings for a protein sequence of length Inline graphic. The features were first projected via a Inline graphic convolution and ReLU activation to a hidden dimension, yielding Inline graphic. Subsequently, Inline graphic was processed through four parallel 1D convolutional layers with varying receptive fields:

graphic file with name DmEquation1.gif (1)

where Inline graphic and Inline graphic denote the kernel sizes corresponding to short, medium, and long functional motifs respectively. The dilated convolution (Inline graphic) specifically accounts for discontinuous sequence motifs that may be proximally correlated in the folded geometry.

To fuse these multi-scale representations, the concatenated tensor was compressed back to Inline graphic via a Inline graphic convolutional fusion layer, producing Inline graphic. Finally, a spatial Squeeze-and-Excitation (SE) Block (Fig. 1B) gate was applied to adaptively calibrate channel-wise importance. The recalibrated sequence representation Inline graphic was obtained via a residual integration [34]:

graphic file with name DmEquation2.gif (2)
graphic file with name DmEquation3.gif (3)

where Inline graphic is the Sigmoid activation and Inline graphic denotes the element-wise product.

Structure branch

To compensate for the absence of explicit 3D spatial geometry in 1D sequence motifs, we established a parallel structure branch to capture macroscopic topological dependencies. We constructed protein-level spatial graphs where nodes, initialized with ESM-1b sequence embeddings, are connected via contact-map edges. Crucially, to comprehensively extract spatial features prior to cross-modal fusion, we employed an internal Dual-Stream Graph Encoder composed of parallel Graph Convolutional Network (GCN) and GAT encoders. The rationale for using both encoders is that they capture complementary aspects of residue-level contact graphs. GCN aggregates information according to the contact-map topology and is suitable for capturing relatively stable structural scaffold information, whereas GAT introduces adaptive attention weights over neighboring residues and can emphasize residue-specific local interactions. This design allows the Dual-Stream Graph Encoder Module to integrate both general topological context and adaptive local neighborhood information before subsequent structural attention and cross-modal fusion. Each encoder consists of two message-passing layers, symmetrically regularized with Batch Normalization, ReLU activation, Dropout, and residual connections. The respective node update functions are defined as follows:

graphic file with name DmEquation4.gif (4)

where Inline graphic denotes the node feature matrix at layer Inline graphic, and Inline graphic represents the adjacency matrix with added self-loops. Inline graphic is the degree matrix of Inline graphic, and Inline graphic is the learnable weight matrix for the Inline graphic-th GCN layer.

Let Inline graphic and Inline graphic denote the outputs of the respective streams. To optimally resolve dimensional redundancy between the two message-passing paradigms, we introduced an intra-modal channel-wise gating mechanism:

graphic file with name DmEquation5.gif (5)
graphic file with name DmEquation6.gif (6)

where Inline graphicdenotes the concatenation along the feature dimension, and Inline graphicis the learnable gating weight.

While GNNs excel at local neighborhood aggregation, assessing the global functional significance of individual residues across the entire macromolecule is non-trivial. We employed a structural attention block (SAB) (Fig. 1C). For each attention head Inline graphic, the node representations Inline graphic were linearly projected to Query (Inline graphic), Key (Inline graphic), and Value (Inline graphic) matrices. The unnormalized attention score for each node Inline graphic was calculated as the self-dot-product scaled by a constant factor Inline graphic:

graphic file with name DmEquation7.gif (7)

Crucially, these scores were normalized across all Inline graphic nodes belonging to the specific protein graph utilizing a softmax function:

graphic file with name DmEquation8.gif (8)

where the updated feature for node Inline graphic was scaled asInline graphic. The multi-head outputs were concatenated, processed through a linear layer, and refined via a Feed-Forward Network (FFN) with LayerNorm and residual connections, ultimately yielding the structural representation:

graphic file with name DmEquation9.gif (9)
graphic file with name DmEquation10.gif (10)

Task-adaptive cross-modal and prediction

To systematically mitigate modality interference particularly the assimilation of spurious structural inductive noise during topology-sensitive predictions, we designed the TACM gating mechanism (Fig. 1D). It computed a specific confidence vector Inline graphic based on the concatenated representations:

graphic file with name DmEquation11.gif (11)

The final multimodal representation Inline graphic was generated via an adaptive, element-wise blockade:

graphic file with name DmEquation12.gif (12)

Consequently, a smaller Inline graphic indicates a greater contribution from the structural representation, whereas a larger Inline graphic indicates a greater contribution from the sequence representation. For topology-sensitive tasks with reliable experimental structures, the model can assign a lower Inline graphic to increase the contribution of native structural topology. Conversely, for sequence-dependent targets or predicted structures in the AFch dataset, where global structural topology may be less informative for certain functional contexts, the model can assign a higher Inline graphic, thereby increasing reliance on sequence representations. Finally, a sum pooling function aggregated the residue-level embeddings into a graph-level representation. These comprehensive vectors were fed into a classification head comprising a Linear Layer, Batch Normalization, GELU Activation [35], and Dropout (P = 0.3), and a final projection layer to output the GO term probabilities.

Model training and parameter settings

To alleviate gradient domination by rare GO terms, the model was optimized utilizing a Positive-Weighted Binary Cross-Entropy with Logits Loss, dynamically re-weighted based on inverse class frequencies. Equation details are in Supplementary Material (S2). The model was trained utilizing a Top-k Checkpoint Ensemble strategy to mitigate stochastic fluctuations. Comprehensive hyperparameters, hardware specifications, and implementation details are available in Supplementary Material (S3).

Baseline methods

To comprehensively evaluate the performance of SSAS-GO, we conducted extensive comparative experiments. Baseline methods were categorized by their input modalities: alignment-based methods (BLAST [7], FunFams [9]), sequence-based deep learning models (DeepGOPlus [10], DeepGO-SE [11]), methods incorporating PPI networks (DeepGO [14]), structure-based methods (DeepFRI [20], GAT-GO [21], GPSFun [36], HEAL [22]), and multimodal fusion methods (TAWFN [26], MAEF-GO [27]). In particular, the multimodal baselines were selected to demonstrate SSAS-GO’s superiority in resolving modality interference and mitigating the modality interference problem inherent in rigid fusion strategies. Detailed descriptions and parameter settings of all baseline methods are provided in the Supplementary Material (S4).

Evaluation metrics

To comprehensively assess predictive performance across various dimensions, we adopted the standard evaluation protocols recommended by the CAFA challenge [37]. All metric calculations were implemented based on the standard Method class. Prior to computing specific metrics, we executed a standard Label Propagation procedure to strictly enforce the hierarchical constraints of the Gene Ontology. Specifically, the predicted scores of child nodes were propagated to all their ancestor nodes. This step ensures the logical consistency of predictions with respect to the GO topology. The detailed calculation method of each indicator is provided in the Supplementary Material (S5).

Results

Experimental results on the PDBch

We compared SSAS-GO against a diverse set of established baselines on the PDBch test set. The quantitative performance across the MF, BP, CC tasks is summarized in Table 1.

Table 1.

AUPR, Fmax, and Smin values of different methods on the PDBch test set, with the highest Fmax and AUPR and the lowest Smin highlighted in bold.

Method AUPR Fmax Smin
MF BP CC MF BP CC MF BP CC
Blast 0.136 0.067 0.096 0.326 0.336 0.443 0.643 0.662 0.632
FunFams 0.370 0.256 0.265 0.573 0.498 0.640 0.542 0.58 0.512
DeepGO 0.391 0.189 0.258 0.576 0.500 0.589 0.475 0.578 0.553
DeepFRI 0.495 0.265 0.274 0.627 0.546 0.617 0.432 0.543 0.530
GAT-GO 0.660 0.381 0.479 0.633 0.492 0.547 0.437 0.521 0.466
GPSFun 0.601 0.203 0.309 0.745 0.456 0.630 0.339 0.564 0.508
DeepGO-SE 0.495 0.233 0.423 0.654 0.566 0.636 0.435 0.530 0.481
HEAL 0.691 0.337 0.467 0.747 0.595 0.687 0.342 0.509 0.458
TAWFN 0.718 0.385 0.488 0.762 0.628 0.693 0.326 0.483 0.454
MAEF-GO 0.758 0.438 0.530 0.787 0.652 0.720 0.298 0.461 0.426
SSAS-GO 0.766 0.463 0.559 0.790 0.654 0.722 0.292 0.456 0.414

As detailed in Table 1, SSAS-GO consistently outperformed all comparative methods across most evaluation metrics. SSAS-GO demonstrated substantial improvements in structure-dependent tasks relative to the leading multimodal baseline, MAEF-GO. For BP prediction, SSAS-GO achieved a 5.7% relative increase in AUPR (0.463 versus 0.438). Similarly, the CC task saw a 5.5% relative AUPR improvement (0.559 versus 0.530). In the sequence-sensitive MF domain, SSAS-GO maintained dominance with the highest AUPR of 0.766. Furthermore, SSAS-GO secured the lowest Inline graphic scores across all three ontologies (0.292, 0.456, and 0.414 for MF, BP, and CC, respectively).

Experimental results on the AFch

Given the rapid expansion of predicted structures, it is crucial to evaluate whether SSAS-GO can generalize well to computationally generated data (Table 2). Notably, the baseline methods evaluated on the AFch test set differ from those used on the PDBch test set because the two evaluations address different structural settings. The PDBch test set provides experimentally resolved structures and is used for a comprehensive comparison with a broad range of baselines. In contrast, the AFch test set consists of AlphaFold2-predicted structures and is used to evaluate model performance in the predicted-structure setting. Therefore, we retained representative sequence-based, structure-based, and multimodal methods that are applicable to AFch under comparable input assumptions and reproducible evaluation settings. Most notably, in the MF task, SSAS-GO achieved an AUPR of 0.639. This represented a substantial 17.7% improvement over the leading baseline MAEF-GO (0.543). For the BP and CC tasks, SSAS-GO similarly secured the highest performance across all metrics, recording AUPR scores of 0.255 and 0.348, and Fmax scores of 0.501 and 0.661, respectively. By integrating structural information, SSAS-GO achieved relative AUPR gains of 25.6% and 30.3% in BP and CC tasks, respectively, compared to DeepGOPlus.

Table 2.

AUPR and Fmax values of different methods on the AFch test set, with the highest Fmax and AUPR highlighted in bold.

Method AUPR
MF
Fmax
MF
AUPR
BP
Fmax
BP
AUPR
CC
Fmax
CC
DeepGOPlus 0.463 0.450 0.203 0.430 0.267 0.567
DeepFRI 0.342 0.398 0.114 0.387 0.192 0.536
HEAL 0.502 0.491 0.200 0.475 0.287 0.614
TAWFN 0.520 0.510 0.223 0.487 0.295 0.626
MAEF-GO 0.543 0.537 0.245 0.495 0.323 0.628
SSAS-GO 0.639 0.631 0.255 0.501 0.348 0.661

Generalization analysis

To rigorously assess the generalization capability of SSAS-GO under increasingly challenging conditions, we evaluated the model using five distinct sequence identity thresholds on the PDBch test set: 30%, 40%, 50%, 70%, and 95% [Supplementary Material (S6)]. The number of test proteins under each sequence identity threshold is provided in Supplementary Figure S1. A comprehensive comparative analysis against established baselines across multiple evaluation metrics, including AUPR, Fmax, and Smin, is illustrated (Fig. 2). As expected in hard generalization scenarios, while the predictive performance of all models naturally declined as sequence identity decreased to the stringent 30% threshold, SSAS-GO consistently maintained a relatively stable and superior performance curve. Quantitatively, the AUPR performance gap between SSAS-GO and the baseline methods observably widened at lower identity thresholds (Supplementary Material Tables S1, S2, and  S3), demonstrating the robust capacity of our framework to capture deep functional dependencies rather than relying on superficial sequence homology.

Figure 2.

(A) AUPR, (B) Fmax, and (C) Smin scores for SSAS-GO and baseline methods. The x-axis represents the maximum sequence identity between the test and training sets.

Performance comparison across varying sequence identity thresholds on the PDBch test set. (A) AUPR, (B) Fmax, and (C) Smin scores for SSAS-GO and baseline methods. The x-axis represents the maximum sequence identity between the test and training sets.

Ablation study

To systematically investigate the synergistic contributions of the core architectural components within SSAS-GO, we conducted a comprehensive ablation study (Table 3). Specifically, five distinct model variants were constructed and evaluated to isolate the functional impact of each module: (i) W/O MSMB: removing the Multi-Scale Motif Block to assess the importance of extracting localized sequence semantic anchors; (ii) W/O GAT: removing the Graph Attention Network pathway to evaluate the role of capturing anisotropic neighbor interactions; (iii) W/O SAB: removing the Structural Attention Block to determine the necessity of assessing the global functional significance of individual residues across the spatial topologies; (iv) W/O GCN: removing the Graph Convolutional Network stream to observe the effect of excluding isotropic structural frameworks and topological aggregation; (v) W/O TACM: replacing the Task-Adaptive Cross-Modal gating mechanism with a standard static concatenation strategy. This final variant was specifically designed to validate the necessity of dynamic routing in systematically mitigating bi-directional modality interference.

Table 3.

Ablation experiment results of SSAS-GO on the PDBch test set, with the highest Fmax and AUPR and the lowest Smin highlighted in bold.

Method AUPR Fmax Smin
MF BP CC MF BP CC MF BP CC
SSAS-GO 0.766 0.463 0.559 0.790 0.654 0.722 0.292 0.456 0.414
SSAS-GO W/O MSMB 0.765 0.448 0.547 0.781 0.647 0.706 0.299 0.468 0.431
SSAS-GO W/O GAT 0.766 0.457 0.564 0.788 0.650 0.713 0.290 0.459 0.422
SSAS-GO W/O SAB 0.774 0.463 0.553 0.789 0.653 0.709 0.292 0.458 0.426
SSAS-GO W/O GCN 0.767 0.437 0.546 0.784 0.644 0.708 0.299 0.470 0.431
SSAS-GO W/O TACM 0.763 0.424 0.539 0.784 0.642 0.709 0.299 0.474 0.428

Table 3 presents the quantitative results of the ablation study on the PDBch test set. The complete SSAS-GO model achieved the highest AUPR in both the BP (0.463) and CC (0.559) domains. A critical performance degradation was observed when the TACM was replaced by simple static concatenation (W/O TACM): in the topology-sensitive BP domain, the AUPR declined precipitously from 0.463 to 0.424. Furthermore, removing the MSMB resulted in a measurable AUPR reduction across all three ontologies, particularly dropping to 0.765 in the MF task. Notably, for the MF task, the variant lacking the SAB marginally outperformed the full model (AUPR 0.774 versus 0.766), whereas the full model-maintained dominance in BP predictions. Finally, the exclusion of topological aggregation (W/O GCN) notably impaired BP performance (AUPR 0.437). In addition, to evaluate robustness across varying levels of functional specificity, GO terms were stratified based on Information Content (IC). In the most challenging subset (IC > 10), SSAS-GO retained substantial predictive capability, reaching AUPR scores of 0.390 (BP), 0.524 (CC), and 0.716 (MF). See Supplementary Material (S6.2) and Supplementary Figure S2 for details.

The architectural rationale for integrating both GCN and GAT lies in the biophysical complementarity of isotropic macroscopic scaffold extraction and anisotropic localized spatial attention. A comprehensive theoretical analysis and an extended single-stream ablation study substantiating this Dual-Stream synergy are detailed in Supplementary Material S7, and the corresponding ablation results are shown in Supplementary Figure S3.

Task-adaptive routing distributions across ontologies

To quantify the internal routing dynamics of SSAS-GO, the sequence modality weights of Inline graphic allocated to targets in both the PDBch and AFch test set were extracted and analyzed. Kruskal–Wallis tests revealed statistically significant divergences in the weight distributions across the BP, MF, and CC tasks (Fig. 3A and B). While sequence representations globally dominated the information flow with mean Inline graphic exceeding 0.85 across all evaluated sets, the explicit structural routing weight (Inline graphic) allocated to topology-sensitive tasks (BP and CC) demonstrated a substantial relative increase compared to the MF task.

Figure 3.

(A–B) Distributions of the learned sequence modality weights (αseq) for BP, MF, and CC tasks on the PDBch and AFch test sets. Kruskal-Wallis tests indicate statistically significant divergences in weight distributions across the three functional ontologies. (C–D) Stratified AUPR performance based on modality reliance. Proteins were partitioned into terciles according to their empirical αseq distributions, with error bars indicating 95% confidence intervals derived from 1000 bootstrap iterations.

Modality routing distributions and stratified performance dynamics. (A–B) Distributions of the learned sequence modality weights (αseq) for BP, MF, and CC tasks on the PDBch and AFch test sets. Kruskal-Wallis tests indicate statistically significant divergences in weight distributions across the three functional ontologies. (C–D) Stratified AUPR performance based on modality reliance. Proteins were partitioned into terciles according to their empirical αseq distributions, with error bars indicating 95% confidence intervals derived from 1000 bootstrap iterations.

Furthermore, to evaluate the impact of these routing strategies on predictive accuracy, proteins were stratified into statistically equivalent terciles (Bottom 33%, Middle 33%, and Top 33%) based on their empirical Inline graphic distributions. AUPR outcomes and their corresponding 95% Confidence Intervals were derived using 1000 bootstrap iterations (Fig. 3C and D). On the PDBch test set, topology-sensitive tasks (BP and CC) exhibited a clear monotonic correlation: the structure-dominated cohort (Bottom 33%) attained the maximal AUPR (0.4658 for BP; 0.6076 for CC), significantly outperforming the sequence-dominated cohort (Top 33%, AUPR 0.4140 for BP; 0.5143 for CC). Conversely, the MF task demonstrated relative invariance to structural routing on the PDBch test set. Notably, the monotonic structural benefit observed for topology-sensitive tasks on the PDBch was severely disrupted within the AFch. In this computationally predicted set, the topology-dependence is completely inverted. Because predicted monomers lack native multimeric contexts, their global topologies introduce inductive noise for complex interactome tasks. Consequently, the sequence-dominated cohort outperformed the structure-reliant cohort for both BP (AUPR 0.3218 versus 0.2953) and markedly for CC (AUPR 0.4783 versus 0.3329). This serves as micro-level validation that the TACM module adaptively suppresses structural noise, dynamically shifting the inferential burden to the sequence modality when spatial geometries are physically unreliable.

Micro-level validation of adaptive modality routing

Two specific cases from the test set that achieved perfect predictive precision of Inline graphic were analyzed (Fig. 4). Figure 4A illustrates the subunit b of the human respiratory complex I (PDB: 5XTD [38]), a BP target. For this target, the cross-modal gate assigned an elevated structural routing weight of Inline graphic. Structural data indicates that this polypeptide chain is a component within a multi-subunit macromolecular assembly. Conversely, Fig. 4B presents the human c-Jun N-terminal kinase 3 (PDB: 1JNK [39]), a MF target. For this protein, the TACM gating mechanism assigned an overwhelming sequence weight of Inline graphic, corresponding to a low structural routing proportion.

Figure 4.

(A) The BP target (PDB: 5XTD). The cross-modal gate assigns a substantial structural weight (1 − αseq = 0.18). Dark blue surfaces indicate high node-level structural attention scores extracted from the structure branch. (B) The MF target (PDB: 1JNK). The network prioritizes the sequence modality (αseq = 0.97). Hot pink spheres highlight crucial localized sequence motifs identified via Gradient-weighted Class Activation Mapping (Grad-CAM) [51] from the MSMB.

Spatial visualization of task-adaptive modality routing. (A) The BP target (PDB: 5XTD). The cross-modal gate assigns a substantial structural weight (1 − αseq = 0.18). Dark blue surfaces indicate high node-level structural attention scores extracted from the structure branch. (B) The MF target (PDB: 1JNK). The network prioritizes the sequence modality (αseq = 0.97). Hot pink spheres highlight crucial localized sequence motifs identified via Gradient-weighted Class Activation Mapping (Grad-CAM) [51] from the MSMB.

Discussion

Our results reveal a fundamental biophysical asymmetry in how proteins execute diverse functions. The routing analysis across the three GO ontologies further indicates that this distinction is not limited to an MF-versus-BP contrast, but reflects a broader ontology-dependent modality pattern: MF is relatively sequence-dominant, whereas BP and CC show stronger dependence on structural information in the experimental PDBch setting. This pattern is plausible. MF annotations are often governed by conserved local motifs, catalytic residues, and binding pockets, whereas BP and CC annotations more frequently involve macromolecular context, subcellular organization, complex membership, or broader structural environments. On the PDBch test set, the routing-stratified analysis supports this interpretation: structure-reliant proteins achieved higher AUPR for both BP and CC, whereas MF showed weaker dependence on structural routing, suggesting that structural information contributes unevenly across GO ontologies.

This ontology-dependent behavior also explains why static multimodal fusion may be suboptimal for protein function prediction. Rigid concatenation inevitably assimilates mutually orthogonal noise distributions, indiscriminately merging the physical topological artifacts inherent in computationally predicted structures [40] with the statistical uncertainties of sequence embeddings. The modality routing distributions (Fig. 3) support the interpretation that TACM adjusts the relative contribution of sequence and structural representations according to ontology type and structural source. As the data indicates, the mean sequence weight is notably lower for targets in the AFch compared to the PDBch test set. This routing shift uncovers a critical limitation of PLMs: the AFch predominantly comprises uncharacterized or orphan proteins with shallow evolutionary depths. For such targets, ESM-1b struggles to extract meaningful co-evolutionary patterns, resulting in highly uncertain semantic embeddings. For BP and CC on the AFch test set, sequence-dominated cohorts outperformed structure-reliant cohorts, suggesting that predicted monomeric structures may provide less reliable global topological context for functions involving assemblies, cellular localization, or interaction-dependent processes. These results suggest that the usefulness of structural information depends on both functional ontology and structural source, when predicted monomeric structures lack native biological context, reducing reliance on global structural topology may improve predictive performance.

The representative structural visualizations provide case-level support for the learned routing patterns, although they should not be interpreted as exhaustive mechanistic validation. For the BP target 5XTD (a subunit of a macromolecular assembly), the SAB dynamically bypassed known local phospholipid (PLX) binding sites (Tyr84/Tyr88) annotated in the BioLiP database [41]. Instead, the attention distribution (score > 90) strictly anchored to the C-terminal region (residues 105–118, centered at Gly107). This pattern suggests that the structure branch may capture global or assembly-related topological cues relevant to process-level annotation, rather than focusing exclusively on localized ligand-binding sites. Conversely, for the MF target 1JNK, the MSMB exhibited extreme local specificity. The core activation peaks (Asp207, score 100.0; Asp189, score 85.4) correspond to the absolute catalytic residues essential for phosphorylation transfer (the DFG and HRD motifs). This case is consistent with the interpretation that localized sequence motifs can provide sufficient information for certain biochemical function predictions.

These routing patterns may also provide useful interpretability signals for identifying proteins whose predicted functions depend more strongly on structural or sequence-derived information. While SSAS-GO effectively addresses spatial modality interference, predicting functional biology strictly from static snapshots remains a fundamental limitation. Moving forward, models for protein function prediction should move beyond static structural snapshots. Because proteins are dynamic molecular systems, future work will focus on integrating conformational dynamics [42] through molecular dynamics simulations or predictive confidence metrics derived from AlphaFold3 [43]. Another promising direction is to combine residue-level structure–sequence representations with heterogeneous interactome networks [44, 45], informed by recent advances in graph representation learning and transformer-powered biomedical network modeling [46–48]. Such extensions may help connect local residue-level representations with broader cellular and interaction-level contexts.

Ultimately, by integrating multi-scale sequence motifs with dynamic, task-adaptive gating, SSAS-GO provides a practical strategy for reducing modality interference and structural noise in multimodal protein function prediction. This architecture advances the paradigm of automated protein annotation from rigid feature aggregation to dynamic, ontology-adaptive modality routing. As structural genomics rapidly expands through computation, SSAS-GO provides a practical framework for ontology-aware multimodal protein annotation, particularly in large-scale settings where experimental and predicted structures differ in reliability and biological context.

Key Points

  • SSAS-GO enriches sequence feature extraction by employing a Multi-Scale Motif Block to capture localized functional pockets and multi-length semantic anchors.

  • SSAS-GO utilizes a Dual-Stream Graph Encoder network and a Structural Attention Block to comprehensively extract macroscopic spatial topologies.

  • SSAS-GO introduces a Task-Adaptive Cross-Modal gating mechanism to dynamically recalibrate modality weights, effectively decoupling bi-directional modality interference.

Supplementary Material

SSAS-GO-supplement_bbag408

Contributor Information

Dong Wang, Yanzhao Electric Power Laboratory, North China Electric Power University, No. 689 Huadian Road, Lianchi District, Baoding 071000, China; Engineering Research Center of Intelligent Computing for Complex Energy Systems, Ministry of Education, No. 689 Huadian Road, Lianchi District, Baoding 071000, China; Hebei Key Laboratory of Knowledge Computing for Energy & Power, North China Electric Power University, No. 689 Huadian Road, Lianchi District, Baoding 071000, China.

Hailong Wang, Yanzhao Electric Power Laboratory, North China Electric Power University, No. 689 Huadian Road, Lianchi District, Baoding 071000, China.

Tao Jiang, Center for Bioinformatics, Faculty of Computing, Harbin Institute of Technology, No. 92 Xidazhi Street, Nangang District, Harbin, Heilongjiang 150001, China.

Bin Lu, Yanzhao Electric Power Laboratory, North China Electric Power University, No. 689 Huadian Road, Lianchi District, Baoding 071000, China; Engineering Research Center of Intelligent Computing for Complex Energy Systems, Ministry of Education, No. 689 Huadian Road, Lianchi District, Baoding 071000, China; Hebei Key Laboratory of Knowledge Computing for Energy & Power, North China Electric Power University, No. 689 Huadian Road, Lianchi District, Baoding 071000, China.

Fujun Xiang, Yanzhao Electric Power Laboratory, North China Electric Power University, No. 689 Huadian Road, Lianchi District, Baoding 071000, China.

Qiang Wang, Yanzhao Electric Power Laboratory, North China Electric Power University, No. 689 Huadian Road, Lianchi District, Baoding 071000, China; Engineering Research Center of Intelligent Computing for Complex Energy Systems, Ministry of Education, No. 689 Huadian Road, Lianchi District, Baoding 071000, China; Hebei Key Laboratory of Knowledge Computing for Energy & Power, North China Electric Power University, No. 689 Huadian Road, Lianchi District, Baoding 071000, China.

Author contributions

D.W. and H.W. developed the method, conceived and designed the experiments, performed the experiments, analyzed the data, and wrote the paper. T.J., B.L., F.X., and Q.W. reviewed and edited the manuscript.

Conflicts of interest

None declared.

Funding

This work was supported by the Fundamental Research Funds for the Central Universities (2024MS128), Hebei Natural Science Foundation (Grant Number F2025502024), and Beijing Natural Science Foundation (Grant Number 4254105).

Data availability

The source codes and datasets of SSAS-GO are available at https://github.com/hlwang613/SSAS-GO.git

References

  • 1. Eisenberg  D, Marcotte  EM, Xenarios  I  et al.  Protein function in the post-genomic era. Nature  2000;405:823–6. 10.1038/35015694 [DOI] [PubMed] [Google Scholar]
  • 2. Ashburner  M, Ball  CA, Blake  JA  et al.  Gene ontology: tool for the unification of biology. Nat Genet  2000;25:25–9. 10.1038/75556 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3. Zhou  N, Jiang  Y, Bergquist  TR  et al.  The CAFA challenge reports improved protein function prediction and new functional annotations for hundreds of genes through experimental screens. Genome Biol  2019;20:244. 10.1186/s13059-019-1835-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4. Wüthrich  K. Protein structure determination in solution by NMR spectroscopy. J Biol Chem  1990;265:22059–62. 10.1016/S0021-9258(18)45665-7 [DOI] [PubMed] [Google Scholar]
  • 5. The UniProt Consortium, Bateman  A, Martin  M-J  et al.  UniProt: the universal protein knowledgebase in 2025. Nucleic Acids Res  2025;53:D609–17. 10.1093/nar/gkae1010 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6. The UniProt Consortium . UniProt: a worldwide hub of protein knowledge. Nucleic Acids Res  2019;47:D506–15. 10.1093/nar/gky1049 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7. Altschul  SF, Gish  W, Miller  W  et al.  Basic local alignment search tool. J Mol Biol  1990;215:403–10. 10.1016/S0022-2836(05)80360-2 [DOI] [PubMed] [Google Scholar]
  • 8. Altschul  S, Madden  TL, Schäffer  AA  et al.  Gapped BLAST and PSI-BLAST: a new generation of protein database search programs. Nucleic Acids Res  1997;25:3389–402. 10.1093/nar/25.17.3389 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9. Littmann  M, Bordin  N, Heinzinger  M  et al.  Clustering FunFams using sequence embeddings improves EC purity. Bioinformatics  2021;37:3449–55. 10.1093/bioinformatics/btab371 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10. Kulmanov  M, Hoehndorf  R. DeepGOPlus: improved protein function prediction from sequence. Bioinformatics  2020;36:422–9. 10.1093/bioinformatics/btz595 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11. Kulmanov  M, Guzmán-Vega  FJ, Duek Roggli  P  et al.  Protein function prediction as approximate semantic entailment. Nat Mach Intell  2024;6:220–8. 10.1038/s42256-024-00795-w [DOI] [Google Scholar]
  • 12. Lin  Z, Akin  H, Rao  R  et al.  Evolutionary-scale prediction of atomic-level protein structure with a language model. Science  2023;379:1123–30. 10.1126/science.ade2574 [DOI] [PubMed] [Google Scholar]
  • 13. Elnaggar  A, Heinzinger  M, Dallago  C  et al.  ProtTrans: toward understanding the language of life through self-supervised learning. IEEE Trans Pattern Anal Mach Intell  2022;44:7112–27. 10.1109/TPAMI.2021.3095381 [DOI] [PubMed] [Google Scholar]
  • 14. Kulmanov  M, Khan  MA, Hoehndorf  R  et al.  DeepGO: predicting protein functions from sequence and interactions using a deep ontology-aware classifier. Bioinformatics  2018;34:660–8. 10.1093/bioinformatics/btx624 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15. You  R, Yao  S, Xiong  Y  et al.  NetGO: improving large-scale protein function prediction with massive network information. Nucleic Acids Res  2019;47:W379–87. 10.1093/nar/gkz388 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16. De Silva  E, Thorne  T, Ingram  P  et al.  The effects of incomplete protein interaction data on structural and evolutionary inferences. BMC Biol  2006;4:39. 10.1186/1741-7007-4-39 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17. De Las  RJ, Fontanillo  C. Protein–protein interactions essentials: key concepts to building and Analyzing Interactome networks. PLoS Comput Biol  2010;6:e1000807. 10.1371/journal.pcbi.1000807 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18. Kipf  TN, Welling  M. Semi-supervised classification with graph convolutional networks. 2017. 10.48550/arXiv.1609.02907 [DOI]
  • 19. Veličković  P, Cucurull  G, Casanova  A  et al. Graph Attention Networks. 2018. https://openreview.net/forum?id=rJXMpikCZ
  • 20. Gligorijević  V, Renfrew  PD, Kosciolek  T  et al.  Structure-based protein function prediction using graph convolutional networks. Nat Commun  2021;12:3168. 10.1038/s41467-021-23303-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21. Lai  B, Xu  J. Accurate protein function prediction via graph attention networks with predicted structure information. Brief Bioinform  2022;23:bbab502. 10.1093/bib/bbab502 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22. Gu  Z, Luo  X, Chen  J  et al.  Hierarchical graph transformer with contrastive learning for protein function prediction. Bioinformatics  2023;39:btad410. 10.1093/bioinformatics/btad410 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23. Varadi  M, Anyango  S, Deshpande  M  et al.  AlphaFold protein structure database: massively expanding the structural coverage of protein-sequence space with high-accuracy models. Nucleic Acids Res  2022;50:D439–44. 10.1093/nar/gkab1061 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24. Peng  J, Zhao  L. A predicted structural interactome reveals binding interference from intrinsically disordered regions. PLoS Comput Biol  2026;22:e1013899. 10.1371/journal.pcbi.1013899 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25. Alderson  TR, Pritišanac  I, Kolarić  Đ  et al.  Systematic identification of conditionally folded intrinsically disordered regions by AlphaFold2. Proc Natl Acad Sci  2023;120:e2304302120. 10.1073/pnas.2304302120 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26. Meng  L, Wang  X. TAWFN: a deep learning framework for protein function prediction. Bioinformatics  2024;40:btae571. 10.1093/bioinformatics/btae571 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27. Liu  M, Wang  S, Luo  Z  et al.  Multistage attention-based extraction and fusion of protein sequence and structural features for protein function prediction. Bioinformatics  2025;41:btaf374. 10.1093/bioinformatics/btaf374 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28. Ding  F, Steinhardt  J. Protein language models are biased by unequal sequence sampling across the tree of life. 2024. 10.1101/2024.03.07.584001 [DOI]
  • 29. Hu  J, Shen  L, Sun  G. Squeeze-and-excitation networks. In: 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 7132–41, 2018.  https://ieeexplore.ieee.org/document/8578843. 10.1109/CVPR.2018.00745. [DOI]
  • 30. Lin  K, Quan  X, Jin  C  et al.  An interpretable double-scale attention model for enzyme protein class prediction based on transformer encoders and multi-scale convolutions. Front Genet  2022;13:885627. 10.3389/fgene.2022.885627 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31. Dana  JM, Gutmanas  A, Tyagi  N  et al.  SIFTS: updated structure integration with function, taxonomy and sequences resource allows 40-fold increase in coverage of structure-based annotations for proteins. Nucleic Acids Res  2019;47:D482–9. 10.1093/nar/gky1114 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32. Rives  A, Meier  J, Sercu  T  et al.  Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences. Proc Natl Acad Sci  2021;118:e2016239118. 10.1073/pnas.2016239118 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33. Suzek  BE, Huang  H, McGarvey  P  et al.  UniRef: comprehensive and non-redundant UniProt reference clusters. Bioinformatics  2007;23:1282–8. 10.1093/bioinformatics/btm098 [DOI] [PubMed] [Google Scholar]
  • 34. He  K, Zhang  X, Ren  S  et al.  Deep residual learning for image recognition. In: 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 770–8. Las Vegas, NV, USA: IEEE, 2016. 10.1109/CVPR.2016.90 [DOI] [Google Scholar]
  • 35. Hendrycks  D, Gimpel  K. Gaussian error linear units (GELUs). 2023. 10.48550/arXiv.1606.08415 [DOI]
  • 36. Yuan  Q, Tian  C, Song  Y  et al.  GPSFun: geometry-aware protein sequence function predictions with language models. Nucleic Acids Res  2024;52:W248–55. 10.1093/nar/gkae381 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37. Radivojac  P, Clark  WT, Oron  TR  et al.  A large-scale evaluation of computational protein function prediction. Nat Methods  2013;10:221–7. 10.1038/nmeth.2340 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38. Guo  R, Zong  S, Wu  M  et al.  Architecture of human mitochondrial respiratory Megacomplex I2III2IV2. Cell  2017;170:1247–1257.e12. 10.1016/j.cell.2017.07.050 [DOI] [PubMed] [Google Scholar]
  • 39. Xie  X, Gu  Y, Fox  T  et al.  Crystal structure of JNK3: a kinase implicated in neuronal apoptosis. Structure  1998;6:983–91. 10.1016/S0969-2126(98)00100-2 [DOI] [PubMed] [Google Scholar]
  • 40. Akdel  M, Pires  DEV, Pardo  EP  et al.  A structural biology community assessment of AlphaFold2 applications. Nat Struct Mol Biol  2022;29:1056–67. 10.1038/s41594-022-00849-w [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41. Yang  J, Roy  A, Zhang  Y. BioLiP: a semi-manually curated database for biologically relevant ligand–protein interactions. Nucleic Acids Res  2012;41:D1096–103. 10.1093/nar/gks966 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42. Hou  C, Zhao  H, Shen  Y. Protein language models trained on biophysical dynamics inform mutation effects. Proc Natl Acad Sci  2026;123:e2530466123. 10.1073/pnas.2530466123 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43. Abramson  J, Adler  J, Dunger  J  et al.  Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature  2024;630:493–500. 10.1038/s41586-024-07487-w [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44. Zhao  B-W, Su  X-R, Yang  Y  et al.  A heterogeneous information network learning model with neighborhood-level structural representation for predicting lncRNA-miRNA interactions. Comput Struct Biotechnol J  2024;23:2924–33. 10.1016/j.csbj.2024.06.032 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45. Zhao  B-W, Su  X-R, Yang  Y  et al.  Regulation-aware graph learning for drug repositioning over heterogeneous biological network. Inf Sci  2025;686:121360. 10.1016/j.ins.2024.121360 [DOI] [Google Scholar]
  • 46. Su  X, Hu  P, Li  D  et al.  Interpretable identification of cancer genes across biological networks via transformer-powered graph representation learning. Nat Biomed Eng  2025;9:371–89. 10.1038/s41551-024-01312-5 [DOI] [PubMed] [Google Scholar]
  • 47. Zhao  B-W, Su  X-R, Hu  P-W  et al.  iGRLDTI: an improved graph representation learning method for predicting drug–target interactions over heterogeneous biological information network. Bioinformatics  2023;39:btad451. 10.1093/bioinformatics/btad451 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48. Zhao  B-W, Su  X-R, Hu  P-W  et al.  A geometric deep learning framework for drug repositioning over heterogeneous information networks. Brief Bioinform  2022;23:bbac384. 10.1093/bib/bbac384 [DOI] [PubMed] [Google Scholar]
  • 49. Vaswani  A, Shazeer  N, Parmar  N  et al. Attention Is All You Need. In: Guyon I, von Luxburg U, Bengio S, Wallach H, Fergus R, Vishwanathan S, Garnett R (eds.), Advances in Neural Information Processing Systems, Vol. 30. Red Hook, NY, USA: Curran Associates, Inc., 2017. [Google Scholar]
  • 50. Wang  W, Shuai  Y, Zeng  M  et al.  DPFunc: accurately predicting protein function via deep learning with domain-guided structure information. Nat Commun  2025;16:70. 10.1038/s41467-024-54816-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51. Selvaraju  RR, Cogswell  M, Das  A  et al.  Grad-CAM: Visual explanations from deep networks via gradient-based localization. In: 2017 IEEE International Conference on Computer Vision (ICCV), pp. 618–26. Venice: IEEE, 2017. 10.1109/ICCV.2017.74 [DOI] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

SSAS-GO-supplement_bbag408

Data Availability Statement

The source codes and datasets of SSAS-GO are available at https://github.com/hlwang613/SSAS-GO.git


Articles from Briefings in Bioinformatics are provided here courtesy of Oxford University Press

RESOURCES