Abstract
The classification of novel protein folds remains a central challenge in structural bioinformatics, particularly as deep learning models like AlphaFold2 dramatically expand the universe of predicted protein structures. In this study, we investigated 664 candidate novel fold (CNF) domains from the TED database that TED and DPAM methods both classified with low confidence. These CNFs span a structurally diverse and largely non-redundant set of domains, most of which lack clear sequence or structural similarity to known folds. Many CNFs appear as insertions into known transmembrane or enzymatic domains, while others occur in modular architectures, co-occurring with interaction or catalytic folds such as β-barrels, zinc fingers, or Rossmann-like domains. Although some CNFs resemble known folds that have undergone topological rearrangements or circular permutations, others result from errors in domain boundary prediction, often due to truncated sequences or tightly packed domain duplications. Our analyses led to the creation of 190 new Pfam families, many classified as domains of unknown function (DUFs), and revealed intriguing cases of zinc-binding and disulfide-rich architectures that contribute to fold space expansion. A small subset of CNFs helped define new superfamilies by linking previously unclassified but structurally related domains. Taken together, this work underscores the importance of integrating structural, evolutionary, and contextual information to resolve challenging fold assignments and provides a roadmap for extending protein classification frameworks into previously uncharted structural territory.
Keywords: Protein domain classification, Candidate novel fold, TED, DPAM, Pfam family, Zinc finger domains, Disulfide-rich domains, Domain of unknown function (DUF)
Summary statement:
Advances in AI, particularly the generation of highly accurate structural models for millions of proteins by tools such as AlphaFold2, have unveiled vast new regions of protein structure space. However, many predicted domains remain challenging to classify. In this study, we investigate hundreds of uncharacterized domains with potential novel folds, uncovering diverse architectures and establishing new protein families and superfamilies. By integrating structural and evolutionary insights, this work expands our understanding of protein fold space and highlights the complexity of interpreting novel domains.
1. Introduction
Identifying and classifying novel protein folds remains one of the most persistent challenges in structural bioinformatics. While many protein domains can be confidently annotated based on sequence or structural homology, the classification of some domains is hindered due to an apparent lack of similarity to any known structural architectures [1]. These domains may represent truly novel folds, characterized by unique topological arrangements of secondary structural elements and no detectable relationship to existing structural families [2]. Alternatively, they may be highly diverged homologs whose sequences and structures have evolved to such an extent that their resemblance to known folds has become less evident [3,4]. The combination of structural divergence, functional ambiguity, and undetectable sequence similarity makes these domains difficult to interpret, categorize, or validate, even when high-confidence structural models are available.
The advent of deep learning-based methods such as AlphaFold2 [5], RoseTTAFold [6], and ESMFold [7] has led to an unprecedented expansion in the availability of predicted protein structures. With the release of over 200 million models in the AlphaFold Protein Structure Database (AFDB) [8], researchers now have an unprecedented opportunity to explore the protein structure universe - including regions that extend well beyond the boundaries of previously known fold space [9–11]. Yet this dramatic increase in structural data also presents a significant challenge: distinguishing truly novel folds from highly divergent homologs or prediction artifacts, particularly in the absence of strong evolutionary or functional signals.
Recently, the TED (The Encyclopedia of Domains) project [10] was developed to identify and classify domain architectures across the entire AFDB, integrating machine learning-based domain segmentation and fast structure-based assignment using tools like Foldseek [12]. This large-scale effort uncovered thousands of domains lacking detectable structural homologs - strong candidates for novel folds. Concurrently, we have developed DPAM, a domain parsing and assignment method that goes beyond TED’s structure-only approach [13]. DPAM combines AlphaFold-derived distance maps (distograms), Predicted Aligned Error (PAE) matrices, and sequence profiles with evolutionary inference using HHpred [14] and high-sensitivity structural alignment via DaliLite [15]. This allows DPAM to detect subtle similarities and assign domains to the ECOD structural hierarchy with improved sensitivity. In addition, recent collaborative efforts of the ECOD and Pfam teams aimed at the harmonization and synchronization of both databases [16,17]. The tighter integration of ECOD with Pfam, as well as the addition of structural predictions, has allowed for more efficient classification of protein domains.
In this study, we focus on the most difficult candidate novel fold domains in TED - specifically, those that TED flagged as novel and DPAM assigned with a low classification probability score. These represent the cases where both structure-based and evolution-informed approaches struggle to provide confident fold-level classification. By closely examining this subset, we aim to determine whether these domains reflect truly novel architectures, extreme structural divergence, limitations of current prediction tools, and boundaries of existing classification systems. Beyond comparing TED and DPAM classifications, our goal is to explore how such challenging domains define the boundaries of current structural classification systems. These domains serve as test cases that reveal both methodological limitations and conceptual ambiguities in defining protein folds and domains. By integrating structural, evolutionary, and contextual analysis, this study seeks to provide a framework for interpreting low-confidence predictions and refining classification strategies across databases. Our analysis offers insight into interpreting low-confidence predictions, refining classification strategies for edge cases, and enhancing understanding of the outer limits of fold space within the predicted protein structure universe.
2. Results and discussion
2.1. Assignment of CNF (Candidate Novel Fold) domains
TED compiled a dataset of 7,427 domains that are structurally well folded yet appear distinct from any known domain folds [10]. We applied our domain-parsing tool, DPAM, to AlphaFold models of the 7232 full-length proteins that contain these TED domains with potential novel folds. A total of 8044 DPAM domains showed significant overlap with these TED domains (see Materials and methods). These DPAM domains were mapped to the ECOD database and classified into four categories: well-assigned domains, unassigned domains, partial domains, and simple-topology domains (see Materials and methods). Among them, 2490 domains (~31%) were well-assigned domains, and small proportions were categorized as partial domains (401 domains, ~5%) and simple-topology domains (333 domains, ~4%). The majority of these DPAM domains (4820 domains, ~60%) were unassigned domains with DPAM confidence scores less than 0.9, indicating limited sequence or structural similarity to known ECOD domains.
We used the overlap fraction as a metric to quantify the consistency between DPAM and TED domain definitions (see Materials and methods). When both DPAM and TED domains exhibit high overlap fractions (above 0.8), their domain definitions are considered consistent. Focusing on the subset of 664 domains with low DPAM confidence scores (<0.5) yet consistent TED/DPAM boundaries, we identified a group of domains that presents the greatest challenge for evolutionary analysis due to the absence of clear sequence or structural similarity. We refer to these as CNF (Candidate Novel Fold) domains. Out of the 664 CNF domains, 428 domains were composed of a single segment, whereas the remaining were defined using sequentially separated parts of the proteins. Nearly 10% of these were models of entries which were made obsolete in the recent Uniprot releases.
Of the 664 CNF domains, only 16 were found to have HHpred hits to ECOD domains with high confidence (HHpred probability scores above 0.9). They received low DPAM confidence scores, likely due to weak structural similarity (reflected by Dali Z-scores). Manual inspection provided further insights into the nature of these apparent inconsistencies. In six cases, including domains such as A0A1E8CLK0_TED02 and A0A315B4N1_TED02, the CNF domain is inserted within larger outer membrane β-barrel domains, complicating structural alignment and likely contributing to the low DPAM confidence. A0A0F6YKA5_TED01 yielded a high-confidence HHpred hit (probability score: 0.95) to PDB entry 6u6g, a small disulfide-rich domain known as the DPY module [18]. Structural analysis confirmed that this TED domain indeed contains a compact DPY module (highlighted in orange) embedded within a larger domain (shown in blue) comprising several β-strands and α-helices (Figure 1A). For A0A820ZPT9_TED02, manual analysis confirmed that the top HHpred hits with probability scores above 0.95 (e.g., PDB: 2xcm, chain E) correspond to Btk/CHORD zinc finger domains [19], which contain two zinc-binding sites. Structural alignment supports this annotation, despite the domain’s overall divergence and the inclusion of several additional α-helices (shown in gray, Figure 1B).
Figure 1. CNF domains with high HHpred probability scores to known structures.

A) A0A0F6YKA5_TED01 contains a DPY module similar to PDB: 1oig. B) A0A820ZPT9_TED02 contains Btk/CHORD zinc finger domains, structurally similar to PDB: 2xcm. C)-H) Six CNF domains contain zinc ribbon motifs (shown in orange) together with other regions (shown in blue).
However, not all high-scoring HHpred hits were supported by structural similarity. For example, A0A6H9LBC1_TED02 returned hits to immunoglobulin-like domains despite clear structural dissimilarity, suggesting a false positive. This example highlights the limitations of relying solely on sequence-based methods and underscores the importance of integrating structural context in novel fold analysis.
Interestingly, six CNF domains with high HHpred scores contain Rubredoxin-like zinc fingers (zinc ribbons) [20]. Two CNF domains (D5SQH0_TED01 and A0A849TGY6_TED01) possess two nested zinc ribbons arranged in a similar fashion (Figure 1, C and D), suggesting they may be remote homologs. The other four (A0A126Q0V0_TED01, A0A3N9Y1R3_TED01, A0A6H1ZRD8_TED01 and A0A101JE48_TED01) contain a single zinc ribbon motif (Figure 1, E–H). These zinc ribbon motifs were inserted into a larger domain that contains additional secondary structural elements. For example, A0A3N9Y1R3_TED01 found an ECOD domain (e2k5cA1) with an HHpred probability of 96.3. Both domains contain a zinc ribbon despite overall structural differences in other parts of their structures.
Detecting evolutionary relationships for small zinc finger domains remains a significant challenge for tools like DPAM. Structural similarity search programs such as DaliLite and Foldseek often perform poorly on these small compact domains, frequently yielding marginal similarity scores due to their limited size and high structural divergence. Zinc finger domains—such as rubredoxin-like folds—pose inherent difficulties for evolutionary analysis because of their minimalistic architecture, modularity, and widespread occurrence across diverse proteins. In some cases, the zinc finger may represent an ancient structural core around which additional secondary structure elements have evolved, gradually giving rise to a more complex fold [20]. In other scenarios, a pre-existing domain may have acquired a zinc-binding motif as a peripheral module to enhance structural rigidity or to enable novel molecular interactions [21]. These diverse modes of integration complicate both sequence- and structure-based attempts to trace their evolutionary origins.
We examined the taxonomic origins of the 664 CNF domains using direct UniProt taxonomy queries. The majority of domains are from bacteria (454; 68.4%) and eukaryotes (141; 21.2%), with smaller contributions from archaea (42; 6.3%) and unclassified sequences derived from metagenomic sources (27; 4.1%). Among eukaryotic entries, most are from unicellular lineages, including Fungi (63 domains) and the SAR supergroup (Stramenopiles, Alveolates, and Rhizaria; 38 domains). Additional minor groups include Amoebozoa (2), Haptophyta (3), Discoba (4), and Choanoflagellata (1). By contrast, metazoan (22 domains) and plant (8 domains) sequences account for only small fractions. Notably, only a single model organism, Dictyostelium discoideum, is represented. This distribution suggests that CNF domains are concentrated in less-characterized prokaryotic and unicellular eukaryotic lineages, rather than in well-studied multicellular model systems.
2.2. Most CNF domains have limited structural similarity to known structures and among themselves
We assessed the structural similarity of CNF domains to known folds using the highest Dali Z-score against the ECOD database (referred to as DALIB). Over half of the CNF domains (386 out of 664) had DALIB scores below 2, indicating minimal structural resemblance to existing ECOD domains. Only 12 domains exhibited DALIB scores above 5. Closer inspection revealed that five of these—such as A0A1Q7WCM7_TED02 (Figure 2A)—matched outer membrane β-barrel structures in DaliLite searches. These domains all feature meander β-sheet architectures. While segments of their β-sheets can be structurally aligned to β-barrel domains, their overall folds differ substantially. This highlights a key limitation: moderately high DALIB scores alone may not be sufficient to infer true homology, particularly when alignment is driven by localized structural features rather than global fold similarity.
Figure 2. Examples of CNF domains with high DALIB scores.

A) A0A1Q7WCM7_TED02 found a DaliLite hit to the outer membrane meander β-barrel domain (PDB: 7cp9, chain I). Two views of different orientations are shown. B) A0A7J4T2G3_TED01 bears structural similarity to an immunoglobulin-like domain (PDB: 3bes, chain R). C) A0A7J4T2G3_TED01 contains a jelly roll fold. D) Two jelly roll modules (one colored rainbow and the other colored purple) defined in the CNF domain A0A3R7BMP6_TED01. PDB IDs and chain IDs are shown for the structural matches.
Inspections of other domains with high DALIB revealed several interesting cases. For example, A0A2U9IC53_TED01 (Figure 2B) contains an immunoglobulin-like domain inserted into and closely packed against a long, wide meander β-sheet (purple in Figure 2B). This domain yields a DALIB score of 7.2 when aligned to an ECOD domain with the outer membrane meander β-barrel fold. However, its top DPAM hit is an immunoglobulin-like domain, supported by a Dali Z-score of 6.3. Other CNF domains with DALIB scores above 5 may reflect instances of remote homology. For example, two CNF domains aligned with ECOD domains containing a jelly roll fold. In one case, A0A7J4T2G3_TED01 (Figure 2C), the CNF domain consists of a canonical jelly roll core with long insertions, suggesting structural elaboration. In the other case, A0A3R7BMP6_TED01, the CNF domain contains a tandem duplication of jelly roll-like modules (shown in cartoon in Figure 2D, one module is rainbow colored and the other is colored in purple), which are tightly packed together and are part of a composite structure with two other jelly roll modules (shown in pink ribbons in Figure 2D). Both TED and DPAM have failed to separate the region into two individual jelly roll-fold domains.
To investigate structural similarities among the 664 CNF domains, we performed an all-against-all structural comparison using DaliLite. The majority of the resulting ~220,000 pairwise comparisons yielded low Dali Z-scores, indicating limited structural similarity across the dataset. Only 14 domain pairs exhibited Dali Z-scores greater than 5, highlighting the overall structural diversity and non-redundancy of these domains. Closer examination of the highest-scoring pairs suggests that some CNF domains may be evolutionarily related. The two top-scoring pairs (A0A6M1NSA2_TED02/A0A7K0E7U2_TED02 with a Dali Z-score of 11.5 and A0A101DGG6_TED02/A0A351RUP4_TED01 with a Dali Z-score of 9.8, Figure 3 A and B) involve CNF domains that are inserted into Rossmann-like fold domains. Given their shared domain architecture, these pairs are likely homologous. Similarly, four additional domain pairs (A0A0S8FXC8_TED03 and A0A382MN27_TED01: Dali Z-score 7.4; A0A315B4N1_TED02 and A0A7X1SQF4_TED02: Dali Z-score 6.4; A0A315B4N1_TED02_nD2 and A0A6G7ZWG8_TED02: Dali Z-score 5.4; A0A4Q3AGZ4_TED01 and A0A6G7ZWG8_TED02: Dali Z-score 5.0) are likely homologous pairs, as they have overall structural similarity and are made up of insertion regions in outer membrane meander β-barrels.
Figure 3. Examples of CNF domain pairs that are structurally similar.

A) A0A6M1NSA2_TED02 and A0A7K0E7U2_TED02 with a Dali Z-score of 11.5. B) A0A101DGG6_TED02 and A0A351RUP4_TED01 with a Dali Z-score of 9.8. C) A0A1I4MD37_TED01 and A0A2V9UWW2_TED01 with a Dali Z-score of 9.7. D) A0A6G1KZ67_TED02 and W6YP97_TED03 with a Dali Z-score of 9.2. E) A0A6N9AP83_TED01 and A0A838GSN0_TED01 with a Dali Z-score of 8.0. F) A0A0C3KVK1_TED02 and A0A4Y9YY98_TED02 with a Dali Z-score of 6.7.
Two CNF domains, A0A1I4MD37_TED01 and A0A2V9UWW2_TED01 (Figure 3C), have a high Dali Z-score of 9.7 and adopt an unusual fold with two kinked β-sheets. This uncommon structural feature and overall structural similarity suggest that they could be homologous. Likewise, A0A6G1KZ67_TED02 and W6YP97_TED03 (Dali Z-score: 9.2, Figure 3D) exhibit similar folds composed of an antiparallel β-sheet formed by two β-hairpins and flanked by several α-helices, suggesting potential homology. Other domain pairs with high Dali-Z-scores and that are potentially homologous include A0A6N9AP83_TED01 and A0A838GSN0_TED01 (Dali Z-score: 8.0, Figure 3E) as well as A0A0C3KVK1_TED02 and A0A4Y9YY98_TED02 (Dali Z-score: 6.7, Figure 3F). A0A2R5FEB3_TED01 and A0A6P1CWQ9_TED01 (Dali Z-score: 6.3) represent a case where one domain (A0A6P1CWQ9_TED01) is likely a sequence fragment.
2.3. CNF domains are often inserted into other domains
According to DPAM domain parsing, 116 out of the 664 CNF domains are inserted within other DPAM domains. Among them, 72 CNF domains are inserted into well-assigned DPAM domains. The most frequent ECOD X-groups of these well-assigned domains that harbor CNF insertions are shown in Figure 4A. The most common host domain of CNF insertions is outer membrane meander β-barrels, with a representative example illustrated in Figure 4B. In this example, the CNF domain (A0A2H0ZAE8_TED02, orange) is embedded within the insertion region of an outer membrane β-barrel domain (blue). β-barrel pore-forming proteins are a structurally diverse and functionally versatile class of membrane proteins that often contain discontinuous domains formed by loop regions connecting β-strands [22,23]. These loops, though not part of the core β-barrel structure, frequently play essential roles in pore function by contributing to substrate selectivity, gating mechanisms, and protein-protein interactions [24].
Figure 4.

A) Distribution of top ECOD X-groups of well-assigned DPAM domains that harbor insertions of CNF domains. B) A CNF domain (orange) inserted into an outer membrane β-barrel domain (blue). C) A CNF domain (orange) inserted into a domain of repetitive α-hairpins (blue). D) A CNF domain (orange) inserted into a domain of Rossmann-like fold (blue).
CNF domains were also identified as insertions in other types of transmembrane domains, including the STT3/PglB/AglB transmembrane domain, voltage-gated ion channels, and type II ABC exporter transmembrane domain. They were also found as insertions into repetitive α-hairpins (Figure 4C) and various enzyme domains such as Rossmann-like (Figure 4D), HAD domain-like, Carbon-nitrogen hydrolase-like, Zincin-like, and P-loop domains-like.
2.4. Domain co-occurrence and interactions of CNF domains
DPAM assignments revealed that 228 CNF domain–containing proteins have only a single DPAM domain. In contrast, 188 proteins contain two DPAM domains, 112 have three, and 136 have four or more domains—including 26 proteins with seven or more. Overall, approximately one-third of the CNF domain-containing proteins (247 out of 664) contain at least one additional unassigned domain. More than half (315 out of 664) contain at least one additional well-assigned DPAM domain, covering 111 ECOD X-groups. The most frequently co-occurring well-assigned domains—each associated with 10 or more CNF domains—belong to these ECOD X-groups: Immunoglobulin-like β-sandwich (50), jelly roll (21), P-loop-like domains (18), helix-turn-helix (HTH) motifs (16), repetitive α-hairpins (14), cradle loop barrel (14), N0-like domains found in phage tail proteins and secretins (12), HAD-like domains (12), outer membrane meander β-barrels (11), and four-helical up-and-down bundles (10).
We assessed the proximity of these domains by defining them as interacting [10] when at least three pairs of Cβ atoms (or Cα atom for glycine residue) from two domains were within 8 Å of each other (see Materials and methods). A total of 422 potential intramolecular domain–domain interactions were identified for 328 CNF domains, indicating that many CNF domains may function jointly with other domains in modular domain architectures. Of these interactions, 178 involve well-assigned DPAM domains, while the remaining 244 involve unassigned DPAM domains. Notably, the most frequent interactions involve CNF domains interacting with common protein–protein interaction modules, including immunoglobulin-like β-sandwiches (20 interactions), jelly roll folds (12 interactions), and repetitive α-hairpins (11 interactions). In addition, CNF domains are also found in association with enzymatic domains such as P-loop NTPase-like (11 interactions) and HAD domain-like folds (seven interactions), suggesting potential roles in catalysis or molecular regulation. These patterns of domain co-occurrence and domain-domain interactions provide valuable insights into the possible functional roles of CNF domains, especially in contexts where structural or sequence homology is insufficient for annotation. As such, examining domain architectures and interaction partners may offer important clues for inferring biological function, molecular mechanism, or evolutionary origin of these novel folds.
2.5. Many CNF domains represent candidate novel domain families.
Limited sequence similarity to known protein families and domains was observed for the CNF domains. InterProScan analysis of the 664 CNF domains showed that only 99 (15%) had detectable matches across major domain databases, including Pfam, PANTHER, NCBIFAM, Gene3D, CDD, SMART and SUPERFAMILY. Of these, only 43 domains (fewer than 7%) matched entries in Pfam. Some Pfam hits correspond to insertions within outer membrane meander β-barrel domains (e.g., PF00593 and PF25183). 15 CNF domains aligned with Pfam entries currently annotated as domains of unknown function (DUFs). These results suggest that the CNF domains represent promising candidates for defining novel protein domain families, thereby expanding the known protein sequence space. Using the sequences of these CNF domains as SEEDs, we constructed 190 new Pfam families. A substantial portion of these families remain uncharacterized and are classified as DUFs (Table S1). Amongst them there were interesting discoveries, which we describe in more detail throughout the next parts of this work.
2.6. Some CNF domains are topological variants of domains with known experimental structure.
A small subset of the CNF domains were identified as variants of known, experimentally characterized structural architectures. Structural comparisons suggest that these variants have probably arisen through diverse evolutionary mechanisms, including insertions, deletions, and circular permutations. Several of these CNF domains present particularly intriguing cases, some of which are described in detail below.
The DUF8237 (PF26707, seed A0A7C6M3Z4_TED03) domain, found in uncharacterized bacterial proteins, exhibits structural similarity to C-type lectin domains — well-known carbohydrate-binding modules (Figure 5, A–B). It shares key features of the canonical C-type lectin fold [25], including a double-looped, two-stranded β-sheet capped by an α-helix and flanked by an additional three-stranded β-sheet. Structural superposition reveals a good alignment with various C-type lectin domains, with root mean square deviation (RMSD) values ranging from 1.88 to 2.5 Å. However, unlike classical C-type lectins, the DUF8237 domain contains a distinctive insertion formed by a six-stranded β-meander (shown in gray in Figure 5B). It also lacks the characteristic calcium-binding motifs typically required for carbohydrate recognition. Notably, the absence of these motifs is not without precedent; for example, the C-terminal domain of Legionella pneumophila Lcl (PDB: 8q4e) also lacks these motifs yet has been shown to recognize sulfated glycosaminoglycans [26]. The DUF8237 domain features several conserved, surface-exposed residues and is frequently found in combination with other domains such as immunoglobulin-like domains Big_2 (PF02368) and Big-like (PF22359), which are commonly associated with bacterial surface proteins. This domain architecture suggests a potential role in adhesion or carbohydrate recognition.
Figure 5.

Side-by-side comparison of A) Structure of the C-terminal domain of L. pneumophila Lcl (PDB: 8q4e, chain C) and B) AlphaFold2 model of a member of DUF8237 (UniProtKB:A0A7C6M3Z4). The distinct β-meander insertion is shown in gray. Side-by-side comparison of C) HemS (PDB:7qxv, chain A), D) ChuX dimer (PDB: 2ovi, chain A) and E) AlphaFold2 model of a member of DUF8240 (UniProtKB:A0A1I1K5F6).
Another interesting finding was the DUF8240 domain (PF26712, seed A0A1I1K5F6_TED01) from uncharacterized proteins from halobacteria. Its predicted β-barrel fold is topologically similar to members of the heme iron utilization protein superfamily (see Figure 5, C–E). This superfamily includes the HemS/ChuS-like family [27], characterized by a tandem duplication of β-barrel domains joined edge-to-edge (Figure 5C), and the ChuX-like family [28] (Figure 5D), whose members form dimeric assemblies resembling individual HemS/ChuS-like protomers. While ChuX homodimers coordinate two heme molecules, HemS and ChuS bind a single heme via their C-terminal domain. The DUF8240 domain lacks the conserved residues required for heme binding. Despite exhibiting greater structural similarity to HemS/ChuS-like domains (with an RMSD of 2.3 Å), DUF8240 domains likely underwent a deletion event that disrupted the canonical heme-binding site.
Circular permutation in protein structures is a versatile evolutionary mechanism where the order of secondary structural elements is sequentially rearranged, effectively creating new N- and C-termini while preserving the overall spatial arrangement [29]. The DUF8212 (PF26640, seed A0A4Q9PEE3_TED03) domain from uncharacterized fungal proteins contains a predicted five-stranded β-barrel, which is structurally similar to the β-barrel of 50S Ribosomal protein L14p/L23e (PF00238) [30] but with a circular permutation (Figure 6, A–C). The C-terminal strand of the β-barrel of DUF8212 occupies the position of the N-terminal strand in the Ribosomal protein L14p/L23e fold. The DUF8212 domain is also structurally similar to the β-barrel domain of Heterokaryon incompatibility protein 6 (HET-6) (PF26639), a key component in the non-self recognition system of filamentous fungi [31]. Thus, these β-barrel domain families define a close evolutionary group that is probably distantly related to the Ribosomal protein L14p/L23e domains.
Figure 6.

Side-by-side comparison of A) Ribosomal protein L14 (PDB:1whi, chain A) and B) A0A4Q9PEE3_TED03 domain. Both are shown in rainbow. C) Topological diagrams showing the circular permutation in these two domains - Ribosomal protein L14 (top) and A0A4Q9PEE3_TED03 (bottom). D) Structure of a classical C2H2 zinc finger domain (PDB: 2eml, chain A). E) The AlphaFold3 model of UniprotKB:A0A7S4IBF0, which predicts the zinc binding sites with a high confidence (ipTM score = 0.82). A0A7S4IBF0 is composed of three structural repeats. F) Topology diagrams showing the location of metal-coordinating residues in classical zinc fingers (left) and in the novel type of CHCC zinc finger (PF26600 and PF26601). Zinc-binding residues are represented by filled black circles. G) A C4 zinc finger domain of FAP1 protein (PDB: 7zw0, chain sh).
Another interesting discovery among the CNF domains is a tandemly repeated structural motif present in uncharacterized eukaryotic proteins, many of which also contain AAA+ ATPase domains. This structural repeat is composed of a β-hairpin followed by an α-helix, with the latter tightly packed against the former. Some repeats contain a large insertion that is observed immediately upstream of the α-helix. Structurally, these repeats are reminiscent of classical C2H2 zinc finger domains (see Figure 6, D–E) and feature highly conserved residues that are potentially involved in metal coordination. However, the spatial arrangement of these residues diverges from that of the canonical C2H2 domains. Three of the putative metal-coordinating residues are located within a single repeat domain, while the fourth residue is contributed by an adjacent repeat (see Figure 6, E–F). This repeat likely represents a new type of CHCC zinc finger, a prediction which is supported by a high-confidence AlphaFold3 model (ipTM score = 0.82). Two new Pfam families, PF26600 and PF26601, have been defined for this repeat architecture, with PF26601 featuring repeats distinguished by a large insertion. Structural relationships of these domains probably extend further and include the structurally analogous zinc finger domain observed in the FAP1 protein [32], in which, however, four cysteine residues coordinate the zinc ion (Figure 6G).
2.7. Sequence fragments hinder protein domain classification
Some CNF domains were falsely classified as novel folds because their sequences, used to build the AlphaFold models, are fragments. In laboratory experiments, protein folding is significantly disrupted when constructs do not encompass the full globular domain or omit residues essential for the hydrophobic core integrity. In contrast, AlphaFold2 can generate structural predictions even for partial sequences. This may result in incomplete domain substructures that do not accurately reflect their true fold. Amongst such fragmented CNF domains, particularly interesting was the discovery of a new family, DUF8202 (PF26628, seed A0A0Q6A750_TED01), of bacterial domains distantly related to ZU5/GAIN domains. ZU5/GAIN domains are adaptable protein-protein interaction modules, participating in different binding contexts by using different surfaces for interaction [33]. The DUF8202 domains superpose well with ZU5/GAIN with RMSD ranging between 1.9–3.0Å and share their common topological features (Figure 7, A–D). Like some ZU5/GAIN domains, ankyrins for example, these domains lack the GPS motif and therefore they are unlikely to carry autoproteolytic activity. The DUF8202 domains appear frequently in combination with domains typical for bacterial adhesins and cell surface proteins, such as DUF11 (PF01345), Big_13 (PF19077) and CHU_C (PF13585), suggesting they could be involved in adhesion.
Figure 7.

Side-by-side comparison of A) A0A0Q6A750_TED01 (UniProtKB:A0A0Q6A750) (partial sequence assigned as a CNF domain), B) DUF8202 domain of UniProtKB:A0A086AVG2 (a full-length domain), C) GAIN domain of CIRL 1/Latrophilin 1 (PDB: 4dlq, chain A) and D) ZU5 domain of human erythrocyte ankyrin (PDB: 3ud1, chain A). Side-by-side comparison of the AlphaFold2 models of members of DUF8188 (PF26603) E) X6Q0D4_TED01 UniProtKB:X6Q0D4 (Prevotella sp., a fragment) and F) UniProtKB:C9MQ94 (Prevotella veroralis, a full-length domain); and G) Pertussis toxin subunit S5 (PDB: 1prt, chain L).
Another similar finding was the DUF8188 (PF26603) family that was identified within the full-length proteins. The DUF8188 domain, present in uncharacterized proteins from gram-negative bacteria, structurally resembles the typical OB-fold β-barrel [34]. It features the characteristic five-stranded β-barrel capped by an α-helix (Figure 7F). In contrast, the DUF8188 domain X6Q0D4_TED01 - though classified as a CNF based on its sequence fragment - lacks the N-terminal β-strand, resulting in an open, four-stranded β-barrel structure (Figure 7E). Full-length DUF8188 domains display greater similarity to the β-barrels domains of Bacterial enterotoxins (Figure 7, E–G). These domains typically occur as standalone units, often preceded by a hydrophobic α-helix that is predicted as a transmembrane segment.
2.8. CNF domains are enriched with zinc fingers and metal-binding domains
In addition to the cases of CNF domains harboring zinc ribbon motifs described above, several interesting cases of zinc finger domains were found. Some domains, especially small domains, could be wrongly classified as novel domains because they are made up of duplicated domains tightly packed against each other. A0A0D2ANI5_TED01, for example, was classified as a single CNF domain despite being composed of a pair of treble clef zinc fingers [35] (Figure 8A). This prediction was confidently supported by AlphaFold3 (ipTM = 0.75 and pTM = 0.7). These two domains are found in uncharacterized fungal proteins and are now members of the new PF26647 and PF26648 families.
Figure 8.

A) AlphaFold3 model of the pair of zinc fingers in A0A0D2ANI5_TED01 colored in rainbow (left). The first treble clef domain (green) of A0A0D2ANI5_TED01 superposed to PDB: 3h15 (magenta) (middle), and the second treble clef domain (green) of A0A0D2ANI5_TED01 superposed to PDB: 3h15 (magenta) (right). B) AlphaFold3 prediction of the A0A167FI29_TED02 bound to zinc (ipTM = 0.97 and pTM = 0.91). C) AlphaFold3 model of A0A820FDY3_TED01 featuring three zinc binding sites.
A different scenario applied to the A0A167FI29_TED02 domain, which is now classified as a member of the Pfam family PF26625. This domain shows remote similarity to the PPR1 yeast transcription factor (PDB: 1pyi) and related Zn2Cys6 DNA-binding domains [36]. The residues responsible for zinc coordination are highly conserved within the PF26625 family, suggesting that this domain likely forms a binuclear zinc cluster, with two zinc atoms coordinated by six cysteine residues (Figure 8B). Compared to its distant homologues, the A0A167FI29_TED02 domain is more elaborate, featuring additional secondary structural elements. Some homologues within the PF26625 family may have lost the second zinc-binding site, potentially impacting their functional properties.
Some of the zinc finger CNF domains could represent genuine novel folds. One such domain is A0A820FDY3_TED01, now classified as a member of the DUF8206 (PF26633) family. It contains multiple highly conserved cysteine and histidine residues, which we hypothesized are involved in zinc ion coordination. This hypothesis is supported by the AlphaFold3 prediction (ipTM = 0.93 and pTM = 0.87), which reveals three putative zinc-binding sites, including a binuclear zinc cluster (Figure 8C). The DUF8206 domain is found in uncharacterized eukaryotic proteins. In many cases, it is preceded by a P-loop-containing domain. Nearly half of these eukaryotic proteins also contain an N-terminal pore-forming MACPF_SNTX domain (PF24674) [37], suggesting a potential functional association.
2.9. CNF domains are enriched with duplications of small disulfide-rich domains
The assignment of duplicated domains as a single domain with a novel fold represents a significant analytical error that could lead to mischaracterization of protein architecture and evolutionary relationships. What appears to be a single, large domain with an unusual fold could actually be composed of two or more duplicated units. The consequence of this misidentification is the artificial creation of a “novel” fold classification that obscures the true evolutionary origin of the protein, potentially leading to incorrect functional predictions and hampering comparative analyses with homologous proteins. Failure of automatic domain parsing programs to separate tightly packed structural units can arise from inadequate sequence analysis and the inability of structure comparison programs to detect subtle similarity within existing structural databases. This error is compounded by the fact that small domains often lack the distinctive structural landmarks that facilitate proper domain boundary identification, leading to the artificial inflation of fold diversity in structural databases and potentially obscuring important evolutionary relationships between protein families that share these duplicated small domain architectures. Proper domain boundary prediction algorithms and careful structural alignment with known domain families are essential to improve domain parsing programs to avoid these misclassifications and ensure accurate protein fold annotation.
We identified various such cases of CNF domains, including duplications of jelly roll fold domains (Figure 2D) and zinc fingers (Figure 1 C–D and Figure 8A). In addition, we found that small disulfide-rich domains are particularly prone to this type of classification error. We analyzed 27 disulfide-rich CNF domains and identified 12 cases of domain duplications or repeats (A0A7J3TU99_TED03, A0A2E8K025_TED03, A0A075WIA8_TED01, A0A2D6MKY8_TED02, A0A835SZW4_TED01, A0A378IH57_TED01, A0A1W0WMH1_TED01, F2UA55_TED01, A0A7S1S8Q0_TED01, A0A850GIF6_TED02, A0A356VJ87_TED01, and A0A2W5YWX4_TED01), as shown in Figure 9A–L.
Figure 9. Disulfide-rich CNF domains consisting of duplications or repeats.

A-I) CNF domains with a duplication of two structurally similar units. The N- and C-terminal units are colored blue and orange, respectively. J-L) CNF domains with three or more structurally similar units. Units are colored blue, orange, yellow, red, and brown from N-terminus to C-terminus. Non-duplicated regions are colored gray. Disulfide bonds are shown in sticks. M) AlphaFold2 model of a member of DUF8241 (UniProtKB:A0A2P6VIE2) with the three repeating units colored blue, orange and yellow from N-terminus to C-terminus. N) A DUF8241 individual repeat (repeat1) colored in rainbow. O) A defensin C-terminal domain (PDB: 2lew, chain A) colored in rainbow. Cysteine residues forming disulfide bonds are shown in sticks. P) An alignment of the three DUF8241 repeats (UniProtKB:A0A2P6VIE2) and defensin (PDB: 2lew, chain A). Cysteine residues forming disulfide bonds are colored in the same way.
Among the identified disulfide-rich CNF domains, one particularly intriguing discovery was the DUF8241 domain (PF26721, seed A0A835SZW4_TED01), found in uncharacterized plant proteins. While this domain in A0A835SZW4_TED01 corresponds to a sequence fragment of two structural repeats (Figure 9E), in related proteins it comprises three internal structural repeats (one example, A0A2P6VIE2, is shown in Figure 9M). Each repeating unit (e.g., repeat1 in Figure 9N) adopts a topology reminiscent of the α-defensin domain [38] (Figure 9O). Each repeat contains two predicted disulfide bonds that are highly conserved across the family and occupy spatial positions analogous to two of the disulfide bonds found in defensin, which is also reflected in their sequence alignment (Figure 9P). The DUF8241 domain is frequently observed in association with other domains, particularly the EGF_Teneurin domain (Pfam: PF23106) and the zf-C3HC4_3 domain (Pfam: PF13920).
2.10. Some CNF domains contribute to the discovery of new superfamilies (clans)
Structural similarities among CNF domains have been rarely observed. Such instances, however, are of particular significance as they enable the identification of novel superfamilies comprising distantly related homologues. A salient example is the identification of a novel superfamily formed by three CNF domains, each displaying significant structural similarity to the bacterial chaperone protein CcmS (PF26619) (Figure 10). CcmS is functionally implicated in the assembly of beta-carboxysomes, mediating interactions with other beta-carboxysome-associated proteins, and specifically engaging with the C-terminal extension of the carboxysome shell protein CcmK1 [39]. CcmS has not been classified in the ECOD and CATH structural databases, as its structures have been relatively recently released in PDB.
Figure 10.

Side-by-side members of the Pfam clan CL0912 - A) CcmS PF26619 (PDB: 8zlz, chain A); B) A0A0C3KVK1_TED02 - PF26617; C) A0A4Y9YY98_TED02 - PF26646; D) A0A409W9H6_TED01 - PF26632; E) Ajm-1 - PF26649 (UniProtKB:C9J069).
Members of this newly defined superfamily adopt a unique α/β fold, characterized by a central, curved, four-stranded mixed β-sheet flanked on both sides by α-helices. The β-strands are arranged in the order 2-1-3-4, with strands 1 and 3 oriented in parallel. A distinguishing structural feature of this superfamily is the presence of a kinked α-helix that closely associates with the β-sheet (Figure 10).
Besides CcmS, this superfamily currently includes the CNF domains A0A0C3KVK1_TED02 (CcmS-like, PF26617), A0A409W9H6_TED01 (DUF8205, PF26632), and A0A4Y9YY98_TED02 (DUF8214, PF26646), as well as the Ajm-1 protein (Ajm-1, PF26649). Ajm-1 protein is essential for the maintenance of adherens junction integrity in Caenorhabditis elegans and is required for appropriate timing and completion of embryonic elongation [40]. All aforementioned families are constituent members of Pfam clan CL0912. Notably, many of these domains are found in combination with the MYND-type zinc finger.
3. Conclusion
This study presents a systematic exploration of candidate novel fold (CNF) domains within the TED dataset using the DPAM domain assignment framework. Our analysis highlights the complexity and diversity of CNF domains, which often resist straightforward classification due to minimal sequence conservation, structural modularity, and frequent embedding within known domain architectures. While many CNFs do not resemble known folds, others were revealed to be variants of established structures, such as zinc fingers, jelly roll β-sandwiches, and β-barrels, frequently complicated by insertions, duplications, or circular permutations. We identified critical sources of misclassification, including domain boundary errors caused by fragmentary sequences and the misidentification of tightly packed repeats as single domains. These findings underscore the necessity of combining multiple approaches—sequence similarity, structural alignment, domain co-occurrence, and manual curation—to accurately interpret ambiguous domain predictions. Importantly, this work contributes to the broader effort of expanding the domain classification landscape by establishing 190 new Pfam families and identifying potential new superfamilies. The CNF domains investigated here enrich our understanding of remote homology, modular protein evolution, and the outer limits of protein fold space. As structure prediction and classification tools continue to improve, integrating contextual, evolutionary, and structural information will remain essential for robust domain annotation. Our findings not only provide a foundation for refining fold classification strategies but also highlight intriguing candidates for experimental validation and functional characterization.
It is important to emphasize that the goal of this study is not merely to compare the TED and DPAM methodologies, but to investigate the subset of protein domains that challenge both approaches. These cases reveal conceptual limits in the current understanding of what constitutes a domain or fold, emphasizing the fluidity of these definitions in the context of AI-based predictions. We also acknowledge that some CNF domains originate from uncharacterized bacterial sequences, including entries later removed from or updated in UniProt. Nevertheless, these examples remain valuable as they highlight how prediction and annotation errors can inform improvements in automated classification. Our findings demonstrate that examining inconsistencies across tools provides critical insights for refining domain parsing algorithms and for guiding curation in databases such as Pfam and ECOD.
4. Materials and methods
4.1. DPAM domain parsing for proteins containing TED domains with putative novel folds
The TED database reported a dataset of 7,427 TED domains with putative novel folds [10]. We applied DPAM [13] to predict domains from the AlphaFold models of the 7232 full-length proteins that contain these TED domains. The method incorporates the predicted aligned error (PAE) provided with AlphaFold structures to exclude regions likely to be disordered or serve as inter-domain linkers. DPAM estimates the probability that two residues belong to the same domain based on their spatial proximity in the 3D structure, the PAE between them, and whether the pair aligns to the same ECOD domain through sequence (HHsearch) or structural (DaliLite) similarity. These pairwise probabilities are then used to cluster 5-residue segments into putative domains. DPAM then utilizes a Neural Network [41] to classify the parsed domains in DPAM partitions and assign domains to ECOD homologous groups.
We consider those globular domains with a DPAM confidence score above 0.9 to an ECOD entry and with a significant number of secondary structural elements (SSEs) as well-assigned domains. There are three categories of domains and regions that cannot be confidently assigned by DPAM to an ECOD homologous group. 1) “Unassigned'' domains are globular domains with a significant number of SSEs that cannot be confidently assigned to a single homologous group (DPAM confidence score < 0.9). Often, these unassigned domains are sufficiently distant from known ECOD domains that additional expert considerations are needed, such as cofactor binding or functional inference from literature, to make an assignment. 2) “Simple topology” domains are composed of one or two SSEs and may contain intrinsically disordered or poorly predicted regions. These domains may have significant stability provided by non-SSEs (e.g., disulfide bonds, cofactors or metal binding sites). These domains may also be components of larger protein complexes and lack a globular structure outside of this broader context. 3) “Partial” domains have confident homology to a much longer reference ECOD domain. These domains are usually the result of errors either in the query protein (i.e., genome annotation) or due to inconsistency within the reference set with respect to repeat and duplication (i.e., a query domain hitting a homologous domain duplication incorrectly annotated as a single domain).
The overlap fraction quantifies the agreement between domain boundaries defined by DPAM and TED. For any DPAM domain and TED domain from the same protein, we calculate two overlap fractions: the DPAM overlap fraction (the number of residues common to DPAM and TED divided by DPAM domain length) and the TED overlap fraction (the number of residues common to DPAM and TED divided by TED domain length). We consider any DPAM domain to be a domain of candidate novel fold if it has a low DPAM confidence score (less than 0.5) and has a consistent domain definition with a TED domain with putative novel fold (both the DPAM overlap fraction and the TED overlap fraction are above 0.8).
We used a similar procedure reported in TED [10] to find the intramolecular interactions of domains: two domains were defined as interacting if at least three Cβ atom pairs (Cα used for glycine residues) were found within 8 Å between them.
4.2. Manual analysis of protein domains
We manually examined a subset of CNF domains by analyzing sequence and structural similarity search results. To aid interpretation, sequence conservation was calculated by AL2CO [42] and mapped to structures with conserved residues and disulfide-bond-forming cysteines highlighted in PyMOL. We inspected the HHpred [14] results against the PDB [43] and Pfam databases [17], paying particular attention to conserved motifs among weak hits. Structural similarity searches were performed by DaliLite [17], and for some proteins also by the FoldSeek server [12]. Functional associations were explored using the STRING web server [44].
4.3. Creation of new Pfam families
The sequences of the CNF domains were used as initial seeds to search the reference proteome database using hmmsearch (HMMER version 3.3.2) [45]. Pfam families were built via repeated iterative searches, which included manual refinement of boundaries, member selection and inclusion thresholds. Models of representative domains from selected families, bound to zinc, were generated with the AlphaFold3 Server [46] and were deposited in ModelArchive (modelarchive.org) with accession codes ma-mtbiy, ma-qhwv4, ma-9n1zk, ma-vho72.
Supplementary Material
Supplementary Table S1. Pfam families created from the CNF set of TED and DPAM domains.
Acknowledgements
The study is supported by grants from the National Institute of General Medical Sciences of the National Institutes of Health GM127390 (to N.V.G.), GM147367 (to R.D.S), 1R35GM160468-01 (to Q.C.), the Welch Foundation I-1505 (to N.V.G.), the Welch Foundation I-2095-20220331 (to Q.C.), the National Science Foundation DBI 2224128 (to N.V.G.), the National Science Foundation MED240004, and the Biotechnology and Biological Sciences Research Council and the National Science Foundation Directorate for Biological Sciences (no. BB/X012492/1 to A.B.). This work was performed in part using computational resources provided by the Texas Advanced Computing Center (TACC). Q.C. is a Southwestern Medical Foundation-endowed scholar. The authors thank Dr. Lisa Kinch for helpful discussions.
Footnotes
Conflict Of Interest Statement: The authors declare no conflicts of interest.
Data availability
The data of well-assigned DPAM domains that overlap with TED novel-fold domains are available in http://conglab.swmed.edu/ted_web/ted_50_80.html. Models are available from ModelArchive (http://www.modelarchive.org/doi/10.5452/ma-vho72 (password:rqnygTi7e6), http://www.modelarchive.org/doi/10.5452/ma-mtbiy (password:Cc6y4GOfvQ), http://www.modelarchive.org/doi/10.5452/ma-qhwv4 (password: ROWVh13kKG), http://www.modelarchive.org/doi/10.5452/ma-9n1zk (password:Pn8z5Jrjpp)
References
- [1].Durairaj J, Waterhouse AM, Mets T, Brodiazhenko T, Abdullah M, Studer G, et al. Uncovering new families and folds in the natural protein universe. Nature 2023;622:646–53. 10.1038/s41586-023-06622-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [2].Waman VP, Bordin N, Alcraft R, Vickerstaff R, Rauer C, Chan Q, et al. CATH 2024: CATH-AlphaFlow Doubles the Number of Structures in CATH and Reveals Nearly 200 New Folds. J Mol Biol 2024;436:168551. 10.1016/j.jmb.2024.168551. [DOI] [PubMed] [Google Scholar]
- [3].Schaeffer RD, Kinch L, Medvedev KE, Pei J, Cheng H, Grishin N. ECOD: identification of distant homology among multidomain and transmembrane domain proteins. BMC Mol Cell Biol 2019;20:18. 10.1186/s12860-019-0204-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [4].Pei J, Schaeffer RD, Cong Q, Grishin NV. Case Studies of Orphan Domain Reclassification in ECOD by Expert Curation. Proteins 2025. 10.1002/prot.26840. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [5].Jumper J, Evans R, Pritzel A, Green T, Figurnov M, Ronneberger O, et al. Highly accurate protein structure prediction with AlphaFold. Nature 2021;596:583–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [6].Humphreys IR, Pei J, Baek M, Krishnakumar A, Anishchenko I, Ovchinnikov S, et al. Computed structures of core eukaryotic protein complexes. Science 2021;374:eabm4805. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [7].Lin Z, Akin H, Rao R, Hie B, Zhu Z, Lu W, et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science 2023;379:1123–30. [DOI] [PubMed] [Google Scholar]
- [8].Varadi M, Anyango S, Deshpande M, Nair S, Natassia C, Yordanova G, et al. AlphaFold Protein Structure Database: massively expanding the structural coverage of protein-sequence space with high-accuracy models. Nucleic Acids Res 2022;50:D439–44. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [9].Durairaj J, Waterhouse AM, Mets T, Brodiazhenko T, Abdullah M, Studer G, et al. Uncovering new families and folds in the natural protein universe. Nature 2023;622:646–53. 10.1038/s41586-023-06622-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [10].Lau AM, Bordin N, Kandathil SM, Sillitoe I, Waman VP, Wells J, et al. Exploring structural diversity across the protein universe with The Encyclopedia of Domains. Science 2024;386:eadq4946. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [11].Pei J, Andreeva A, Chuguransky S, Pinto BL, Paysan-Lafosse T, Schaeffer RD, et al. Bridging the gap between sequence and structure classifications of proteins with AlphaFold Models. J Mol Biol 2024;436:168764. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [12].Van Kempen M, Kim SS, Tumescheit C, Mirdita M, Lee J, Gilchrist CL, et al. Fast and accurate protein structure search with Foldseek. Nat Biotechnol 2024;42:243–6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [13].Zhang J, Schaeffer RD, Durham J, Cong Q, Grishin NV. DPAM: A domain parser for AlphaFold models. Protein Sci 2023;32:e4548. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [14].Gabler F, Nam S, Till S, Mirdita M, Steinegger M, Söding J, et al. Protein sequence analysis using the MPI bioinformatics toolkit. Curr Protoc Bioinforma 2020;72:e108. [DOI] [PubMed] [Google Scholar]
- [15].Holm L, Park J. DaliLite workbench for protein structure comparison. Bioinformatics 2000;16:566–7. [DOI] [PubMed] [Google Scholar]
- [16].Schaeffer RD, Medvedev KE, Andreeva A, Chuguransky SR, Pinto BL, Zhang J, et al. ECOD: integrating classifications of protein domains from experimental and predicted structures. Nucleic Acids Res 2025;53:D411–8. 10.1093/nar/gkae1029. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [17].Paysan-Lafosse T, Andreeva A, Blum M, Chuguransky SR, Grego T, Pinto BL, et al. The Pfam protein families database: embracing AI/ML. Nucleic Acids Res 2025;53:D523–34. 10.1093/nar/gkae997. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [18].Wilkin MB, Becker MN, Mulvey D, Phan I, Chao A, Cooper K, et al. Drosophila dumpy is a gigantic extracellular protein required to maintain tension at epidermal-cuticle attachment sites. Curr Biol CB 2000;10:559–67. 10.1016/s0960-9822(00)00482-6. [DOI] [PubMed] [Google Scholar]
- [19].Kaur G, Subramanian S. Evolutionary relationship between the cysteine and histidine rich domains (CHORDs) and Btk-type zinc fingers. Bioinformatics 2018;34:1981–5. [DOI] [PubMed] [Google Scholar]
- [20].Krishna SS, Majumdar I, Grishin NV. Structural classification of zinc fingers: survey and summary. Nucleic Acids Res 2003;31:532–50. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [21].Bochkareva E, Korolev S, Lees-Miller SP, Bochkarev A. Structure of the RPA trimerization core and its role in the multistep DNA-binding mechanism of RPA. EMBO J 2002. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [22].Mayse LA, Movileanu L. Gating of β-Barrel Protein Pores, Porins, and Channels: An Old Problem with New Facets. Int J Mol Sci 2023;24:12095. 10.3390/ijms241512095. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [23].Wzorek JS, Lee J, Tomasek D, Hagan CL, Kahne DE. Membrane integration of an essential β-barrel protein prerequires burial of an extracellular loop. Proc Natl Acad Sci U S A 2017;114:2598–603. 10.1073/pnas.1616576114. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [24].Tamm LK, Hong H, Liang B. Folding and assembly of beta-barrel membrane proteins. Biochim Biophys Acta 2004;1666:250–63. 10.1016/j.bbamem.2004.06.011. [DOI] [PubMed] [Google Scholar]
- [25].McMahon SA, Miller JL, Lawton JA, Kerkow DE, Hodes A, Marti-Renom MA, et al. The C-type lectin fold as an evolutionary solution for massive sequence variation. Nat Struct Mol Biol 2005;12:886–92. [DOI] [PubMed] [Google Scholar]
- [26].Rehman S, Antonovic AK, McIntire IE, Zheng H, Cleaver L, Baczynska M, et al. The Legionella collagen-like protein employs a distinct binding mechanism for the recognition of host glycosaminoglycans. Nat Commun 2024;15:4912. 10.1038/s41467-024-49255-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [27].Lyles KV, Eichenbaum Z. From Host Heme To Iron: The Expanding Spectrum of Heme Degrading Enzymes Used by Pathogenic Bacteria. Front Cell Infect Microbiol 2018;8:198. 10.3389/fcimb.2018.00198. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [28].Suits MDL, Lang J, Pal GP, Couture M, Jia Z. Structure and heme binding properties of Escherichia coli O157:H7 ChuX. Protein Sci Publ Protein Soc 2009;18:825–38. 10.1002/pro.84. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [29].Bliven S, Prlić A. Circular permutation in proteins. PLoS Comput Biol 2012;8:e1002445. 10.1371/journal.pcbi.1002445. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [30].Davies C, Gerchman SE, Kycia JH, McGee K, Ramakrishnan V, White SW. Crystallization and preliminary X-ray diffraction studies of bacterial ribosomal protein L14. Acta Crystallogr D Biol Crystallogr 1994;50:790–2. 10.1107/S0907444994004117. [DOI] [PubMed] [Google Scholar]
- [31].Saupe SJ. Molecular genetics of heterokaryon incompatibility in filamentous ascomycetes. Microbiol Mol Biol Rev MMBR 2000;64:489–502. 10.1128/MMBR.64.3.489-502.2000. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [32].Li S, Ikeuchi K, Kato M, Buschauer R, Sugiyama T, Adachi S, et al. Sensing of individual stalled 80S ribosomes by Fap1 for nonfunctional rRNA turnover. Mol Cell 2022;82:3424–3437.e8. 10.1016/j.molcel.2022.08.018. [DOI] [PubMed] [Google Scholar]
- [33].Liao Y, Pei J, Cheng H, Grishin NV. An ancient autoproteolytic domain found in GAIN, ZU5 and Nucleoporin98. J Mol Biol 2014;426:3935–45. 10.1016/j.jmb.2014.10.011. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [34].Murzin AG. OB(oligonucleotide/oligosaccharide binding)-fold: common structural and functional solution for non-homologous sequences. EMBO J 1993;12:861–7. 10.1002/j.1460-2075.1993.tb05726.x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [35].Grishin NV. Treble clef finger--a functionally diverse zinc-binding structural motif. Nucleic Acids Res 2001;29:1703–14. 10.1093/nar/29.8.1703. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [36].Marmorstein R, Carey M, Ptashne M, Harrison SC. DNA recognition by GAL4: structure of a protein-DNA complex. Nature 1992;356:408–14. 10.1038/356408a0. [DOI] [PubMed] [Google Scholar]
- [37].Ellisdon AM, Reboul CF, Panjikar S, Huynh K, Oellig CA, Winter KL, et al. Stonefish toxin defines an ancient branch of the perforin-like superfamily. Proc Natl Acad Sci U S A 2015;112:15360–5. 10.1073/pnas.1507622112. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [38].Andersson HS, Figueredo SM, Haugaard-Kedström LM, Bengtsson E, Daly NL, Qu X, et al. The α-defensin salt-bridge induces backbone stability to facilitate folding and confer proteolytic resistance. Amino Acids 2012;43:1471–83. 10.1007/s00726-012-1220-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [39].Cheng J, Li C-Y, Meng M, Jian-Xun Li, Liu S-J, Cao H-Y, et al. Molecular interactions of the chaperone CcmS and carboxysome shell protein CcmK1 that mediate β-carboxysome assembly. Plant Physiol 2024;196:1778–87. 10.1093/plphys/kiae438. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [40].Köppen M, Simske JS, Sims PA, Firestein BL, Hall DH, Radice AD, et al. Cooperative regulation of AJM-1 controls junctional integrity in Caenorhabditis elegans epithelia. Nat Cell Biol 2001;3:983–91. 10.1038/ncb1101-983. [DOI] [PubMed] [Google Scholar]
- [41].Schaeffer RD, Zhang J, Kinch LN, Pei J, Cong Q, Grishin NV. Classification of domains in predicted structures of the human proteome. Proc Natl Acad Sci U S A 2023;120:e2214069120. 10.1073/pnas.2214069120. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [42].Pei J, Grishin NV. AL2CO: calculation of positional conservation in a protein sequence alignment. Bioinforma Oxf Engl 2001;17:700–12. 10.1093/bioinformatics/17.8.700. [DOI] [PubMed] [Google Scholar]
- [43].Burley SK, Bhikadiya C, Bi C, Bittrich S, Chao H, Chen L, et al. RCSB Protein Data Bank (RCSB.org): delivery of experimentally-determined PDB structures alongside one million computed structure models of proteins from artificial intelligence/machine learning. Nucleic Acids Res 2023;51:D488–508. 10.1093/nar/gkac1077. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [44].Szklarczyk D, Kirsch R, Koutrouli M, Nastou K, Mehryary F, Hachilif R, et al. The STRING database in 2023: protein-protein association networks and functional enrichment analyses for any sequenced genome of interest. Nucleic Acids Res 2023;51:D638–46. 10.1093/nar/gkac1000. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [45].Potter SC, Luciani A, Eddy SR, Park Y, Lopez R, Finn RD. HMMER web server: 2018 update. Nucleic Acids Res 2018;46:W200–4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [46].Abramson J, Adler J, Dunger J, Evans R, Green T, Pritzel A, et al. Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature 2024;630:493–500. 10.1038/s41586-024-07487-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Supplementary Table S1. Pfam families created from the CNF set of TED and DPAM domains.
Data Availability Statement
The data of well-assigned DPAM domains that overlap with TED novel-fold domains are available in http://conglab.swmed.edu/ted_web/ted_50_80.html. Models are available from ModelArchive (http://www.modelarchive.org/doi/10.5452/ma-vho72 (password:rqnygTi7e6), http://www.modelarchive.org/doi/10.5452/ma-mtbiy (password:Cc6y4GOfvQ), http://www.modelarchive.org/doi/10.5452/ma-qhwv4 (password: ROWVh13kKG), http://www.modelarchive.org/doi/10.5452/ma-9n1zk (password:Pn8z5Jrjpp)
