Skip to main content
Protein Science : A Publication of the Protein Society logoLink to Protein Science : A Publication of the Protein Society
. 2026 May 20;35(6):e70636. doi: 10.1002/pro.70636

Sketching microprotein portraits

Gabriel Diaz 1, Philippe Valenti 2, Marc Gueroult 1, Simon Marques‐Prieto 2, Kenza Benachenhou 1, Jennifer Zanet 2,, Matthieu Chavent 1,
PMCID: PMC13240139  PMID: 42159250

Abstract

The illustrations of intricate molecular machineries inside cells created by David Goodsell continue to inspire the scientific community. Here, we aim to extend his artworks to include microproteins, a newly recognized class of small proteins with less than 100 amino acids, encoded by small open reading frames. Given the rapidly expanding number of identified microproteins, potentially exceeding the number of canonical proteins, we highlight, in this perspective article, diverse computational approaches to classify these proteins. By predicting localization, assessing structural homology, and modeling environments and dynamics of microproteins, these methods could provide clues about the subcellular localization of these microproteins and their structural domain homology, guiding further investigation into their biological functions in living systems.

Keywords: membrane, microproteins, modeling, organelles, structure prediction, subcellular localization

1. INTRODUCTION

David Goodsell has painted stunning molecular landscapes that have inspired scientists from molecular biologists to computational physicists. His precise renderings of cellular panoramas (Goodsell, 2011) serve as gateways into the nanoworld, offering a powerful means to grasp the complexity and intertwining of biological systems and molecules. At the intersection of art and science (Goodsell, 2021), the work of David Goodsell perpetuates a rich visual tradition pioneered by scientist‐artists like Ernst Haeckel (Haeckel, 1899), Santiago Ramon y Cajal (Gomez‐Marin, 2022), Roger Hayward* (Pauling et al., 1954), and Irving Geis (Dickerson, 1997). Here, we wish to pay homage to his legacy by adopting his artistic style to illuminate recent advances in biology and computational science, particularly in charting the universe of microproteins, a field that is rapidly reshaping our understanding of the coding potential of our genomes and, consequently, of cell biology (Callaway, 2025).

Microproteins are encoded by small open reading frames (small ORFs or smORFs), typically defined as having less than 100 codons. Due to their short length, these smORFs were historically excluded from annotation pipelines and classified as non‐coding to avoid the annotation of false ORFs (Oliver et al., 1992). Therefore, these microproteins were largely ignored until pioneering works identified some of them thanks to developmental genetic approaches in arthropods (Galindo et al., 2007; Kondo et al., 2007; Savard et al., 2006). Following these discoveries, the development of ribosome profiling, which allows capturing ribosome‐bound RNA sequences, and advances in mass spectrometry and bioinformatics unveiled the coding potential of genomes and the translation of thousands of smORFs into microproteins (Chen et al., 2020; Chothani et al., 2026; Ingolia et al., 2012; Prensner et al., 2023). Of note, due to the various methodologies used to search for coding smORFs and the different backgrounds of the associated research teams, this resulted in a disparate nomenclature used to name these microproteins. Indeed, a variety of terms, such as miniproteins, non‐canonical ORFs (ncORFs), alternative ORFs (AltORFs), micropeptides, SEP (Small ORF Encoded Peptides) or smORF peptides, have been used to describe them. Here, we use the term microproteins, which has emerged as the most widely used term within the scientific community. The family of microproteins constitutes the microproteome, also named the ghost or dark proteome, as it refers to the historically ignored part of the coding genome (Wright et al., 2022).

Currently, the number of microproteins continues to increase rapidly, and the most recent estimated number of microproteins approaches, or even exceeds, the number of canonical proteins (i.e., proteins composed of more than 100 amino acids) (Chothani et al., 2026; Prensner et al., 2023). To date, researchers have only scratched the surface of this huge pool of putative bioactive molecules. Some of these have been shown to regulate numerous cellular processes controlling development (Chanut‐Delalande & Zanet, 2024), immunity (Prins & Billerbeck, 2025), mitochondrial function (Liang et al., 2022), metabolism in eukaryotes (Zheng & Xiang, 2022) as well as in bacteria (Fesenko et al., 2025) or the development of diseases, such as cancer (Posner et al., 2023; Ruiz‐Orera & Hübner, 2025). This highlights the microproteome as an unexplored pool of new potential biological regulators, posing the challenge of determining whether these microproteins have functional roles and, if so, uncovering the mechanisms by which they act.

To yield insights into the functions of microproteins, functional screens can be carried out in cellulo (Chen et al., 2020; Sandmann et al., 2023; Schlesinger et al., 2025) or in vivo (Markus et al., 2023; Treichel & Bazzini, 2022). Once a microprotein with potential biological activity has been identified, comprehensive functional analyses based on genetic, molecular, and cellular biology are necessary to determine its role in vivo. This top‐down approach is time consuming, expensive, and cannot be scaled up to study numerous microproteins at once. To address these challenges, emerging computational methods, including machine learning‐based approaches, bioinformatics, and molecular modeling, offer the possibility to invert the traditional experimental trajectory by inferring functions based on protein sequences or structures (Middendorf et al., 2024; Whited et al., 2024). Nevertheless, these computational approaches may not be applicable for the entire pool of microproteins as some of them might be unstable as seen in other organisms (Cuevas et al., 2021; Kesner et al., 2023).

In this perspective, we would like to highlight computational approaches to predict subcellular localization and three‐dimensional (3D) structures of proteins to hypothesize potential microprotein functions. We will use here annotated microproteins from D. melanogaster to illustrate our strategy on a tractable dataset.

2. PREDICTING SUBCELLULAR LOCALIZATIONS AS A FIRST STEP TO INFER MICROPROTEIN FUNCTION

We obtained two sets of annotated microproteome of D. melanogaster from the Flybase (release 6.61) (Öztürk‐Çolak et al., 2024) and UniProt (UniprotKB, January 2025) (The Uniprot Consortium et al., 2025) sites. An R script was used to merge the two files and to extract a list of all proteome entries with Uniprot entry, gene name, parent transcript, parent gene, protein ID, amino acid sequence and sequence length (see Data Availability Statement). The resulting table was then checked for unique amino acid sequences regardless of transcript sequence. We obtained a list of 1200 microprotein sequences, including 1017 sequences up to 100 amino acids, the classical microprotein length limit, and 183 sequences up to 150 amino acid, a threshold also used for eukaryotic microproteins (Table S1). There exist nowadays several bioinformatic tools that can help to investigate microprotein properties based on the amino acid sequence (see Table 1). We first used DeepLoc 2.1 (Ødum et al., 2024) to propose a first classification of the subcellular localization of these 1200 microproteins (Figure 1a, inset and Table S1). The majority of microproteins were classified as soluble (865) while the remainder (335) were associated with membranes: either transmembrane (233), peripheral (80), or lipid‐anchored (22) microproteins. A little over a quarter of microproteins we have referenced were cytoplasmic while another quarter were extracellular. We should point out that current prediction software may have limitations for identifying functional motifs in very small proteins, given that most of these tools have been trained on canonical (i.e., larger) proteins. For example, we observed a microprotein that was predicted to reside in the plastid (Figure 1a), an organelle which does not exist in Drosophila. To limit mislocalization artifacts and improve reliability of these predictions, a consensus approach, combining other tools such as TargetP or MitoFates (see Table 1), would be beneficial. We can also compare these first results with the set of microproteins identified by SignalP 6.0 (Teufel et al., 2022). SignalP identified 493 microproteins containing signal peptides, a majority with high confidence (Figure 1b, inset), hence with a high probability to be exported. The size of these signal peptides is around 20 amino acids (Figure 1b). While SignalP predicts the presence of an N‐terminal signal peptide directing a protein to the secretory pathway, DeepLoc predicts the final subcellular localization, classifying 290 proteins as extracellular (Figure 1a). This apparent discrepancy can be explained by the fact that not all the proteins entering the secretory pathway will be secreted. The remaining microproteins (203) were predicted to localize mainly in membranes and organelles. Indeed, numerous microproteins are predicted to be distributed in diverse cellular organelles such as the nucleus, mitochondria, the endoplasmic reticulum (ER), and the plasma membrane (PM) (Figure 1a).

TABLE 1.

Non‐exhaustive list of useful bioinformatics tools developed to predict biological features of proteins of interest based on amino acid sequences.

Name Type of prediction Website References
SignalP‐6.0 Signal peptide https://services.healthtech.dtu.dk/services/SignalP‐6.0/ (Teufel et al., 2022)
DeepMito Mitochondrial peptide adressing motif https://busca.biocomp.unibo.it/deepmito/ (Savojardo et al., 2019)
MitoFates Mitochondrial targeting sequences and cleavage sites https://mitf.cbrc.pj.aist.go.jp/MitoFates/cgi‐bin/top.cgi (Fukasawa et al., 2015)
DeepTMHMM‐1.0 Transmembrane domain https://services.healthtech.dtu.dk/services/DeepTMHMM‐1.0/ (Hallgren et al., 2022)
DeepLoc‐2.1 Cellular localization https://services.healthtech.dtu.dk/services/DeepLoc‐2.1/ (Ødum et al., 2024)
TargetP‐2.0 N‐terminal sorting signal https://services.healthtech.dtu.dk/services/TargetP‐2.0/ (Armenteros et al., 2019)
ELMdatabase Short linear motif (slim) http://elm.eu.org/ (Kumar et al., 2023)
InterPro Protein domain family https://www.ebi.ac.uk/interpro/ (Blum et al., 2024)
Human PPI Human interactome http://prodata.swmed.edu/humanPPI/ (Zhang et al., 2025)
Flypredictome Interactors for Drosophila https://www.flyrnai.org/tools/fly_predictome/web/ (Kim et al., 2024)
HHpred Protein sequence analysis https://toolkit.tuebingen.mpg.de/tools/hhpred (Gabler et al., 2020)
IntAct Interactors https://www.ebi.ac.uk/intact/home (Orchard et al., 2014)
DeepFRI GO term prediction https://beta.deepfri.flatironinstitute.org/AboutUs (Gligorijević et al., 2021)
PICNIC Condensate‐forming proteins https://picnic.cd‐code.org (Hadarovich et al., 2024)

FIGURE 1.

FIGURE 1

Microprotein cellular localizations in Drosophila. (a) Microprotein localization predicted by DeepLoc. (b) Signal peptide length predicted by SignalP 6.0. (c) Textbook‐style rendering of cellular landscape with large empty regions. Depicted are the plasma membrane (PM, orange), the cytoplasm (yellow), the extracellular space (gray), the nucleus (pink) with nuclear pore (blue), a mitochondrion (green), the endoplasmic reticulum (ER, blue), lipid droplets (LDs, light pink), and the Golgi apparatus (purple). Overlapping with this scheme are more realistic representations of crowded cellular regions inspired by David Goodsell's paintings (Goodsell, 2011, 2010) and presented in detail in panels (d–h). (d) PM (orange) and extracellular material (gray). (e) Nuclear membrane (pink) with nuclear pore (blue). (f) Mitochondrion (green). (g) ER (blue), cytoplasm (yellow), and LD (light pink). (h) Golgi apparatus (purple) and cytoplasm (yellow). Soluble microproteins are displayed in green while transmembrane microproteins are displayed in red.

Although the annotated Drosophila microproteome is still relatively small, very few microproteins have been experimentally assigned to a subcellular localization. Besides, several recent reviews summarized what is known experimentally (Deng et al., 2023) or computationally (Whited et al., 2024) about microprotein localization in organisms other than Drosophila, and especially in humans and mammals. Based on DeepLoc analyses (Figure 1a) and combined with these reviews, we have reproduced David Goodsell's way of rendering cellular landscapes (Goodsell, 2016) to illustrate putative microprotein cellular localizations (Figure 1d–h). DeepLoc 2.1 was trained on a data set constituted by proteins longer than 40 amino acids. Therefore, DeepLoc may not be entirely suited to predict the location of all of the microproteins. Thus, the degree of confidence of predicted localization can vary from one protein to another (see Table S1 and Figure S1). Beyond DeepLoc predictions, some of these localizations were inferred based on a subset of experimentally validated cases (Table 2); however, the majority remain highly putative due to the limited knowledge currently available for most of the microproteins. Even if speculative, in our opinion, this rendering gives a more realistic description of the crowded cellular environment in which microproteins are embedded than the simplistic “empty cellular” view relayed by textbook representations (Figure 1c). Furthermore, our illustrative landscapes may give new ideas to readers to reflect upon the localization and function of their microproteins of interest. For instance, the microprotein CIP2A‐BP is localized in the cytoplasm where it interacts with CIP2A (cancerous inhibitor of PP2A cancer inhibitory factor) to prevent its interaction with PP2A, an inhibitor of the PI3K/AKT/NFkB signaling pathways (Guo et al., 2020). Adipogenin is a transmembrane microprotein localized in the ER membrane in adipocytes where it controls LD formation by binding to the seipin complex (Li et al., 2025).

TABLE 2.

Diverse subcellular localizations of microproteins reflect functional diversity. Here is a non‐exhaustive list of microproteins investigated at the molecular level in Drosophila or in vertebrates.

Drosophila microprotein Vertebrate microprotein Length (aa) Cellular localization Physiological significance DOI
Pri/tal 11 Cytoplasmic and nuclear Regulation of E3 ubiquitin ligase and epidermal differentiation (Zanet et al., 2015)
Würmchen 2 57 Plasmic membrane Regulation of epithelium polarity (Königsmann et al., 2020)
Pegasus 80 Extracellular space Regulation of Wg (Magny et al., 2021)
Sarcolamban Sarcolipin/phospholamban 28/31/52 Sarco/endoplasmic reticulum Regulation of ATPase Ca++ pump SERCA in muscle (Magny et al., 2013)
Sloth1 SMIM4 79/70 Inner mitochondrial membrane Regulation of respiratory complex III assembly (Bosch et al., 2022)
Sloth2 Brawnin 61/71 Inner mitochondrial membrane Regulation of respiratory complex III assembly (Bosch et al., 2022)
Hemotin Stannin 88/88 Endosome membrane regulation of phagocytosis (Pueyo et al., 2016)
CG34210 PIGBOS1 74/54 Mitochondria Regulation of ER stress (Chu et al., 2019)
Gm15781 80 Nucleus Regulation of nuclear actin (Na et al., 2022)
CASIMO1 83 Endosome membrane Lipid homeostasis regulation (Polycarpou‐Schwarz et al., 2018)
Elabela 53 Secreted Binds to apelin receptor (Chng et al., 2013)
NoBody 68 P‐Body mRNA metabolism (D'Lima et al., 2017)
Myomixer 84 Plasma membrane Regulation of muscle cell fusion (Bi et al., 2017)
APPLE 90 Within ER Regulation of translation (Sun et al., 2021)
SERTM2 89 Plasma membrane Regulation of neurone acticity (Lisi et al., 2025)
FXYD3 83 ER membrane Regulation of SERCA2 activity (Yang et al., 2025)
Adipogenin 80 ER membrane Promoting the development of lipid droplet (Li et al., 2025)
ASDURF 96 Cytoplasmic Belongs to the PAQosome complex (Cloutier et al., 2020)
CIP2A‐BP 52 Cytoplasmic Inhibits migration and invasion of triple‐negative breast cancer cells (Guo et al., 2020)
pTINCR 87 Nucleus and plasma membrane Induction of epidermal differentiation (Boix et al., 2022)
SMIM26 95 Mitochondria Suppression of renal cancer cell growth (Meng et al., 2023)
C16orf74 76 Plasma membrane Promotes thermogenis in brown adipose tissue (Dinh et al., 2025)
Myoregulin 46 Sarcoplasmic reticulum Regulation of SERCA pump activity in muscle (Anderson et al., 2015)
Dworf 35 Sarcoplasmic reticulum Regulation of SERCA pump activity in muscle (Nelson et al., 2016)

The illustrations presented here (Figure 1d–h) propose a vision in which microproteins, like canonical proteins, are distributed across all cellular compartments and can interact with both cytoplasmic and membrane proteins (see also Figure 4). We have subjectively positioned them within various cellular compartments to illustrate the pervasiveness of microproteins. We hypothesize that these microproteins participate in the regulation of all cellular processes and that their function most commonly involves binding to and modulating the function of canonical proteins. To date, the number of microproteins actually encoded by genomes and how their abundance compares to canonical proteins remains open questions.

FIGURE 4.

FIGURE 4

Transmembrane microproteins. (a) Number of Transmembrane Domains (TMD) containing microproteins. (b) Size of single TMDs identified in microproteins. (c) AlphaFold model of Sarcolamban (in space‐filling representation) in interaction with the SERCA protein (in cartoon representation) from Drosophila. Figure made with the Mol* program (Sehnal et al., 2021, 2018). (d) Type of interactions between Sarcolamban and SERCA found in the AlphaFold model. (e) pLDDT score of the model (dark blue: PLDDT >90; light blue: 90> pLDDT >70; yellow: 70> pLDDT >50; orange: 50> pLDDT) showing a high confidence in the core of the SERCA structure but a low confidence for the Sarcolamban structure. Figure made with the Mol*program (Sehnal et al., 2021, 2018). (f) Localization of the microprotein CG12617::V5 (green) to mitochondria in Drosophila S2 cells is revealed by co‐immunostaining (white dots) with the mitochondrial marker ATP5A (magenta). The nucleus is stained with DAPI (DNA, blue). (g) Coarse‐Grained simulations of the CG12617 microprotein (surface colored by residue type) embedded in models of inner and outer mitochondrial membranes from Drosophila (DPPC: Dipalmitoylphosphatidylcholine; POPE: Palmitoyloleoylphosphatidylethanolamine; POPS: Palmitoyloleoylphosphatidylserine; POPI: Palmitoyloleoylphosphatidylinositol; CL: Cardiolipin).

3. 3D STRUCTURAL PREDICTION TO ASSIGN FUNCTIONAL DOMAINS TO MICROPROTEINS

The 3D conformation of structured protein domains is critical for protein function, as these features underlie conformational dynamics and mediate interactions with regulatory partners. In contrast, most microproteins lack identifiable functional domains, largely because their amino acid sequences are too short or too poorly conserved to enable domain annotation by traditional sequence‐based methods. Here, we propose to infer putative functional domains directly from the predicted 3D structures of microproteins and not from the amino acid sequence. Indeed, it was shown that de novo microproteins, that is, microproteins which have recently emerged, can localize to diverse subcellular compartments, where they interact with canonical proteins to modulate their functions, even in the absence of sequence conservation (Sandmann et al., 2023). As another example, the microproteins Sarcolamban in Drosophila and Phospholamban in humans are functionally conserved despite having only 15.4% of sequence identity and 28.8% of sequence similarity (Magny et al., 2013). Both microproteins localize to the Sarco‐Endoplasmic Reticulum and bind to the calcium pump SERCA to regulate its activity and muscle contractility (see also Figure 4c,d). These Drosophila and mammalian microproteins display conserved secondary structures that allow them to be functionally interchangeable between the species (Magny et al., 2013). Thus, detecting similarity between proteins by superimposing their 3D structures may offer higher sensitivity for identifying proteins of similar function than only comparing their sequences (Illergård et al., 2009; Iyer et al., 2022).

With the advent of AI‐boosted protein prediction methods (Heinzinger & Rost, 2025; Park et al., 2025), it is now possible to easily and quickly predict structures of thousands of proteins as well as their interactions with partners. As an example, researchers recently employed AlphaFold2 (AF2) (Jumper et al., 2021) to predict protein–protein interactions among components of the nuage (Kawaguchi et al., 2025). This germline‐specific membraneless structure is localized in the perinuclear region and is composed of proteins and RNAs essential for piRNA production in Drosophila. We have used AF2 to generate structures of our pool of 1200 microproteins (see Data Availability Statement). We then applied the program Define Secondary Structure of Proteins (DSSP) (Hekkelman et al., 2025; Kabsch & Sander, 1983) to quantify the structuration of these proteins (Figure 2a). We divided our dataset of microproteins into four categories according to their size: from very small (less than 25 amino acids) to larger (from 75 to 100 amino acids) microproteins. The majority of our pool of microproteins belong to the latter category. Overall, around half the proteins in each category displayed large unstructured regions (i.e., less than 50% of a protein is predicted as having a secondary structure region). This trend seemed to change for the largest microproteins (i.e., 76–100 amino acids long) with a larger proportion of structured protein regions. Focusing on the population of more organized proteins (more than 50% of the protein displaying secondary structure) for each category (Figure 2b), we observed a large propensity of alpha‐helix folds in smaller proteins, which diminished with increasing size of the microprotein. We note that AF2 was trained on structures of larger proteins from the Protein Data Bank (see statistics of protein size: https://www.rcsb.org/stats/distribution-atom-count) and has a tendency to overestimate the propensity to form alpha‐helices instead of beta‐sheets and loop regions (Stevens & He, 2022). AF2 may therefore be less suitable for predicting the structure of microproteins. That being said, AF2 predicted beta‐sheet regions in microproteins with a higher occurrence in larger microproteins (Figure 2b). For larger microproteins, these structure predictions may be accurate as illustrated by two examples. As a first example, we chose the gene CG33672 that encodes a microprotein with a structure predicted as a succession of beta‐sheets and alpha‐helices with an αββααβα topology (Figure 2c), which is typical of BolA‐like proteins. Indeed, this microprotein shares about 50% of sequence identity with a BolA‐like protein from mice (PDB: 1V9J) (Kasai et al., 2004). This protein type was recently identified by genomics and proteomics for microproteins in cyanobacteria which might interact with monothiol glutaredoxins to mediate iron homeostasis (Peng et al., 2025). This function and interaction are also conserved in eukaryotic cells (Uzarska et al., 2016). For the second example, the microprotein encoded by CG44882, AF2 predicted a structure constituted of only beta‐sheets (Figure 2d). This fold is close to the zinc‐ion‐containing FLYWCH domain determined by NMR (PDB: 2RPR) (Figure 2d). Even if the sequence identity of these two proteins is relatively low (24%), the zinc binding site is well conserved (Figure 2d). The FLYWCH domain is found in canonical protein encoded by the Mod(mdg4) gene in Drosophila (Dorn & Krauss, 2003). This protein is typically localized in the nucleus and can interact with various DNA binding proteins, and plays a role in chromatin organization and insulator function (Melnikova et al., 2017). This led us to hypothesize that this microprotein may localize in the nucleus where it could compete with the canonical Mod(mdg4) protein, which also contains several copies of the FLYWCH domain. Last but not least, visualizing the structure of these microproteins at the same scale as larger molecular machineries, such as ATP synthase (Figure 2d,e), helps illustrate their markedly smaller size relative to canonical proteins. This reduced size may enable microproteins to access regions that are less accessible to larger proteins.

FIGURE 2.

FIGURE 2

Microprotein structures. (a) Distribution of microproteins by sequence‐length ranges (1–25, 26–50, 51–75, 76–100 amino acids) using stacked counts of two categories: (blue) Microproteins with ≥50% of their amino acids located in structured regions (alpha‐helix or beta‐sheet) and (red) Microproteins with <50% of their amino acids located in structured regions (predominantly unstructured). The numbers inside the bars denote the count for each category. (b) Average secondary structure composition of the subset of microproteins with ≥50% of their amino acids located in structured regions (as defined in a). For each microprotein, the fractions of residues in alpha‐helix, beta‐sheet, and unstructured conformations were calculated (summing to 100% per sequence). The bars show the mean percentage of each secondary structure type for four length ranges (1–25, 26–50, 51–75, 76–100 amino acids). (c) CG33672 encodes for the human ortholog of Bol2A and is predicted to be involved in intracellular iron ion homeostasis. (d) CG44882 encodes a pre‐mod(mdg4)‐J microprotein which is a FLYWCH zinc finger, a protein domain found in several canonical Drosophila and human proteins. (e) ATP synthase beside the microprotein encoded by the gene CG44882 (see d) at the same size scale and rendered using David Goodsell's representation for proteins presented in his book: The Machinery of Life (Goodsell, 2009).

We further capitalized on AF2 predictions by clustering the microprotein structures into families to identify common functional domains. To do so, we structurally aligned a pool of 555 microproteins containing more than 25 residues and with more than 50% of structured region, using TM‐align (Zhang & Skolnick, 2005) (Figure 3). Their structural similarities were quantified based on Root Mean Square Deviation (RMSD) values (Carugo, 2003) between pairs of microproteins (Figure 3). Beyond RMSD, other metrics can be used for scoring short proteins such as TM‐score (Zhang & Skolnick, 2004) or US‐align‐based (Zhang et al., 2022) measures. The resulting RMSD matrix was used to perform hierarchical clustering using the Weighted Pair Group Method with Arithmetic mean (WPGMA). This procedure revealed distinct structural groups among the microproteins. In Figure 3 (see also a high‐resolution version in Figure SI2), we present six different protein families, hereafter referred to as clusters. In Cluster 1, nine microproteins share a structure constituted by one short alpha‐helix followed by three‐stranded anti‐parallel beta‐sheets. This structure is close to Kazal‐type modules of the human serine protease HtrA1 domain (PDB: 3TJQ) (Eigenbrot et al., 2012) or serine protease inhibitor of the gland secretion of Coptotermes formosanus Shiraki, a subterranean termite (PDB: 2N17) (Negulescu et al., 2015). The Kazal family represents one of the well‐known families of serine protease inhibitors. Kazal domains generally comprise 40–60 amino acid residues and display a relatively conserved structure between vertebrates and invertebrates. The Kazal type serine protease inhibitors in invertebrate possess anti‐microbial, anti‐coagulational or anti‐gelatinolytic activities and consequently regulate immunity, blood feeding or reproduction (Rimphanitchayakit & Tassanakajon, 2010). Six of the nine microproteins in Cluster 1 are annotated in FlyBase as containing a Kazal domain, inferred from primary sequence analysis. The structure‐based approach developed here enables the remaining three microproteins in this cluster to be confidently assigned to the same domain family. Cluster 2 is constituted of 15 structures of three anti‐parallel beta‐sheets surrounded by two short alpha‐helices. Structurally speaking, this cluster is close to Cluster 1 (Figure 3). This fold may also be associated with serine protease inhibitors but, in this case, related to the Kunitz‐type (Mishra, 2020). This structure is close to the Kunitz inhibitor domain amyloid precursor protein (APPI) (PDB: 6GFI) (Naftaly et al., 2018). It could also be related to a selective antagonist of the vasopressin type 2 receptor, a G‐protein‐coupled receptor (PDB: 5M4V) (Ciolek et al., 2017). The representative structure of Cluster 3, comprising 10 members, is composed of three interlaced alpha‐helices concluded by two short beta‐sheets. This structure is related to the H1 module of histone (PDB 7K5Y) (Zhou et al., 2021). Interestingly, nothing is known about the microproteins in Cluster 3. Given their structural homology with the histone family, we hypothesize that they might be located in the nucleus and play a role in chromatin remodeling. Interestingly, Cluster 5 gathers genes identified as DRS and DRSL1 to 6. This group corresponds to drosomycins, antifungal peptides identified in Drosophila (Deng et al., 2009; Hanson et al., 2019). Indeed, the fold corresponds to the structure of drosomycin identified by NMR (PDB 1MYN) (Landon et al., 1997). It also corresponds to plant defensins which may bind to and alter biological membranes (Baxter et al., 2015; Ong et al., 2020). The representative structure of Cluster 6 is composed of five anti‐parallel beta strands close to the structure found in human Sm‐like protein LSm8 (PDB: 6QX9) (Charenton et al., 2019). Sm and Sm‐like proteins are RNA‐binding proteins that contribute to the formation of the spliceosome complex involved in mRNA splicing. Interestingly, Cluster 4 gathers a large family of 24 proteins which share structural features with numerous microproteins, as attested by the low RMSD values (Figure 3). This cluster regroups microproteins having a single alpha‐helix. As discussed before, this might partly be explained by AF2's tendency to overestimate the alpha‐helix content but could represent realistic structural features seen in single‐pass transmembrane proteins (Anderson et al., 2025; Pogozheva & Lomize, 2018). This aspect is discussed in the following section and illustrated in Figure 4d.

FIGURE 3.

FIGURE 3

Microprotein families. RMSD matrix highlighting microprotein families based on structural similarities (center). On the left, tree based on hierarchical clustering. On the right and bottom, representative structures of the six clusters are shown: Cluster 1: Kazal domains (cyan). Cluster 2: Kunitz‐type domains (red). Cluster 3: H1 module (yellow). Cluster 4: Alpha‐helix structures (purple). Cluster 5: Drosomycins (orange). Cluster 6: Sm‐like proteins (pink). A high resolution version is available in Figure SI2.

Finally, many microproteins can be considered as weakly structured with a low proportion of their residues participating in secondary structures (Figure 2a). These proteins can be considered as Intrinsically Disordered Proteins (IDPs) and are not characterized by a single structure but by an ensemble of structures (Tompa, 2011). Furthermore, these IDPs present specific characteristics such as a complex landscape of interactions with other molecular partners, notably through the presence of Short Linear Motifs (SLiMs), which mediates interactions with folded proteins (Orand & Jensen, 2025). Recent work has harnessed AF2 to generate an ensemble of structures of IDPs (Schnapka et al., 2025). It is also possible to generate sets of IDPs at atomistic resolution by integrating all‐atom molecular dynamics (MD) simulations and experimental data (Borthakur et al., 2025). Other models, referred to as Coarse‐Grained (CG) models, propose simplified representations to study larger IDPs in different environmental conditions (Rauh et al., 2025) or to predict interactions with molecular partners (Ginell et al., 2025). Recently, CG models specifically designed to study IDP dynamics have been proposed such as CALVADOS (Cao et al., 2024; Tesei et al., 2021; Tesei & Lindorff‐Larsen, 2022) or MARTINI3 models (Wang et al., 2025). These models are also particularly useful to study how disordered proteins may self‐interact to form biomolecular condensates (Alberti et al., 2025) as described in a recent review (Schäfer & Stelzl, 2025). With the advances in Machine Learning, new CG models are rapidly emerging (Majewski et al., 2023), potentially proposing new and more efficient solutions for the study of IDPs. Finally, in silico design of new proteins to bind and structure IDPs (Liu et al., 2025; Wu et al., 2025) constitutes another promising strategy for their study.

4. A POOL OF MICROPROTEINS PREDICTED AS TRANSMEMBRANE DOMAINS (TMD)

A large amount of microproteins in our dataset is predicted to localize in membrane regions such as the ER, mitochondria, or the PM (Figure 1a). Indeed, numerous transmembrane microproteins have already been identified to reside in biological membranes to modulate membrane proteins (Hassel et al., 2023). Using the algorithms TMHMM 2.0 (Krogh et al., 2001; Sonnhammer et al., 1998) and DeepTMHMM (Hallgren et al., 2022) to detect TMD, we identified 226 membrane‐associated microproteins from our pool of 1200 microproteins. A similar value was obtained using DeepLoc (Figure 1a). Analyzing the structure of these proteins, integrating secondary structure analysis and amino acid properties, led to the classification presented in Figure 4a. Of the 226 membrane‐associated microproteins, 10 were excluded from this classification as 6 were beta‐barrel proteins (not shown in this analysis), and 4 had transmembrane regions predicted to be unstructured by AF2. Most of the remaining 216 proteins were identified as single‐pass TMD proteins constituted of a single alpha‐helical region spanning the membrane. Some of these microproteins are predicted to contain two to three alpha‐helical TMDs. The size of the TMDs was mainly found to span 15–25 residues (Figure 4b). TMD length may give some clues as to the membrane localization, as it is now clear that transmembrane proteins have evolved concurrently with the thickness and compositions of the membranes they fit in (Levental & Lyman, 2023; Lorent et al., 2020, 2025). Furthermore, these single‐pass TMDs can interact with other membrane proteins. Predictive approaches such as AlphaFold can provide a first clue of this type of interaction. For example, we have modeled the interaction of sarcolamban with SERCA in Drosophila (Figure 4c). This predicted interface is mainly formed by hydrophobic interactions within the membrane, complemented by polar and charged interactions (Figure 4c,d). This model is consistent with the identified protein complex between Sarcolipin and SERCA1a (Figure S3A) (Magny et al., 2013). As seen in our AlphaFold model, the Sarcolipin‐SERCA1a interaction involved hydrophobic as well as polar residues (Figure S3B). Alternative models were also generated positioning sarcolamban at different locations around SERCA while remaining in the membrane region (Figure S3C,D). Further structural analyses are now necessary to characterize the accuracy of these different models. A simple analysis is the display of the predicted local distance difference test (pLDDT), a per‐residue measure of local confidence, showing a relatively good confidence for the SERCA folding but a poor folding/modeling of the Sarcolamban (Figure 4e). pLDDT scores are also available for all the models of microproteins (see Data Availability Statement). To further characterize these membrane proteins interactions and refine the predicted complexes, MD simulations represent a valuable complementary method (Jackson et al., 2022).

MD simulations is a computational approach especially useful to study membrane proteins and their interactions with lipids constituting biological membranes (Corradi et al., 2019; Enkavi et al., 2019; Marrink et al., 2019; Muller et al., 2019). Combining MD simulations with in cellulo experiments can provide molecular details about the association of microproteins with biological membranes. As an illustrative example, we have identified a microprotein in our dataset, encoded by the gene CG12617, which is localized in the mitochondria (Figure 4f). This organelle possesses two membranes, an inner and an outer one (Figure 1f), which cannot be distinguished using classical confocal fluorescence microscopy, even though their biophysical properties and lipid composition may differ (Decker & Funai, 2024; Konar et al., 2023). Using MD simulations, it is possible to model these two membranes and assess how this microprotein can adapt to the two membrane types (Figure 4g). Furthermore, this approach can assess the preference of the microproteins for specific lipids (Corey et al., 2019, 2022; Hedger et al., 2016). That is why MD simulations are used in combination with experiments to study how microproteins can modulate the functions of other protein, as recently shown for the adipogenin microprotein (Li et al., 2025). Thus, based on the proof of concept presented here, we believe that integrating different computational approaches with experiments to create models that can then be simulated in realistic environments, such as a crowded cytoplasm or a complex biological membrane, will shed new light on the comprehension of microprotein functions.

5. CONCLUSION

In this perspective, we highlighted tools and methods to predict both cellular localization and protein structure to infer microprotein function, using annotated Drosophila microproteins as a manageable and well‐characterized dataset. This combination of computational approaches can help to prioritize candidates for in vivo studies and significantly reduce the time and cost required for experimental validation. It would be highly interesting and valuable to apply this approach to the numerous unannotated microproteins identified by mass spectrometry or ribosome profiling, which may number in the tens of thousands, to help distinguish between translational noise and potentially bioactive microproteins. As illustrated in our molecular landscapes, microproteins could spread out into all the subcellular regions and potentially interact with numerous partners. It is thus tempting to speculate that microproteins constitute an additional, widespread layer of regulation of cellular protein activity, by modulating many canonical proteins. Modeling methods presented in this perspective may further help decipher the interactome of microproteins. To illustrate our point of view, we have humbly reinterpreted David Goodsell's beautiful renderings from protein structure to cellular landscapes. The drawing and coloring of the latter places us in the footsteps of David Goodsell's work and reminds us that even simple lines and colors can truly bring clarity to very complex biological systems. In the era of AI, where Art, as well as Science, never stop accelerating, let's keep David Goodsell's molecular landscape, fruit of a long labor and thinking, as an example to follow to also take time out to simply enjoy the Science.

AUTHOR CONTRIBUTIONS

Matthieu Chavent: Conceptualization; funding acquisition; visualization; writing – review and editing; writing – original draft; supervision. Gabriel Diaz: Investigation; methodology; visualization; writing – review and editing. Jennifer Zanet: Conceptualization; supervision; funding acquisition; writing – original draft; writing – review and editing. Kenza Benachenhou: Methodology; visualization; writing – review and editing. Marc Gueroult: Methodology; investigation; visualization; writing – review and editing. Simon Marques‐Prieto: Methodology; visualization; investigation; writing – review and editing. Philippe Valenti: Investigation; methodology; visualization; writing – review and editing.

CONFLICT OF INTEREST STATEMENT

The authors declare no conflicts of interest.

Supporting information

Figure S1. Distribution of predicted localization probabilities obtained with DeepLoc 2.1. Probability value for each microprotein sequence is available in Table S1.

Figure S2. Distribution of pLDDT values for microproteins grouped by sequence length (1–25, 26–50, 51–75, and 76–100 amino acids). The figure shows the distribution of pLDDT scores within each length group, with the horizontal black line indicating the mean pLDDT value for each group. The proportion of microproteins with high‐confidence structural predictions (pLDDT ≥70%) is 86.0% for the 1–25 aa group, 57.4% for 26–50 aa, 49.7% for 51–75 aa, and 60.6% for the 76–100 aa group.

Figure S3. A‐ Sarcolipin‐SERCA1a interaction (PDB: 3W5A) side by side with its closest AlphaFold model of the Sarcolipin (in space‐filling representation) in interaction with the SERCA protein (in cartoon representation) from Drosophila. Figure made with the Mol* program (Sehnal et al., 2021, 2018). B‐ Type of interactions between Sarcolipin and SERCA found in the pdb structure. C and D‐ Left, AlphaFold alternative models model of the Sarcolamban (in space‐filling representation) in interaction with the SERCA protein (in cartoon representation) from Drosophila. Figure made with the Mol* program (Sehnal et al., 2021, 2018). Right, type of interactions between Sarcolamban and SERCA found in the respective AlphaFold models.

PRO-35-e70636-s002.pdf (1.1MB, pdf)

Figure SI2. High resolution of Figure 3.

PRO-35-e70636-s001.pdf (4.8MB, pdf)

Table S1. Table of the 1200 microproteins with cellular localization and Deeploc probability.

PRO-35-e70636-s003.csv (154.1KB, csv)

ACKNOWLEDGMENTS

We thank Hélène Chavent for her work on drawings in Figure 1. We also thank Ludovic Autin, Evert Haanappel, and Hélène Chanut‐Delalande for helpful comments. M.C. and J.Z. are grateful to the CBI transversal program for their support. This work was granted access to the TGCC Joliot‐Curie supercomputer (under the GENCI allocations A0180716209). Open access publication funding provided by COUPERIN CY26.

Diaz G, Valenti P, Gueroult M, Marques‐Prieto S, Benachenhou K, Zanet J, et al. Sketching microprotein portraits. Protein Science. 2026;35(6):e70636. 10.1002/pro.70636

Review Editor: Nir Ben‐Tal

Footnotes

Contributor Information

Jennifer Zanet, Email: jennifer.zanet@utoulouse.fr.

Matthieu Chavent, Email: matthieu.chavent@utoulouse.fr.

DATA AVAILABILITY STATEMENT

Some of the data presented in this perspective are deposited at: https://doi.org/10.5281/zenodo.19450794. This comprises: Raw data files and R script to create the list of 1200 microproteins. Cellular landscapes presented in Figure 1d–h. Structure of the 1200 microproteins with pLDDT score in beta‐factor. AlphaFold models for Sarcolamban‐SERCA interactions presented Figure 4c–e and Figure S3. Table with peptide signal probabilities.

REFERENCES

  1. Alberti S, Arosio P, Best RB, Boeynaems S, Cai D, Collepardo‐Guevara R, et al. Current practices in the study of biomolecular condensates: a community comment. Nat Commun. 2025;16:7730. [DOI] [PMC free article] [PubMed] [Google Scholar]
  2. Anderson DM, Anderson KM, Chang C‐L, Makarewich CA, Nelson BR, McAnally JR, et al. A micropeptide encoded by a putative long noncoding RNA regulates muscle performance. Cell. 2015;160:595–606. [DOI] [PMC free article] [PubMed] [Google Scholar]
  3. Anderson SM, Choi J, Cushman EM, Leander M, Raman S, Senes A. High‐throughput discovery of transmembrane helix dimers from human single‐pass membrane proteins with TOXGREEN sort‐seq. PNAS Nexus. 2025;4:pgaf305. [DOI] [PMC free article] [PubMed] [Google Scholar]
  4. Armenteros JJA, Salvatore M, Emanuelsson O, Winther O, von Heijne G, Elofsson A, et al. Detecting sequence signals in targeting peptides using deep learning. Life Sci Alliance. 2019;2:e201900429. [DOI] [PMC free article] [PubMed] [Google Scholar]
  5. Baxter AA, Richter V, Lay FT, Poon IKH, Adda CG, Veneer PK, et al. The tomato Defensin TPP3 binds phosphatidylinositol (4,5)‐bisphosphate via a conserved dimeric cationic grip conformation to mediate cell lysis. Mol Cell Biol. 2015;35:1964–1978. [DOI] [PMC free article] [PubMed] [Google Scholar]
  6. Bi P, Ramirez‐Martinez A, Li H, Cannavino J, McAnally JR, Shelton JM, et al. Control of muscle formation by the fusogenic micropeptide myomixer. Science. 2017;356:323–327. [DOI] [PMC free article] [PubMed] [Google Scholar]
  7. Blum M, Andreeva A, Florentino LC, Chuguransky SR, Grego T, Hobbs E, et al. InterPro: the protein sequence classification resource in 2025. Nucleic Acids Res. 2024;53:D444–D456. [DOI] [PMC free article] [PubMed] [Google Scholar]
  8. Boix O, Martinez M, Vidal S, Giménez‐Alejandre M, Palenzuela L, Lorenzo‐Sanz L, et al. pTINCR microprotein promotes epithelial differentiation and suppresses tumor growth through CDC42 SUMOylation and activation. Nat Commun. 2022;13:6840. [DOI] [PMC free article] [PubMed] [Google Scholar]
  9. Borthakur K, Sisk TR, Panei FP, Bonomi M, Robustelli P. Determining accurate conformational ensembles of intrinsically disordered proteins at atomic resolution. Nat Commun. 2025;16:9036. [DOI] [PMC free article] [PubMed] [Google Scholar]
  10. Bosch JA, Ugur B, Pichardo‐Casas I, Rabasco J, Escobedo F, Zuo Z, et al. Two neuronal peptides encoded from a single transcript regulate mitochondrial complex III in drosophila. elife. 2022;11:e82709. [DOI] [PMC free article] [PubMed] [Google Scholar]
  11. Callaway E. ‘Dark proteins’ hiding in our cells could hold clues to cancer and other diseases. Nature. 2025;637:1038–1040. [DOI] [PubMed] [Google Scholar]
  12. Cao F, von Bülow S, Tesei G, Lindorff‐Larsen K. A coarse‐grained model for disordered and multi‐domain proteins. Protein Sci. 2024;33:e5172. [DOI] [PMC free article] [PubMed] [Google Scholar]
  13. Carugo O. How root‐mean‐square distance (r.m.s.d.) values depend on the resolution of protein structures that are compared. J Appl Crystallogr. 2003;36:125–128. [Google Scholar]
  14. Chanut‐Delalande H, Zanet J. Small ORFs, big insights: drosophila as a model to unraveling microprotein functions. Cells. 2024;13:1645. [DOI] [PMC free article] [PubMed] [Google Scholar]
  15. Charenton C, Wilkinson ME, Nagai K. Mechanism of 5′ splice site transfer for human spliceosome activation. Science. 2019;364:362–367. [DOI] [PMC free article] [PubMed] [Google Scholar]
  16. Chen J, Brunner A‐D, Cogan JZ, Nuñez JK, Fields AP, Adamson B, et al. Pervasive functional translation of noncanonical human open reading frames. Science. 2020;367:1140–1146. [DOI] [PMC free article] [PubMed] [Google Scholar]
  17. Chng SC, Ho L, Tian J, Reversade B. ELABELA: a hormone essential for heart development signals via the apelin receptor. Dev Cell. 2013;27:672–680. [DOI] [PubMed] [Google Scholar]
  18. Chothani S, Ruiz‐Orera J, Tierney JAS, Swirski MI, Tjeldnes H, Kok LW, et al. An expanded reference catalog of translated open reading frames for biomedical research. Nucleic Acids Res. 2026;54:gkag234. [DOI] [PMC free article] [PubMed] [Google Scholar]
  19. Chu Q, Martinez TF, Novak SW, Donaldson CJ, Tan D, Vaughan JM, et al. Regulation of the ER stress response by a mitochondrial microprotein. Nat Commun. 2019;10:4883. [DOI] [PMC free article] [PubMed] [Google Scholar]
  20. Ciolek J, Reinfrank H, Quinton L, Viengchareun S, Stura EA, Vera L, et al. Green mamba peptide targets type‐2 vasopressin receptor against polycystic kidney disease. Proc Natl Acad Sci. 2017;114:7154–7159. [DOI] [PMC free article] [PubMed] [Google Scholar]
  21. Cloutier P, Poitras C, Faubert D, Bouchard A, Blanchette M, Gauthier M‐S, et al. Upstream ORF‐encoded ASDURF is a novel prefoldin‐like subunit of the PAQosome. J Proteome Res. 2020;19:18–27. [DOI] [PubMed] [Google Scholar]
  22. Corey RA, Harrison N, Stansfeld PJ, Sansom MSP, Duncan AL. Cardiolipin, and not monolysocardiolipin, preferentially binds to the interface of complexes III and IV. Chem Sci. 2022;13:13489–13498. [DOI] [PMC free article] [PubMed] [Google Scholar]
  23. Corey RA, Vickery ON, Sansom MSP, Stansfeld PJ. Insights into membrane protein–lipid interactions from free energy calculations. J Chem Theory Comput. 2019;15:5727–5736. [DOI] [PMC free article] [PubMed] [Google Scholar]
  24. Corradi V, Sejdiu BI, Mesa‐Galloso H, Abdizadeh H, Noskov SY, Marrink SJ, et al. Emerging diversity in lipid–protein interactions. Chem Rev. 2019;119:5775–5848. [DOI] [PMC free article] [PubMed] [Google Scholar]
  25. Cuevas MVR, Hardy M‐P, Hollý J, Bonneil É, Durette C, Courcelles M, et al. Most non‐canonical proteins uniquely populate the proteome or immunopeptidome. Cell Rep. 2021;34:108815. [DOI] [PMC free article] [PubMed] [Google Scholar]
  26. D'Lima NG, Ma J, Winkler L, Chu Q, Loh KH, Corpuz EO, et al. A human microprotein that interacts with the mRNA decapping complex. Nat Chem Biol. 2017;13:174–180. [DOI] [PMC free article] [PubMed] [Google Scholar]
  27. Decker ST, Funai K. Mitochondrial membrane lipids in the regulation of bioenergetic flux. Cell Metab. 2024;36:1963–1978. [DOI] [PMC free article] [PubMed] [Google Scholar]
  28. Deng J, Xu W, Jie Y, Chong Y. Subcellular localization and relevant mechanisms of human cancer‐related micropeptides. FASEB J. 2023;37:e23270. [DOI] [PubMed] [Google Scholar]
  29. Deng X‐J, Yang W‐Y, Huang Y‐D, Cao Y, Wen S‐Y, Xia Q‐Y, et al. Gene expression divergence and evolutionary analysis of the drosomycin gene family in Drosophila melanogaster. Biomed Res Int. 2009;2009:315423. [DOI] [PMC free article] [PubMed] [Google Scholar]
  30. Dickerson RE. Irving Geis, molecular artist, 1908‐1997. Protein Sci. 1997;6:2483–2484. [Google Scholar]
  31. Dinh J, Yi D, Lin F, Xue P, Holloway ND, Xie Y, et al. The microprotein C16orf74/MICT1 promotes thermogenesis in brown adipose tissue. EMBO J. 2025;44:3381–3412. [DOI] [PMC free article] [PubMed] [Google Scholar]
  32. Dorn R, Krauss V. The modifier of mdg4 locus in drosophila: functional complexity is resolved by trans splicing. Genetica. 2003;117:165–177. [DOI] [PubMed] [Google Scholar]
  33. Eigenbrot C, Ultsch M, Lipari MT, Moran P, Lin SJ, Ganesan R, et al. Structural and functional analysis of HtrA1 and its subdomains. Structure. 2012;20:1040–1050. [DOI] [PubMed] [Google Scholar]
  34. Enkavi G, Javanainen M, Kulig W, Róg T, Vattulainen I. Multiscale simulations of biological membranes: the challenge to understand biological phenomena in a living substance. Chem Rev. 2019;119:5607–5774. [DOI] [PMC free article] [PubMed] [Google Scholar]
  35. Fesenko I, Sahakyan H, Dhyani R, Shabalina SA, Storz G, Koonin EV. The hidden bacterial microproteome. Mol Cell. 2025;85:1024–1041.e6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  36. Fukasawa Y, Tsuji J, Fu S‐C, Tomii K, Horton P, Imai K. MitoFates: improved prediction of mitochondrial targeting sequences and their cleavage sites*[S]. Mol Cell Proteomics. 2015;14:1113–1126. [DOI] [PMC free article] [PubMed] [Google Scholar]
  37. Gabler F, Nam S, Till S, Mirdita M, Steinegger M, Söding J, et al. Protein sequence analysis using the MPI bioinformatics toolkit. Curr Protoc Bioinformatics. 2020;72:e108. [DOI] [PubMed] [Google Scholar]
  38. Galindo MI, Pueyo JI, Fouix S, Bishop SA, Couso JP. Peptides encoded by short ORFs control development and define a new eukaryotic gene family. PLoS Biol. 2007;5:e106. [DOI] [PMC free article] [PubMed] [Google Scholar]
  39. Ginell GM, Emenecker RJ, Lotthammer JM, Keeley AT, Plassmeyer SP, Razo N, et al. Sequence‐based prediction of intermolecular interactions driven by disordered regions. Science. 2025;388:eadq8381. [DOI] [PMC free article] [PubMed] [Google Scholar]
  40. Gligorijević V, Renfrew PD, Kosciolek T, Leman JK, Berenberg D, Vatanen T, et al. Structure‐based protein function prediction using graph convolutional networks. Nat Commun. 2021;12:3168. [DOI] [PMC free article] [PubMed] [Google Scholar]
  41. Gomez‐Marin A. Drawing the mind, one neuron at a time the brain in search of itself: Santiago Ramón y Cajal and the story of the neuron Benjamin Ehrlich Farrar, Straus and Giroux. Science. 2022;375:1237. [Google Scholar]
  42. Goodsel DS. Mitochondrion. Biochem. Mol. Biol. Educ. 2010;38:134–140. [DOI] [PubMed] [Google Scholar]
  43. Goodsell DS. Art as a tool for science. Nat Struct Mol Biol. 2021;28:402–403. [DOI] [PubMed] [Google Scholar]
  44. Goodsell DS. Cellular landscapes in watercolor. J Biocommun. 2016;40:e6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  45. Goodsell DS. Eukaryotic cell panorama. Biochem Mol Biol Educ. 2011;39:91–101. [DOI] [PubMed] [Google Scholar]
  46. Goodsell DS. Life and Death. The machinery of life. New York: Springer New York; 2009. p. 108–125. [Google Scholar]
  47. Guo B, Wu S, Zhu X, Zhang L, Deng J, Li F, et al. Micropeptide CIP2A‐BP encoded by LINC00665 inhibits triple‐negative breast cancer progression. EMBO J. 2020;39:EMBJ2019102190. [DOI] [PMC free article] [PubMed] [Google Scholar]
  48. Hadarovich A, Singh HR, Ghosh S, Scheremetjew M, Rostam N, Hyman AA, et al. PICNIC accurately predicts condensate‐forming proteins regardless of their structural disorder across organisms. Nat Commun. 2024;15:10668. [DOI] [PMC free article] [PubMed] [Google Scholar]
  49. Haeckel E. Kunstformen der Natur. Leipzig und Wien: Verlag des Bibliographischen Instituts ; 1899. [Google Scholar]
  50. Hallgren J, Tsirigos KD, Pedersen MD, Armenteros JJA, Marcatili P, Nielsen H, et al. DeepTMHMM predicts alpha and beta transmembrane proteins using deep neural networks. BioRXiv. 2022. 10.1101/2022.04.08.487609 [DOI] [Google Scholar]
  51. Hanson MA, Dostálová A, Ceroni C, Poidevin M, Kondo S, Lemaitre B. Synergy and remarkable specificity of antimicrobial peptides in vivo using a systematic knockout approach. elife. 2019;8:e44341. [DOI] [PMC free article] [PubMed] [Google Scholar]
  52. Hassel KR, Brito‐Estrada O, Makarewich CA. Microproteins: overlooked regulators of physiology and disease. iScience. 2023;26:106781. [DOI] [PMC free article] [PubMed] [Google Scholar]
  53. Hedger G, Rouse SL, Domański J, Chavent M, Koldsø H, Sansom MSP. Lipid‐loving ANTs: molecular simulations of Cardiolipin interactions and the organization of the adenine nucleotide translocase in model mitochondrial membranes. Biochemistry. 2016;55:6238–6249. [DOI] [PMC free article] [PubMed] [Google Scholar]
  54. Heinzinger M, Rost B. Teaching AI to speak protein. Curr Opin Struct Biol. 2025;91:102986. [DOI] [PubMed] [Google Scholar]
  55. Hekkelman ML, Salmoral DÁ, Perrakis A, Joosten RP. DSSP 4: FAIR annotation of protein secondary structure. Protein Sci. 2025;34:e70208. [DOI] [PMC free article] [PubMed] [Google Scholar]
  56. Illergård K, Ardell DH, Elofsson A. Structure is three to ten times more conserved than sequence—A study of structural response in protein cores. Proteins: Struct, Funct, Bioinf. 2009;77:499–508. [DOI] [PubMed] [Google Scholar]
  57. Ingolia NT, Brar GA, Rouskin S, McGeachy AM, Weissman JS. The ribosome profiling strategy for monitoring translation in vivo by deep sequencing of ribosome‐protected mRNA fragments. Nat Protoc. 2012;7:1534–1550. [DOI] [PMC free article] [PubMed] [Google Scholar]
  58. Iyer M, Jaroszewski L, Sedova M, Godzik A. What the protein data bank tells us about the evolutionary conservation of protein conformational diversity. Protein Sci. 2022;31:e4325. [DOI] [PMC free article] [PubMed] [Google Scholar]
  59. Jackson V, Hermann J, Tynan CJ, Rolfe DJ, Corey RA, Duncan AL, et al. The guidance and adhesion protein FLRT2 dimerizes in cis via dual small‐X3‐small transmembrane motifs. Structure. 2022;30:1354–1365.e5. [DOI] [PubMed] [Google Scholar]
  60. Jumper J, Evans R, Pritzel A, Green T, Figurnov M, Ronneberger O, et al. Highly accurate protein structure prediction with AlphaFold. Nature. 2021;596:583–589. [DOI] [PMC free article] [PubMed] [Google Scholar]
  61. Kabsch W, Sander C. Dictionary of protein secondary structure: pattern recognition of hydrogen‐bonded and geometrical features. Biopolymers. 1983;22:2577–2637. [DOI] [PubMed] [Google Scholar]
  62. Kasai T, Inoue M, Koshiba S, Yabuki T, Aoki M, Nunokawa E, et al. Solution structure of a BolA‐like protein from Mus musculus. Protein Sci. 2004;13:545–548. [DOI] [PMC free article] [PubMed] [Google Scholar]
  63. Kawaguchi S, Xu X, Soga T, Yamaguchi K, Kawasaki R, Shimouchi R, et al. In silico screening by AlphaFold2 program revealed the potential binding partners of nuage‐localizing proteins and piRNA‐related proteins. elife. 2025;13:RP101967. [DOI] [PMC free article] [PubMed] [Google Scholar]
  64. Kesner JS, Chen Z, Shi P, Aparicio AO, Murphy MR, Guo Y, et al. Noncoding translation mitigation. Nature. 2023;617:395–402. [DOI] [PMC free article] [PubMed] [Google Scholar]
  65. Kim A‐R, Hu Y, Comjean A, Rodiger J, Mohr SE, Perrimon N. FlyPredictome: A structural atlas of predicted protein‐protein interactions in Drosophila. bioRxiv. 2026, 2026.04.14.71852. [Google Scholar]
  66. Konar S, Arif H, Allolio C. Mitochondrial membrane model: lipids, elastic properties, and the changing curvature of cardiolipin. Biophys J. 2023;122:4274–4287. [DOI] [PMC free article] [PubMed] [Google Scholar]
  67. Kondo T, Hashimoto Y, Kato K, Inagaki S, Hayashi S, Kageyama Y. Small peptide regulators of actin‐based cell morphogenesis encoded by a polycistronic mRNA. Nat Cell Biol. 2007;9:660–665. [DOI] [PubMed] [Google Scholar]
  68. Königsmann T, Parfentev I, Urlaub H, Riedel D, Schuh R. The bicistronic gene würmchen encodes two essential components for epithelial development in drosophila. Dev Biol. 2020;463:53–62. [DOI] [PubMed] [Google Scholar]
  69. Krogh A, Larsson B, von Heijne G, Sonnhammer EL. Predicting transmembrane protein topology with a hidden Markov model: application to complete genomes. J Mol Biol. 2001;305:567–580. [DOI] [PubMed] [Google Scholar]
  70. Kumar M, Michael S, Alvarado‐Valverde J, Zeke A, Lazar T, Glavina J, et al. ELM—the eukaryotic linear motif resource—2024 update. Nucleic Acids Res. 2023;52:D442–D455. [DOI] [PMC free article] [PubMed] [Google Scholar]
  71. Landon C, Sodano P, Hetru C, Hoffmann J, Ptak M. Solution structure of drosomycin, the first inducible antifungal protein from insects. Protein Sci. 1997;6:1878–1884. [DOI] [PMC free article] [PubMed] [Google Scholar]
  72. Levental I, Lyman E. Regulation of membrane protein structure and function by their lipid nano‐environment. Nat Rev Mol Cell Biol. 2023;24:107–122. [DOI] [PMC free article] [PubMed] [Google Scholar]
  73. Li C, Sun X‐N, Funcke J‐B, Vanharanta L, Prasanna X, Gov K, et al. Adipogenin promotes the development of lipid droplets by binding a dodecameric seipin complex. Science. 2025;390:eadr9755. [DOI] [PMC free article] [PubMed] [Google Scholar]
  74. Liang C, Zhang S, Robinson D, Ploeg MV, Wilson R, Nah J, et al. Mitochondrial microproteins link metabolic cues to respiratory chain biogenesis. Cell Rep. 2022;40:111204. [DOI] [PubMed] [Google Scholar]
  75. Lisi M, Santini T, D'Andrea T, Salvatori B, Setti A, Paiardini A, et al. SERTM2: a neuroactive player in the world of micropeptides. EMBO Rep. 2025;26:2044–2076. [DOI] [PMC free article] [PubMed] [Google Scholar]
  76. Liu C, Wu K, Choi H, Han HL, Zhang X, Watson JL, et al. Diffusing protein binders to intrinsically disordered proteins. Nature. 2025;644:809–817. [DOI] [PMC free article] [PubMed] [Google Scholar]
  77. Lorent JH, Cabrera‐Jojoa A, Levental KR, Levental I, Lyman E. Asymmetric membrane properties through a protein lens. Faraday Discuss. 2025;259:597–613. [DOI] [PMC free article] [PubMed] [Google Scholar]
  78. Lorent JH, Levental KR, Ganesan L, Rivera‐Longsworth G, Sezgin E, Doktorova M, et al. Plasma membranes are asymmetric in lipid unsaturation, packing and protein shape. Nat Chem Biol. 2020;16:644–652. [DOI] [PMC free article] [PubMed] [Google Scholar]
  79. Magny EG, Platero AI, Bishop SA, Pueyo JI, Aguilar‐Hidalgo D, Couso JP. Pegasus, a small extracellular peptide enhancing short‐range diffusion of wingless. Nat Commun. 2021;12:5660. [DOI] [PMC free article] [PubMed] [Google Scholar]
  80. Magny EG, Pueyo JI, Pearl FMG, Cespedes MA, Niven JE, Bishop SA, et al. Conserved regulation of cardiac calcium uptake by peptides encoded in small open Reading frames. Science. 2013;341:1116–1120. [DOI] [PubMed] [Google Scholar]
  81. Majewski M, Pérez A, Thölke P, Doerr S, Charron NE, Giorgino T, et al. Machine learning coarse‐grained potentials of protein thermodynamics. Nat Commun. 2023;14:5739. [DOI] [PMC free article] [PubMed] [Google Scholar]
  82. Markus D, Pelletier A, Boube M, Port F, Boutros M, Payre F, et al. The pleiotropic functions of Pri smORF peptides synchronize leg development regulators. PLoS Genet. 2023;19:e1011004. [DOI] [PMC free article] [PubMed] [Google Scholar]
  83. Marrink SJ, Corradi V, Souza PCT, Ingólfsson HI, Tieleman DP, Sansom MSP. Computational modeling of realistic cell membranes. Chem Rev. 2019;119:6184–6226. [DOI] [PMC free article] [PubMed] [Google Scholar]
  84. Melnikova L, Kostyuchenko M, Molodina V, Parshikov A, Georgiev P, Golovnin A. Multiple interactions are involved in a highly specific association of the mod(mdg4)‐67.2 isoform with the Su(Hw) sites in drosophila. Open Biol. 2017;7:170150. [DOI] [PMC free article] [PubMed] [Google Scholar]
  85. Meng K, Lu S, Li Y, Hu L, Zhang J, Cao Y, et al. LINC00493‐encoded microprotein SMIM26 exerts anti‐metastatic activity in renal cell carcinoma. EMBO Rep. 2023;24:EMBR202256282. [DOI] [PMC free article] [PubMed] [Google Scholar]
  86. Middendorf L, Iyengar BR, Eicholt LA. Sequence, structure, and functional space of drosophila de novo proteins. Genome Biol Evol. 2024;16:evae176. [DOI] [PMC free article] [PubMed] [Google Scholar]
  87. Mishra M. Evolutionary aspects of the structural convergence and functional diversification of Kunitz‐domain inhibitors. J Mol Evol. 2020;88:537–548. [DOI] [PubMed] [Google Scholar]
  88. Muller MP, Jiang T, Sun C, Lihan M, Pant S, Mahinthichaichan P, et al. Characterization of lipid–protein interactions and lipid‐mediated modulation of membrane protein function through molecular simulation. Chem Rev. 2019;119:6086–6161. [DOI] [PMC free article] [PubMed] [Google Scholar]
  89. Na Z, Dai X, Zheng S‐J, Bryant CJ, Loh KH, Su H, et al. Mapping subcellular localizations of unannotated microproteins and alternative proteins with MicroID. Mol Cell. 2022;82:2900–2911.e7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  90. Naftaly S, Cohen I, Shahar A, Hockla A, Radisky ES, Papo N. Mapping protein selectivity landscapes using multi‐target selective screening and next‐generation sequencing of combinatorial libraries. Nat Commun. 2018;9:3935. [DOI] [PMC free article] [PubMed] [Google Scholar]
  91. Negulescu H, Guo Y, Garner TP, Goodwin OY, Henderson G, Laine RA, et al. A Kazal‐type serine protease inhibitor from the defense gland secretion of the subterranean termite Coptotermes formosanus Shiraki. PLoS One. 2015;10:e0125376. [DOI] [PMC free article] [PubMed] [Google Scholar]
  92. Nelson BR, Makarewich CA, Anderson DM, Winders BR, Troupes CD, Wu F, et al. A peptide encoded by a transcript annotated as long noncoding RNA enhances SERCA activity in muscle. Science. 2016;351:271–275. [DOI] [PMC free article] [PubMed] [Google Scholar]
  93. Ødum MT, Teufel F, Thumuluri V, Armenteros JJA, Johansen AR, Winther O, et al. DeepLoc 2.1: multi‐label membrane protein type prediction using protein language models. Nucleic Acids Res. 2024;52:W215–W220. [DOI] [PMC free article] [PubMed] [Google Scholar]
  94. Oliver SG, van der AQJM, Agostoni‐Carbone ML, Aigle M, Alberghina L, Alexandraki D, et al. The complete DNA sequence of yeast chromosome III. Nature. 1992;357:38–46. [DOI] [PubMed] [Google Scholar]
  95. Ong ST, Bajaj S, Tanner MR, Chang SC, Krishnarjuna B, Ng XR, et al. Modulation of lymphocyte potassium channel KV1.3 by membrane‐penetrating, joint‐targeting immunomodulatory plant defensin. ACS Pharmacol Transl Sci. 2020;3:720–736. [DOI] [PMC free article] [PubMed] [Google Scholar]
  96. Orand T, Jensen MR. Binding mechanisms of intrinsically disordered proteins: insights from experimental studies and structural predictions. Curr Opin Struct Biol. 2025;90:102958. [DOI] [PubMed] [Google Scholar]
  97. Orchard S, Ammari M, Aranda B, Breuza L, Briganti L, Broackes‐Carter F, et al. The MIntAct project—IntAct as a common curation platform for 11 molecular interaction databases. Nucleic Acids Res. 2014;42:D358–D363. [DOI] [PMC free article] [PubMed] [Google Scholar]
  98. Öztürk‐Çolak A, Marygold SJ, Antonazzo G, Attrill H, Goutte‐Gattat D, Jenkins VK, et al. FlyBase: updates to the drosophila genes and genomes database. Genetics. 2024;227:iyad211. [DOI] [PMC free article] [PubMed] [Google Scholar]
  99. Park S, Myung S, Baek M. Advancing protein structure prediction beyond AlphaFold2. Curr Opin Struct Biol. 2025;90:102985. [DOI] [PubMed] [Google Scholar]
  100. Pauling L, Corey RB, Hayward R. The structure of protein molecules. Sci Am. 1954;191:51–59. [Google Scholar]
  101. Peng Z, Xiao Q, Wan C. Identification and validation of SmORF‐encoded peptides by genomics and proteomics in five cyanobacteria. J Proteome Res. 2025;24:5818–5829. [DOI] [PubMed] [Google Scholar]
  102. Pogozheva ID, Lomize AL. Evolution and adaptation of single‐pass transmembrane proteins. Biochim Biophys Acta (BBA)‐Biomembr. 2018;1860:364–377. [DOI] [PubMed] [Google Scholar]
  103. Polycarpou‐Schwarz M, Groß M, Mestdagh P, Schott J, Grund SE, Hildenbrand C, et al. The cancer‐associated microprotein CASIMO1 controls cell proliferation and interacts with squalene epoxidase modulating lipid droplet formation. Oncogene. 2018;37:4750–4768. [DOI] [PubMed] [Google Scholar]
  104. Posner Z, Yannuzzi I, Prensner JR. Shining a light on the dark proteome: non‐canonical open reading frames and their encoded miniproteins as a new frontier in cancer biology. Protein Sci. 2023;32:e4708. [DOI] [PMC free article] [PubMed] [Google Scholar]
  105. Prensner JR, Abelin JG, Kok LW, Clauser KR, Mudge JM, Ruiz‐Orera J, et al. What can Ribo‐Seq, immunopeptidomics, and proteomics tell us about the noncanonical proteome? Mol Cell Proteomics. 2023;22:100631. [DOI] [PMC free article] [PubMed] [Google Scholar]
  106. Prins RC, Billerbeck S. Small proteins and peptides conferring protection against antimicrobial compounds. Trends Microbiol. 2025;33:586–602. [DOI] [PubMed] [Google Scholar]
  107. Pueyo JI, Magny EG, Sampson CJ, Amin U, Evans IR, Bishop SA, et al. Hemotin, a regulator of phagocytosis encoded by a small ORF and conserved across metazoans. PLoS Biol. 2016;14:e1002395. [DOI] [PMC free article] [PubMed] [Google Scholar]
  108. Rauh AS, Tesei G, Lindorff‐Larsen K. A coarse‐grained model for disordered proteins under crowded conditions. Protein Sci: A Publ Protein Soc. 2025;34:e70232. [DOI] [PMC free article] [PubMed] [Google Scholar]
  109. Rimphanitchayakit V, Tassanakajon A. Structure and function of invertebrate Kazal‐type serine proteinase inhibitors. Dev Comp Immunol. 2010;34:377–386. [DOI] [PubMed] [Google Scholar]
  110. Ruiz‐Orera J, Hübner N. The non‐canonical proteome: a novel contributor to cancer proliferation. Cell Res. 2025;35:155–156. [DOI] [PMC free article] [PubMed] [Google Scholar]
  111. Sandmann C‐L, Schulz JF, Ruiz‐Orera J, Kirchner M, Ziehm M, Adami E, et al. Evolutionary origins and interactomes of human, young microproteins and small peptides translated from short open reading frames. Mol Cell. 2023;83:994–1011.e18. [DOI] [PMC free article] [PubMed] [Google Scholar]
  112. Savard J, Marques‐Souza H, Aranda M, Tautz D. A segmentation gene in tribolium produces a polycistronic mRNA that codes for multiple conserved peptides. Cell. 2006;126:559–569. [DOI] [PubMed] [Google Scholar]
  113. Savojardo C, Bruciaferri N, Tartari G, Martelli PL, Casadio R. DeepMito: accurate prediction of protein sub‐mitochondrial localization using convolutional neural networks. Bioinformatics. 2019;36:56–64. [DOI] [PMC free article] [PubMed] [Google Scholar]
  114. Schäfer LV, Stelzl LS. Deciphering driving forces of biomolecular phase separation from simulations. Curr Opin Struct Biol. 2025;92:103026. [DOI] [PubMed] [Google Scholar]
  115. Schlesinger D, Dirks C, Navarro C, Lafranchi L, Spinner A, Raja GL, et al. A large‐scale sORF screen identifies putative microproteins involved in cancer cell fitness. iScience. 2025;28:111884. [DOI] [PMC free article] [PubMed] [Google Scholar]
  116. Schnapka V, Morozova TI, Sen S, Bonomi M. Atomic resolution ensembles of intrinsically disordered proteins with Alphafold. Nat Commun. 2026;17:2399. [DOI] [PMC free article] [PubMed] [Google Scholar]
  117. Sehnal D, Bittrich S, Deshpande M, Svobodová R, Berka K, Bazgier V, et al. Mol* viewer: modern web app for 3D visualization and analysis of large biomolecular structures. Nucleic Acids Res. 2021;49:W431–W437. [DOI] [PMC free article] [PubMed] [Google Scholar]
  118. Sehnal D, Rose A, Koca J, Burley S, Velankar S. Mol*: towards a common library and tools for web molecular graphics. 2018.
  119. Sonnhammer EL, von Heijne G, Krogh A. A hidden Markov model for predicting transmembrane helices in protein sequences. Proc Int Conf Intell Syst Mol Biol. 1998;6:175–182. [PubMed] [Google Scholar]
  120. Stevens AO, He Y. Benchmarking the accuracy of AlphaFold 2 in loop structure prediction. Biomolecules. 2022;12:985. [DOI] [PMC free article] [PubMed] [Google Scholar]
  121. Sun L, Wang W, Han C, Huang W, Sun Y, Fang K, et al. The oncomicropeptide APPLE promotes hematopoietic malignancy by enhancing translation initiation. Mol Cell. 2021;81:4493–4508.e9. [DOI] [PubMed] [Google Scholar]
  122. Tesei G, Lindorff‐Larsen K. Improved predictions of phase behaviour of intrinsically disordered proteins by tuning the interaction range. Open Res Eur. 2022;2:94. [DOI] [PMC free article] [PubMed] [Google Scholar]
  123. Tesei G, Schulze TK, Crehuet R, Lindorff‐Larsen K. Accurate model of liquid–liquid phase behavior of intrinsically disordered proteins from optimization of single‐chain properties. Proc Natl Acad Sci. 2021;118:e2111696118. [DOI] [PMC free article] [PubMed] [Google Scholar]
  124. Teufel F, Armenteros JJA, Johansen AR, Gíslason MH, Pihl SI, Tsirigos KD, et al. SignalP 6.0 predicts all five types of signal peptides using protein language models. Nat Biotechnol. 2022;40:1023–1025. [DOI] [PMC free article] [PubMed] [Google Scholar]
  125. The Uniprot Consortium , Bateman A, Martin M‐J, Orchard S, Magrane M, Adesina A, et al. UniProt: the universal protein knowledgebase in 2025. Nucleic Acids Res. 2025;53:D609–D617. [DOI] [PMC free article] [PubMed] [Google Scholar]
  126. Tompa P. Unstructural biology coming of age. Curr Opin Struct Biol. 2011;21:419–425. [DOI] [PubMed] [Google Scholar]
  127. Treichel AJ, Bazzini AA. Casting CRISPR‐Cas13d to fish for microprotein functions in animal development. iScience. 2022;25:105547. [DOI] [PMC free article] [PubMed] [Google Scholar]
  128. Uzarska MA, Nasta V, Weiler BD, Spantgar F, Ciofi‐Baffoni S, Saviello MR, et al. Mitochondrial Bol1 and Bol3 function as assembly factors for specific iron‐sulfur proteins. elife. 2016;5:e16673. [DOI] [PMC free article] [PubMed] [Google Scholar]
  129. Wang L, Brasnett C, Borges‐Araújo L, Souza PCT, Marrink SJ. Martini3‐IDP: improved martini 3 force field for disordered proteins. Nat Commun. 2025;16:2874. [DOI] [PMC free article] [PubMed] [Google Scholar]
  130. Whited AM, Jungreis I, Allen J, Cleveland CL, Mudge JM, Kellis M, et al. Biophysical characterization of high‐confidence, small human proteins. Biophys Rep. 2024;4:100167. [DOI] [PMC free article] [PubMed] [Google Scholar]
  131. Wright BW, Yi Z, Weissman JS, Chen J. The dark proteome: translation from noncanonical open reading frames. Trends Cell Biol. 2022;32:243–258. [DOI] [PMC free article] [PubMed] [Google Scholar]
  132. Wu K, Jiang H, Hicks DR, Liu C, Muratspahić E, Ramelot TA, et al. Design of intrinsically disordered region binding proteins. Science. 2025;389:eadr8063. [DOI] [PMC free article] [PubMed] [Google Scholar]
  133. Yang W, Xue Y, Zhu P, Jiang Z, He R, Tian H, et al. Goblet cell‐expressed microprotein FXYD3 determines gut homeostasis by maintaining mucus barrier integrity. Cell Rep. 2025;44:116502. [DOI] [PubMed] [Google Scholar]
  134. Zanet J, Benrabah E, Li T, Pélissier‐Monier A, Chanut‐Delalande H, Ronsin B, et al. Pri sORF peptides induce selective proteasome‐mediated protein processing. Science. 2015;349:1356–1358. [DOI] [PubMed] [Google Scholar]
  135. Zhang C, Shine M, Pyle AM, Zhang Y. US‐align: universal structure alignments of proteins, nucleic acids, and macromolecular complexes. Nat Methods. 2022;19:1109–1115. [DOI] [PubMed] [Google Scholar]
  136. Zhang J, Humphreys IR, Pei J, Kim J, Choi C, Yuan R, et al. Predicting protein‐protein interactions in the human proteome. Science. 2025;390:eadt1630. [DOI] [PMC free article] [PubMed] [Google Scholar]
  137. Zhang Y, Skolnick J. Scoring function for automated assessment of protein structure template quality. Proteins: Struct, Funct, Bioinf. 2004;57:702–710. [DOI] [PubMed] [Google Scholar]
  138. Zhang Y, Skolnick J. TM‐align: a protein structure alignment algorithm based on the TM‐score. Nucleic Acids Res. 2005;33:2302–2309. [DOI] [PMC free article] [PubMed] [Google Scholar]
  139. Zheng X, Xiang M. Mitochondrion‐located peptides and their pleiotropic physiological functions. FEBS J. 2022;289:6919–6935. [DOI] [PubMed] [Google Scholar]
  140. Zhou B‐R, Feng H, Kale S, Fox T, Khant H, de Val N, et al. Distinct structures and dynamics of chromatosomes with different human linker histone isoforms. Mol Cell. 2021;81:166–182. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Figure S1. Distribution of predicted localization probabilities obtained with DeepLoc 2.1. Probability value for each microprotein sequence is available in Table S1.

Figure S2. Distribution of pLDDT values for microproteins grouped by sequence length (1–25, 26–50, 51–75, and 76–100 amino acids). The figure shows the distribution of pLDDT scores within each length group, with the horizontal black line indicating the mean pLDDT value for each group. The proportion of microproteins with high‐confidence structural predictions (pLDDT ≥70%) is 86.0% for the 1–25 aa group, 57.4% for 26–50 aa, 49.7% for 51–75 aa, and 60.6% for the 76–100 aa group.

Figure S3. A‐ Sarcolipin‐SERCA1a interaction (PDB: 3W5A) side by side with its closest AlphaFold model of the Sarcolipin (in space‐filling representation) in interaction with the SERCA protein (in cartoon representation) from Drosophila. Figure made with the Mol* program (Sehnal et al., 2021, 2018). B‐ Type of interactions between Sarcolipin and SERCA found in the pdb structure. C and D‐ Left, AlphaFold alternative models model of the Sarcolamban (in space‐filling representation) in interaction with the SERCA protein (in cartoon representation) from Drosophila. Figure made with the Mol* program (Sehnal et al., 2021, 2018). Right, type of interactions between Sarcolamban and SERCA found in the respective AlphaFold models.

PRO-35-e70636-s002.pdf (1.1MB, pdf)

Figure SI2. High resolution of Figure 3.

PRO-35-e70636-s001.pdf (4.8MB, pdf)

Table S1. Table of the 1200 microproteins with cellular localization and Deeploc probability.

PRO-35-e70636-s003.csv (154.1KB, csv)

Data Availability Statement

Some of the data presented in this perspective are deposited at: https://doi.org/10.5281/zenodo.19450794. This comprises: Raw data files and R script to create the list of 1200 microproteins. Cellular landscapes presented in Figure 1d–h. Structure of the 1200 microproteins with pLDDT score in beta‐factor. AlphaFold models for Sarcolamban‐SERCA interactions presented Figure 4c–e and Figure S3. Table with peptide signal probabilities.


Articles from Protein Science : A Publication of the Protein Society are provided here courtesy of The Protein Society

RESOURCES