Skip to main content

This is a preprint.

It has not yet been peer reviewed by a journal.

The National Library of Medicine is running a pilot to include preprints that result from research funded by NIH in PMC and PubMed.

bioRxiv logoLink to bioRxiv
[Preprint]. 2025 Aug 17:2025.07.17.665375. Originally published 2025 Jul 19. [Version 2] doi: 10.1101/2025.07.17.665375

Protein structure alignment significance is often exaggerated

Robert C Edgar, Harutyun Sahakyan
PMCID: PMC12312179  PMID: 40747427

Abstract

Machine learning has generated millions of high-quality predicted protein structures, creating a need for computationally efficient structure search algorithms and robust estimates of statistical significance at this scale. We show that unrelated proteins have a universal tendency towards convergent evolution of secondary and tertiary motifs, causing an excess of high-scoring false positive alignments. We investigate popular structure search and alignment algorithms, finding that previous methods routinely overestimate significance by up to six orders of magnitude. To address these issues, and to accommodate recent innovations in search algorithm design, we describe a novel method for estimating statistical significance. We show that its E -values are accurate, scale successfully with database size, and are robust against the (generally unknown) diversity of folds in the database.

We implement our approach in an online structure search service based on Reseek at https://reseek.online .

Full Text

The Full Text of this preprint is available as a PDF (3.1 MB). The Web version will be available soon.


Articles from bioRxiv are provided here courtesy of Cold Spring Harbor Laboratory Preprints

RESOURCES