ABSTRACT
Tandem repeats play an important role in centromere structure, subtelomeric regions, DNA methylation, recombination and the regulation of gene activity. Analysis of their distribution in genomes offers a potential means for predicting putative centromere locations, which continues to be a challenge for genome annotation. Here we present RepeatOBserver (https://github.com/celphin/RepeatOBserverV1), a new tool for visualising repeat patterns and identifying putative centromere locations, using a Fourier transform of DNA walks. RepeatOBserver can identify and visualise a broad range of perfect and imperfect repeats (3–5000 bp long) in genome assemblies without any a priori knowledge of repeat sequences or the need for optimising parameters. RepeatOBserver heatmaps can distinguish between tandem and retrotransposon repeats. We analysed 159 chromosomes with experimentally‐verified centromere positions from 12 plant and animal species. We find that 93% of experimentally‐verified tandem repeat centromeres occur in regions of low sequence diversity and 97% of retrotransposon centromeres occur in regions with a high abundance of repeat lengths. Depending on the centromere type predicted by the heatmaps, putative centromere locations can be predicted using either a genomic Shannon diversity index or a repeat abundance sum. RepeatOBserver can also locate other regions of interest including potential neocentromeres and gene copy variation. Split and inverted tandem repeats at inversion boundaries suggest that chromosomal inversions or mis‐assemblies can also be located. RepeatOBserver is a flexible tool for comprehensive characterisation of repeat patterns that can be used to visualise and identify a variety of regions of interest in genome assemblies.
Keywords: centromere, DNA periodicity, DNA walk, Fourier transform, inversions, neocentromere, Shannon diversity, subtelomeric regions, tandem repeats, telomeres
Short abstract
see also the Perspective by Amanda M. Larracuente and John S. Sproul
1. Introduction
The increasing availability of chromosome‐scale genome assemblies for a wide range of non‐model species, as well as fully haplotype‐resolved assemblies and pan‐genomes for well‐studied species (e.g., Homo sapiens , Solanum spp., Sorghum bicolor , Gallus gallus domesticus; Li et al. 2022; Gong et al. 2023; Liao et al. 2023; Li et al. 2023; Ruperao et al. 2021; Wang et al. 2021; Wlodzimierz, Rabanal, et al. 2023b), is opening a world of possibilities for looking at non‐genic regions (including tandem repeats and centromeres) and genome structural evolution. Along with improving long read and phasing sequencing technologies, bioinformatics tools are needed to better visualise genome‐wide patterns of sequence and structural variation in these genome assemblies (Miga 2020). For example, being able to predict putative centromeres for species with no direct experimental detection data available (e.g., fluorescent in situ hybridisation, FISH or chromatin immunoprecipitation sequencing, ChIP‐Seq), accurate genome annotations or information on DNA methylation patterns would be extremely useful for understanding genome structure and evolution (Hartley and O'Neill 2019; Wlodzimierz, Hong, et al. 2023a). Likewise, visualising tandem repeat patterns throughout the genome would facilitate identification of other regions of potential interest (e.g., inversions, neocentromeres and gene copy number variants).
1.1. Centromeres
Centromeres play a crucial role in chromosome segregation during cell division and knowing their locations helps predict recombination rates across chromosomes. Centromeres are difficult to reliably identify since they undergo rapid evolution. There are multiple hypotheses as to why. One is centromeric drive (Henikoff et al. 2001; Kursel and Malik 2018; Talbert and Henikoff 2022; Bracewell et al. 2019; Lampson and Black 2017; Fishman and Saunders 2008; Melters et al. 2012; Wlodzimierz, Rabanal, et al. 2023b). Centromere drive occurs when centromeres act selfishly and trigger unequal segregation during female meiosis such that the chromosome with the selfish centromere is more likely than its homologue to be inherited.
Centromeres are made of a combination of epigenetic (e.g., histone methylation) and genomic (often AT‐rich, satellite DNA) features (Saha 2019; Naish et al. 2021; Gent et al. 2012; Talbert and Henikoff 2020; Naish and Henderson 2024). Chromosomes can have one centromere (point, regional or satellite), two centromeres (dicentric) or even centromeres distributed across the whole chromosome (polycentric or holocentric) (Scott and Sullivan 2014; Melters et al. 2012; Talbert and Henikoff 2020; Naish and Henderson 2024). De novo or neo‐centromeres can form due to a shifted epigenetic signature, often to a position without the same underlying sequence; in humans, this can lead to some types of cancer and other extreme phenotypes (Amor and Choo 2002). Typically, centromeres are identified in the lab using FISH probes, ChIP‐Seq, CENP‐A footprinting or chromatin fibre analysis (Saha 2019; Nietzel et al. 2001). Centromeres are known to be enriched with tandem repeats or retrotransposon clusters (Melters et al. 2013; Hartley and O'Neill 2019; Naish and Henderson 2024). The number of repeats at the centromere is thought to strengthen centromere binding during meiosis, leading to centromeric drive (Talbert and Henikoff 2020). However, in many newly formed or artificial centromeres, there are no large tandem repeat arrays, implying that these repeats form later in centromere development (Hartley and O'Neill 2019). In some domesticated species with recent centromere shifts, such as Zea mays , centromeres may lack tandem repeats (Talbert and Henikoff 2020). Past studies have explored and classified centromeric repeat sequence similarities across many species (Melters et al. 2013; Talbert and Henikoff 2020).
Putative centromere locations have only been predicted in a few species using only bioinformatics. One approach used wavelet analyses of gene densities and DNA methylation patterns to detect putative centromeres in Populus trichocarpa (Weighill et al. 2019). However, this analysis requires methylation data and an accurate genome annotation, which are not available for all non‐model genome assemblies. Another proposed method uses patterns in HiC maps to predict putative centromere locations (Varoquaux et al. 2015). Other studies have shown links between CG isochores (Costantini et al. 2006; Zhang and Zhang 2004; Jabbari and Bernardi 2017; Bernardi 2021; Woody et al. 2012; Lynch et al. 2010) and putative centromere positions in some specific species. In plants, some recent new methods (quarTeT and CentIER) identify putative centromere locations using k‐mer analyses (Lin et al. 2023; Xu et al. 2024). Overall, these programs often require complex datasets (e.g., methylation, HiC or gene/TE annotations), multiple input parameters, typically lack simple chromosome‐scale visualisation of repeat structure and sometimes predict multiple or very large regions of potential putative centromere locations.
Other bioinformatics tools annotate tandem repeat and/or transposable elements (TEs) for model reference genomes and their centromeres (TRASH (Wlodzimierz, Hong, et al. 2023a), CentromereArchitect (Dvorkina et al. 2021), HiCAT (Gao et al. 2023), HORmon (Kunyavskaya et al. 2021) and Repeat Modeler (Flynn et al. 2020)). These programs do not predict putative centromere positions but annotate repeat structures within known centromeres. Most rely on knowing the tandem (centromeric) repeats for a given species, limiting these analyses to model species with known repeats (HiCAT and CentromereArchitect). One exception is TRASH, which does not need the sequence of Higher Order Repeats (HORs) to be known but uses a k‐mer technique that requires the tandem repeats it detects to be close to exact matches, preventing annotation of some imperfect repeats (Wlodzimierz, Hong, et al. 2023a).
1.2. Tandem Repeats
Tandem repeats also appear in telomeres and many subtelomeric regions (Greider and Blackburn 1985; Kwapisz and Morillon 2020) and may modulate recombination rates (Bzymek and Lovett 2001). In addition, they can contribute to genetic disorders by affecting gene regulation (Horton et al. 2023). Many assembled chromosomes include nearly perfect tandem telomeric repeats (TTTAGGG in plants (Richards and Ausubel 1988), TTGGGG in protists and TTAGGG in animals (Greider and Blackburn 1985; Kilburn et al. 2001)) that can span a few thousand base pairs. Adjacent to these telomeric repeats, there are often semi‐regular subtelomeric repeats, which cover a larger portion of the chromosome (sometimes a few Mbp) (Kwapisz and Morillon 2020) and can be conserved between species. Coding and noncoding RNAs are transcribed from these subtelomeric regions and are believed to help with the maintenance of the telomere (Kwapisz and Morillon 2020).
Although most repetitive DNA sequences (e.g., long terminal repeat retrotransposons) tend to accumulate in regions of chromosomes that have lower rates of recombination (Tiley and Burleigh 2015), some tandem repeats (specifically GT‐rich ones) have been suggested to be recombination hotspots (Majewski and Ott 2000; Bzymek and Lovett 2001). Exchanges between non‐homologous repeats can also lead to expansions and contractions of repeat arrays (Smith 1976; Pâques et al. 1998; Read et al. 2004). Similarly, non‐homologous repeats can result in chromosomal inversions (Flores et al. 2007), and inversions have been found to be occasionally bracketed by split tandem repeats (Hirabayashi and Owens ).
Finally, increases in the abundance of short tandem repeats (STR), in particular trinucleotide repeats, are associated with some genetic disorders in both H. sapiens (Paulson 2018) and Arabidopsis thaliana (Sureshkumar et al. 2009). Gene regulation is also affected by STRs that can bind transcription factors (Horton et al. 2023). Overall, both short and long tandem repeats play an important role in many aspects of genome organisation and function.
Several heuristic programs have been developed to identify tandem repeats of a wide range of lengths in whole‐genome assemblies. One of the first, and most frequently used, is Tandem Repeat Finder, which uses Bernoulli trials of k‐mers to statistically detect candidate tandem repeats (Benson 1999). This program requires the user to set input parameters that allow for varying amounts of mutation or error in the repeats they can identify. More recently, other programs such as TROLL (Castelo et al. 2002), mreps (Kolpakov et al. 2003), Spectral Repeat Finder (SRF) (Sharma et al. 2004), ATRHunter (Wexler et al. 2005), RepeatMasker (Tarailo‐Graovac and Chen 2009), TRStalker (Pellegrini et al. 2010), RepeatExplorer (Novák et al. 2013) and Tide Hunter (Gao et al. 2019) have been designed to better search for fuzzy or imperfect tandem repeats. K‐Seek is a tool that has been used in Drosophila genomes to find common tandem repeats in Illumina short reads (Wei et al. 2014). There are also databases of common tandem repeats that can be used to compare between species and aid in annotation (Navajas‐Pérez and Paterson 2009). None of these tandem repeat detection programs predict putative centromere locations based on these repeats and most programs require the user to adjust input parameters to detect all types of repeats.
Visualising repeating sequences across the genome can be difficult. One recent method for visualising and comparing repeating sequences across a genome is called StainedGlass plots (Vollger et al. 2022). It builds upon an older method of using dotplots to visualise the locations of a particular repeat length (k‐mer). However, visualising the locations of many different repeat lengths across the genome in one plot remains challenging.
1.3. RepeatOBserver
We developed a new repeat visualisation tool, RepeatOBserver, to explore chromosome‐wide patterns of repeats and predict putative centromeric regions. Starting from only a fasta file for a given chromosome, RepeatOBserver produces Fourier spectra showing locations of repeats (including their length and how perfectly they repeat). Fourier transforms and wavelet analyses have previously been used to study DNA, including exploration of DNA periodicity (Elloumi et al. 2012; Nagai et al. 2001, 2020), DNA palindromes (Qi et al. 2012), sequence comparison (MAFFT) (Katoh et al. 2002), sequence evolution (Machado 2013), exon/intron identification (Haimovich et al. 2006), visualisation of regular features (Dodin et al. 2000) and even tandem repeats (Sharma et al. 2004; Brodzik 2007; Buchner and Janjarasjitt 2003; Yadav et al. 2022). DNA sequence variations in Fourier transforms have previously been shown to have biological importance (Haimovich et al. 2006) and are known to show periodicity and tandem repeats (Elloumi et al. 2012; Sharma et al. 2004; Brodzik 2007; Buchner and Janjarasjitt 2003; Yadav et al. 2022). Fourier transforms specifically detect repeats that occur near each other in the genome while missing repeats (such as rare TEs) that occur only once in each Fourier window. This gives them a unique advantage when searching for highly repetitive regions of the genome such as centromeric and pericentromeric regions.
Here, we apply Fourier transforms to chromosome‐scale AT or CG DNA walks to provide an accessible tool for visualising repeat patterns and predicting putative centromere regions based on repeat abundance and diversity in chromosome‐level genome assemblies. RepeatOBserver also allows for the easy visualisation of positions of potential neocentromeres, gene copy number variation, repeats around inversions and other short and long tandem repeat patterns. We chose to use Fourier transforms of DNA walks since this allows for a high degree of imperfection in the repeats that are detected, as well as not requiring any preset k‐mer lengths that are otherwise required for k‐mer methods.
The advantages of RepeatOBserver when compared to other tandem repeat detection software (Table S0) derive from its ability to: (1) predict putative regions containing centromeric and pericentromeric repeats; (2) find tandem repeats without any prior knowledge; (3) detect a wide range of repeat lengths (3–5000 bp); (4) identify imperfect tandem repeats, since it does not rely on k‐mers; (5) visualise whole chromosomes; and (6) perform these functions without a user needing to optimise input parameters.
2. Methods
2.1. R Package
The R package RepeatOBserver can be found on Github: https://github.com/celphin/RepeatOBserverV1 and can be run locally (1 cpu) or with multiple cpu locally or on a server. Scripts for running all the code described in this paper are included in the repository.
2.2. 2D DNA Walks
The R package RepeatOBserver first converts the chromosome fasta file input into a numerical DNA walk. Beginning at (0,0) and stepping through each nucleotide in the chromosome, the 2D DNA walk moves horizontally along the x axis by +1 for each A and by −1 for each T in the genome (Figure 1A). Similarly, the walk moves vertically along the y axis by +1 for C and −1 for G. Plots of this numerical DNA walk allow the overall chromosome structure and large‐scale tandem repeats to be viewed (Berger et al. 2004). Dominance in any nucleotide or set of repeating nucleotides can be viewed as a slow movement along a vector in the 2D walk.
FIGURE 1.

Example 40 kbp sequences from the A. thaliana genome (COLCEN) (Naish et al. 2021) chromosome 4 showing a non‐repetitive section of the genome (left) and a repetitive sequence (right). (A) Plot of a 2D DNA walk with the Araport11 annotation showing forward genes plotted in blue and reverse genes plotted in pink (Cheng et al. 2017). (B) 1D AT DNA walk with a smoothing spline (red line), again forward genes plotted in blue and reverse genes plotted in pink. (C) Smoothed sequence prior to the Fourier Transform. (D) Fast Fourier Transform output showing the relative abundance of the different frequencies (with frequency represented by 1/repeat length in bp). The largest repeat length (smallest frequency) is the fundamental and the length of the repeating DNA sequence; the harmonics are also shown and describe the complexity of the repeat.
2.3. 1D DNA Walks
One‐dimensional DNA walks are plotted and used for the Fourier transform input. A one‐dimensional AT walk starts at (0,0) and steps along the x‐axis for each base pair in the genome sequence. On the y‐axis, each A in the sequence increments the plot by +1, each T by −1 and both C and G increment by 0 (Figure 1B). Similarly, the CG walk starts at (0,0) and steps along the x‐axis for each base pair in the genome sequence. On the y‐axis, each C in the sequence increments the plot by +1, each G by −1 and both A and T increment by 0. Gene annotations can be plotted on 1D DNA walks in RepeatOBserver, with forward genes (blue) most often appearing in the AT walk as negative slopes and reverse genes (pink) as positive slopes. This pattern likely results from a higher frequency of codons enriched in T and explains some of the DNA walk movement outside of the repetitive regions.
2.4. Smoothing Spline
Over large regions of the genome there can be a trend in base pair composition that is not directly related to the shorter repeats of interest. This large‐scale trend often has a large amplitude that will dominate the repeat abundance in the Fourier transform. To remove this large‐scale trend, we first fit a cubic spline curve to the original 1D DNA walk (Figure 1B). We then smooth this trend by subtracting this curve from the original 1D DNA walk (Figure 1C). This is done using the smooth.spline function in the R stats package (smooth.spline function—RDocumentation). This allows the Fourier transform to focus on the local repeats and not on the large‐scale trends/waves in the DNA walks.
2.5. Fourier Transform
A Fast Fourier Transform (FFT) is run on the smoothed 1D DNA walk. The FFT is run on windows of 5000 bp at a time to allow for the detection of a wide range of repeat lengths (from 2 to 2500 bp). We used the fft function in the R stats package (RPubs—Introduction Fast Fourier Transform in R). The FFT outputs complex values. We plot the modulus of these values to produce a strength for each discrete frequency (1/repeat length) (Figure 1D). Large strengths imply a relatively high abundance of the given repeat length in that 5000 bp window. To visualise repeats across a chromosome, these 5000 bp windows can be joined and relative abundances of each repeat length can be plotted as a heat map (Figure 2D).
FIGURE 2.

Arabidopsis thaliana genome (COLCEN) chromosome 1 (Naish et al. 2021) (A) Full chromosome 1D AT DNA walk. (B) Full chromosome 1D CG DNA walk. (C) Full chromosome 2D DNA walks. DNA walks are coloured based on the rainbow; starting at red going to violet and changing colour every 5 Mbp to help visualise how the plots relate to each other. The long green linear regions correspond to centromeric repeats. (D) Heat map showing all the 5000 bp frequency spectra (e.g., from Figure 1D) joined together to plot the entire COLCEN chromosome 1, with the x axis showing genome position, the y axis plotting the frequency (1/repeat length) and the colour showing the relative abundance of the repeat at that frequency in that position of the genome. The relative repeat abundance colour does not have units. (E) Heat map of the older A. thaliana TAIR10.1 assembly chromosome 1 (Berardini et al. 2015). (F) EMBOSS plots of CG isochores across chromosome 1 in the new COLCEN and older TAIR10.1 assemblies. Comparison between the older A. thaliana TAIR10.1 assembly (Berardini et al. 2015) and the newer COLCEN genome assembly (Naish et al. 2021) shows the improvements in assembly quality for repeat sequences in recent years.
These Fourier transforms of 1D AT walks can infer tandem repeats not only in AT but also other nucleotide combinations (e.g., CG, AG, GT, etc.) based on gaps in the AT repeats. There are options within the R package to run the Fourier transform on any pair of base pairs, although AT is the default. Fourier transforms of AT walks were compared to Fourier transforms of CG walks. The same fundamental repeat lengths at the same positions in the genome were detected in both transforms. However, occasionally the abundance of harmonics can vary between the Fourier transforms run on the AT and CG walks. Running 2D wavelet analyses results in complex 2D transformations that are more difficult to interpret and visualise. Here we use 1D walks since we can show that they are able to detect the same fundamental repeat lengths and are easier to understand and visualise. In rare cases where repeats occur only in CG, running the CG Fourier transform would allow for the specific CG repeat lengths to be determined. Note that these regions will still appear in the AT Fourier transform as dark vertical bars (with no obvious repeat lengths).
2.6. Genomic Shannon Diversity
To explore repeat diversity across the chromosome, we introduce a genomic form of the Shannon diversity index (Shannon 1948; Good 1953) often used in ecology to describe species diversity (i.e., high values of the Shannon diversity index mean that many species are present at similar abundance in a certain geographic region). In our case, each repeat length is the equivalent of a species, and each window of the genome is the equivalent of a geographic location. The abundance of each repeat output from the Fourier transform is the equivalent of the number of individuals of a given species in a location. Shannon diversity was calculated in R on the Fourier transform output using the diversity function in the vegan package (Oksanen 2022; diversity function—RDocumentation). The Fourier transform abundance matrix was normalised to values between zero and one by dividing each value in the matrix by the total summed abundance for each repeat length. The Shannon diversity index can be run on the default 5 kbp Fourier windows to find small regions containing densely clustered repeats or averaged over larger (up to 5 Mbp) windows to detect large‐scale changes in repeat diversity. We find that tandem repeat clusters are defined as regions of low repeat diversity (i.e., a low Shannon diversity index).
2.7. Repeat Abundance Sum
Along with repeat diversity, total repeat abundance changes over the chromosome. To calculate repeat abundance across a chromosome, we sum the Fourier abundance of all repeat lengths in each 5 kbp Fourier window. Then we run a rolling sum across these 5 kbp summed values to determine the repeat abundance over window sizes ranging from 500 kbp to 5 Mbp. The 500 kbp sums show detailed smaller features while the larger windows show large‐scale patterns in repeat abundance over the chromosomes. We find that clusters of such repetitive sequences, not in tandem (such as retrotransposons and other transposable elements), appear clearly as maxima in the repeat abundance sums, while clusters of tandem repeats often show up as minima.
2.8. Putative Centromere Prediction
The centromere and pericentromeric regions are usually characterised by either the presence of a single or a couple of dominant tandem repeats, or a cluster of retrotransposons or other non‐tandemly repeating sequences (Hartley and O'Neill 2002; Naish and Henderson 2024). Regions of dense tandem repeats appear as minima in Shannon diversity plots and clusters of non‐tandem repetitive elements (e.g., retrotransposons) appear as maxima in the repeat abundance plots.
To determine what type of centromere exists in each species, we can examine the heatmaps of the Fourier spectra (Figure 3 heatmaps). Major types of centromeres have been described in Naish and Henderson (2024). A common type of centromere is one with a tandem repeating sequence appearing over a large region and will show up in the Fourier heatmaps as a wide dark vertical bar with a few bright horizontal lines representing the repeat lengths present (e.g., Figure 3A–C). In Naish and Henderson (2024) this is referred to as a monocentric satellite centromere. We hypothesise that chromosomes with clear clusters of tandem repeats have centromeres predicted by the minimum in the Shannon diversity plots.
FIGURE 3.

Five example species illustrating various centromere types. For each species, one chromosome's Fourier heatmap, a diagram of the ChIPseq experimentally‐verified centromere positions (top row, blue) compared to RepeatOBserver's predicted centromere locations (bottom row, yellow) and predictive RepeatOBserver plot are shown. Below this, more diagrams comparing the known and predicted centromere locations in other chromosomes of a similar length from the same species are shown. Similar plots for all species and chromosomes listed in Table 1 can be found in Figures S1–S11. (A) Oryza sativa (rice) Chr 1, (B) Brassica rapa (turnip) Chr 1 and (C) H. sapiens (human) Chr 10, showing monocentric satellite or tandem repeat centromeres with a minimum in the Shannon diversity plot at the centromere. (D) Gossypium barbadense (sea cotton) Chr 14 showing a monocentric retrotransposon centromere with a maximum in the repeat abundance plot at the centromere. (E) Morus notabilis (mulberry) holocentric Chr 1 showing multiple tandem repeat clusters in the heatmaps and associated local minima in the Shannon diversity index.
Another less common but very distinct type of centromere is made up of retrotransposons and is referred to in Naish and Henderson (2024) as a monocentric retrotransposon type centromere. Repetitive sequences that vary in length will appear as blurs in the Fourier heatmaps. These blurs are picking up the repeated insertions of retrotransposons or other repeating sequences that are near each other but not perfectly in tandem; hence there is variation in detected repeat lengths. Thus, retrotransposon centromeres should appear in the heatmaps as a bright vertically blurred region (Figure 3D). We hypothesise that chromosomes that show clusters of retrotransposons in the heatmaps will have centromeres predicted by the maximum in the repeat abundance plots.
Other even rarer centromere types are found on metapolycentric and holocentric chromosomes (Figure 3E). These chromosomes should show many regions with the same tandem repeat pattern appearing in multiple places across each chromosome. We hypothesise that the multiple centromere locations will be identified as local minima in the Shannon diversity index.
3. Results and Discussion
RepeatOBserver can be run on any chromosome‐scale assembly to visualise tandem repeats, predict putative centromeric regions (including the centromeric and pericentromeric repeats) and identify other regions of potential interest (e.g., neocentromeres, historic centromeres and inversions). Entire chromosomes are represented as 1D and 2D DNA walks, which can be useful for visualising the dominant base pairs in each region of the genome (Figure 2A–C). Fourier transforms of AT DNA walks are then used to visualise the repeat spectra in 5 kbp windows as a heatmap of the abundance of repeats of different lengths across each chromosome (Figure 2D,E). Large regions of tandem repeats appear clearly as bright horizontal lines, against a dark background, (e.g., Figure 2D 15–17 Mbp). These bars describe a particular repeat spanning many 5 kbp Fourier transform windows along the chromosome. The fundamental frequency detected by the transform is the true repeat length and the lower bars on the plot represent its harmonics. These plots make it easy to visualise repeat patterns across the chromosome.
In what follows, we show how RepeatOBserver can be used to: (1) predict regions that may contain centromeric and pericentromeric repeats, (2) visualise repeat patterns across the chromosome, (3) explore variation in repeat patterns across species and (4) investigate repeat patterns at neocentromeres, areas with gene copy number variation and inversion boundaries, using a wide range of published plant and animal genomes (Table S1).
TABLE 1.
Comparison of experimentally‐verified centromere ranges with RepeatOBserver predicted centromere locations.
| Chromosome type | Species | Reference | Centromere count | Absolute min/max overlap | Range overlap |
|---|---|---|---|---|---|
| Monocentric satellite (tandem repeat) | 103 | 76 (74%) | 96 (93%) | ||
| Mouse‐ear cress Arabidopsis thaliana (COLCEN assembly) | Naish et al. (2021) | 5 | 5 | 5 | |
| Human Homo sapiens | Nassar et al. (2023) and Logsdon et al. (2024) | 23 | 15 | 22 | |
| Maize Zea mays | Gent et al. (2015), Jiao et al. (2017) and Wolfgruber et al. (2009) | 10 | 0 | 5 | |
| Fruit fly Drosophila pseudoobscura lowei | Bracewell et al. (2019) | 4 | 4 | 4 | |
| Rice Oryza sativa (Japonica Group) | Zhang et al. (2013), Nawaz et al. (2014) and Song et al. (2021) | 12 | 11 | 12 | |
| Turnip Brassica rapa | Lim et al. (2005) and Zhang et al. (2023) | 10 | 10 | 10 | |
| Mouse Mus musculus | Hughes et al. (2007) and Nassar et al. (2023) | 19 | 16 | 19 | |
| Soybean Glycine max | Liu et al. (2023) | 20 | 15 | 19 | |
| Monocentric retrotransposon | 33 | 30 (91%) | 32 (97%) | ||
| Einkorn Wheat Triticum monococcum | Ahmed et al. (2023) | 7 | 6 | 6 | |
| Sea Cotton Gossypium barbadense | Chang et al. (2024) | 26 | 24 | 26 | |
| Holocentric | 40 | NA | 38 (95%) | ||
| Mulberry Morus notabilis | Ma et al. (2023) | 40 | NA | 38 | |
| No clear repeats | 17 | 3 (18%) | 12 (70%) | ||
| Sunflower Helianthus annuus (HA412) | Nagaki et al. (2015) and Huang et al. (2022) | 17 | 3 | 12 | |
3.1. Comparing Putative Centromere Predictions With Experimentally‐Verified Positions
To assess RepeatOBserver's ability to detect centromeric regions from the Fourier transform of a DNA walk, we compare predicted putative centromere locations to experimentally verified known centromere locations in 159 chromosomes across 12 species (Table S2, Figure 3). Only species with ChIP‐seq verified centromere locations were included (all species are listed in Table 1). To compare our estimates to the experimentally‐verified locations, we compare how many of the specific locations predicted by RepeatOBserver lie within the ranges experimentally verified by ChIP‐seq data. This method of comparison works even though centromeres are often multiple mega‐base‐pairs (Mbp) in length and can vary tremendously in size between individuals of the same species (Hallast et al. 2023; Logsdon et al. 2024). In some cases, we had to infer the known centromere positions from figures in papers and use slightly different genomes than those described in the papers, introducing further uncertainty in the exact known positions.
3.1.1. Shannon Diversity Minimum and Tandem Repeat Centromeres
We hypothesised that, for species with clear tandem repeat patterns in the heatmaps (Figures 3A–C and 4A–D,H), a minimum in the Shannon diversity index (Shannon 1948) should predict the putative centromere locations, assuming no retrotransposon clusters are apparent. The Shannon diversity index is run on each 5 kbp Fourier window (that can then be averaged across varying window sizes). In the program, a wide range of rolling mean window sizes from 50 kbp up to 5 Mbp are plotted for each chromosome to visualise repeat patterns at various scales. Pericentromeric tandem repeats generally occur on the scale of 2–3 Mbp so we used a rolling window size of 2.5 Mbp to predict putative centromere locations in what follows. We also calculate a putative centromere range based on where in the genome the Shannon diversity index is (somewhat arbitrarily) two standard deviations below the mean. Often this is a range directly around the absolute minimum, but in some cases, this also detects other local minima elsewhere on the chromosome (e.g., multiple yellow bars in Figure 3B B. rapa Chr 4 and Figure 3C H. sapiens Chr 2).
FIGURE 4.

Heat maps of the relative abundance of each repeat length for the centromeric regions in a wide range of species. (A) Drosophila pseudoobscura lowei (Fruit fly) Chr AD showing metacentric tandem repeats ~21 bp long, (B) Solanum pimpinellifolium (tomato) Chr1 showing metacentric tandem repeats ~53 bp long, (C) Danio rerio (zebrafish) Chr 3 showing metacentric tandem repeats ~95 and ~190 bp long, (D) H. sapiens (human) Chr 22 showing telocentric tandem repeats ~171 bp long and other subtelomeric repeats, (E) Z. mays (maize) Chr 4 showing a retrotransposon cluster at the centromere, (F) Triticum monococcum (einkorn wheat) Chr 4A showing a retrotransposon cluster at the centromere, (G) Ceratopteris richardii (fern) Chr 2 showing a retrotransposon cluster at the centromere and an example of the 4 bp repeats found across the genome in small clusters, (H) Brassica rapa (turnip) Chr 5 showing metacentric tandem repeats ~176 bp long and (I) Taxus wallichiana var. yunnanensis (yew) Chr 3 with metacentric blurred (imperfect) tandem repeats that are ~7 and ~18 bp long.
The heatmaps for B. rapa (with 10 chromosomes), A. thaliana (5), Drosophilia pseudoobscura lowei (4), Oryza sativa (12), H. sapiens (23), Glycine max (20), Mus musculus (19) and Z. mays (10) show consistent clear tandem repeat patterns and thus the Shannon diversity method with 2.5 Mbp windows was used to predict the putative centromere positions. In 96 (93%) of the 103 chromosomes, the predicted centromere range overlapped with the ChIP‐Seq experimentally‐verified range. However, only in 76 (74%) of the 103 chromosomes of these species, we observed that the absolute minima in repeat diversity overlapped with the experimentally‐verified centromere range (Figure 3A–C, Figures S1–S3, S6–S9, Table 1 and Table S2). The offset of the centre of the absolute minimum window and the experimentally‐verified centromere range was in most cases very small < 1 Mbp (except for Z. mays discussed below). This offset only occurred in two species ( H. sapiens and G. max ). It may be a result of the genome assemblies we used not being the exact genome assemblies that the ChIP‐Seq data had been mapped to, or because the large rolling windows (2.5 Mbp) prevented the exact minimum position from being determined. It is also possible that the centromeres in some species fall right next to but not at the minimum repeat diversity location.
Multiple centromere positions in polycentric and holocentric chromosomes were also predicted using the Shannon diversity index ranges. Although the program outputs a single putative centromere location for each chromosome at the absolute minimum repeat diversity, the ranges defined by two standard deviations from the mean were able to identify 38 of the 40 ChIP‐Seq identified centromere locations across the six chromosomes of Morus notabilis (Figure 3E, Figure S5).
3.1.2. Repeat Abundance Maximum and Retrotransposon Centromeres
We hypothesised that for species with a distinct cluster of retrotransposons (appearing as brightened vertical ‘blurs’ in the Fourier heatmaps) (Figures 3D and 4E–G, Figure S12), we could predict the putative centromere location based on where the repeat abundance sum maximised. Calculating the repeat abundance involves running a rolling sum of the abundances for all repeat lengths in the Fourier transform across a given window. RepeatOBserver calculates this rolling sum for window sizes ranging from 50 kbp to 5 Mbp. Based on the heatmaps, we found that clusters of retrotransposons varied in size from 0.5 to 1 Mbp. Thus, we used a rolling window of 500 kbp to determine the location with the maximum repeat abundance across the chromosome. In what follows, the putative centromere position in these retrotransposon centromeres was predicted as the centre of the 500 kbp window with the maximum summed repeat abundance. A range was predicted arbitrarily as any windows more than two and a half standard deviations above the mean repeat abundance. These ranges give a sense of the quality of the result. A single large block range suggests a single putative centromere position and multiple spread‐out ranges imply other possible putative centromere positions may exist.
Gossypium barbadense (sea cotton, 26 chromosomes, Chang et al. 2024) and Triticum monococcum (einkorn wheat, 7 chromosomes, Ahmed et al. 2023) are examples of species that had chromosomes with a blurred bright bar in the heatmaps indicating a cluster of retrotransposons (Figures 3D and 4E–G, Figures S4, S12, S17). In these species, the window with the absolute maximum repeat abundance aligned with the ChIP‐Seq identified centromere locations in 30 out of 33 chromosomes (and the predicted and known ranges overlapped in 97% of chromosomes).
3.1.3. Repeat Abundance Minimums and Sunflower Centromeres
Of the 159 chromosomes and 12 species with experimentally‐verified centromeres that we explored, only one species' chromosomes, Helianthus annuus (sunflower), did not have clearly visible repeating patterns in their heatmaps (i.e., no clear consistent tandem repeats or retrotransposon clusters across chromosomes) (Figure S10). We noticed that the absolute minimum of the repeat abundance sums aligned with centromere locations in H. annuus and we propose that this may also work in other species without clear repeat patterns in their heatmaps. However, we have no other similar species to test this in and thus have no evidence if this would work in other species as well. We also do not have an explanation for why the centromeric regions in H. annuus have such a low abundance of most repeat lengths since we do not observe any one repeat length dominating as we normally do in tandem repeat satellite type centromeres.
3.1.4. Limitations
All predictions outputted by RepeatOBserver are putative centromere ranges based on only the repetitive nature of the DNA sequence. RepeatOBserver does not consider other important epigenetic features that may be essential in determining the true centromere. The program identifies both putative centromeric and pericentromeric repeats, so the region predicted is often large. Users of RepeatOBserver should be aware that predicted centromere locations are putative positions given sequence data but that only experimental evidence such as ChIP‐Seq can confirm these positions. Comparisons of our estimates with experimentally‐verified data show that patterns in the heatmaps provide information suggesting a type of centromere, a putative centromere location and how much confidence a user should have regarding the prediction's accuracy. Single clear tandem repeat or retrotransposon clusters are likely centromeres, while chromosomes with no clear pattern in the heatmaps may have significantly lower precision and accuracy in their centromere prediction.
The heatmaps should support the locations predicted based on the minima or maxima in the Shannon diversity or repeat abundance plots. In some cases, the location predicted by RepeatOBserver may be less accurate than the positions that appear obvious in the heatmaps (e.g., Figure 3A O. sativa Chr 1). In these cases, using the heatmap locations and potential associated local minima or maxima will be a better estimate. For example, in G. barbadense chromosome 26 (Figure 3D), although the absolute maximum does not lie in the experimentally‐verified range, another local maximum does. Another example, in Z. mays (Dawe et al. 2023), shows that although tandem repeat and retrotransposon centromeres are visibly clear in the Fourier heatmaps, only 5 of the 10 centromere locations in Z. mays were accurately predicted using the Shannon diversity index alone (Figure S11). This may be due to the Z. mays centromeres having moved in recent evolutionary time (Talbert and Henikoff 2020) or due to them having both tandem repeat and retrotransposon type centromeres, resulting in the detection of multiple locations in the genome (Figure S11).
3.2. Comparing RepeatOBserver Output to Known Repeat Patterns
To confirm that Fourier transforms of AT DNA walks can accurately describe genome‐wide repeat patterns, we ran RepeatOBserver on chromosomes containing well‐studied patterns of tandem repeats. In D. pseudoobscura chromosomes (Bracewell et al. 2019), we find a clear signal at 20–21 bp, matching the known centromeric repeat (Figure 4A). Similarly, we can detect the ~52 bp repeat known to be associated with putative centromeric regions in the Solanum lycopersicum genome assembly (Jo et al. 2009) (Figure 4B). In A. thaliana , RepeatOBserver plots permit clear identification of the 178 bp repeats (CEN178) that represent the sites where centromere‐specific histones (CENTROMERIC HISTONE3, CENH3) are loaded (Wlodzimierz, Rabanal, et al. 2023b). We also detect the pericentromeric ~150 and ~500 bp repeats that precede and follow CEN178 repeats, respectively (Figure 2D) (Wlodzimierz, Hong, et al. 2023a). Interestingly, comparisons between an older A. thaliana genome assembly (TAIR 10.01; Berardini et al. 2015) and a more recent assembly (Col‐CEN; Naish et al. 2021) highlight the improvements in the assembly of structural variants and repeated regions that can be obtained with long read technologies (Figure 2D,E). When applied to the H. sapiens Y chromosome, RepeatOBserver can easily identify and visualise the boundaries of the centromeric, euchromatin and heterochromatin regions (Rhie et al. 2023) (Figure 5).
FIGURE 5.

Homo sapiens (human) Y chromosome (Rhie et al. 2023) heat map (rotated) showing the relative abundance of each repeat length at each chromosome position. Transitions in repeat structure match known positions in the chromosome for pseudoautosomal region (PAR), euchromatin, heterochromatin and the centromere.
RepeatOBserver, if EMBOSS is installed, will plot the CG isochore patterns to compare with the repeat patterns in each chromosome. EMBOSS plots (Rice et al. 2000) of CG isochores provide a simple means for detecting the centromeres in many species (e.g., fungi) with AT‐rich centromeric sequences. However, in some plant and animal genome assemblies it is less clear if there is always an isochore signal at the centromere (e.g., H. annuus ). We were not able to determine whether this lack of signal is related to the quality of the reference genome (e.g., newer vs. older A. thaliana chromosomes in Figure 2D–F) or to the actual composition of repeats throughout that genome. We hypothesise that it may be related to the type of centromere, with satellite‐rich, AT‐dense old centromeres appearing clearly in CG isochore plots and recently originated centromeres characterised by retrotransposons not appearing at all (e.g., in Z. mays or H. annuus ).
3.3. Variation in Repeat Patterns
By plotting the Fourier transforms as heatmaps in a wide variety of species with and without experimentally‐verified centromere positions, we can begin to visualise the diversity in repeat patterns that exist across chromosomes and species (Figure 4).
3.3.1. Tandem Repeats
Drosophila pseudoobscura lowei, B. rapa and S. pimpinellifolium are examples of species that are predicted to have monocentric satellite centromeres on most chromosomes. These centromeres are consistently represented by a known repeat of length 21 bp in D. pseudoobscura lowei (Bracewell et al. 2019), 52 bp in S. pimpinellifolium (Jo et al. 2009) and 176 bp in B. rapa (Zhang et al. 2023) (Figure 4A,B,H).
Some species' chromosomes contain tandem repeats of varying lengths (Figure 5C). These different tandem repeats appear in the Fourier spectra as bright horizontal bars at different positions (the top bar is the fundamental frequency or true repeat length and the bars under it show higher frequency harmonics that describe the complexity of the specific repeat). Each fundamental frequency and its harmonics represent a particular kind of repeat pattern. The same horizontal bar pattern in two locations implies the same repeat at two different positions in the genome (e.g., a fundamental repeat that is 90 bp long with a harmonic at 45 bp can be seen at both 47 and 57 Mbp in Danio rerio (zebrafish) chromosome 3 in Figure 4C). In D. rerio , it is evident that these matching repeats are separated by a distinct repeat with a different fundamental length (180 bp long) and harmonics at 1/2 (90 bp), 1/3 (60 bp), 1/4 (45 bp) and 1/5 (36 bp). If there is a lower fundamental frequency (large repeat) that is too long for the 5000 bp window that we used, then this pattern would be the fraction of the true fundamental repeat length. The true fundamental could be calculated (or determined using a larger Fourier window, for example, run_long_repeats function using a 20 kbp window option in R package).
3.3.2. Subtelomeric Repeats
RepeatOBserver can help with the visualisation and identification of subtelomeric repeats. In H. sapiens chromosomes 13, 14, 15, 21, 22, X and Y, an extended set of complex repeats marks the subtelomeric region and centromere (e.g., H. sapiens Chr 22, Figure 4D). These subtelomeric repeats can be even less diverse (using Shannon diversity) than the centromeric repeats, making it difficult to consistently differentiate between centromeres and subtelomeric repeats. However, it is usually still possible to determine the approximate centromere location since the subtelomeric repeats are right next to the centromeres in these cases.
3.3.3. Retrotransposon Clusters
In T. monococcum , clusters of retrotransposons previously identified as the centromere (Ahmed et al. 2023) can be recognised as blurs in the Fourier spectra that start and end at the boundaries of the retrotransposon clusters (Figure 4F). Similar patterns can be seen in some Z. mays centromeres (Figure 4E). Although no ChIP‐Seq data is available, clear retrotransposon clusters indicating putative monocentric retrotransposon centromeres appear in Ceratopteris richardii (fern, Figure 4G) and Rubus idaeus (raspberry, Figure S12) chromosomes. In Taxus wallichiana (yew, Figure 4I) we show a putative centromere location that has somewhat blurred repeats likely due to imperfections (mutations, insertions and deletions) in the tandem repeats and not retrotransposons.
3.3.4. Periodicity
Chromosome‐scale repetitive patterns, termed DNA periodicity here, are also evident when looking at the Fourier heatmaps. In all chromosomes there is a well‐known 3 bp periodicity in exons that we could observe in the Fourier heatmaps (Figure S13; Eskesen et al. 2004; Nagai et al. 2001). In C. richardii chromosomes, there is also a 4 bp tandem repeat that appears at semi‐regular intervals throughout all the chromosomes (Figure 4G). Future studies should explore if the appearance and disappearance of these repeats mark any important boundaries in the genome.
In some chromosomes, we also observe imperfect repeats that occur across the entire chromosome and appear as horizontal blurs in the Fourier heatmaps. Often these blurs seem to occur at approximately the same repeat lengths as the centromeric repeat (e.g., ~135 bp in Anolis sagrei , ~170 bp in H. sapiens ); with a very clear example in Secale cereale (rye) Chr 8 (Figures S14 and S15). These transposon‐like repeats may cluster to form tandem repeat arrays over time (Sharma et al. 2013; Vondrak et al. 2020). Alternatively, this pattern may be caused by many mutations, insertions and deletions introduced to the tandem repeats over time making them appear as retrotransposon‐like repeats.
3.3.5. Neocentromeres, Variable Number Tandem Repeats (VNTRs) and Gene Copy Number Variation
RepeatOBserver can identify a wide range of perfect and imperfect repeating patterns in DNA sequences throughout the chromosome. While the dominant and often the clearest tandem repeats in a chromosome are often found at the centromere, it is possible to visualise other repetitive sections of the genome including subtelomeric repeats, neocentromeres, variable number tandem repeats, variation in short tandem repeats and regions containing gene copy number variation. To explore the visualisation of these non‐centromeric repeats, we chose to look at H. sapiens chromosome 8 (Figure 6), which contains a variable number tandem repeat (VNTR) that has been labelled as a neocentromere (85–87 Mbp) as well as three regions (7–7.6, 11.5–12.2 and 12.2–12.8 Mbp) that contain copy number variation for beta‐defensin (a gene involved in flu immune response) (Logsdon et al. 2021). Repetitive sequences at all these locations were visible in the Fourier spectra of this chromosome and appeared as weaker bands than the ‘pure’ tandem repeats seen in many centromeres but not as blurred as retrotransposon clusters.
FIGURE 6.

(A) Homo sapiens (human) chromosome 8 heat map showing the relative abundance of each repeat length at each chromosome position. The centromere appears as the dominant tandem repeat at ~45 Mbp. A variable number tandem repeat that has been observed to act as a neocentromere appears at 86 Mbp. Gene copy variation of beta‐defensin appears from 7 to 13 Mbp. (B) This region (7–13 Mbp) contains an inversion that we can visualise in the DNA walks of these sequences. Tandem repeats start at the bottom right and move to the top left in the region from 7 to 7.6 Mbp and are reversed in the following region (11.5 Mbp) starting in the top left and moving to the bottom right. DNA walks are coloured based on the rainbow; starting at red going to violet and changing colour to help visualise directionality.
3.3.6. Inversions Bounded by Tandem Repeats
The gene copy variation regions in the H. sapiens chromosome 8 described above are known to be split by an inversion from position 7.5–11.6 Mbp (Logsdon et al. 2021). This inversion can be visualised in the DNA walks of these sequences (Figure 6B,C) since the tandem repeats in these walks are inverted relative to each other on either side of the inversion. Similarly, in H. annuus there is a known inversion on chromosome 5 that differs between two cultivars (HA412 and HA89) for which genome assemblies are available (Huang et al. 2022). The repeat spectra in HA89 show the same repeat on both sides of the inversion while in HA412 the repeats are not split by the inversion (Figure 7A,B). When these split repeats were plotted in a 2D DNA walk, it was apparent that the first set of repeats in HA89 were in the same direction as the repeats in HA412. However, the second set of repeats in HA89 were inverted relative to those in HA412 and the first set of repeats in HA89 (like the inverted repeats in H. sapiens chromosome 8, Figure 6), supportive of the original repeat having been split in two by the inversion (Figure 7C).
FIGURE 7.

Inversion between sunflower ( H. annuus ) lines HA89 and HA412 on chromosome 5. (A, B) Zoomed in Fourier spectra to look closely at tandem repeats found at 35‐40 Mbp in both sunflower haplotypes (HA412 and HA89). There is a known inversion that varies between haplotypes at this location (shown by the rectangles above each heatmap). Repeats are split at the inversion boundaries in HA89. (C) Repeats shown in the DNA walks for HA412, HA89 part1 and then same repeats inverted for HA89 part2. DNA walks are coloured based on the rainbow; starting at red going to violet and changing colour to help visualise directionality.
3.4. Main Features to Look for in RepeatOBserver Output
When looking at a newly assembled chromosome using RepeatOBserver, it is helpful to initially confirm that the 6–7 bp telomerase repeats are found in at least some chromosome ends. Interstitial telomere repeats may be identified in the center of chromosomes in both plants and animals (Lin and Yan 2008; Maravilla et al. 2021) and may be associated with chromosomal rearrangements (Rosas Bringas et al. 2024). The next step is to identify the centromere type for the species and chromosomes one is working with. In the heat map plots, there will often be a dark vertical bar with bright clear horizontal bars, which is indicative of a metacentric satellite centromere (example in Figure 2D). The minima of the Shannon repeat diversity across the chromosome will best predict a putative centromere in these types of chromosomes (Figure 3A–C). The top bar will usually be the fundamental or true repeat length and the bars below it are the higher frequency harmonics. To confirm the fundamental, one should check the next harmonic below the fundamental. This should be half of the fundamental repeat length. If not, it can be worth running the Fourier transform with a larger window size (e.g., 20 kbp) to check for larger repeats. The fundamental repeat length can also be calculated from the harmonics, if needed. Specific tandem repeats that show up clearly in the heatmaps can be investigated further using a script available on the RepeatOBserver GitHub. The chromosome position and fundamental repeat length are input to the script and the DNA sequence of the longest tandem repeat (up to 600 bp long) will be returned. This repeat can then be found across all chromosomes using BLAST or other sequence mapping software. Blurred bright bars (e.g., Figures 3D and 4E–G, Figure S12) in the heat maps indicate a potentially retrotransposon‐dense region and a metacentric retrotransposon centromere. In the chromosomes we explored, we found that there is only one such region in most chromosomes, and if it exists this is the centromere, even if another region contains more clear tandem repeats (Figure 5F). Finally, polycentric or holocentric chromosomes will appear as multiple dark vertical bars with similar repeat structure throughout the chromosome (Kuo et al. 2022; Figure 8C,D).
FIGURE 8.

Example repeat spectra showing clear tandem repeats, split repeat regions, different fundamental repeat lengths, potential dense regions of retrotransposons and potential neocentromeres from (A) Brassica rapa (turnip) Chr 9 showing clear tandem repeats in centromeres and blurred retrotransposons, which resemble the gene copy number variation and neocentromere in human Chr 8 (B) B. rapa Chr 1 showing split repeats. (C) holocentric chromosomes in Luzula sylvatica Chr 4 and (D) Chionographis japonica Chr 4 showing the spread out but still clearly visible tandem repeats (Kuo et al. 2022).
Not only can RepeatOBserver work well for exploring reference genome assemblies, but it also provides a simple method to visualise repeat patterns in many haplotypes within a species. This allows users to spot mis‐assemblies and quickly determine changes in repeat composition and position between haplotypes. Other features such as gene copy variation and some neocentromeres can appear as anything from perfect to semi‐blurred fundamental and harmonic repeat lengths in the heat maps. Matching but separated tandem repeats can potentially indicate an inverted region (Figures 7B HA89 and 8B B. rapa ), especially if the split repeats are inverted relative to each other in a 2D DNA walk (Figures 6B and 7C). Using the Shannon diversity index to measure repeat diversity across each chromosome, it is possible to not only identify putative centromere positions, but also visualise other changes in sequence repeat diversity that could be caused by small repeat clusters (e.g., trinucleotide repeats) to much larger tandem repeat clusters (occasionally associated with gene copy number variation and inversion boundaries). The amount of imperfection in repeats due to insertions and deletions can be determined for a specific repeat based on the difference in repeat abundance between the fundamental repeat length and the next closest repeat lengths (e.g., for a 50 bp repeat the abundance of 50 bp compared to 51 and 49 bp gives a relative imperfection amount). For base pair changes in the sequence itself, these can be spotted in the similar ‘blur’ of the harmonics but are harder to interpret without finding the specific repeating sequence in the fasta file. Using the Shannon diversity index, we were also able to quickly spot small changes in repeats between varieties of the same species (e.g., a triplet expansion just ~1000 bp long in one line (Bur‐0) of A. thaliana (Lian et al. 2024)).
3.5. What Evolutionary Forces Could Be Driving the Lack of Repeat Diversity at Centromeres?
Using RepeatOBserver we can visually observe the different kinds of repeats and repeat patterns characteristic of centromeres that have been reported in many previous studies (e.g., Melters et al. 2012; Logsdon et al. 2024). Haig (2022) states that centromeres must be ‘hospitable’ environments for repeats. Using 159 chromosomes across 12 species with experimentally‐verified centromere locations, we observe that for monocentric tandem repeat centromeres the diversity of repeat types is lowest at the centromere implying that one or a few pure repeats typically play an important role in centromere structure. We can also see, as noted before (Shepelev et al. 2009), that the clearest repeats are often in the middle of the centromeric region and sequences that are less conserved ‘blur’ as you go farther from this central point (Figure S14). We also observe that most species have either centromeres composed of tandem repeats (usually a similar repeat length across many, if not all, chromosomes, Figure S16) or a cluster of retrotransposons (Figure 3D, Figure S12) and relatively few species have a mixture of centromere types (e.g., Z. mays , Figure S11).
Which evolutionary forces are driving the lack of repeat diversity at centromeres within chromosomes and species while allowing for rapid evolution of centromeres between species? Many past studies have looked at the evolution of repetitive elements in centromeres (reviewed in Hartley and O'Neill 2019). It was initially suggested that unequal crossing over generated the varying abundances of repeats around centromeres (Smith 1976). However, replication slippage and unequal crossing over are unlikely to be the driving factor since recombination is extremely limited around centromeres (Walsh 1987). Gene conversion (Shi et al. 2010) and interchromosomal and mitotic recombination (Jaco et al. 2008) have been suggested as factors that may be contributing to the rapidly varying copy numbers and similarities seen between some centromeres in the same species. A lack of methylation causing a change in mitotic recombination resulted in a decrease in the repeat copy numbers in mice (Jaco et al. 2008). It is still unclear why centromeric repeats match between some chromosomes and not others within a given species.
There is also uncertainty regarding the origin of centromeric repeats. One of the older hypotheses is the library hypothesis (Salser et al. 1976). It suggests that every genome contains a wide variety of repetitive sequences. If one of these sequences ends up in the centromere, due to a chromosomal rearrangement or some other shift, this repeat will increase in abundance due to the processes described above (Hartley and O'Neill 2019). Another hypothesis is that centromeric repeats derive from subtelomeric repeats (Villasante et al. 2007). Although this may be the case in some species (especially those starting with telocentric centromeres), many studies have reported very fast rates of evolution in centromeres that do not resemble the subtelomeric repeats in that species. The centromere drive hypothesis (Malik 2009) suggests that some repeat sequences provide a transmission advantage. In response to this, centromeric proteins (e.g., CEN‐A) evolve rapidly to re‐balance meiosis. Recently formed centromeres do not always have clear, large tandem repeat arrays suggesting these repeats form later, once the new centromere has reached a certain stability in the population (Hartley and O'Neill 2019). Throughout all the genomes that we explored (including recently formed centromeres in Drosophila), we were mostly able to detect the centromere as the genomic location with the lowest diversity of repeats (not necessarily implying a single dominant repeat length but implying a lower genome complexity than found in the rest of the genome). In some chromosomes in both Triticum and Z. mays , there is a high abundance of retrotransposons but no clear tandem repeats making up the centromere (Ahmed et al. 2023). Interestingly, in some species (e.g., S. cereale (rye) Chr 8, Figure S14) we can see that the length of the clear tandem repeat arrays closely matches that of more blurred repeats found elsewhere in the chromosomes. Satellite centromeric repeats have been proposed to have originated from pre‐existing retrotransposon clusters (Sharma et al. 2013; Vondrak et al. 2020). Given the many types of centromeric repeats and their structure, a pluralistic explanation of the origin and evolution of centromeres is required.
4. Conclusions
With the recent advances in long read sequencing technology, genome assemblies more frequently include long stretches of tandem repeats, creating an opportunity to visualise and compare chromosome‐scale tandem repeat patterns between species and individuals. Across plants and animals there is enormous variation in tandem repeat structure. RepeatOBserver plots allow for easy visualisation of fundamental tandem repeat lengths, repeat blur/clarity and repeat positions throughout the genome without the user needing to choose any specific input parameters. For chromosomes with a single satellite centromere, the centromere appears to consistently be found where all but a few repeat lengths have minimal abundance due to the dominance of a particular repeat and its harmonics. Detection of centromeres composed of retrotransposons show up as the maximum repeat abundance across the chromosome. Clear tandem repeat patterns are also associated with subtelomeric repeats, gene copy variation, neocentromeres and potential inversion boundaries. These genomic features appear clearly in the repeat spectra plots highlighting the potential for this tool to visualise and search for these features in newly assembled genomes of non‐model species. Even for top model species like A. thaliana , current reference assemblies contain large stretches of tandem repeats that were missing from genomes assembled only a few years earlier, making inferences from these patterns much easier now.
RepeatOBserver is a simple and accessible programto detect tandem repeat patterns and putative centromeric regions in a chromosome‐scale genome assembly. It requires no input parameters or prior knowledge of the repeats and can also detect imperfect repeats. Although it is ideally run on a server (e.g., ~20 min on 15 CPU per 100 Mbps or less of chromosome), it works on a personal computer with a few hours per 100 Mbp of chromosome. This program provides a quick way to visualise whole chromosome structures, helping bioinformaticians and biologists locate putative centromeres and other repetitive regions of interest like retrotransposon clusters, changes in repeat diversity and other tandem repeat regions.
Author Contributions
C.E., R.E. and L.R. designed research. C.E. and R.E. performed research. C.E. analysed data used in the manuscript. C.E., M.T. and L.R. wrote the paper.
Conflicts of Interest
The authors declare no conflicts of interest.
Benefit‐Sharing Statement
We are committed to making our code as accessible and open source as possible. We did not generate any new genomic data for this manuscript but hope that our code will allow many people around the world without a lot of bioinformatics or lab expertise to determine centromere positions in their genomes of interest. We have supplied all code for regenerating our results on Github including code to download the same genomes directly from NCBI, https://github.com/celphin/RepeatOBserverV1.
Supporting information
Appendix S1.
Acknowledgements
This paper is written in memory of Rob Elphinstone. His endless passion and curiosity for mathematics, physics and biology inspired all this work. Thank you also to Debbie Hearn for all her help and support in writing the paper. Funding to support this paper came from the Weston Family Doctoral Award in Northern Research and the NSERC Vanier (to C.E.). We would also like to thank Korbinian Scheeberger and Qichao Lian for providing us access to the unpublished BUR0 Arabidopsis thaliana assembly to explore its triplet repeat expansions.
Handling Editor: Jason Bragg
Funding: This study was supported by Weston Family Foundation and Natural Sciences and Engineering Research Council of Canada.
Data Availability Statement
The data that support the findings of this study are available in Tables S1 and S2. These data were derived from the following genome assemblies available in the public domain on NCBI with URLs listed in the data citations below and Table S1. Known centromere positions used in this manuscript were obtained from the papers listed in Table S2 and the Literature cited below.
References
References
- Ahmed, H. I. , Heuberger M., Schoen A., et al. 2023. “Einkorn Genomics Sheds Light on History of the Oldest Domesticated Wheat.” Nature 620, no. 7975: 830–838. 10.1038/s41586-023-06389-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Amor, D. J. , and Choo K. A.. 2002. “Neocentromeres: Role in Human Disease, Evolution, and Centromere Study.” American Journal of Human Genetics 71: 695–714. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Benson, G. 1999. “Tandem Repeats Finder: A Program to Analyze DNA Sequences.” Nucleic Acids Research 27: 573–580. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Berardini, T. Z. , Reiser L., Li D., et al. 2015. “The Arabidopsis Information Resource: Making and Mining the ‘Gold Standard’ Annotated Reference Plant Genome.” Genesis 53: 474–485. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Berger, J. A. , Mitra S. K., Carli M., and Neri A.. 2004. “Visualization and Analysis of DNA Sequences Using DNA Walks.” Journal of the Franklin Institute 341: 37–53. [Google Scholar]
- Bernardi, G. 2021. “The ‘Genomic Code’: DNA Pervasively Moulds Chromatin Structures Leaving No Room for ‘Junk’.” Life (Basel) 11: 342. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bracewell, R. , Chatla K., Nalley M. J., and Bachtrog D.. 2019. “Dynamic Turnover of Centromeres Drives Karyotype Evolution in Drosophila .” eLife 8: e49002. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Brodzik, A. K. 2007. “Quaternionic Periodicity Transform: An Algebraic Solution to the Tandem Repeat Detection Problem.” Bioinformatics 23: 694–700. [DOI] [PubMed] [Google Scholar]
- Buchner, M. , and Janjarasjitt S.. 2003. “Detection and Visualization of Tandem Repeats in DNA Sequences.” IEEE Transactions on Signal Processing 51: 2280–2287. [Google Scholar]
- Bzymek, M. , and Lovett S. T.. 2001. “Instability of Repetitive DNA Sequences: The Role of Replication in Multiple Mechanisms.” Proceedings of the National Academy of Sciences of the United States of America 98: 8319–8325. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Castelo, A. T. , Martins W., and Gao G. R.. 2002. “TROLL—Tandem Repeat Occurrence Locator.” Bioinformatics 18: 634–636. [DOI] [PubMed] [Google Scholar]
- Chang, X. , He X., Li J., et al. 2024. “High‐Quality Gossypium hirsutum and Gossypium barbadense Genome Assemblies Reveal the Landscape and Evolution of Centromeres.” Plant Communications 5, no. 2: 100722. 10.1016/j.xplc.2023.100722. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Cheng, C. Y. , Krishnakumar V., Chan A. P., Thibaud‐Nissen F., Schobel S., and Town C. D.. 2017. “Araport11: A Complete Reannotation of the Arabidopsis thaliana Reference Genome.” Plant Journal 89: 789–804. [DOI] [PubMed] [Google Scholar]
- Costantini, M. , Clay O., Auletta F., and Bernardi G.. 2006. “An Isochore Map of Human Chromosomes.” Genome Research 16: 536–541. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Dawe, R. K. , Gent J. I., Zeng Y., et al. 2023. “Synthetic Maize Centromeres Transmit Chromosomes Across Generations.” Nature Plants 9: 433–441. [DOI] [PubMed] [Google Scholar]
- “Diversity Function—RDocumentation.”.
- Dodin, G. , Vandergheynst P., Levoir P., Cordier C., and Marcourt L.. 2000. “Fourier and Wavelet Transform Analysis, a Tool for Visualizing Regular Patterns in DNA Sequences.” Journal of Theoretical Biology 206: 323–326. [DOI] [PubMed] [Google Scholar]
- Dvorkina, T. , Kunyavskaya O., Bzikadze A. V., Alexandrov I., and Pevzner P. A.. 2021. “CentromereArchitect: Inference and Analysis of the Architecture of Centromeres.” Bioinformatics 37: i196–i204. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Elloumi, A. , Messaoudi I., Lachiri Z., and Ellouze N.. 2012. “Spectral Analysis of Global Behaviour of C. elegans Chromosomes.” In Fourier Transform Applications, edited by Salih S.. InTech. 10.5772/36493. [DOI] [Google Scholar]
- Eskesen, S. T. , Eskesen F. N., Kinghorn B., and Ruvinsky A.. 2004. “Periodicity of DNA in Exons.” BMC Molecular Biology 5: 12. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Fishman, L. , and Saunders A.. 2008. “Centromere‐Associated Female Meiotic Drive Entails Male Fitness Costs in Monkeyflowers.” Science 322: 1559–1562. [DOI] [PubMed] [Google Scholar]
- Flores, M. , Morales L., Gonzaga‐Jauregui C., et al. 2007. “Recurrent DNA Inversion Rearrangements in the Human Genome.” Proceedings of the National Academy of Sciences of the United States of America 104: 6099–6106. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Flynn, J. M. , Hubley R., Goubert C., et al. 2020. “RepeatModeler2 for Automated Genomic Discovery of Transposable Element Families.” Proceedings of the National Academy of Sciences of the United States of America 117: 9451–9457. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Gao, S. , Yang X., Guo H., Zhao X., Wang B., and Ye K.. 2023. “HiCAT: A Tool for Automatic Annotation of Centromere Structure.” Genome Biology 24: 58. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Gao, Y. , Liu B., Wang Y., and Xing Y.. 2019. “TideHunter: Efficient and Sensitive Tandem Repeat Detection From Noisy Long‐Reads Using Seed‐and‐Chain.” Bioinformatics 35: i200–i207. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Gent, J. I. , Dong Y., Jiang J., and Dawe R. K.. 2012. “Strong Epigenetic Similarity Between Maize Centromeric and Pericentromeric Regions at the Level of Small RNAs, DNA Methylation and H3 Chromatin Modifications.” Nucleic Acids Research 40: 1550–1560. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Gent, J. I. , Wang K., Jiang J., and Dawe R. K.. 2015. “Stable Patterns of CENH3 Occupancy Through Maize Lineages Containing Genetically Similar Centromeres.” Genetics 200: 1105–1116. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Gong, Y. , Li Y., Liu X., Ma Y., and Jiang L.. 2023. “A Review of the Pangenome: How It Affects Our Understanding of Genomic Variation, Selection and Breeding in Domestic Animals?” Journal of Animal Science and Biotechnology 14: 73. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Good, I. J. 1953. “The Population Frequencies of Species and the Estimation of Population Parameters.” Biometrika 40: 237–264. [Google Scholar]
- Greider, C. W. , and Blackburn E. H.. 1985. “Identification of a Specific Telomere Terminal Transferase Activity in Tetrahymena Extracts.” Cell 43: 405–413. [DOI] [PubMed] [Google Scholar]
- Haig, D. 2022. “Paradox Lost: Concerted Evolution and Centromeric Instability.” BioEssays 44: 2200023. [DOI] [PubMed] [Google Scholar]
- Haimovich, A. D. , Byrne B., Ramaswamy R., and Welsh W. J.. 2006. “Wavelet Analysis of DNA Walks.” Journal of Computational Biology: A Journal of Computational Molecular Cell Biology 13: 1289–1298. [DOI] [PubMed] [Google Scholar]
- Hallast, P. , Ebert P., Loftus M., et al. 2023. “Assembly of 43 Human Y Chromosomes Reveals Extensive Complexity and Variation.” Nature 621: 355–364. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hartley, G. , and O'Neill R. J.. 2019. “Centromere Repeats: Hidden Gems of the Genome.” Genes (Basel) 10: 223. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Henikoff, S. , Ahmad K., and Malik H. S.. 2001. “The Centromere Paradox: Stable Inheritance With Rapidly Evolving DNA.” Science 293: 1098–1102. [DOI] [PubMed] [Google Scholar]
- Horton, C. A. , Alexandari A. M., Hayes M. G., et al. 2023. “Short Tandem Repeats Bind Transcription Factors to Tune Eukaryotic Gene Expression.” Science 381: eadd1250. [DOI] [PubMed] [Google Scholar]
- Huang, K. , Ostevik K. L., Elphinstone C., et al. 2022. “Mutation Load in Sunflower Inversions Is Negatively Correlated With Inversion Heterozygosity.” Molecular Biology and Evolution 39, no. 5: msac101. 10.1093/molbev/msac101. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hughes, E. D. , Qu Y. Y., Genik S. J., et al. 2007. “Genetic Variation in C57BL/6 ES Cell Lines and Genetic Instability in the Bruce4 C57BL/6 ES Cell Line.” Mammalian Genome 18: 549–558. [DOI] [PubMed] [Google Scholar]
- Jabbari, K. , and Bernardi G.. 2017. “An Isochore Framework Underlies Chromatin Architecture.” PLoS One 12: e0168023. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Jaco, I. , Canela A., Vera E., and Blasco M. A.. 2008. “Centromere Mitotic Recombination in Mammalian Cells.” Journal of Cell Biology 181: 885–892. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Jiao, Y. , Peluso P., Shi J., et al. 2017. “Improved Maize Reference Genome With Single‐Molecule Technologies.” Nature 546: 524–527. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Jo, S. H. , Koo D. H., Kim J. F., et al. 2009. “Evolution of Ribosomal DNA‐Derived Satellite Repeat in Tomato Genome.” BMC Plant Biology 9: 42. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Katoh, K. , Misawa K., Kuma K. I., and Miyata T.. 2002. “MAFFT: A Novel Method for Rapid Multiple Sequence Alignment Based on Fast Fourier Transform.” Nucleic Acids Research 30: 3059–3066. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kilburn, A. E. , Shea M. J., Sargent R. G., and Wilson J. H.. 2001. “Insertion of a Telomere Repeat Sequence Into a Mammalian Gene Causes Chromosome Instability.” Molecular and Cellular Biology 21: 126–135. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kolpakov, R. , Bana G., and Kucherov G.. 2003. “Mreps: Efficient and Flexible Detection of Tandem Repeats in DNA.” Nucleic Acids Research 31: 3672–3678. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kunyavskaya, O. , Dvorkina T., Bzikadze A. V., Alexandrov I. A., and Pevzner P. A.. 2021. “HORmon: Automated Annotation of Human Centromeres.” bioRxiv. 10.1101/2021.10.12.464028. [DOI] [PMC free article] [PubMed]
- Kuo, Y. T. , Câmara A. S., Schubert V., et al. 2022. “Plasticity in Centromere Organization: Holocentromeres Can Consist of Merely a Few Megabase‐Sized Satellite Arrays.” bioRxiv. 10.1101/2022.11.23.516916. [DOI] [PMC free article] [PubMed]
- Kursel, L. E. , and Malik H. S.. 2018. “The Cellular Mechanisms and Consequences of Centromere Drive.” Current Opinion in Cell Biology 52: 58–65. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kwapisz, M. , and Morillon A.. 2020. “Subtelomeric Transcription and Its Regulation.” Journal of Molecular Biology 432: 4199–4219. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lampson, M. A. , and Black B. E.. 2017. “Cellular and Molecular Mechanisms of Centromere Drive.” In Cold Spring Harbor Symposia on Quantitative Biology, vol. 82, 249–257. Cold Spring Harbor Laboratory Press. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Li, N. , He Q., Wang J., et al. 2023. “Super‐Pangenome Analyses Highlight Genomic Diversity and Structural Variation Across Wild and Cultivated Tomato Species.” Nature Genetics 55: 852–860. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Li, W. , Liu J., Zhang H., et al. 2022. “Plant Pan‐Genomics: Recent Advances, New Challenges, and Roads Ahead.” Journal of Genetics and Genomics 49: 833–846. [DOI] [PubMed] [Google Scholar]
- Lian, Q. , Huettel B., Walkemeier B., et al. 2024. “A Pan‐Genome of 69 Arabidopsis thaliana Accessions Reveals a Conserved Genome Structure Throughout the Global Species Range.” Nature Genetics 56: 982–991. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Liao, W. W. , Asri M., Ebler J., et al. 2023. “A Draft Human Pangenome Reference.” Nature 617: 312–324. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lim, K. B. , De Jong H., Yang T. J., et al. 2005. “Characterization of rDNAs and Tandem Repeats in the Heterochromatin of Brassica rapa .” Molecules and Cells 19: 436–444. [PubMed] [Google Scholar]
- Lin, K. W. , and Yan J.. 2008. “Endings in the Middle: Current Knowledge of Interstitial Telomeric Sequences.” Mutation Research 658: 95–110. [DOI] [PubMed] [Google Scholar]
- Lin, Y. , Ye C., Li X., et al. 2023. “quarTeT: A Telomere‐to‐Telomere Toolkit for Gap‐Free Genome Assembly and Centromeric Repeat Identification.” Horticulture Research 10: uhad127. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Liu, Y. , Yi C., Fan C., et al. 2023. “Pan‐Centromere Reveals Widespread Centromere Repositioning of Soybean Genomes.” Proceedings of the National Academy of Sciences of the United States of America 120: e2310177120. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Logsdon, G. A. , Rozanski A. N., Ryabov F., et al. 2024. “The Variation and Evolution of Complete Human Centromeres.” Nature 629: 136–145. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Logsdon, G. A. , Vollger M. R., Hsieh P., et al. 2021. “The Structure, Function and Evolution of a Complete Human Chromosome 8.” Nature 593: 101–107. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lynch, D. B. , Logue M. E., Butler G., and Wolfe K. H.. 2010. “Chromosomal G+ C Content Evolution in Yeasts: Systematic Interspecies Differences, and GC‐Poor Troughs at Centromeres.” Genome Biology and Evolution 2: 572–583. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ma, B. , Wang H., Liu J., et al. 2023. “The Gap‐Free Genome of Mulberry Elucidates the Architecture and Evolution of Polycentric Chromosomes.” Horticulture Research 10: uhad111. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Machado, J. A. T. 2013. “Fractional‐Order Fourier Analysis of the DNA.” IFAC Proceedings Volumes 46: 248–253. [Google Scholar]
- Majewski, J. , and Ott J.. 2000. “GT Repeats Are Associated With Recombination on Human Chromosome 22.” Genome Research 10: 1108–1114. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Malik, H. S. 2009. “The Centromere‐Drive Hypothesis: A Simple Basis for Centromere Complexity.” Progress in Molecular and Subcellular Biology 48: 33–52. [DOI] [PubMed] [Google Scholar]
- Maravilla, A. J. , Rosato M., and Rosselló J. A.. 2021. “Interstitial Telomeric‐Like Repeats (ITR) in Seed Plants as Assessed by Molecular Cytogenetic Techniques: A Review.” Plants 10: 2541. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Melters, D. P. , Bradnam K. R., Young H. A., et al. 2013. “Comparative Analysis of Tandem Repeats From Hundreds of Species Reveals Unique Insights Into Centromere Evolution.” Genome Biology 14: R10. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Melters, D. P. , Paliulis L. V., Korf I. F., and Chan S. W.. 2012. “Holocentric Chromosomes: Convergent Evolution, Meiotic Adaptations, and Genomic Analysis.” Chromosome Research 20: 579–593. [DOI] [PubMed] [Google Scholar]
- Miga, K. H. 2020. “Centromere Studies in the Era of ‘Telomere‐to‐Telomere’ Genomics.” Experimental Cell Research 394: 112127. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Nagai, N. , Kuwata K., Hayashi T., Kuwata H., and Era S.. 2001. “Evolution of the Periodicity and the Self‐Similarity in DNA Sequence: A Fourier Transform Analysis.” Japanese Journal of Physiology 51: 159–168. [DOI] [PubMed] [Google Scholar]
- Nagai, M. , Kurokawa M., and Ying B.‐W.. 2020. “The Highly Conserved Chromosomal Periodicity of Transcriptomes and the Correlation of Its Amplitude with the Growth Rate in Escherichia coli .” DNA Research 27, no. 3: dsaa018. 10.1093/dnares/dsaa018. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Nagaki, K. , Tanaka K., Yamaji N., Kobayashi H., and Murata M.. 2015. “Sunflower Centromeres Consist of a Centromere‐Specific LINE and a Chromosome‐Specific Tandem Repeat.” Frontiers in Plant Science 6: 912. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Naish, M. , Alonge M., Wlodzimierz P., et al. 2021. “The Genetic and Epigenetic Landscape of the Arabidopsis Centromeres.” Science 374, no. 6569: eabi7489. 10.1126/science.abi7489. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Naish, M. , and Henderson I. R.. 2024. “The Structure, Function, and Evolution of Plant Centromeres.” Genome Research 34, no. 2: 161–178. 10.1101/gr.278409.123. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Nassar, L. R. , Barber G. P., Benet‐Pagès A., et al. 2023. “The UCSC Genome Browser Database: 2023 Update.” Nucleic Acids Research 51: D1188–D1195. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Navajas‐Pérez, R. , and Paterson A. H.. 2009. “Patterns of Tandem Repetition in Plant Whole Genome Assemblies.” Molecular Genetics and Genomics 281: 579–590. [DOI] [PubMed] [Google Scholar]
- Nawaz, Z. , Kakar K. U., Saand M. A., and Shu Q. Y.. 2014. “Cyclic Nucleotide‐Gated Ion Channel Gene Family in Rice, Identification, Characterization and Experimental Analysis of Expression Response to Plant Hormones, Biotic and Abiotic Stresses.” BMC Genomics 15: 853. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Nietzel, A. , Rocchi M., Starke H., et al. 2001. “A New Multicolor‐FISH Approach for the Characterization of Marker Chromosomes: Centromere‐Specific Multicolor‐FISH (cenM‐FISH).” Human Genetics 108: 199–204. [DOI] [PubMed] [Google Scholar]
- Novák, P. , Neumann P., Pech J., Steinhaisl J., and Macas J.. 2013. “RepeatExplorer: A Galaxy‐Based Web Server for Genome‐Wide Characterization of Eukaryotic Repetitive Elements From Next‐Generation Sequence Reads.” Bioinformatics 29: 792–793. [DOI] [PubMed] [Google Scholar]
- Oksanen, J. 2022. “Vegan: Ecological Diversity.”
- Pâques, F. , Leung W. Y., and Haber J. E.. 1998. “Expansions and Contractions in a Tandem Repeat Induced by Double‐Strand Break Repair.” Molecular and Cellular Biology 18: 2045–2054. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Paulson, H. 2018. “Chapter 9—Repeat Expansion Diseases.” In Handbook of Clinical Neurology, Neurogenetics, Part I, edited by Geschwind D. H., Paulson H. L., and Klein C., 105–123. Elsevier. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Pellegrini, M. , Renda M. E., and Vecchio A.. 2010. “TRStalker: An Efficient Heuristic for Finding Fuzzy Tandem Repeats.” Bioinformatics 26: i358–i366. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Qi, Y. , Jin N., and Ai D.. 2012. “Wavelet Analysis of DNA Walks on the Human and Chimpanzee MAGE/CSAG‐Palindromes Genomics.” Proteomics & Bioinformatics 10: 230–236. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Read, L. R. , Raynard S. J., Rukść A., and Baker M. D.. 2004. “Gene Repeat Expansion and Contraction by Spontaneous Intrachromosomal Homologous Recombination in Mammalian Cells.” Nucleic Acids Research 32: 1184–1196. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Rhie, A. , Nurk S., Cechova M., et al. 2023. “The Complete Sequence of a Human Y Chromosome.” Nature 621: 344–354. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Rice, P. , Longden I., and Bleasby A.. 2000. “EMBOSS: the European Molecular Biology Open Software Suite.” Trends in Genetics 16: 276–277. [DOI] [PubMed] [Google Scholar]
- Richards, E. J. , and Ausubel F. M.. 1988. “Isolation of a Higher Eukaryotic Telomere From Arabidopsis thaliana .” Cell 53: 127–136. [DOI] [PubMed] [Google Scholar]
- Rosas Bringas, F. R. , Yin Z., Yao Y., Boudeman J., Ollivaud S., and Chang M.. 2024. “Interstitial Telomeric Sequences Promote Gross Chromosomal Rearrangement via Multiple Mechanisms.” Proceedings of the National Academy of Sciences of the United States of America 121, no. 49: e2407314121. 10.1073/pnas.2407314121. [DOI] [PMC free article] [PubMed] [Google Scholar]
- “RPubs—Introduction Fast Fourier Transform in R.”.
- Ruperao, P. , Thirunavukkarasu N., Gandham P., et al. 2021. “Sorghum Pan‐Genome Explores the Functional Utility for Genomic‐Assisted Breeding to Accelerate the Genetic Gain.” Frontiers in Plant Science 12: 666342. 10.3389/fpls.2021.666342. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Saha, A. K. 2019. “The Genetics and Epigenetics of Centromeres in Cancer.”
- Salser, W. , Bowen S., Browne D., et al. 1976. “Investigation of the Organization of Mammalian Chromosomes at the DNA Sequence Level.” Federation Proceedings 35: 23–35. [PubMed] [Google Scholar]
- Shannon, C. E. 1948. “A Mathematical Theory of Communication.” Bell System Technical Journal 27: 379–423. [Google Scholar]
- Sharma, A. , Wolfgruber T. K., and Presting G. G.. 2013. “Tandem Repeats Derived From Centromeric Retrotransposons.” BMC Genomics 14: 1–11. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Sharma, D. , Issac B., Raghava G. P., and Ramaswamy R.. 2004. “Spectral Repeat Finder (SRF): Identification of Repetitive Sequences Using Fourier Transformation.” Bioinformatics 20: 1405–1412. [DOI] [PubMed] [Google Scholar]
- Shepelev, V. A. , Alexandrov A. A., Yurov Y. B., and Alexandrov I. A.. 2009. “The Evolutionary Origin of Man Can be Traced in the Layers of Defunct Ancestral Alpha Satellites Flanking the Active Centromeres of Human Chromosomes.” PLoS Genetics 5: e1000641. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Shi, J. , Wolf S. E., Burke J. M., Presting G. G., Ross‐Ibarra J., and Dawe R. K.. 2010. “Widespread Gene Conversion in Centromere Cores.” PLoS Biology 8: e1000327. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Smith, G. P. 1976. “Evolution of Repeated DNA Sequences by Unequal Crossover.” Science 191: 528–535. [DOI] [PubMed] [Google Scholar]
- “smooth.spline function—RDocumentation.”.
- Song, J. M. , Xie W. Z., Wang S., et al. 2021. “Two Gap‐Free Reference Genomes and a Global View of the Centromere Architecture in Rice.” Molecular Plant 14: 1757–1767. [DOI] [PubMed] [Google Scholar]
- Scott, K. C. , and Sullivan, B. A. 2014. “Neocentromeres: A Place for Everything and Everything in Its Place.” Trends in Genetics 30: 66–74. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Sureshkumar, S. , Todesco M., Schneeberger K., Harilal R., Balasubramanian S., and Weigel D.. 2009. “A Genetic Defect Caused by a Triplet Repeat Expansion in Arabidopsis thaliana .” Science 323: 1060–1063. [DOI] [PubMed] [Google Scholar]
- Talbert, P. , and Henikoff S.. 2022. “Centromere Drive: Chromatin Conflict in Meiosis.” Current Opinion in Genetics & Development 77: 102005. [DOI] [PubMed] [Google Scholar]
- Talbert, P. B. , and Henikoff S.. 2020. “What Makes a Centromere?” Experimental Cell Research 389: 111895. [DOI] [PubMed] [Google Scholar]
- Tarailo‐Graovac, M. , and Chen N.. 2009. “Using RepeatMasker to Identify Repetitive Elements in Genomic Sequences.” Current Protocols in Bioinformatics 25: 4–10. [DOI] [PubMed] [Google Scholar]
- Tiley, G. P. , and Burleigh J. G.. 2015. “The Relationship of Recombination Rate, Genome Structure, and Patterns of Molecular Evolution Across Angiosperms.” BMC Evolutionary Biology 15: 194. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Varoquaux, N. , Liachko I., Ay F., et al. 2015. “Accurate Identification of Centromere Locations in Yeast Genomes Using Hi‐C.” Nucleic Acids Research 43: 5331–5339. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Villasante, A. , Abad J. P., and Méndez‐Lago M.. 2007. “Centromeres Were Derived From Telomeres During the Evolution of the Eukaryotic Chromosome.” Proceedings of the National Academy of Sciences of the United States of America 104: 10542–10547. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Vollger, M. R. , Kerpedjiev P., Phillippy A. M., and Eichler E. E.. 2022. “StainedGlass: Interactive Visualization of Massive Tandem Repeat Structures With Identity Heatmaps.” Bioinformatics 38: 2049–2051. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Vondrak, T. , Ávila Robledillo L., Novák P., Koblížková A., Neumann P., and Macas J.. 2020. “Characterization of Repeat Arrays in Ultra‐Long Nanopore Reads Reveals Frequent Origin of Satellite DNA From Retrotransposon‐Derived Tandem Repeats.” Plant Journal 101: 484–500. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Walsh, J. B. 1987. “Persistence of Tandem Arrays: Implications for Satellite and Simple‐Sequence DNAs.” Genetics 115: 553–567. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wang, K. , Hu H., Tian Y., et al. 2021. “The Chicken Pan‐Genome Reveals Gene Content Variation and a Promoter Region Deletion in IGF2BP1 Affecting Body Size.” Molecular Biology and Evolution 38: 5066–5081. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wei, K. H. , Grenier J. K., Barbash D. A., and Clark A. G.. 2014. “Correlated Variation and Population Differentiation in Satellite DNA Abundance Among Lines of Drosophila melanogaster .” Proceedings of the National Academy of Sciences of the United States of America 111: 18793–18798. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Weighill, D. , Macaya‐Sanz D., DiFazio S. P., et al. 2019. “Wavelet‐Based Genomic Signal Processing for Centromere Identification and Hypothesis Generation.” Frontiers in Genetics 10: 487. 10.3389/fgene.2019.00487. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wexler, Y. , Yakhini Z., Kashi Y., and Geiger D.. 2005. “Finding Approximate Tandem Repeats in Genomic Sequences.” Journal of Computational Biology: A Journal of Computational Molecular Cell Biology 12: 928–942. [DOI] [PubMed] [Google Scholar]
- Wlodzimierz, P. , Hong M., and Henderson I. R.. 2023a. “TRASH: Tandem Repeat Annotation and Structural Hierarchy.” Bioinformatics 39: btad308. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wlodzimierz, P. , Rabanal F. A., Burns R., et al. 2023b. “Cycles of Satellite and Transposon Evolution in Arabidopsis Centromeres.” Nature 618: 557–565. [DOI] [PubMed] [Google Scholar]
- Wolfgruber, T. K. , Sharma A., Schneider K. L., et al. 2009. “Maize Centromere Structure and Evolution: Sequence Analysis of Centromeres 2 and 5 Reveals Dynamic Loci Shaped Primarily by Retrotransposons.” PLoS Genetics 5: e1000743. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Woody, J. L. , Beavis W., and Shoemaker R. C.. 2012. “Large Homogeneous Genome Regions (Isochores) in Soybean [Glycine max (L.) Merr.].” Frontiers in Genetics 3: 98. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Xu, D. , Yang J., Wen H., et al. 2024. “CentIER: Accurate Centromere Identification for Plant Genome.” Plant Communications 5: 101046. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Yadav, Y. , Sharma S. N., and Shakya D. K.. 2022. “Detection of Tandem Repeats in DNA Sequences Using Short‐Time Ramanujan Fourier Transform.” IEEE/ACM Transactions on Computational Biology and Bioinformatics 19: 1583–1591. [DOI] [PubMed] [Google Scholar]
- Zhang, L. , Liang J., Chen H., Zhang Z., Wu J., and Wang X.. 2023. “A Near‐Complete Genome Assembly of Brassica rapa Provides New Insights Into the Evolution of Centromeres.” Plant Biotechnology Journal 21: 1022–1032. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zhang, R. , and Zhang C.‐T.. 2004. “Isochore Structures in the Genome of the Plant Arabidopsis thaliana .” Journal of Molecular Evolution 59: 227–238. [DOI] [PubMed] [Google Scholar]
- Zhang, T. , Talbert P. B., Zhang W., et al. 2013. “The CentO Satellite Confers Translational and Rotational Phasing on cenH3 Nucleosomes in Rice Centromeres.” Proceedings of the National Academy of Sciences of the United States of America 110: E4875–E4883. [DOI] [PMC free article] [PubMed] [Google Scholar]
Data Citations
- Ahmed, H. I. , Heuberger M., Schoen A., et al. 2023. “Einkorn Genomics Sheds Light on History of the Oldest Domesticated Wheat.” Nature 620, no. 7975: 830–838. 10.1038/s41586-023-06389-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Chang, X. 2023. “TM‐1 HAU v2 and 3‐79 HAU v3 genomes.” [Dataset]. Figshare. 10.6084/m9.figshare.22682833.v2. [DOI]
- Geneva, A. J. , Park S., Bock D., et al. 2022. “Anolis sagrei Genome Assembly and Annotation [Dataset].” Harvard Dataverse. 10.7910/DVN/TTKBFU. [DOI]
- Huang, K. , Jahani M., Gouzy J., et al. 2023. “The Genomics of Linkage Drag in Inbred Lines of Sunflower.” Proceedings of the National Academy of Sciences of the United States of America 120, no. 14: e2205783119. 10.1073/pnas.2205783119. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ma, B. , Wang H., Liu J., et al. 2023. “The Gap‐Free Genome of Mulberry Elucidates the Architecture and Evolution of Polycentric Chromosomes.” Horticulture Research 10, no. 7: uhad111. 10.1093/hr/uhad111. [DOI] [PMC free article] [PubMed] [Google Scholar]
- MorusDB . 2023. “M. notabilis Mnot‐SWU.” [Dataset]. https://morus.biodb.org/downloads.
- NCBI . 2015. “Oryza sativa Japonica Group Genome Assembly IRGSP‐1.0.” [Dataset]. https://www.ncbi.nlm.nih.gov/data‐hub/assembly/GCF_001433935.1/.
- NCBI . 2018. “Arabidopsis thaliana Genome Assembly TAIR10.1.” [Dataset]. https://www.ncbi.nlm.nih.gov/data‐hub/assembly/GCF_000001735.4/.
- NCBI . 2018. “Solanum lycopersicum Genome Assembly SL3.1.” [Dataset]. https://www.ncbi.nlm.nih.gov/data‐hub/assembly/GCF_000188115.5/.
- NCBI . 2019. “Drosophila lowei Genome Assembly UCBerk_Dlow_1.0.” [Dataset]. https://www.ncbi.nlm.nih.gov/data‐hub/assembly/GCA_008121275.1/.
- NCBI . 2020. “Solanum pimpinellifolium Genome Assembly ASM1496433v1.” [Dataset]. https://www.ncbi.nlm.nih.gov/data‐hub/assembly/GCA_014964335.1/.
- NCBI . 2021. “Taxus wallichiana var. yunnanensis Genome Assembly ASM1834077v1.” [Dataset]. https://www.ncbi.nlm.nih.gov/data‐hub/assembly/GCA_018340775.1/.
- NCBI . 2021. “Ceratopteris richardii Genome Assembly C. richardii_v2.” [Dataset]. https://www.ncbi.nlm.nih.gov/data‐hub/assembly/GCA_020310875.1/.
- NCBI . 2021. “Secale cereale Genome Assembly Rye_Lo7_2018_v1p1p1.” [Dataset]. https://www.ncbi.nlm.nih.gov/data‐hub/assembly/GCA_902687465.1/.
- NCBI . 2021. “Glycine max Genome Assembly Glycine_max_v4.0.” [Dataset]. https://www.ncbi.nlm.nih.gov/data‐hub/assembly/GCF_000004515.6/.
- NCBI . 2021. “Homo sapiens Genome Assembly T2T‐CHM13v1.1—NCBI—NLM.” [Dataset]. https://www.ncbi.nlm.nih.gov/datasets/genome/GCA_009914755.3/.
- NCBI . 2022. “Zea mays Genome Assembly Zm‐Mo17‐REFERENCE‐CAU‐T2T‐Assembly.” [Dataset]. https://www.ncbi.nlm.nih.gov/data‐hub/assembly/GCA_022117705.1/.
- NCBI . 2022. “Danio rerio Genome Assembly fDanRer4.1.” [Dataset]. https://www.ncbi.nlm.nih.gov/data‐hub/assembly/GCA_944039275.1/.
- NCBI . 2022. “Luzula sylvatica Genome Assembly lpLuzSylv1.1.” [Dataset]. https://www.ncbi.nlm.nih.gov/data‐hub/assembly/GCA_946800325.1/.
- NCBI . 2023. “Rubus idaeus Genome Assembly RiMJ.” [Dataset]. https://www.ncbi.nlm.nih.gov/data‐hub/assembly/GCA_030142095.1/.
- NCBI . 2023. “Mus musculus Genome Assembly NEI_Mmus_1.0.” [Dataset]. https://www.ncbi.nlm.nih.gov/data‐hub/assembly/GCA_030265425.1/.
- NCBI . 2023. “Chionographis japonica Genome Assembly Cjaponica_Koshiki_v1.0.” [Dataset]. https://www.ncbi.nlm.nih.gov/data‐hub/assembly/GCA_947650365.1/.
- NCBI . 2023. “Vitis vinifera Genome Assembly ASM3070453v1.” [Dataset]. https://www.ncbi.nlm.nih.gov/data‐hub/assembly/GCF_030704535.1/.
- The Arabidopsis Information Resource (TAIR) . 2024. “TAIR—Arabidopsis.” [Dataset]. https://www.arabidopsis.org/download/list?dir=Sequences%2FAssemblies.
- Todesco, M. , Owens G. L., Bercovich N., et al. 2020. “Massive Haplotypes Underlie Ecotypic Differentiation in Sunflowers.” Nature 584, no. 7822: 602–607. 10.1038/s41586-020-2467-6. [DOI] [PubMed] [Google Scholar]
- Zhang, L. , Liang J., Chen H., Zhang Z., Wu J., and Wang X.. 2023. “A Near‐Complete Genome Assembly of Brassica Rapa Provides New Insights Into the Evolution of Centromeres Plant.” Biotechnology Journal 21, no. 5: 1022–1032. 10.1111/pbi.14015. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Appendix S1.
Data Availability Statement
The data that support the findings of this study are available in Tables S1 and S2. These data were derived from the following genome assemblies available in the public domain on NCBI with URLs listed in the data citations below and Table S1. Known centromere positions used in this manuscript were obtained from the papers listed in Table S2 and the Literature cited below.
