Skip to main content
Biomolecules logoLink to Biomolecules
. 2026 Aug 27;16(9):1247. doi: 10.3390/biom16091247

Active Human Transposable Elements: Long-Read Sequencing Technologies, Computational Analysis, and Implications for Human Disease

Dániel Vörösvácki 1, Nikolett Szakállas 1,2, Alexandra Kalmár 1, István Takács 1, Béla Molnár 1,*
Editor: Kouji Hirota
PMCID: PMC13604091  PMID: 42793079

Abstract

Transposable elements (TEs) account for nearly half of the human genome and shape chromatin organization, gene regulation, and genome evolution. However, their contributions to human physiology and disease remain incompletely understood. The most active elements in humans, LINE-1 (L1), Alu, and SVA, retain some copies with the ability to evade epigenetic repression and mobilize via target-primed reverse transcription (TPRT), whereas copies become inactive through various fragmentations and mutations. TE activity contributes to genomic instability and has been implicated in aging, cancer, neurological disorders, chromatin organization, and epigenetic regulation. Studying TE is challenging due to their repetitive and polymorphic nature. Recent advances in sequencing technologies and short- and long-read sequencing platforms, combined with specialized bioinformatic pipelines, currently enable more comprehensive characterization of TE insertions, deletions, expression, and epigenetic status. Computational approaches vary in sensitivity, specificity, and resource requirements, and their performance is influenced by sequencing modality, coverage, and the reference genome used. Assembly-based and read-based methods, as well as integrating methylation data or single-cell data, provide complementary insights into TE biology. This review summarizes the biology of active human TE, surveys state-of-the-art short- and long-read pipelines for TE analysis, and highlights their applications in studies of aging, cancer, and other complex diseases. We also provide practical guidance for selecting appropriate sequencing strategies and tools for TE-focused projects, and discuss emerging approaches and open questions in the field.

Keywords: transposable elements, LINE-1, Alu, SVA, long-read sequencing, Oxford Nanopore Technologies, Pacific Biosciences, bioinformatics, epigenetics, structural variation

1. Key Messages

  • Transposable elements are not merely “junk DNA”; some remain active and mobilize via target-primed reverse transcription, while others contribute to genome regulation and evolution.

  • Long-read sequencing technologies, including Oxford Nanopore Technologies and Pacific Biosciences, provide improved resolution of repetitive genomic regions and enable characterization of complex TE insertions and structural rearrangements.

  • Modern TE analysis pipelines integrate read-based, assembly-based, methylation-aware, and single-cell approaches to investigate TE insertions, expression, epigenetic regulation, and disease associations.

  • Selection of an appropriate TE analysis workflow depends on sequencing modality, computational resources, the biological question, and the balance between sensitivity and specificity.

2. Introduction

Repetitive sequences account for a substantial fraction of eukaryotic genomes [1], with transposable elements (TEs) as the primary contributors [2,3]. In humans, approximately 45% of the genome is derived from TEs, and some estimates suggest that up to two-thirds of the genome may originate from highly fragmented ancestral TE insertions whose origins are no longer readily identifiable [4].

Recent advances in high-throughput sequencing technologies, particularly long-read sequencing platforms, together with increasingly sophisticated bioinformatic pipelines, have enabled more comprehensive characterization of the TE landscape [5]. These developments provide important insights into the evolutionary history, regulation, and functional impact of TEs in humans and other organisms.

Multiple TE classes have been identified in genomes analyzed to date and can be classified based on their propagation mechanisms and sequence homology [6]. A comprehensive overview of TE classification is beyond the scope of this review, which focuses primarily on active human TEs and their analysis using long-read sequencing approaches. Such information can be found in the reviews of human TEs [7]. For researchers new to the field, inconsistent TE nomenclature and overlapping classification systems may present a substantial challenge. Community resources such as TE Hub [8] provide useful summaries of current classification frameworks and TE-related reviews. Table 1 summarizes the major retrotransposon families that remain active in the human genome.

Table 1.

Short summary of active mobile elements in the human genome. Length values are approximate and may vary between different subfamilies. Estimates for some values differ substantially between methods. For the Non-AUG translation of SVA, the evidence is limited. * The retrotransposon-competent SVA number is still not conclusively determined; the value of 1927, representing the number of full-length SVA elements detected in GRCh38 by Chong et al. [9], should be treated as an upper bound without accounting for mutations, epigenetic suppression and other possible mechanisms preventing SVA TPRT.

Name L1 Alu SVA [9]
Length (bp) 6000 280 700–4000
Copies in genome 500,000 [10] −919,967 [11] 1,100,000 [12] 3000 [13]–5100 [9]
Retrotransposition competent elements 80 [10]–114 [14] ≥852 [12] 1927 *
Autonomy Autonomous Dependent on L1 Dependent on L1
Transcribed Yes Yes Yes
Translated ORF0, ORF1, ORF2 No Non-AUG translation [15]
Required for transposition ORF1p, ORF2p ORF2p ORF1p, ORF2p

Recent telomere-to-telomere sequencing efforts and large-scale analyses of repetitive DNA suggest that most major human TE families have now been identified [16,17]. Accurate TE detection remains challenging, especially in short-read data, because the high copy number and sequence similarity of TEs complicate unambiguous mapping to their true genomic loci [18]. This review focuses on recent computational approaches and bioinformatic tools developed for TE analysis in human genomic datasets, with particular emphasis on long-read sequencing applications.

2.1. LINE-1

LINE-1 (L1) elements constitute approximately 17% of the human genome [3], and more than 30% of the genome reflects L1-driven activity through other retrotransposons that exploit the L1-encoded retrotransposition machinery [19,20,21]. L1 is the only currently active autonomous TE family in humans [19,20,21], with 500,000–900,000 L1 copies in various states of truncation and fragmentation according to different reference genomes and annotation methods. L1 mobilizes through target-primed reverse transcription (TPRT), as explained in Figure 1.

Figure 1.

Figure 1

The schematic overview of L1-mediated TPRT. (A) L1 RNA is transcribed and exported from the nucleus. (B) L1 mRNA is translated to produce ORF1p and ORF2p, which form a ribonucleoprotein (RNP) complex, which is transported back to the nucleus. (C) The endonuclease activity of ORF2p nicks a genomic target site and the resulting 3′-OH group primes reverse transcription of L1 RNA leading to the integration of a new copy. Created in BioRender. Linkner, T. (2026) https://BioRender.com/g57plm7.

Nearly all active L1 copies characterized to date belong to the L1Pa1 lineage, also referred to as L1HS (human-specific) or L1TA (transcriptionally active). A small subset of these L1 elements, often termed “hot L1s” (3–9 copies per genome), accounts for approximately 80% of retrotransposition events [20]. Loss of L1 activity arises through multiple mechanisms, including 5′ truncation frequently caused by premature termination of TPRT, point mutations [18], nested insertions [22], and epigenetic repression such as hypermethylation. TPRT events frequently generate truncated L1 copies incapable of retrotransposition. On average, two human genomes differ by around 285 L1 insertions [23].

An intact L1Pa1 element is approximately 6 kb in length, comprises a 5′ untranslated region (UTR) containing an internal RNA polymerase II promoter, followed by three open reading frames (ORFs). ORF0 (213 bp) encodes a short peptide [24], ORF1 (1122 bp) encodes a 40-kDa RNA-binding protein, and ORF2 (3852 bp) encodes a 150-kDa protein with endonuclease and reverse transcriptase activity. ORF1p and ORF2p are essential for L1 retrotransposition, whereas ORF0p enhances retrotransposition frequency. Non-autonomous elements such as Alu and SVA rely on L1-encoded ORF2p for mobilization. L1 elements end with a 3′ UTR enriched in adenine and thymine bases.

Current estimates place the L1 germline retrotransposition rate at approximately one new insertion per 20–200 births, while short-read whole-genome sequencing in three-generation pedigrees suggests a rate of one insertion per roughly 63 births [25]. Insertions often occur at motifs that deviate from the canonical TTAAAA L1 endonuclease target site by up to two mismatches.

Because L1 is subject to strong epigenetic repression, most retrotransposition events occur in the germline when the genome is hypomethylated. New insertions are frequently generated by aberrant reverse transcription, as the transcription terminates before reaching the 5′ end of the L1 mRNA. Another phenomenon that reduces intact L1 is nesting, in which a new L1 copy integrates into a pre-existing element, damaging the pre-existing element. When L1 activity bypasses host repression, somatic retrotransposition occurs [18,26,27]. Somatic insertions typically feature a truncated 5′ end, an intact 3′ poly(A) tail, and flanking target site duplications (TSDs) of approximately 4–20 bp generated by L1 endonuclease during integration. Many bioinformatic pipelines identify somatic insertions by their absence in matched normal tissue DNA [28], although adjacent tissues may already exhibit early molecular alterations prior to overt tumorigenesis, potentially confounding detection [29].

2.2. Alu

Alu elements are non-autonomous, primate-specific retrotransposons [30] and account for approximately 11% of the human genome. With over one million copies present in various states of fragmentation, Alu elements represent the most numerous TE family by copy number. Intact Alu elements are approximately 280 bp long, are transcribed by RNA polymerase III, and rely on L1-encoded ORF2p for reverse transcription. Alu elements can undergo target-primed reverse transcription using ORF2p alone, whereas L1 elements require ORF1p. Alu likely derive from non-coding 7SL RNA, distinguishing it from most other SINEs, which generally originate from tRNA ancestors. Alu elements emerged approximately 65 million years ago [31], and current retrotransposition estimates range from one new insertion per approximately 20 births (phylogenetic inference) to one per approximately 40 births (short-read whole-genome sequencing) [3,12,25].

Alu retrotransposition has significantly shaped human evolution. At least 5% of alternatively spliced internal exons in the human genome originate from Alu sequences [32]. Like L1, Alu elements comprise hierarchical subfamilies in which older lineages (AluJ, AluS) have become mostly immobile due to truncation, mutation, or host repression, allowing younger subfamilies (AluY) to dominate current activity.

2.3. SVA

SVA (SINE–VNTR–Alu) elements are composite, non-autonomous retrotransposons [9] and constitute the youngest active TE family in humans. SVAs consist of a (CCCTCT)n hexameric tandem repeat, a reversed Alu-like region, a GC-rich variable number tandem repeat (VNTR) domain, a SINE-R domain, and a poly(A) tail [9]. SVA elements range from roughly 700 bp to several kilobases in length, and some copies remain capable of retrotransposition. Their mobilization requires an intact L1 ORF2p and RNA polymerase II for transcription. The necessity of L1 ORF1p for TPRT remains contested and is likely subfamily- and assay-dependent. SVAs account for only approximately 0.15% of the human genome, with roughly 5100 annotated copies according to latest studies [9]. However, due to their variable-length structure, the number of full-length, potentially active SVAs remains difficult to determine accurately [26].

2.4. HERV-K

Human endogenous retrovirus K (HERV-K) represents the youngest endogenous retrovirus family in the human genome and retains several biologically active loci [33]. Although its current retrotransposition activity remains debated, and no confirmed disease-causing de novo HERV-K insertion has been reported to date [34], polymorphic insertions and transcriptional activity have been associated with several human diseases. Because this review focuses primarily on currently active non-LTR retrotransposons (L1, Alu, and SVA) and the computational methods developed for their detection, the biology of HERV-K is not discussed in detail. Some of the tools mentioned later natively handle HERV-K or can be configured to detect them.

3. TE Contributions to Human Health and Disease

One of the first disease-causing TEs identified in humans was an L1 insertion that caused insertional mutagenesis within exon 14 of the F8 gene, resulting in hemophilia [35]. The earliest disease-causing Alu insertion was later detected in an intron of the NF1 gene, where it disrupted normal gene function and caused neurofibromatosis [36]. TE insertions within exons or promoter regions often have deleterious effects on gene expression. As illustrated in Figure 1, TPRT involves endonuclease-mediated cleavage and can be associated with DNA damage and somatic insertions, providing a potential mechanism by which uncontrolled TE activity may contribute to genomic instability.

Some studies have shown that increased L1 ORF1p levels correlates with several types of cancer [37], ORF2p may also be associated with cancer because it is a protein with DNA nicking and TPRT activity, and any genomic instability could be carcinogenic. However, L1 ORF2p is difficult to measure, and only recently has it become possible to measure it reliably [38]. Some cancer tissues also exhibit a higher L1 insertion burden [18,39].

In vivo studies in mice have shown that L1-mediated activation of the cGAS–STING (cyclic GMP-AMP synthase–stimulator of interferon genes) pathway can accelerate aging in cardiovascular tissue [40]. Whether similar mechanisms contribute to human aging remains an open question.

Despite these detrimental effects, TEs also provide beneficial regulatory functions [41], shaping chromatin structure, gene expression, and genome evolution.

Higher-order chromatin organization is critical for long-range promoter–enhancer interactions and complex gene regulatory networks. CTCF (CCCTC-binding factor), a conserved architectural protein, binds directly to specific repetitive elements [42], and L1 elements have been implicated in topologically associating domain (TAD) organization and other aspects of three-dimensional genome architecture [43].

TE activity also contributes to somatic genome mosaicism in human brain tissue. For example, one study estimated an average of 13.7 L1 insertions per hippocampal neuron [44]. In contrast, another study estimated a somatic L1 insertion rate of 0.19 per cell and found no significant differences between hippocampal neurons and glia, cortical neurons, or AGS-derived hippocampal neurons [45]. While the exact number is contested, it has been hypothesized that this activity underlies aspects of neural plasticity required for synapse formation. In vivo experiments in mouse embryos have shown that L1 is involved in neural progenitor cell differentiation [46]. L1 elements may also take part in human brain development. A study has shown that increased L1 silencing results in reduced cerebral organoid development [47].

TE involvement is suggested in tissue regeneration and homeostasis: dental pulp stem cell-derived osteoblasts show lower TE methylation levels [48], and in axolotls, L1 reactivation occurs at the onset of limb regeneration [49], suggesting that loosening TE repression may accompany regenerative programs.

It is important to clarify that TE expression and hypomethylation correlations with cancer have been experimentally validated in human patients [6,37], while the beneficial effects listed above are supported by less robust evidence. The likely physiological functions shown here are still in the basic research phase and have potential applications. TEs appear to be part of complex regulatory networks. Such biological systems are likely easier to disrupt and damage than improve if they are carelessly interfered with.

Regulation of TE Activity

As the previous segment explained, carefully controlled and regulated TEs can have many beneficial effects. This could be one of the reasons they occupy such a large percentage of the human genome. On the other hand, uncontrolled TEs pose a threat to genomic stability due to their potential for unchecked proliferation, illegitimate recombination, and the generation of DNA cuts during TPRT [18,37,50,51,52,53,54]. Consequently, host organisms have evolved multiple mechanisms to suppress TE activity. Additionally, immobile TEs can be domesticated, serving as scaffolds for chromatin organization and stability.

Genome-wide screens in the K562 cancer cell line identified numerous factors modulating L1 activity [55]. Hypermethylation of CpG islands within TE promoters is a major repression mechanism observed in differentiated tissues [50,56]. Tumor suppressors, such as P53, directly inhibit L1 transcription [54], while DNA repair factors (e.g., RAD51C, BRCA1/2, RAD54L) suppress TPRT [57]. The Human Silencing Hub (HUSH) complex, comprising TASOR, MPP8, and Periphilin, mediates L1 silencing in somatic cells via methylation-dependent chromatin condensation [58]. Genome-wide CRISPR knockout screens in K562 cells identified CBX1 and GA8PA as the strongest candidates for L1 activators and SRSF1 and RAD5 as the most likely candidates for L1 suppressors [55].

The Krüppel-associated box zinc finger protein (KRAB-ZFP) family provides another major host defense mechanism. Approximately two-thirds of KRAB-ZFPs recognize multiple sequences across diverse repetitive element families (LTRs, LINEs, SINEs, SVAs, simple repeats) [59]. Upon binding, KRAB-ZFPs recruit KAP1 (TRIM28) to induce sequence-specific transcriptional silencing through histone marking, maintaining genomic integrity. In germline cells, PIWI proteins and PIWI-interacting RNAs (piRNAs) guide cleavage of complementary TE transcripts, preventing transcription and mobilization [53].

Genes regulating nucleic acid metabolism, such as TREX1, SAMHD1, and ADAR1 have also been implicated in suppressing L1 mobilization [60]. TREX1 has been hypothesized to metabolize ssDNA products of TEs as part of the innate immune system [61]. SAMHD1 has several proposed models for restricting TE activity. SAMHD1 may restrict L1 retrotransposition by inducing depletion of L1 ORF2p. Nucleocytoplasmic shuttling is important for SAMHD1-mediated L1 suppression [62], and suggesting the nuclear transport needed for TPRT is also one possible point of L1 and TE regulation. SAMHD1 could also assist in the sequestration of the LINE-1 in cytoplasmic stress granules [63]. Furthermore some studies show ADAR1 can bind to L1 RNP and inhibit L1 RNP activity locally [64].

These intricate layers of TE regulation underscore the need for robust computational tools capable of accurately detecting, annotating, and interpreting TE insertions across diverse human tissues and diseases. Given these complex regulatory mechanisms, TEs may also hold promise as therapeutic targets or biomarkers, pending deeper mechanistic understanding.

4. Sequencing Techniques for TE Analysis

As discussed in the previous section, TEs play an important role in the genome and possess a complex regulatory system, and they could be used as prognostic or predictive biomarkers upon further understanding. This section reviews sequencing techniques, bioinformatic tools, and analytical pipelines developed to uncover TE insertions, deletions, genomic locations, disease associations, regulatory effects, and chromatin structure, and to support the planning and execution of future research.

Since the completion of the Human Genome Project [65], DNA sequencing has become significantly faster, cheaper, and more accessible. This rapid technological leap also allowed more specialized experimental methods to acquire more information not deducible from Sanger sequencing. The most widely used platforms in TE-focused sequencing studies are Illumina (IL), Oxford Nanopore Technologies (ONT), and Pacific Biosciences (PB). As many laboratories are bottlenecked by the type of sequencing platform available, this section presents the most used platforms.

Currently, the most widespread next-generation sequencing technology is based on sequencing by synthesis with fluorescently labelled nucleotides on the Illumina platform [66]. It has well-documented difficulties at properly mapping long repetitive sequences including TEs [67]. The difficulties arise from the fact that Illumina reads most commonly uses 2 × 150 bp paired-end reads. Thus, large repetitive sequences cannot be mapped unambiguously. This has led to a plethora of specialized techniques and tools. An independent comprehensive benchmark for TE detection in Illumina short-read data has been performed [68].

For choosing sequencing protocols, reagents and informatics pipelines combined with Illumina, SequenceEnG [69] could be a useful starting point. It is an interactive database of Illumina sequencing-based pipelines that is very helpful in selecting an analysis method for a specific task from a long list of possibilities, complete with methods and references. In the effort of making the bioinformatics of Illumina require less IT knowledge, many widely used tools are bundled together in the Galaxy online platform [70]. This makes TE research more accessible than approaches based on other sequencing platforms.

Third-generation sequencing (TGS) technologies have substantially improved TE analysis by enabling long-read sequencing, real-time signal acquisition, and detection of selected epigenetic modifications. The repetitive and fragmented nature of TEs creates substantial challenges for alignment and assembly algorithms [6]. Because many TE copies share near-identical sequences, short-reads frequently cannot be mapped unambiguously to a single genomic locus [71]. Long-read sequencing has been particularly transformative for the analysis of active human TE families such as L1, Alu, and SVA. These elements frequently generate polymorphic insertions, truncated copies, inversions, and nested retrotransposition events [22] that are difficult to reconstruct using short-read sequencing alone. Long-reads can span entire insertion loci together with their flanking genomic regions, enabling improved breakpoint resolution, more accurate genotyping, and characterization of insertion hallmarks including poly(A) tails, target-site duplications and complex structural rearrangements.

ONT platforms generate long-reads typically on the range of 10 kb or more [72], with ultra-long protocols producing reads exceeding hundreds of kilobases and, in some cases, reaching megabase scale. These long-reads enable spanning of entire TE insertions and other repetitive regions, reducing mapping ambiguity and improving structural variant (SV) detection. ONT sequencing operates by measuring changes in ionic current as native DNA or RNA molecules pass through nanopores, and the absence of PCR amplification also reduces amplification bias. This allows for preservation and direct detection of base modifications such as DNA methylation [73], making ONT particularly valuable for studying both TE insertions and their epigenetic regulation. Drawbacks include a higher raw error rate compared with short-read platforms [74], which can be mitigated by increased coverage, consensus polishing, or hybrid assembly strategies combining ONT and Illumina data.

The Pacific Biosciences (PacBio, PB) platform employs single-molecule real-time (SMRT) sequencing to generate long-reads, typically 10–25 kb, with highly accurate circular consensus reads (HiFis) reaching 99.9% accuracy [75]. PB is particularly advantageous for TE analysis because the combination of read length and accuracy allows confident detection of full-length TE insertions, SV, and complex rearrangements. SMRT sequencing can also indirectly detect epigenetic modifications, such as DNA methylation; combining SMRT with bisulfite sequencing [76] enabled studies on TE regulation. Limitations include higher per-sample cost and lower throughput compared with Illumina, which may be a bottleneck for large population studies.

Although ONT and PB both overcome many limitations of short-read sequencing for TE analysis, their strengths differ substantially. ONT provides substantially longer reads and direct detection of base modifications, making it particularly suitable for methylation-aware TE studies, resolving large repetitive regions, and detecting complex insertions spanning multiple kilobases [77]. In contrast, PacBio HiFi sequencing offers superior per-base accuracy, which improves breakpoint resolution and characterization of highly similar TE subfamilies. Consequently, the optimal platform depends on the primary biological question, sequencing budget, and required balance between read length, throughput, and nucleotide-level accuracy.

Despite their advantages, long-read technologies retain several limitations relevant to TE analysis. Ultra-long-sequencing protocols require high-molecular-weight DNA and careful sample preparation, which may not be feasible for archived or degraded samples [78]. In addition, long-read datasets typically require substantially greater storage capacity and computational resources during alignment, assembly, and polishing. For ONT data, sequencing chemistry and basecalling models may also influence methylation detection and insertion accuracy, complicating reproducibility between studies.

5. Bioinformatic Pipelines for TE Analysis

As described in the previous section, TEs play important roles in genome regulation, disease, and evolution. This section reviews sequencing techniques and computational pipelines developed to detect TE insertions and deletions, assess their genomic locations, quantify expression, and explore epigenetic regulation, providing guidance for experimental design and bioinformatic analysis.

Fundamental questions in TE research include identifying the genomic locations of insertions, determining whether they disrupt genes or regulatory elements, and assessing their potential effects on gene expression. Addressing these questions is challenging because TEs are highly repetitive, many genomic copies are fragmented, and newly inserted elements are frequently structurally incomplete or aberrant [6]. While the main premise of these tools is highly similar, the differences lie in the types of questions the tools address and in where each tool is positioned within the analytical pipeline, as summarized in Table 2. It is also worth noting that while some of these tools do not directly answer these questions, it may be possible to use specialized R packages or other downstream analysis tools, measurements to answer questions such as: Are these insertions near specific genes? How many of these L1 insertion are full-length or have intact ORFs? How methylated are the insertions? Figure 2 shows the framework of TE analysis pipelines and the many possible changes, which makes comparison between studies challenging.

Table 2.

Conceptual classification of long-read TE detection tools according to workflow architecture, preprocessing requirements, computational complexity, and organism specificity.

Tool ONT PB IL Strategy Scope Preprocessing Complexity Organism Coverage Key Strategies/Primary Application
Alignment and read-based approaches
PALMER [79] ✓ ✓ × Read-based Detection only Alignment Medium Human-focused Detection of human TE families (L1, Alu, SVA, HERV-K) using long-read sequencing. Detects hallmark features of retrotransposition.
xTea [80] ✓ ✓ ✓ Read-based Detection + genotyping Alignment Medium Adaptable Multi-platform TE insertion detection supporting short-read and long-read sequencing with machine learning-based genotyping.
sTELLeR [81] ✓ ✓ × Read-based VCF annotation Alignment + SV calling Low Adaptable Lightweight TE insertion detection utilizing DBSCAN with low computational requirements and fast runtimes.
TraDetIONS [82] ✓ × × Read-based Somatic TE insertion detection Alignment + SV calling Medium Human-focused Detection of somatic and germline TE insertions in ONT tumor-normal paired samples with TPRT hallmark detection.
MELT [83] × × ✓ Read-based Detection Alignment Low Human-focused Benchmark detection tool for short-read sequencing of L1, Alu and SVA.
TraFiC-mem [84] × × ✓ Read-based Somatic TE insertion detection Alignment Low–Medium Human-focused Pan-cancer somatic TE insertion detection tool for Illumina data.
MEIGA-PAV [85] ✓ ✓ × Hybrid SV analysis Complex rearrangement analysis Alignment + Variant calling High Human-focused Detection of complex TE-associated rearrangements including inversions and potentially active L1 elements with intact ORFs.
Assembly-based approaches
TELR [86] ✓ ✓ × Assembly-based Detection + local assembly Alignment High Adaptable High-precision reconstruction and annotation of non-reference TE insertions using local assembly and polishing.
TrEMOLO [87] ✓ ✓ × Assembly-based Integrated TE analysis Assembly + alignment High Adaptable Assembly-supported TE characterization. Distinguishes “insider” and “outsider” TE insertions using assembled genomes and aligned reads with graphical summaries.
Hybrid and end-to-end workflows
GraffiTE [88] ✓ ✓ × Hybrid Full pipeline Alignment + optional assembly Medium–High Adaptable Flexible TE-associated SV detection and genotyping pipeline supporting batch analysis and multiple input types.
Retroinspector [89] ✓ × × Read-based Full workflow + visualization Raw reads or alignment High Human-focused Integrated workflow combining alignment, SV calling, TE annotation and built-in visualization/report generation.
TLDR [90] ✓ × × Read-based Methylation-aware detection Alignment Medium–High Adaptable Simultaneous TE insertion and methylation analysis using ONT long-read sequencing.

Figure 2.

Figure 2

Conceptual framework for TE analysis using modern sequencing technologies, showing the major analytical decision points in TE analysis pipelines.

TE analysis must take the vast amount of fractured, no longer mobile, or otherwise inactive segments into account. The goal is to distinguish potentially active or polymorphic TE insertions and deletions from the large background of ancient, shared, and inactive TE-derived sequences present in the human genome. A common strategy in TE analysis is to identify sequence variants relative to a reference genome and subsequently annotate variants overlapping known TE sequences. By retaining only insertions absent from the reference genome and classified as TE-derived, the candidate search space can be substantially reduced. Although most TE detection pipelines follow this general principle, they differ considerably in how candidate insertions are identified, filtered, reconstructed, and validated.

5.1. Read-Based TE Insertion Detection

Read-based approaches operate directly on aligned sequencing reads and infer TE insertions from discordant alignments, clipped reads, split reads, or characteristic hallmarks of retrotransposition such as poly(A) tails and target site duplications. Because these methods avoid computationally intensive genome assembly, they are typically faster and require fewer computational resources. They are particularly suitable for studies involving large cohorts or heterogeneous tumor samples in which detecting low-frequency insertions may take priority over reconstructing full insertion architecture. However, increased sensitivity often comes at the expense of specificity, and read-based approaches may be more susceptible to false-positive predictions in highly repetitive genomic regions. Tools such as PALMER [79], xTea [80], sTELLeR [81], and TraDetIONS [82] primarily follow this strategy.

5.2. Assembly-Assisted TE Reconstruction

Assembly-based approaches attempt to locally or globally reconstruct the genomic sequence prior to TE annotation. By rebuilding insertion loci directly from sequencing reads, these methods can achieve improved breakpoint resolution and more accurate reconstruction of complex insertions, truncations, inversions, and nested retrotransposition events [22]. Such approaches are advantageous when validating potentially deleterious TE insertions or studying complex structural rearrangements. However, assembly quality is strongly dependent on sequencing depth and read length, and assembly-based workflows are generally computationally demanding due to the additional assembly and polishing steps. TELR [86] and TrEMOLO [87] are examples of assembly-assisted TE detection pipelines. Notably, TrEMOLO distinguishes between “insider” insertions incorporated into the genome assembly and “outsider” insertions supported only by aligned reads, highlighting the limitations of assembly completeness in repetitive regions and the differences between varied approaches.

5.3. Integrated End-to-End Workflows

A third category includes integrated, hybrid workflows that combine multiple analytical stages including alignment, SV calling, TE annotation, and visualization into unified pipelines. These integrated pipelines emphasize workflow standardization and simplified execution, although this may reduce flexibility for highly customized analyses. For example, Retroinspector [89] implements alignment, SV detection, annotation, and graphical summarization within a unified, script-driven workflow, whereas GraffiTE [88] combines SV discovery with TE-focused genotyping and annotation in a Nextflow [91] pipeline. Such integrated pipelines can reduce technical barriers for laboratories without specialized bioinformatics expertise but may provide less flexibility for custom-tailored analysis strategies.

Most current TE detection methods remain fundamentally library-based. Since the development of RepeatMasker [92], curated TE consensus databases, such as Repbase [17] and Dfam [93], have formed the basis of TE annotation pipelines. These approaches rely on sequence similarity and are highly effective for identifying previously characterized TE families. However, their performance depends heavily on the completeness and quality of the underlying TE libraries. This limitation becomes particularly important in non-human organisms where TE catalogs remain incomplete. Although machine learning approaches are increasingly being explored for TE classification and insertion detection, their adoption remains substantially more limited than library-based approaches. sTELLeR [81] is one example for complementing DBSCAN machine learning approaches with sequence similarity.

Despite the advantages of long-read sequencing, short-read sequencing data remain substantially more widely available in many research settings. Consequently, TE detection from short-read datasets remains highly relevant, particularly for retrospective analyses of existing clinical cohorts [71]. Large-scale benchmarking studies have demonstrated that specialized short-read pipelines can achieve robust TE detection performance in specific settings. For example, MELT [83] showed strong performance in exome sequencing datasets [68], whereas xTea demonstrated improved detection accuracy in whole-genome sequencing data [80]. These comparisons highlight that optimal pipeline selection is highly dependent on experimental design, sequencing modality, biological question being addressed, and the value of independent benchmarking.

Long-read sequencing has allowed the easier integration of TE analysis with epigenetic studies. ONT-based methylation-aware sequencing allows simultaneous characterization of TE insertions and their epigenetic state. Tools such as TLDR [90] integrate insertion detection with methylation profiling, enabling investigation of whether newly inserted elements remain epigenetically repressed or transcriptionally active. For those who have chosen a different pipeline and have methylation information available, using Modkit, Methylartist [94] or other methylation analysis in downstream analysis could complement insertion data with epigenetic information.

5.4. TE Expression Analysis Workflows

While the previous segment focused on the detection of TE insertional burden, deletions, and other SV, transcriptional profiling represents an equally important aspect of TE biology. Bulk RNA sequencing has been widely used to quantify TE expression, with tools including TEtranscripts [95], L1EM [96], Telescope [97] and SQuIRE [98]. A major computational challenge is the already mentioned extensive sequence similarity among TE copies. This is especially relevant to the youngest and most active segments, which causes many sequencing reads to map equally well to multiple genomic loci. Consequently, accurate locus-specific quantification remains substantially more difficult than family-level expression analysis, and benchmarking studies have demonstrated that multi-mapping can result in considerable locus-level false discovery rates. Comprehensive reviews and methodological comparisons of TE expression pipelines have recently been published and are therefore not discussed in detail here [99,100].

The use of long reads could substantially improve TE expression analysis by enhancing the reassignment of multimapping reads. LocusMasterTE is one such approach [101]. The authors reported improved mapping of the youngest and most active TEs compared with short-read RNA-seq tools. However, no independent benchmarking study of LocusMasterTE has yet to be completed. Because TE insertion detection has benefited from long-read technologies, expression-analysis pipelines may also achieve important methodological advances. At present, more advancements are needed for potential clinical applications.

Similarly, single-cell sequencing approaches are beginning to address cell-type-specific TE activity and somatic mosaicism. In heterogeneous tissues such as tumors or neuronal populations, bulk sequencing may obscure low-frequency TE insertions restricted to specific cell populations. CELLO-seq [102] extends long-read sequencing to single-cell TE expression analysis, while MATES [103] applies machine learning approaches to TE quantification in single-cell datasets. These methods remain computationally demanding but represent an important future direction for understanding TE dynamics at cellular resolution.

5.5. Computational Considerations

Computational requirements vary substantially between TE analysis strategies and are influenced by factors including alignment complexity, genome assembly, variant calling, methylation analysis, and machine learning integration. Read-based heuristic filtering approaches generally require fewer computational resources than assembly-based reconstruction methods [86,87]. Likewise, workflows incorporating methylation-aware basecalling, genome polishing, or single-cell analyses introduce additional computational overhead. Importantly, the computational complexity classifications provided in Table 2 represent qualitative methodological estimates rather than direct benchmarking measurements, as runtime and memory usage remain highly dependent on sequencing depth, read length, sample number, and available hardware infrastructure.

Most of the discussed tools are open source and typically installed from repositories through command line environment managers such as Conda, Mamba, or through the use of containers. Their use requires computational expertise, and analyses can be computationally demanding, particularly for long-read WGS data. A detailed summary of insertion detection tool availability, pipeline type, and organism scope is provided in Table 2, with download links listed in Data Availability Statement Section.

6. Practical Considerations and Guidelines for TE-Focused Studies

6.1. Literature and Tool Selection

This article is a narrative review and was not conducted as a systematic review. The relevant literature was identified through searches of PubMed and Google Scholar using combinations of terms including “Transposable Element detection,” “Mobile element detection,” “TE insertion detection,” “TE structural variation,” “TE methylation,” “TE expression analysis”. Tool selection was based on relevance to the detection and characterization of human TE insertions from sequencing data, with emphasis on methods published up to March 2026. Tools were included when they provided a distinct analytical strategy, supported long-read data, or represented an important comparator for TE detection in short-read data. As a narrative review, the presented tool collection is intended to illustrate the current methodological landscape rather than provide an exhaustive systematic inventory of all available software.

6.2. Selection of Appropriate TE Analysis Pipelines

This section presents a couple of possible scenarios and research questions and discusses what would be the optimal TE pipeline for that particular situation. Table 3 lists the different research objectives, their key considerations and some recommended tools for that task. These recommendations are intended as general guidance. Multiple tools may be suitable for a given study depending on sequencing technology, sample quality, computational resources, and desired balance between sensitivity and specificity. Table 2 highlights the necessary preprocessing steps needed for different tools, indicating the pipelines that use hard-wired TE consensus, which of them can be configured by the user, and most importantly, which sequencing platforms they support. While not an exhaustive list, this table should help new research groups entering the field.

Table 3.

Practical guidance table on which TE analysis tools are useful for different research objectives. It is very likely that several other questions can be answered by these tools. At present, clinical validation has no gold standard and any such study should complement it with experimental validation.

Research Objective Key Requirement Suitable Approaches
Population cohorts Scalability, multi-platform support xTea [80], TraFiC-mem [84]
Clinical validation High confidence Human-focused tool with PCR or other experimental validation
Tumor-normal-pairing Somatic comparison TraDetIONS [82], TraFiC-mem [84]
Using existing Illumina data Short-read support MELT [83], xTea [80]
Methylation analysis Native ONT reads TLDR [90], other ONT compatible tool combined with Methylartist [94] or Modkit
Bulk expression analysis RNA expression data TEtranscripts [95], SQuIRE [98]
Single-cell expression analysis Single-cell level expression data CELLO-seq [102], MATES [103]
Structural characterization Complex insertions MEIGA-PAV [85]
Exploratory analysis High sensitivity, built-in downstream analysis Retroinspector [89], GraffiTE [88]

6.2.1. Retrospective Analysis of Different Cohorts

Research question: Which non-reference TE insertions are associated with a disease across hundreds or thousands of genomes?

Methodological considerations:

Large cohorts such as The Cancer Genome Atlas or the Pan-Cancer Analysis of Whole Genomes [14,104] often contain sequencing data generated using different platforms and library preparations. Therefore, compatibility across sequencing technologies, scalability, and standardized output become more important than reconstructing every insertion in maximal detail.

Recommended analytical strategy:

When both long-read and short-read samples are available, tools supporting multiple sequencing technologies, such as xTea, are well suited for these analyses because they can process both Illumina and long-read datasets using a unified analytical framework. When only short-read data were available, MELT demonstrated strong performance in independent benchmarking of exome sequencing datasets [68]. There have been successful breakthroughs with large-scale genomics studies involving TE; Rodriguez-Martin et al. [84] have performed one of the largest-scale TE detection studies to date in the framework of the Pan-Cancer Analysis of Whole Genomes project using TraFiC-mem. Their methodology could also serve as guidance.

6.2.2. Clinical Validation

Research question: Does a patient carry a pathogenic TE insertion affecting a disease-associated gene?

Methodological considerations: In clinical investigations, confidence in individual insertion calls is generally more important than processing speed. Assembly-assisted reconstruction can improve breakpoint resolution and facilitate downstream validation. Extra attention must be paid for experimental verification as well.

Recommended analytical strategy: The most important caveat here is that no established gold standard or formally validated diagnostic pipeline exists as of the writing of this article for TE detection pipelines. If such study were undertaken, a haplotype-resolved assembly on an enriched long-read sample, annotated with a human-specific annotation tool and combined with experimental validation such as PCR, would be a good place to start.

6.2.3. Tumor-Normal Comparison

Research question: Which TE insertions are somatically acquired during tumor development?

Methodological considerations: The analytical workflow should distinguish germline from somatic insertions while accounting for tumor heterogeneity.

Recommended analytical strategy: TraFiC-mem and TraDetIONS were specifically developed for paired tumor-normal analyses, and therefore, directly address this experimental design.

6.2.4. Epigenetic Regulation

Research question: Are the TE insertions epigenetically silenced?

Methodological considerations: DNA methylation information must be available in addition to the sequence.

Recommended analytical strategy: TLDR was designed for this research question. Furthermore, combining other tools with downstream methylation analysis using Modkit or Methylartist [94] could also address the research question.

6.2.5. Single-Cell Biology

Research question: Which cell populations express TEs?

Difficulty: Bulk sequencing averages signals across cells. Cell-level expression data is lost.

Recommended analytical strategy: Single-cell approaches that preserve cell identity are recommended, particularly single-cell barcoding combined with expression analysis. CELLO-seq or MATES are appropriate pipelines developed to solve this question.

6.2.6. Structural Characterization

Research question: Are the TE insertions truncated, inverted, or associated with 3′ transduction events? Do they have intact open reading frames? Are they potentially active or aberrant?

Methodological consideration: Beyond localization, structural characterization is required.

Recommended analytical strategy: MEIGA-PAV was designed to address this research question.

6.2.7. Low Tumor-Purity Analysis

Research question: Where are the TE insertions located in a liquid biopsy sample or other low tumor-purity sample?

Methodological considerations: The true signal, like cancer-relevant insertions, need to be filtered out from a potentially very small portion of the reads. Careful compromises must be made between precision and sensitivity.

Recommended analytical strategy: Similarly to clinical validation, no currently adopted standard practice exists. Every option to reduce the error rate and to increase the signal-to-noise ratio should be considered. Options such as ONT adaptive sequencing, to enrich the target sequences [105], or using Unique Molecular Identifiers (UMIs) [106], to reduce error rate, have the potential to make such studies more effective. Combining those with an adjustable-read-based tool could detect an insertion that would not appear in a tool not sensitive enough, like one based on de novo assembly.

No single workflow is optimal for all research objectives. Pipeline selection should primarily be driven by the biological question, sequencing modality, desired sensitivity, available computational resources, and downstream analyses rather than by overall tool popularity.

Finally, it is important to remember that the accuracy of TE detection depends strongly on sequencing quality. Read length, sequencing depth, N50, alignment quality, and basecalling accuracy influence both sensitivity and precision. Consequently, appropriate quality control should precede any biological interpretation.

6.3. Current Limitations of TE Analysis

Despite major advances in sequencing technologies and computational methods, several important limitations continue to constrain TE analysis. These limitations affect detection accuracy, reproducibility, computational requirements, and biological interpretation.

In human genomes, curated databases such as Dfam 3.9 and Repbase are generally assumed to cover the vast majority of currently recognized human TE families in the era of telomere-to-telomere sequencing. However, Repbase is now under commercial license, which may remain a practical barrier for some laboratories and institutions, especially in resource-limited regions. In non-human organisms, incomplete TE annotation and limited availability of curated TE libraries remain substantial challenges and may reduce detection sensitivity [107,108]. Consequently, researchers must carefully select the TE consensus libraries used during analysis, particularly when using tools restricted to predefined TE families. TE insertions absent from the selected database cannot be detected. Conversely, tools hardwired to broad databases, such as Retroinspector with the Dfam 3.9 human library, may recover large numbers of inactive or highly fragmented TE copies that are not biologically relevant for a given study, increasing downstream filtering requirements and computational burden.

Another important limitation is that many TE analysis pipelines were developed for highly specific research questions. Default parameters are therefore not universally applicable, as hyperparameters such as minimum read support, alignment thresholds, or SV filtering criteria can substantially influence sensitivity and precision depending on sequencing depth, tumor purity, and experimental design [80,81]. For example, highly heterogeneous tumor samples or samples with low tumor purity may require relaxed filtering thresholds to recover low-frequency somatic insertions, whereas clinical diagnostic workflows may prioritize specificity and reproducibility over maximal sensitivity. Sensitivity estimates reported by different tools are often difficult to compare directly because benchmarking datasets, sequencing depth, read lengths, and validation criteria vary substantially between studies.

A major unresolved issue in the field is the lack of comprehensive independent benchmarking studies across recently developed tools. Existing evaluations are often performed by the developers themselves and commonly compare new methods only against older approaches rather than against the most recent alternatives [80,81,88]. Although early independent benchmarking efforts for long-read TE detection tools such as PALMER, TLDR, sTELLeR, and xTea have begun to emerge, the rapidly expanding number of pipelines highlights the need for larger standardized benchmarking studies. Such efforts would improve reproducibility, clarify optimal use cases for each tool, and facilitate more objective selection of analysis workflows.

Computational reproducibility also remains challenging. Although most TE analysis tools are publicly available and open source, practical implementation often requires substantial bioinformatic expertise. Installation may be complicated by dependency conflicts, outdated packages, variable documentation quality, or incompatibilities between software versions. In addition, many pipelines rely on whole-genome sequencing data and therefore require substantial storage capacity, memory allocation, and wall-clock runtime. These technical barriers may limit adoption in smaller laboratories lacking dedicated computational infrastructure.

Reference genome selection introduces an additional source of bias. As illustrated in Table 4, substantial differences exist between commonly used human reference genomes such as GRCh38 and T2T assemblies. TE insertions identified relative to one reference genome may represent common polymorphisms absent from another assembly rather than novel insertion events. This issue is particularly relevant in underrepresented populations, where population-specific germline variants may be incorrectly interpreted as rare or disease-associated insertions if they are not represented in the reference genome. Potential solutions to this problem include the Human Pangenome reference [109]. Using their data could help in deciding if a detected non-reference TE insertion is a true somatic insertion or a subpopulation-specific TE.

Table 4.

TE content of the GRCh38 and T2T reference genomes. Data from [16] are rounded. Almost half of the genome is measurably taken up by TEs. Some estimates hypothesize that up to two-thirds of the genome consisted of TEs, but became so fragmented that they no longer readily recognizable as TEs [4]. The T2T reference genome contains around 200 Mb more base pairs in difficult to map regions. These differences highlight the importance of reference selection in experiments as some regions detected in T2T might not show up if the sample was aligned against the GRCh38 reference. Care must be taken when evaluating or comparing results using different reference genomes.

TE GRCh38% GRCh38 Mbp T2T% T2T Mbp
L1 17.36 507 16.77 512
Other LINE 4.07 118 3.90 119
Alu 10.43 304 10.09 308
Other SINE 2.80 81 2.68 82
SVA 0.15 4.50 0.15 4.65
LTR 9.16 267 8.84 270
DNA transposon 3.71 108 3.58 109

7. Future Directions and Open Questions

Understanding the contributions of currently active and replicating TEs requires careful investigation of their role in disease. Transcript levels of Alu and L1 ORF1p [30,51] correlate with disease severity and may serve as biomarkers of poor prognosis. These observations suggest that TE activation may contribute to disease progression, although causality remains incompletely established.

The broader adoption of long-read sequencing, along with improved bioinformatic pipelines and experimental approaches specifically designed for WGS and TE detection, is expected to improve our understanding of the mechanistic links between TEs and human diseases in the near future. Integrating long-read sequencing, functional studies, and standardized analysis pipelines will be essential to fully understand TE biology and their roles in human health and disease.

At the time of writing, many of the listed TE analysis tools reviewed here have not been independently benchmarked. Such benchmarking could establish community standards for TE insertion detection, facilitate reproducibility, and provide developers with critical feedback for improving and harmonizing analysis pipelines.

Pangenome references and graph-based genome representations are already available [109,110,111], but their application in TE analysis remains underutilized. Other advancements that could support TE analysis include graph genome representations, standardized benchmark datasets, improved TE annotations, increasingly integrated long-read sequencing workflows, direct native RNA sequencing [112], advances in machine learning, and single-cell multiomics approaches. While several of these technologies are already established, their broader adoption and integration into TE research represent important opportunities for future methodological development. Together, these advances could help shift TE research beyond the identification of insertion sites toward addressing broader biological and clinical questions.

8. Conclusions

TEs have profoundly shaped the human genome, participating in an ongoing evolutionary dynamic as hosts evolve mechanisms to suppress their activity. Although most TE subfamilies have lost their ability to mobilize and now persist as genomic fossils, their continued presence underscores their biological significance. Future studies may clarify how TE activity contributes to aging, elucidate the mechanisms underlying the association between L1 expression and cancer, and determine whether TE activity can be therapeutically targeted.

Acknowledgments

We thank Linkner Tamás for his help in creating the BioRender figure featured in Figure 1. Latest versions of ChatGPT 5.5 mini and Perplexity Max were used for LaTeX formatting, language checking, and initial review. We have carefully reviewed and edited the output and take full responsibility for the content of this publication.

Abbreviations

The following abbreviations are used in this manuscript:

HERV-K Human endogenous retrovirus K
HiFi Highly accurate circular consensus reads
IL Illumina
L1 Long interspersed nuclear element-1
ONT Oxford Nanopore Technologies
ORF Open reading frame
PB Pacific Biosciences
RNP Ribonucleoprotein
SMRT Single-molecule real-time
SV Structural variants
SVA SINE-VNTR-Alu
TAD Topologically associating domain
TGS Third-generation sequencing
TE Transposable element
TPRT Target-primed reverse transcription
TSD Target site duplication
UMI Unique molecular identifiers
UTR Untranslated region
VNTR Variable number tandem repeat
WGS Whole-genome sequencing

Author Contributions

D.V. and B.M. conceived the idea for the review and wrote the manuscript. N.S. and A.K. contributed their previous expertise in TE methylation and its analysis. I.T. reviewed the article. All authors have read and agreed to the published version of the manuscript.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

No new data were created or analyzed in this study. Data sharing is not applicable to this article. The source codes for most bioinformatic tools discussed in this review are publicly available at the following repositories: CELLO-seq: https://github.com/MarioniLab/CELLOseq (accessed on 24 August 2026); GraffiTE: https://github.com/cgroza/GraffiTE (accessed on 24 August 2026); L1EM: https://github.com/FenyoLab/L1EM (accessed on 24 August 2026); LocusMasterTE: https://github.com/jasonwong-lab/LocusMasterTE (accessed on 24 August 2026); MATES: https://github.com/mcgilldinglab/MATES (accessed on 24 August 2026); MEIGA-PAV: https://github.com/MEIGA-tk/MEIGA-PAV (accessed on 24 August 2026); MELT: https://melt.igs.umaryland.edu/downloads.php (accessed on 24 August 2026); PALMER: https://github.com/WeichenZhou/PALMER (accessed on 24 August 2026); Retroinspector: https://github.com/javiercguard/retroinspector (accessed on 24 August 2026); sTELLeR: https://github.com/kristinebilgrav/sTELLeR (accessed on 24 August 2026); SQuIRE: https://github.com/wyang17/SQuIRE (accessed on 24 August 2026); TEtranscripts: https://github.com/mhammell-laboratory/TEtranscripts (accessed on 24 August 2026); Telescope: https://github.com/mlbendall/telescope (accessed on 24 August 2026); TELR: https://github.com/bergmanlab/TELR (accessed on 24 August 2026); tldr: https://github.com/adamewing/tldr (accessed on 24 August 2026); TraDetIONS: https://github.com/panummi/TraDetIONS (accessed on 24 August 2026); TraFiC-mem: https://gitlab.com/mobilegenomesgroup/TraFiC (accessed on 24 August 2026); TrEMOLO: https://github.com/DrosophilaGenomeEvolution/TrEMOLO (accessed on 24 August 2026); xTea: https://github.com/parklab/xTea (accessed on 24 August 2026).

Conflicts of Interest

The authors declare no conflicts of interest.

Funding Statement

This study was funded by the Hungarian Scientific Research Funds NRDI-FK0201NEPE/TKPNKTA-47 and RRF-2.3.1-21-2022-00003.

Footnotes

Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

References

  • 1.Liehr T. Repetitive elements in humans. Int. J. Mol. Sci. 2021;22:2072. doi: 10.3390/ijms22042072. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Mangiavacchi A., Liu P., Della Valle F., Orlando V. New insights into the functional role of retrotransposon dynamics in mammalian somatic cells. Cell. Mol. Life Sci. 2021;78:5245–5256. doi: 10.1007/s00018-021-03851-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Hancks D.C., Kazazian H.H. Active human retrotransposons: Variation and disease. Curr. Opin. Genet. Dev. 2012;22:191–203. doi: 10.1016/j.gde.2012.02.006. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.de Koning A.P., Gu W., Castoe T.A., Batzer M.A., Pollock D.D. Repetitive elements may comprise over two-thirds of the human genome. PLoS Genet. 2011;7:e1002384. doi: 10.1371/journal.pgen.1002384. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Szakállas N., Kalmár A., Rada K.R., Kucarov M., Linkner T.R., Barták B.K., Takács I., Molnár B. Methodological comparison of short-read and long-read sequencing methods on colorectal cancer samples. Int. J. Mol. Sci. 2025;26:9254. doi: 10.3390/ijms26189254. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Snowbarger J., Koganti P., Spruck C. Evolution of repetitive elements, their roles in homeostasis and human disease, and potential therapeutic applications. Biomolecules. 2024;14:1250. doi: 10.3390/biom14101250. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Wells J.N., Feschotte C. A field guide to eukaryotic transposable elements. Annu. Rev. Genet. 2020;54:539–561. doi: 10.1146/annurev-genet-040620-022145. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Elliott T.A., Heitkam T., Hubley R., Quesneville H., Suh A., Wheeler T.J. TE Hub: A Community-Oriented Space for Sharing and Connecting Tools, Data, Resources, and Methods for Transposable Element Annotation. Mob. DNA. 2021;12:16. doi: 10.1186/s13100-021-00244-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Chu C., Lin E.W., Tran A., Jin H., Ho N.I., Veit A., Cortes-Ciriano I., Burns K.H., Ting D.T., Park P.J. The landscape of human SVA retrotransposons. Nucleic Acids Res. 2023;51:11453–11465. doi: 10.1093/nar/gkad821. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Brouha B., Schustak J., Badge R.M., Lutz-Prigge S., Farley A.H., Moran J.V., Kazazian H.H. Hot l1s account for the bulk of retrotransposition in the human population. Proc. Natl. Acad. Sci. USA. 2003;100:5280–5285. doi: 10.1073/pnas.0831042100. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Yang L., Metzger G.A., Padilla Del Valle R., Delgadillo Rubalcaba D., McLaughlin R.N. Evolutionary Insights from Profiling LINE-1 Activity at Allelic Resolution in a Single Human Genome. EMBO J. 2023;43:6. doi: 10.1038/s44318-023-00007-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Bennett E.A., Keller H., Mills R.E., Schmidt S., Moran J.V., Weichenrieder O., Devine S.E. Active alu retrotransposons in the human genome. Genome Res. 2008;18:1875–1883. doi: 10.1101/gr.081737.108. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Wang H., Xing J., Grover D., Hedges D.J., Han K., Walker J.A., Batzer M.A. SVA Elements: A Hominid-specific Retroposon Family. J. Mol. Biol. 2005;354:994–1007. doi: 10.1016/j.jmb.2005.09.085. [DOI] [PubMed] [Google Scholar]
  • 14.Aaltonen L.A., Abascal F., Abeshouse A., Aburatani H., Adams D.J., Agrawal N., Ahn K.S., Ahn S.M., Aikata H., Akbani R., et al. Pan-Cancer Analysis of Whole Genomes. Nature. 2020;578:82–93. doi: 10.1038/s41586-020-1969-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Reyes C.J., Asano K., Todd P.K., Klein C., Rakovic A. Repeat-Associated Non-AUG Translation of AGAGGG Repeats That Cause X-Linked Dystonia-Parkinsonism. Mov. Disord. 2022;37:2284–2289. doi: 10.1002/mds.29183. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Hoyt S.J., Storer J.M., Hartley G.A., Grady P.G., Gershman A., de Lima L.G., Limouse C., Halabian R., Wojenski L., Rodriguez M., et al. From telomere to telomere: The transcriptional and epigenetic state of human repeat elements. Science. 2022;376:eabk3112. doi: 10.1126/science.abk3112. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Kojima K.K. Human transposable elements in Repbase: Genomic footprints from fish to humans. Mob. DNA. 2018;9:2. doi: 10.1186/s13100-017-0107-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Lee E., Iskow R., Yang L., Gokcumen O., Haseley P., Luquette L.J., Lohr J.G., Harris C.C., Ding L., Wilson R.K., et al. Landscape of somatic retrotransposition in human cancers. Science. 2012;337:967–971. doi: 10.1126/science.1222077. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Jurka J. Subfamily Structure and evolution of the human L1 family of repetive sequences. J. Mol. Evol. 1989;29:496–503. doi: 10.1007/bf02602921. [DOI] [PubMed] [Google Scholar]
  • 20.Beck C.R., Collier P., Macfarlane C., Malig M., Kidd J.M., Eichler E.E., Badge R.M., Moran J.V. Line-1 retrotransposition activity in human genomes. Cell. 2010;141:1159–1170. doi: 10.1016/j.cell.2010.05.021. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Wang P.J. Tracking Line1 retrotransposition in the Germline. Proc. Natl. Acad. Sci. USA. 2017;114:7194–7196. doi: 10.1073/pnas.1709067114. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.SanMiguel P., Tikhonov A., Jin Y.K., Motchoulskaia N., Zakharov D., Melake-Berhan A., Springer P.S., Edwards K.J., Lee M., Avramova Z., et al. Nested retrotransposons in the intergenic regions of the maize genome. Science. 1996;274:765–768. doi: 10.1126/science.274.5288.765. [DOI] [PubMed] [Google Scholar]
  • 23.Lanciano S., Philippe C., Sarkar A., Pratella D., Domrane C., Doucet A.J., van Essen D., Saccani S., Ferry L., Defossez P.A., et al. Locus-level L1 DNA methylation profiling reveals the epigenetic and transcriptional interplay between L1s and their integration sites. Cell Genom. 2024;4:100498. doi: 10.1016/j.xgen.2024.100498. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Denli A.M., Narvaiza I., Kerman B.E., Pena M., Benner C., Marchetto M.C., Diedrich J.K., Aslanian A., Ma J., Moresco J.J., et al. Primate-Specific ORF0 Contributes to Retrotransposon-Mediated Diversity. Cell. 2015;163:583–593. doi: 10.1016/j.cell.2015.09.025. [DOI] [PubMed] [Google Scholar]
  • 25.Feusier J., Watkins W.S., Thomas J., Farrell A., Witherspoon D.J., Baird L., Ha H., Xing J., Jorde L.B. Pedigree-based estimation of human mobile element retrotransposition rates. Genome Res. 2019;29:1567–1577. doi: 10.1101/gr.247965.118. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Helman E., Lawrence M.S., Stewart C., Sougnez C., Getz G., Meyerson M. Somatic retrotransposition in human cancer revealed by whole-genome and exome sequencing. Genome Res. 2014;24:1053–1063. doi: 10.1101/gr.163659.113. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Belancio V.P., Deininger P.L., Roy-Engel A.M. Line dancing in the human genome: Transposable elements and disease. Genome Med. 2009;1:97. doi: 10.1186/gm97. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Goerner-Potvin P., Bourque G. Computational tools to unmask transposable elements. Nat. Rev. Genet. 2018;19:688–704. doi: 10.1038/S41576-018-0050-X. [DOI] [PubMed] [Google Scholar]
  • 29.Braakhuis B.J., Leemans C.R., Brakenhoff R.H. Using tissue adjacent to carcinoma as a normal control: An obvious but questionable practice. J. Pathol. 2004;203:620–621. doi: 10.1002/path.1549. [DOI] [PubMed] [Google Scholar]
  • 30.Ade C., Roy-Engel A.M., Deininger P.L. Alu elements: An intrinsic source of human genome instability. Curr. Opin. Virol. 2013;3:639–645. doi: 10.1016/j.coviro.2013.09.002. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Terreros M.C., Alfonso-Sánchez M.A., Novick G.E., Luis J.R., Lacau H., Lowery R.K., Regueiro M., Herrera R.J. Insights on Human Evolution: An Analysis of Alu Insertion Polymorphisms. J. Hum. Genet. 2009;54:603–611. doi: 10.1038/jhg.2009.86. [DOI] [PubMed] [Google Scholar]
  • 32.Sorek R., Lev-Maor G., Reznik M., Dagan T., Belinky F., Graur D., Ast G. Minimal conditions for exonization of intronic sequences. Mol. Cell. 2004;14:221–231. doi: 10.1016/s1097-2765(04)00181-9. [DOI] [PubMed] [Google Scholar]
  • 33.Garcia-Montojo M., Doucet-O’Hare T., Henderson L., Nath A. Human Endogenous Retrovirus-K (HML-2): A Comprehensive Review. Crit. Rev. Microbiol. 2018;44:715–738. doi: 10.1080/1040841X.2018.1501345. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Xue B., Sechi L.A., Kelvin D.J. Human Endogenous Retrovirus K (HML-2) in Health and Disease. Front. Microbiol. 2020;11:1690. doi: 10.3389/fmicb.2020.01690. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Kazazian H.H., Wong C., Youssoufian H., Scott A.F., Phillips D.G., Antonarakis S.E. Haemophilia a resulting from de novo insertion of L1 sequences represents a novel mechanism for mutation in man. Nature. 1988;332:164–166. doi: 10.1038/332164a0. [DOI] [PubMed] [Google Scholar]
  • 36.Wallace M.R., Andersen L.B., Saulino A.M., Gregory P.E., Glover T.W., Collins F.S. A de novo alu insertion results in neurofibromatosis type 1. Nature. 1991;353:864–866. doi: 10.1038/353864a0. [DOI] [PubMed] [Google Scholar]
  • 37.Scott E., Devine S. The role of somatic L1 retrotransposition in human cancers. Viruses. 2017;9:131. doi: 10.3390/v9060131. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Nielsen M.I., Wolters J.C., Bringas O.G.R., Jiang H., Di Stefano L.H., Oghbaie M., Hozeifi S., Nitert M.J., van Pijkeren A., Smit M., et al. Targeted Detection of Endogenous LINE-1 Proteins and ORF2p Interactions. Mob. DNA. 2025;16:3. doi: 10.1186/s13100-024-00339-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Pradhan B., Cajuso T., Katainen R., Sulo P., Tanskanen T., Kilpivaara O., Pitkänen E., Aaltonen L.A., Kauppi L., Palin K. Detection of subclonal L1 transductions in colorectal cancer by long-distance inverse-PCR and nanopore sequencing. Sci. Rep. 2017;7:14521. doi: 10.1038/s41598-017-15076-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Yang C., Du H., Liu S., Xu P., Wang Y., Zhou Y., Yuan H., Li Y., Shen J., Yuan X., et al. Targeting age-related line-1 activation alleviates cardiac aging. Nat. Aging. 2026;6:414–429. doi: 10.1038/s43587-025-01056-0. [DOI] [PubMed] [Google Scholar]
  • 41.Ilık E.A., Yang X., Zhang Z.Z., Aktaş T. Transcriptional and post-transcriptional regulation of transposable elements and their roles in development and disease. Nat. Rev. Mol. Cell Biol. 2025;26:759–775. doi: 10.1038/s41580-025-00867-8. [DOI] [PubMed] [Google Scholar]
  • 42.Zhu K., Zhuo J., Zhen Y. Evolution of CTCF Binding Sites in the Human Genome. Mol. Biol. Evol. 2026;43:msag167. doi: 10.1093/molbev/msag167. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Hong Y., Bie L., Zhang T., Yan X., Jin G., Chen Z., Wang Y., Li X., Pei G., Zhang Y., et al. SAFB Restricts Contact Domain Boundaries Associated with L1 Chimeric Transcription. Mol. Cell. 2024;84:1637–1650.e10. doi: 10.1016/j.molcel.2024.03.021. [DOI] [PubMed] [Google Scholar]
  • 44.Upton K., Gerhardt D., Jesuadian J., Richardson S., Sánchez-Luque F., Bodea G., Ewing A., Salvador-Palomeque C., vanderKnaap M., Brennan P., et al. Ubiquitous L1 mosaicism in hippocampal neurons. Cell. 2015;161:228–239. doi: 10.1016/j.cell.2015.03.026. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Evrony G.D., Lee E., Park P.J., Walsh C.A. Resolving Rates of Mutation in the Brain Using Single-Neuron Genomics. eLife. 2016;5:e12966. doi: 10.7554/eLife.12966. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Mangoni D., Simi A., Lau P., Armaos A., Ansaloni F., Codino A., Damiani D., Floreani L., Di Carlo V., Vozzi D., et al. LINE-1 Regulates Cortical Development by Acting as Long Non-Coding RNAs. Nat. Commun. 2023;14:4974. doi: 10.1038/s41467-023-40743-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Garza R., Atacho D.A.M., Adami A., Gerdes P., Vinod M., Hsieh P., Karlsson O., Horvath V., Johansson P.A., Pandiloski N., et al. LINE-1 Retrotransposons Drive Human Neuronal Transcriptome Complexity and Functional Diversification. Sci. Adv. 2023;9:eadh9543. doi: 10.1126/sciadv.adh9543. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Prucksakorn T., Mutirangura A., Pavasant P., Subbalekha K. Altered methylation levels in line-1 in dental pulp stem cell–derived osteoblasts. Int. Dent. J. 2025;75:1269–1276. doi: 10.1016/j.identj.2024.09.009. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Zhu W., Kuo D., Nathanson J., Satoh A., Pao G.M., Yeo G.W., Bryant S.V., Voss S.R., Gardiner D.M., Hunter T. Retrotransposon long interspersed nucleotide element-1 (line-1) is activated during salamander limb regeneration. Dev. Growth Differ. 2012;54:673–685. doi: 10.1111/j.1440-169x.2012.01368.x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Chuong E.B., Elde N.C., Feschotte C. Regulatory activities of transposable elements: From conflicts to benefits. Nat. Rev. Genet. 2017;18:71–86. doi: 10.1038/NRG.2016.139. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Ardeljan D., Taylor M.S., Ting D.T., Burns K.H. The Human Long Interspersed Element-1 Retrotransposon: An Emerging Biomarker of Neoplasia. Clin. Chem. 2017;63:816–822. doi: 10.1373/CLINCHEM.2016.257444. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Burns K.H. Transposable elements in cancer. Nat. Rev. Cancer. 2017;17:415–424. doi: 10.1038/NRC.2017.35. [DOI] [PubMed] [Google Scholar]
  • 53.Ozata D.M., Gainetdinov I., Zoch A., O’Carroll D., Zamore P.D. Piwi-interacting RNAS: Small RNAS with big functions. Nat. Rev. Genet. 2018;20:89–108. doi: 10.1038/s41576-018-0073-3. [DOI] [PubMed] [Google Scholar]
  • 54.Tiwari B., Jones A.E., Caillet C.J., Das S., Royer S.K., Abrams J.M. P53 directly represses human LINE1 transposons. Genes Dev. 2020;34:1439–1451. doi: 10.1101/gad.343186.120. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Liu N., Lee C.H., Swigut T., Grow E., Gu B., Bassik M.C., Wysocka J. Selective silencing of euchromatic l1s revealed by genome-wide screens for L1 regulators. Nature. 2017;553:228–232. doi: 10.1038/nature25179. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Weisenberger D.J. Analysis of repetitive element DNA methylation by methylight. Nucleic Acids Res. 2005;33:6823–6836. doi: 10.1093/nar/gki987. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57.Bona N., Crossan G.P. Fanconi anemia DNA crosslink repair factors protect against line-1 retrotransposition during Mouse Development. Nat. Struct. Mol. Biol. 2023;30:1434–1445. doi: 10.1038/s41594-023-01067-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58.Tunbak H., Enriquez-Gasca R., Tie C.H., Gould P.A., Mlcochova P., Gupta R.K., Fernandes L., Holt J., van der Veen A.G., Giampazolias E., et al. The Hush Complex is a gatekeeper of type I interferon through epigenetic regulation of line-1s. Nat. Commun. 2020;11:5387. doi: 10.1038/s41467-020-19170-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59.Yang P., Wang Y., Macfarlan T.S. The role of Krab-ZFPs in transposable element repression and mammalian evolution. Trends Genet. 2017;33:871–881. doi: 10.1016/j.tig.2017.08.006. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60.Gázquez-Gutiérrez A., Witteveldt J., Heras S.R., Macias S. Sensing of Transposable Elements by the Antiviral Innate Immune System. RNA. 2021;27:735–752. doi: 10.1261/rna.078721.121. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 61.Stetson D.B., Ko J.S., Heidmann T., Medzhitov R. Trex1 Prevents Cell-Intrinsic Initiation of Autoimmunity. Cell. 2008;134:587–598. doi: 10.1016/j.cell.2008.06.032. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62.Du J., Peng Y., Wang S., Hou J., Wang Y., Sun T., Zhao K. Nucleocytoplasmic Shuttling of SAMHD1 Is Important for LINE-1 Suppression. Biochem. Biophys. Res. Commun. 2019;510:551–557. doi: 10.1016/j.bbrc.2019.02.009. [DOI] [PubMed] [Google Scholar]
  • 63.Hu S., Li J., Xu F., Mei S., Le Duff Y., Yin L., Pang X., Cen S., Jin Q., Liang C., et al. SAMHD1 Inhibits LINE-1 Retrotransposition by Promoting Stress Granule Formation. PLoS Genet. 2015;11:e1005367. doi: 10.1371/journal.pgen.1005367. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64.Orecchini E., Doria M., Antonioni A., Galardi S., Ciafrè S.A., Frassinelli L., Mancone C., Montaldo C., Tripodi M., Michienzi A. ADAR1 Restricts LINE-1 Retrotransposition. Nucleic Acids Res. 2017;45:155–168. doi: 10.1093/nar/gkw834. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65.Nurk S., Koren S., Rhie A., Rautiainen M., Bzikadze A.V., Mikheenko A., Vollger M.R., Altemose N., Uralsky L., Gershman A., et al. The Complete Sequence of a Human Genome. Science. 2022;376:44–53. doi: 10.1126/science.abj6987. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 66.Uhlen M., Quake S.R. Sequential Sequencing by Synthesis and the Next-Generation Sequencing Revolution. Trends Biotechnol. 2023;41:1565–1572. doi: 10.1016/j.tibtech.2023.06.007. [DOI] [PubMed] [Google Scholar]
  • 67.Witherspoon D.J., Zhang Y., Xing J., Watkins W.S., Ha H., Batzer M.A., Jorde L.B. Mobile Element Scanning (ME-Scan) Identifies Thousands of Novel Alu Insertions in Diverse Human Populations. Genome Res. 2013;23:1170–1181. doi: 10.1101/gr.148973.112. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 68.Wijngaard R., Demidov G., O’Gorman L., Corominas-Galbany J., Yaldiz B., Steyaert W., De Boer E., Vissers L.E.L.M., Kamsteeg E.J., Pfundt R., et al. Mobile Element Insertions in Rare Diseases: A Comparative Benchmark and Reanalysis of 60,000 Exome Samples. Eur. J. Hum. Genet. 2024;32:200–208. doi: 10.1038/s41431-023-01478-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 69.Zhang Y., Manjunath M., Kim Y., Heintz J., Song J.S. SequencEnG: An Interactive Knowledge Base of Sequencing Techniques. Bioinformatics. 2019;35:1438–1440. doi: 10.1093/bioinformatics/bty794. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 70.Jalili V., Afgan E., Gu Q., Clements D., Blankenberg D., Goecks J., Taylor J., Nekrutenko A. The Galaxy Platform for Accessible, Reproducible and Collaborative Biomedical Analyses: 2020 Update. Nucleic Acids Res. 2020;48:8205–8207. doi: 10.1093/nar/gkaa554. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 71.Keane T.M., Wong K., Adams D.J. RetroSeq: Transposable Element Discovery from next-Generation Sequencing Data. Bioinformatics. 2013;29:389–390. doi: 10.1093/bioinformatics/bts697. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 72.Delahaye C., Nicolas J. Sequencing DNA with Nanopores: Troubles and Biases. PLoS ONE. 2021;16:e0257521. doi: 10.1371/journal.pone.0257521. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 73.Rand A.C., Jain M., Eizenga J.M., Musselman-Brown A., Olsen H.E., Akeson M., Paten B. Mapping DNA Methylation with High-Throughput Nanopore Sequencing. Nat. Methods. 2017;14:411–413. doi: 10.1038/nmeth.4189. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 74.Rang F.J., Kloosterman W.P., De Ridder J. From Squiggle to Basepair: Computational Approaches for Improving Nanopore Sequencing Read Accuracy. Genome Biol. 2018;19:90. doi: 10.1186/s13059-018-1462-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 75.Veselovsky V., Romanov M., Zoruk P., Larin A., Babenko V., Morozov M., Strokach A., Zakharevich N., Khamidova S., Danilova A., et al. Comparative Evaluation of Sequencing Platforms: Pacific Biosciences, Oxford Nanopore Technologies, and Illumina for 16S rRNA-based Soil Microbiome Profiling. Front. Microbiol. 2025;16:1633360. doi: 10.3389/fmicb.2025.1633360. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 76.Yang Y., Scott S.A. DNA Methylation Profiling Using Long-Read Single Molecule Real-Time Bisulfite Sequencing (SMRT-BS) In: Kaufmann M., Klinger C., Savelsbergh A., editors. Functional Genomics. Volume 1654. Springer; New York, NY, USA: 2017. pp. 125–134. [DOI] [PubMed] [Google Scholar]
  • 77.Logsdon G.A., Vollger M.R., Eichler E.E. Long-Read Human Genome Sequencing and Its Applications. Nat. Rev. Genet. 2020;21:597–614. doi: 10.1038/s41576-020-0236-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 78.Amarasinghe S.L., Su S., Dong X., Zappia L., Ritchie M.E., Gouil Q. Opportunities and Challenges in Long-Read Sequencing Data Analysis. Genome Biol. 2020;21:30. doi: 10.1186/s13059-020-1935-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 79.Zhou W., Emery S.B., Flasch D.A., Wang Y., Kwan K.Y., Kidd J.M., Moran J.V., Mills R.E. Identification and characterization of occult human-specific line-1 insertions using long-read sequencing technology. Nucleic Acids Res. 2019;48:1146–1163. doi: 10.1093/nar/gkz1173. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 80.Chu C., Borges-Monroy R., Viswanadham V.V., Lee S., Li H., Lee E.A., Park P.J. Comprehensive identification of transposable element insertions using multiple sequencing technologies. Nat. Commun. 2021;12:3836. doi: 10.1038/s41467-021-24041-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 81.Bilgrav Saether K., Eisfeldt J. Detecting transposable elements in long-read genomes using sTELLeR. Bioinformatics. 2024;40:btae686. doi: 10.1093/bioinformatics/btae686. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 82.Nummi P., Cajuso T., Norri T., Taira A., Kuisma H., Välimäki N., Lepistö A., Renkonen-Sinisalo L., Koskensalo S., Seppälä T.T., et al. Structural features of somatic and germline retrotransposition events in humans. Mob. DNA. 2025;16:20. doi: 10.1186/s13100-025-00357-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 83.Gardner E.J., Lam V.K., Harris D.N., Chuang N.T., Scott E.C., Pittard W.S., Mills R.E., Devine S.E. The Mobile Element Locator Tool (melt): Population-scale mobile element discovery and biology. Genome Res. 2017;27:1916–1929. doi: 10.1101/gr.218032.116. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 84.Rodriguez-Martin B., Alvarez E.G., Baez-Ortega A., Zamora J., Supek F., Demeulemeester J., Santamarina M., Ju Y.S., Temes J., Garcia-Souto D., et al. Pan-Cancer Analysis of Whole Genomes Identifies Driver Rearrangements Promoted by LINE-1 Retrotransposition. Nat. Genet. 2020;52:306–319. doi: 10.1038/s41588-019-0562-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 85.Zumalave S., Santamarina M., P. Espasandín N., Zamora J., Garcia-Souto D., Temes J., Baker T.M., Rodríguez-Castro J., Otero P., Pequeño-Valtierra A., et al. Concurrent L1 Retrotransposition Events Promote Reciprocal Translocations in Human Tumorigenesis. Science. 2026;392:eaee4513. doi: 10.1126/science.aee4513. [DOI] [PubMed] [Google Scholar]
  • 86.Han S., Dias G.B., Basting P.J., Viswanatha R., Perrimon N., Bergman C. Local Assembly of long reads enables phylogenomics of transposable elements in a polyploid cell line. Nucleic Acids Res. 2022;50:e124. doi: 10.1093/nar/gkac794. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 87.Mohamed M., Sabot F., Varoqui M., Mugat B., Audouin K., Pélisson A., Fiston-Lavier A.S., Chambeyron S. Tremolo: Accurate transposable element allele frequency estimation using long-read sequencing data combining assembly and mapping-based approaches. Genome Biol. 2023;24:63. doi: 10.1186/s13059-023-02911-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 88.Groza C., Chen X., Wheeler T.J., Bourque G., Goubert C. A unified framework to analyze transposable element insertion polymorphisms using graph genomes. Nat. Commun. 2024;15:8915. doi: 10.1038/s41467-024-53294-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 89.Cuenca-Guardiola J., de la Morena-Barrio B., Corral J., Fernández-Breis J.T. Advanced analysis of retrotransposon variation in the human genome with nanopore sequencing using retroinspector. Sci. Rep. 2025;15:14489. doi: 10.1038/s41598-025-98847-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 90.Ewing A.D., Smits N., Sanchez-Luque F.J., Faivre J., Brennan P.M., Richardson S.R., Cheetham S.W., Faulkner G.J. Nanopore sequencing enables comprehensive transposable element epigenomic profiling. Mol. Cell. 2020;80:915–928. doi: 10.1016/j.molcel.2020.10.024. [DOI] [PubMed] [Google Scholar]
  • 91.Di Tommaso P., Chatzou M., Floden E.W., Barja P.P., Palumbo E., Notredame C. Nextflow enables reproducible computational workflows. Nat. Biotechnol. 2017;35:316–319. doi: 10.1038/nbt.3820. [DOI] [PubMed] [Google Scholar]
  • 92.Tarailo-Graovac M., Chen N. Using RepeatMasker to identify repetitive elements in genomic sequences. Curr. Protoc. Bioinform. 2009;25:4.10.1–4.10.14. doi: 10.1002/0471250953.bi0410s25. [DOI] [PubMed] [Google Scholar]
  • 93.Storer J., Hubley R., Rosen J., Wheeler T.J., Smit A.F. The DFAM community resource of transposable element families, sequence models, and genome annotations. Mob. DNA. 2021;12:2. doi: 10.1186/s13100-020-00230-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 94.Cheetham S.W., Kindlova M., Ewing A.D. Methylartist: Tools for Visualizing Modified Bases from Nanopore Sequence Data. Bioinformatics. 2022;38:3109–3112. doi: 10.1093/bioinformatics/btac292. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 95.Jin Y., Tam O.H., Paniagua E., Hammell M. TEtranscripts: A Package for Including Transposable Elements in Differential Expression Analysis of RNA-seq Datasets. Bioinformatics. 2015;31:3593–3599. doi: 10.1093/bioinformatics/btv422. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 96.McKerrow W., Fenyö D. L1EM: A Tool for Accurate Locus Specific LINE-1 RNA Quantification. Bioinformatics. 2020;36:1167–1173. doi: 10.1093/bioinformatics/btz724. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 97.Bendall M.L., de Mulder M., Iñiguez L.P., Lecanda-Sánchez A., Pérez-Losada M., Ostrowski M.A., Jones R.B., Mulder L.C., Reyes-Terán G., Crandall K.A., et al. Telescope: Characterization of the retrotranscriptome by accurate estimation of transposable element expression. PLoS Comput. Biol. 2019;15:e1006453. doi: 10.1371/journal.pcbi.1006453. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 98.Yang W.R., Ardeljan D., Pacyna C.N., Payer L.M., Burns K.H. SQuIRE Reveals Locus-Specific Regulation of Interspersed Repeat Expression. Nucleic Acids Res. 2019;47:e27. doi: 10.1093/nar/gky1301. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 99.Lanciano S., Cristofari G. Measuring and Interpreting Transposable Element Expression. Nat. Rev. Genet. 2020;21:721–736. doi: 10.1038/s41576-020-0251-y. [DOI] [PubMed] [Google Scholar]
  • 100.Schwarz R., Koch P., Wilbrandt J., Hoffmann S. Locus-Specific Expression Analysis of Transposable Elements. Brief. Bioinform. 2022;23:bbab417. doi: 10.1093/bib/bbab417. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 101.Lee S., Barbour J.A., Tam Y.M., Yang H., Huang Y., Wong J.W. Locusmasterte: Integrating long-read RNA sequencing improves locus-specific quantification of transposable element expression. Genome Biol. 2025;26:72. doi: 10.1186/s13059-025-03522-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 102.Berrens R.V., Yang A., Laumer C.E., Lun A.T.L., Bieberich F., Law C.T., Lan G., Imaz M., Bowness J.S., Brockdorff N., et al. Locus-Specific Expression of Transposable Elements in Single Cells with CELLO-seq. Nat. Biotechnol. 2022;40:546–554. doi: 10.1038/s41587-021-01093-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 103.Wang R., Zheng Y., Zhang Z., Song K., Wu E., Zhu X., Wu T.P., Ding J. MATES: A Deep Learning-Based Model for Locus-Specific Quantification of Transposable Elements in Single Cell. Nat. Commun. 2024;15:8798. doi: 10.1038/s41467-024-53114-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 104.The Cancer Genome Atlas Research Network. Weinstein J.N., Collisson E.A., Mills G.B., Shaw K.R.M., Ozenberger B.A., Ellrott K., Shmulevich I., Sander C., Stuart J.M. The Cancer Genome Atlas Pan-Cancer Analysis Project. Nat. Genet. 2013;45:1113–1120. doi: 10.1038/ng.2764. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 105.Chevrier S., Richard C., Mille M., Bertrand D., Boidot R. Nanopore Adaptive Sampling Accurately Detects Nucleotide Variants and Improves the Characterization of Large-Scale Rearrangement for the Diagnosis of Cancer Predisposition. Clin. Transl. Med. 2025;15:e70138. doi: 10.1002/ctm2.70138. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 106.Smith T., Heger A., Sudbery I. UMI-tools: Modeling Sequencing Errors in Unique Molecular Identifiers to Improve Quantification Accuracy. Genome Res. 2017;27:491–499. doi: 10.1101/gr.209601.116. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 107.Ou S., Su W., Liao Y., Chougule K., Agda J.R.A., Hellinga A.J., Lugo C.S.B., Elliott T.A., Ware D., Peterson T., et al. Benchmarking Transposable Element Annotation Methods for Creation of a Streamlined, Comprehensive Pipeline. Genome Biol. 2019;20:275. doi: 10.1186/s13059-019-1905-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 108.Flynn J.M., Hubley R., Goubert C., Rosen J., Clark A.G., Feschotte C., Smit A.F. RepeatModeler2 for Automated Genomic Discovery of Transposable Element Families. Proc. Natl. Acad. Sci. USA. 2020;117:9451–9457. doi: 10.1073/pnas.1921046117. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 109.Liao W.W., Asri M., Ebler J., Doerr D., Haukness M., Hickey G., Lu S., Lucas J.K., Monlong J., Abel H.J., et al. A Draft Human Pangenome Reference. Nature. 2023;617:312–324. doi: 10.1038/s41586-023-05896-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 110.Wang T., Antonacci-Fulton L., Howe K., Lawson H.A., Lucas J.K., Phillippy A.M., Popejoy A.B., Asri M., Carson C., Chaisson M.J.P., et al. The Human Pangenome Project: A Global Resource to Map Genomic Diversity. Nature. 2022;604:437–446. doi: 10.1038/s41586-022-04601-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 111.Paten B., Novak A.M., Eizenga J.M., Garrison E. Genome Graphs and the Evolution of Genome Inference. Genome Res. 2017;27:665–676. doi: 10.1101/gr.214155.116. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 112.Garalde D.R., Snell E.A., Jachimowicz D., Sipos B., Lloyd J.H., Bruce M., Pantic N., Admassu T., James P., Warland A., et al. Highly Parallel Direct RNA Sequencing on an Array of Nanopores. Nat. Methods. 2018;15:201–206. doi: 10.1038/nmeth.4577. [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

No new data were created or analyzed in this study. Data sharing is not applicable to this article. The source codes for most bioinformatic tools discussed in this review are publicly available at the following repositories: CELLO-seq: https://github.com/MarioniLab/CELLOseq (accessed on 24 August 2026); GraffiTE: https://github.com/cgroza/GraffiTE (accessed on 24 August 2026); L1EM: https://github.com/FenyoLab/L1EM (accessed on 24 August 2026); LocusMasterTE: https://github.com/jasonwong-lab/LocusMasterTE (accessed on 24 August 2026); MATES: https://github.com/mcgilldinglab/MATES (accessed on 24 August 2026); MEIGA-PAV: https://github.com/MEIGA-tk/MEIGA-PAV (accessed on 24 August 2026); MELT: https://melt.igs.umaryland.edu/downloads.php (accessed on 24 August 2026); PALMER: https://github.com/WeichenZhou/PALMER (accessed on 24 August 2026); Retroinspector: https://github.com/javiercguard/retroinspector (accessed on 24 August 2026); sTELLeR: https://github.com/kristinebilgrav/sTELLeR (accessed on 24 August 2026); SQuIRE: https://github.com/wyang17/SQuIRE (accessed on 24 August 2026); TEtranscripts: https://github.com/mhammell-laboratory/TEtranscripts (accessed on 24 August 2026); Telescope: https://github.com/mlbendall/telescope (accessed on 24 August 2026); TELR: https://github.com/bergmanlab/TELR (accessed on 24 August 2026); tldr: https://github.com/adamewing/tldr (accessed on 24 August 2026); TraDetIONS: https://github.com/panummi/TraDetIONS (accessed on 24 August 2026); TraFiC-mem: https://gitlab.com/mobilegenomesgroup/TraFiC (accessed on 24 August 2026); TrEMOLO: https://github.com/DrosophilaGenomeEvolution/TrEMOLO (accessed on 24 August 2026); xTea: https://github.com/parklab/xTea (accessed on 24 August 2026).


Articles from Biomolecules are provided here courtesy of Multidisciplinary Digital Publishing Institute (MDPI)

RESOURCES