Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2026 Apr 1.
Published in final edited form as: Fungal Genet Biol. 2026 Jan 25;183:104057. doi: 10.1016/j.fgb.2026.104057

The dynamics of transposable element content in the genome of the human pathogen Histoplasma

Tania Kurbessoian 1, David A Turissini 1, Patrick W Kelly 1, Oliver Kompathoum 1, Jonathan A Rader 2, Gaston I Jofre 3, Jingbaoyi Li 1, McKenna Sutherland 1, Victoria E Sepúlveda 1, Daniel R Matute 1,*
PMCID: PMC13036692  NIHMSID: NIHMS2144232  PMID: 41592676

Abstract

Histoplasma is a genus of human fungal pathogens that frequently affects immunosuppressed patients. Previous genetic surveys have largely focused on nucleotide-level variation, but much less attention has been given to more complex forms of mutation. Among these, transposable elements (TEs) represent an important class of mobile genetic elements that can alter genome size and play key roles in adaptation and speciation. In this study, we address this gap by examining the content and evolutionary dynamics of TEs in the human pathogen Histoplasma. Using previously published Histoplasma genome assemblies, we quantified TE content across eight phylogenetic species within the genus. Our analyses reveal heterogeneity in the evolutionary patterns of different TE families. The majority of TE orders and superfamilies show strong phylogenetic signal suggesting that phylogenetic relatedness significantly constrains the content of mobile genetic elements. We find no correlation between RNA or DNA TEs and genome size. Together, our results highlight the diverse landscape of TEs in Histoplasma and suggest that future studies should investigate their impact on genome evolution, fitness, and virulence.

Keywords: Histoplasma, genome size, transposable elements, comparative phylogenetics

INTRODUCTION

Histoplasmosis is a fungal pulmonary infection that is common in immunosuppressed patients (Azar et al., 2020; Sepúlveda et al., 2024). The infection typically occurs in the lungs, where inhaled spores or mycelium particles convert to yeast, causing symptoms ranging from mild flu-like illness to severe respiratory distress. Some patients can develop disseminated histoplasmosis as a severe, potentially life-threatening infection that spreads from the lungs to multiple organs (Adenis et al., 2014a; Linder and Kauffman, 2019; Nacher et al., 2020). Immunocompromised patients are the most susceptible—especially individuals living with AIDS, where up to 25% require prolonged antifungal therapy and the disease leads to over 700,000 deaths per year (Adenis et al., 2014b; Bongomin et al., 2019; Ca et al., 2013; Myint et al., 2020; Pasqualotto and Quieroz-Telles, 2018; Sepúlveda et al., 2024). Diagnosis often involves histopathology, culture, or antigen detection, and treatment may require administration of antifungals such as itraconazole or amphotericin B in severe cases. The disease is considered both an emergent and neglected disease, and Histoplasma was recently included in the World Health Organization’s Fungal Priority Pathogens list (Organization, 2022).

Histoplasma spp., the causal agents of histoplasmosis, are a genus of dimorphic fungi, which are characterized by their ability to transition from a saprophytic mold state to a parasitic yeast phase in a transition that is thermally regulated (Caballero Van Dyke et al., 2019; Höft et al., n.d.; Klein and Tebbets, 2007; Nemecek et al., 2006). At 25°C, Histoplasma lives as a filamentous organism in the soil, but upon entering an animal or human system where the surrounding temperature is above 37°C, the fungus converts to a virulent yeast phase (Beyhan and Sil, 2019; Sepúlveda et al., 2024; Sil, 2019). The genus is globally distributed and comprises four named species with at least six additional, as yet unnamed, phylogenetic species (Almeida-Silva et al., 2021; Li et al., 2025; Mapengo et al., 2025; Sepúlveda et al., 2017). Single isolates that remain to be classified, but seem to be divergent enough to be different species, have also been identified. Histoplasma has been collected as hyphae from bird droppings and guano-heavy soils (Mulec et al., 2020; Sudnik et al., 2024; Taylor et al., 1962), old buildings (Allard et al., 2014; Morgan et al., 2003), caves (Bernardes et al., 2025; Gugnani et al., 1994; LOTTENBERG et al., 1979; Lyon et al., 2004), and chicken coops (Benedict and Mody, 2016; de Perio et al., 2021; Norkaew et al., 2013). The fungus is additionally associated with animals such as horses (Al Mheiri et al., 2024; Al-Ani, 1999; Ameni, 2006; Rezabek et al., 1993), bats (Bernardes et al., 2025; Gugnani and Denning, 2023; Vite-Garín et al., 2021), dogs (Brömel and Sykes, 2005a, 2005b; Clinkenbeard et al., 1988), skunks (Emmons et al., 1949), and others (Guillot et al., 2018; Wilmes et al., 2022).

Several surveys have quantified genetic diversity in Histoplasma, and the majority of them have focused on detecting nucleotide variants. Yet structural variants can play important roles in evolution and remain largely underexplored in human fungal pathogens. Transposable elements, nucleic acid sequences that can change their position within a genome, comprise a major class of structural variants that are important in evolution (Oliver and Greene, 2009). TEs can affect gene diversity by introducing new chromosomal variants (Araya et al., 2025; Cáceres et al., 1999; Sharma et al., 2021), altering gene expression (Drongitis et al., 2019; Gebrie, 2023; Hirsch and Springer, 2017), potentially creating new genes (Nekrutenko and Li, 2001; Volff, 2006), and generating new adaptive variants (Casacuberta and González, 2013; Schrader et al., 2014; Schrader and Schmitz, 2019). TEs have also been hypothesized to alter genome length ((Jing et al., 2024; Sun et al., 2012) cf. (Elliott and Gregory, 2015; Kidwell, 2002)) and even to drive increased diversification rates ((Böhne et al., 2008; Carleton et al., 2020) cf. (Serrato-Capuchina and Matute, 2018)).

The proportional representation of TEs in fungal genomes shows a large range within and across species (Abraham et al., 2024; Amselem et al., 2015; McTaggart et al., 2016; Muszewska et al., 2011, 2019; Nakamoto et al., 2023; Torres et al., 2021). For example, non-pathogenic yeast species, the first to be surveyed for TEs in fungi, harbor retrotransposons (~3% of their genome, (Kim et al., 1998)) but have few DNA transposons ((Bleykasten-Grosshans et al., 2013; Chang et al., 2023) reviewed in (Bleykasten-Grosshans and Neuvéglise, 2011)). In contrast, other fungi have extremely large genomes composed of over 85% TEs (e.g., Entomophtora, (Stajich et al., 2024). TEs are key players in genome evolution and virulence determination among plant pathogens (reviewed in (Fouché et al., 2022; Möller and Stukenbrock, 2017). However, the effects of TEs on fungal human pathogens have been less directly explored. Genome surveys have indicated the existence of TEs in Histoplasma (Voorhies et al., 2022), Candida (Oggenfuss et al., 2025; Potocki et al., 2019), Coccidioides (Dubin et al., 2022; Kirkland et al., 2018a; Neafsey et al., 2010) and Cryptococcus (Goodwin et al., 2003; Gusa et al., 2023, 2020). In this latter species, TE mobilization can occur rapidly and is conditioned by temperature and genetic background (Gusa et al., 2023; Mackey et al., 2025).

In the case of Histoplasma, chromosomal assemblies of five isolates from four different species indicated differences in RNA TE content between isolates (Voorhies et al., 2022). Notably, transcripts in transposon-rich regions of the genome are more expressed in the virulent yeast phase compared to the mycelial stage (Voorhies et al., 2022). In addition to these assemblies, we also assembled available sequence data used to infer the phylogenetic history of Histoplasma (Almeida-Silva et al., 2021; Ly et al., 2025; Mapengo et al., 2025; Sepúlveda et al., 2017), making Histoplasma a promising candidate to assess the extent of TE diversity within a fungal genus of human pathogens. Here we address gaps in knowledge of human fungal pathogen TE content by studying the TE content in eight phylogenetic species of the genus Histoplasma. We report patterns of TE variation across the Histoplasma genus by analyzing 58 assembled genomes: 57 from Histoplasma and one from Blastomyces parvus as an outgroup. Our results indicate differences in TE proportion inference depending on the bioinformatic tool. We also study two hypotheses: 1) whether the presence and content of TEs follow a phylogenetic pattern, and 2) whether the variation in genome size in Histoplasma is correlated with the genomic percentage of TEs across the genus. We find support for both ideas. Our work indicates that TEs are an important source of variation and suggests the need to incorporate their characterization and scoring in surveys of genetic diversity in human fungal pathogens.

MATERIALS AND METHODS

Data

Our study leverages previously generated sequences of Histoplasma. First, we used five highly contiguous assemblies from four different species of Histoplasma (Voorhies et al., 2022). These genomes were generated with long reads and are assembled into full chromosomes. Second, we used unassembled short read sequences from 57 isolates. This sample included isolates from eight Histoplasma species–H. ohiense, H. mississippiense, H. suramericanum, Amazon-III, Africa, H. capsulatum sensu stricto (ss), mz5-like, India–and one differentiated isolate yet to be included in a phylogenetic species (B05821). We also included short reads from Blastomyces parvus (Ep_130_s_7; previously classified as Emmonsia parva) which we used as an outgroup for the comparative phylogenetic analyses. Table S1 lists the Sequence Read Archive (SRA) accession numbers for all the reads used in this study and the available metadata for each isolate.

Genome assembly and annotation

We assembled the short read sequences de novo into genomes with the AAFTF pipeline v.0.2.3 (Palmer and Stajich, 2022). This pipeline performs read QC and filtering with BBTools BBDuk (v.38.86) (Bushnell, 2014), assembles each genome using the default parameters for SPAdes (v.3.15.2) (Bankevich et al., 2012), and removes short contigs < 200 bp and contamination using NCBI’s VecScreen (Schäffer et al., 2018). Each assembled genome was then scaffolded to strains H. mississippiense WU24, H. ohiense G217B, and H. capsulatum var. dubosii H88 (Histoplasma Africa lineage, (Mapengo et al., 2025)) using Ragtag (v.1.0.2) (Alonge et al., 2022), which uses minimap2 (v. 2.17-r941) (Li, 2018) to further link scaffolds based on shared co-linearity of these isolates’ genomes. We used the ‘assess’ tool in AAFTF (https://github.com/stajichlab/AAFTF) to calculate summary statistics (L50, N50), and we estimated genome completeness with BUSCO (v5.4.4) (Manni et al., 2021) and the eurotiomycetes_odb10 database of 3,546 marker genes. Since the five chromosomal assemblies (Voorhies et al., 2022) were also sequenced with short reads in a different study (Sepúlveda et al., 2017), we determined whether the inferred size with our short read assemblies was correlated with the genome size inferred with long read assemblies using a Pearson’s correlation test (function cor.test, library stats, (R Core Team, 2018)).

We predicted the genes of 57 Histoplasma and one Blastomyces parvus short-read assemblies using Funannotate v1.8.1 (Palmer and Stajich, 2020). Masked genomes were created by first generating libraries of sequence repeats with the RepeatModeler v2.0.3 pipeline (Flynn et al., 2020). These species-specific predicted repeats were combined with fungal repeats in RepBase (Bao et al., 2015) to identify and mask repetitive regions in the genome assembly with RepeatMasker v.4-1-1 (Smit, 2004). To predict genes, ab initio gene predictors SNAP (v.2013_11_29) (Korf, 2004) and AUGUSTUS (v.3.3.3) (Stanke and Morgenstern, 2005) were trained using the Funannotate ‘train’ command, which was based on full-length transcripts constructed by a genome-guided run of Trinity v.2.11.0 (Grabherr et al., 2011) using RNA-Seq from published Histoplasma capsulatum data (SRR18230644). The assembled transcripts were aligned with PASA v.2.4.1 (Haas et al., 2008) to produce full-length spliced alignments and predicted open reading frames, which served both for training the ab initio predictors and as informant data for gene predictions.

Additional gene models were predicted by GeneMark.HMM-ES v.4.62_lic (Bruuna et al., 2020), and GlimmerHMM v.3.0.4 (Majoros et al., 2004), both of which utilize a self-training procedure to optimize gene predictions. Additional exon evidence was generated from scans of SwissprotDB proteins with DIAMOND BLASTX, and polished by Exonerate v.2.4.0 (Slater and Birney, 2005). Finally, EvidenceModeler v.1.1.1 (Haas et al., 2008) generated consensus gene models in Funannotate using default evidence weights. Non-protein-coding tRNA genes were predicted by tRNAscan-SE (v.2.0.9) (Chan et al., 2021; Lowe and Chan, 2016).

Each annotated genome was processed with antiSMASH v.5.1.1(Blin et al., 2021) to predict secondary metabolite biosynthesis gene clusters. These annotations were also incorporated into the functional annotation by Funannotate. Putative protein functions were assigned to genes based on sequence similarity to InterProScan5 v.5.51-85.0 (Mulder and Apweiler, 2007), Pfam v.35.0 (Bateman et al., 2004; Paysan-Lafosse et al., 2025), Eggnog v.2.1.6-d35afda (Huerta-Cepas et al., 2019), dbCAN2 (v.9.0) (Zhang et al., 2018) and MEROPS v.12.0 (Rawlings et al., 2014) databases relying on NCBI BLAST (v.2.9.0+) (Johnson et al., 2008) and HMMer v.3.3.2 (Potter et al., 2018). Gene Ontology terms were assigned to protein products based on the inferred homology from these sequence similarity analyses. Identification of telomeric repeat sequences was performed using the FindTelomeres.py script (https://github.com/JanaSperschneider/FindTelomeres). Briefly, this pipeline searches for chromosomal assembly with a regular expression pattern for telomeric sequences at the 5’ and 3’ ends of each scaffold. Telomere repeat sequences were also predicted using the Telomere Identification toolkit (tidk) v.0.1.5 with the “explore” option (https://github.com/tolkit/telomeric-identifier). Genome assemblies and annotation files can be found under BioProject PRJNA1118793.

Detection of Transposable elements (TEs)

Next, we annotated TEs in Histoplasma with the explicit goal of determining what percentage of the genome was composed of TEs. TE classification is dynamic and varies depending on the source (e.g., (Kapitonov and Jurka, 2008; Wicker et al., 2007). Table 1 lists the classification we used which follows the assignments from (Wicker et al., 2007). RNA Transposable Elements (RNA TEs) are split into five orders: Long Terminal Repeats (LTR), Dictyostelium intermediate repeat sequence (DIRS), Penelope-like elements (PLE), Long Interspersed Nuclear Elements (LINEs), and Short Interspersed Nuclear Elements (SINEs). The last two categories are often referred to as non-LTR retrotransposons.

TABLE 1. TE classification used in this manuscript. This classification follows (Wicker et al., 2007).

Note that MITEs are not a taxonomic classification but instead are miniaturized and non-autonomous of DNA transposons.

Class Subclass Order Superfamily
RNA LTR Copia
Ty3
Bel-Pao
Endogenous Retrovirus LTR retrotransposon (ERV)
Retrovirus LTR retrotransposon
TRIM
Solo
PLE Penelope
DIRS Ngaro
Viper
DIRS
LINE R2
Jockey
L1
I
RTE
SINE tRNA
7SL
5S
DNA Subclass 1 TIR DNA/DTA (hAT)
DNA/DTB (PiggyBac)
DNA/DTC (CACTA)
DNA/DTE (Merlin)
DNA/DTH (PIF-Harbinger)
DNA/DTM (Mutator)
DNA/DTP (P-element)
DNA/DTR (Transib)
DNA/DTT (Tc1-Mariner)
Crypton (DYC) -
Subclass 2 Helitron -
Maverick (Polinton) -
MITE - MITE/DTA
MITE/DTB
MITE/DTC
MITE/DTE
MITE/DTH
MITE/DTM
MITE/DTR
MITE/DTT

LTRs and their flanking elements have a large size range (reviewed in (Havecker et al., 2004)). LTRs usually start with 5′-TG-3′ and end with 5′-CA-3′ and typically contain open reading frames (ORFs) for GAG, a structural protein for virus-like particles, and for POL. Pol encodes an aspartic proteinase (AP), DDE integrase (INT), reverse transcriptase (RT), and RNase H (RH); POL can also contain additional ORFs. Within LTR retrotransposons we looked for five superfamilies: LTR/Ty3, LTR/Copia, LTR/Bel-Pao, LTR/ERV1, and Retrovirus superfamilies, but also report LTRs that could not be classified as ‘unknown’.

DIRS also contain a RT but instead of an INT, they have a tyrosine recombinase gene; their termini resemble either split direct repeats (SDR) or inverted repeats. PLEs have a notoriously patchy taxonomic distribution (Arkhipova, 2006), and encode a RT that is more closely related to telomerase than to the RTs of LTR elements or LINEs, along with an endonuclease related to both intron-encoded endonucleases, and the bacterial repair protein UvrC.

LINEs lack LTRs, but they can be several kilobases long. Autonomous LINEs encode at least an RT and a nuclease in their pol ORF for transposition but often carry additional genes. LINEs are common in some animals and might be prevalent in plants as well. SINEs are small and non-autonomous, and they are thought to originate from instances of retrotransposition of polymerase III (Pol III) transcripts. We also looked for Terminal-Repeat Retrotransposons in Miniature (TRIM) which are non-autonomous LTR retrotransposons with two terminal repeats and a short internal domain. Finally, we looked for LTR/Solo elements which are single LTRs where, due to homologous recombination, one LTR and the internal coding regions have been excised. Although there have been several proposals to divide LINEs into clades and families (Biedler and Tu, 2003), we did not attempt any such classification and report only the superfamily.

DNA Transposable Elements are split into two subclasses. Subclass 1 includes mostly TEs that mobilize by excising themselves from one location in the genome and inserting it into another and are characterized by terminal inverted repeats (TIRs). The nine known superfamilies are distinguished by their TIR sequences and the sequences within the TIRs. This latter category includes nine superfamilies: DNA/DTA, DNA/DTB, DNA/DTC, DNA/DTE, DNA/DTH, DNA/DTM, DNA/DTP, DNA/DTR, and DNA/DTT. Note that some methods separately list TIR/TC1_mariner, which we combined with DNA/DTT in all downstream analyses. Within Subclass 1, Wicker et al. also included DYC (Crypton) which contain a tyrosine recombinase, but lack a reverse transcriptase domain, which suggests they transpose via a DNA intermediate. DNA subclass 2 includes DNA TEs that undergo a transposition with replication but no double-stranded cleavage. Within subclass 2, we looked for autonomous Helitrons (DNA/Helitron) and Mavericks (DNA/Maverick). Helitron replicates through a rolling-circle mechanism that cuts only one DNA strand. Helitron termini are marked by simple TC or CTRR (R being a purine) motifs and a small hairpin just upstream of the 3′ end, though the hairpin is not strictly diagnostic. Autonomous Helitrons encode a Y2-type tyrosine recombinase—similar to the one in IS91 elements—along with helicase domain and replication-initiator activity. Within this category we also looked for non-autonomous Helitrons, those that lack the gene that encodes the functional recombinase/helicase transposase. Mavericks (also known as Polintons) are large 10–20 kb TEs flanked by long TIRs. They encode up to 11 proteins, though gene content and order vary, and some show limited similarity to DNA virus proteins. Mavericks carry DNA polymerase B, and a c-int–type integrase, but lack reverse transcription ability, implying a replicative transposition mechanism without an RNA intermediate—likely via single-strand excision, extrachromosomal replication, and reintegration.

Within DNA transposons, we also inspected for the presence of Miniature inverted–repeat transposable elements (MITEs) which are truncated derivatives of autonomous DNA transposons (Feschotte and Mouches 2000; Zhang et al. 2000; Yang et al. 2001; Feschotte et al. 2003; Yang and Hall 2003).

Given the variation among different TE annotation methods (Gozashti and Hoekstra, 2024; Loreto et al., 2023), we used four pipelines to determine whether our results were consistent regardless of the methodology.

1. EDTA-only.

First, we built a TE library using the extensive de novo TE annotator (EDTA) v.2.1.3 (Ou et al., 2019). EDTA is a computational suite that integrates a variety of computational packages (LTR_FINDER, LTRharvest, LTR_retriever, Generic Repeat Finder, TIR Learner, HelitronScanner, and RepeatMasker) to produce a nonredundant TE library for each genome. Using this library as an input, we ran each of the 58 genome assemblies with the following parameters: “--species other, --step all, --sensitive 1, --force 1, --anno 1, --evaluate 1 --cds $CDS” where $CDS was a fasta file containing CDS sequences for Histoplasma capsulatum G186A. EDTA provides a variety of outputs including TE sequences in FASTA format, and TE positions in GFF3 format.

2. EDTA-RepeatMasker (EDTA-RM).

We ran EDTA to build the TE library and then ran RepeatMasker (RM, (Smit, 2004)) to identify putative TEs. The goal of this additional step was to determine whether RepeatMasker provided more resolution than EDTA. RepeatMasker (v.4.1.5) (Flynn et al., 2020; Smit, 2004) was run using the library generated by EDTA with the following options: “-lib (EDTA generated library) -a -s -e rmblst”.

3. dnaPipeTE RM lib.

We used an additional method to identify TEs, dnaPipeTE v1.4c (Goubert et al., 2015). This approach differs from EDTA-based methods (i.e., EDTA-only and EDTA-RM) in that dnaPipeTE does not require an assembled genome. Instead dnaPipeTE does a de novo assembly of raw sequencing reads at a very low genomic coverage and then maps the assembled contigs to a TE library. We used the recommended RepeatMasker (v.4.1.5) TE library. We ran dnaPipeTE for each sample using the length of our assembled genome for that isolate (Table S2) and a range of different genomic coverage parameters to determine the most appropriate parameters for our analyses. We found that, regardless of TE library, the inferred genomic TE percentage increased as the genome coverage parameter increased but dropped slightly for genome coverage = 2x, and that the average number of unique TE subclasses identified increased as the genome coverage increased (Table S3). Subsequent analyses used the following dnaPipeTE parameters:s “-genome_coverage 1 -sample_number 5 -RM_t 0.2”.

4. dnaPipeTE EDTA lib.

We used the same approach described in dnaPipeTE RM lib but instead of using the library generated by RepeatMasker, we used the TE libraries generated by EDTA for each isolate. Otherwise, the parameters were the same as for dnaPipeTE RM lib.

Phylogenetic signal and models of evolution of genome length and TE content

Genome size has been hypothesized to strongly reflect phylogenetic history (Kidwell, 2002), however recent TE invasion can lead to the loss of this phylogenetic structure (Amselem et al., 2015b). Comparative phylogenetics provides the tools to determine whether TE content in the genome follows the genealogical relationships of Histoplasma isolates or deviates therefrom. Please note that given the diversity in the quality of our genome assemblies, we focused on the amount of TE in each genome and not on the genealogical history of each type of element. We describe each of the steps as follows.

i). Phylogenetic tree

We generated a phylogenetic tree for Histoplasma following the methodology described in Jofre et al. (2022) to map sequencing reads and identify single-nucleotide polymorphisms (SNPs). Raw FASTQ files were aligned to the H. mississippiense WU24 reference genome (Voorhies et al. 2022) using BWA v0.7.15 (Li and Durbin 2010). Resulting BAM files were sorted, cleaned, and duplicates marked using Samtools v1.11 (Li et al. 2009). SNPs were called with GATK HaplotypeCaller v4.1.7.0, with the ploidy parameter set to 1 (McKenna et al. 2010). The resulting gVCF files were merged using GATK GenomicsDBImport and joint-genotyped with GenotypeGVCFs with ploidy set to 1. We retained only biallelic SNPs using bcftools (Li 2011). Variant filtering was performed with GATK VariantFiltration using the following thresholds: QD < 2.0, FS > 60.0, MQ < 30.0, MQRankSum < −12.5, and ReadPosRankSum < −8.0. A single multi-sample VCF file was generated from all sequences. We then constructed a concatenated whole-genome maximum likelihood (ML) phylogeny. We converted our multi-sample VCF file to a genome-wide concatenated alignment using the Python script vcf2phylip (Ortiz 2019). The ML tree was then inferred with IQ-TREE 2 (Nguyen et al. 2015; Minh et al. 2020) using the GTR+G substitution model. Branch support was evaluated with 1,000 ultrafast bootstrap replicates (Hoang et al. 2018). We used the resulting tree for all comparative approaches described below.

ii). Comparisons among species

We used two different approaches to compare genome size and TE content among species. First, we used factorial ANOVAs (function lm, library stats, (R Core Team, 2018)) in which the genome size, DNA or RNA TE content were the response, Histoplasma species was a fixed effect, and the BUSCO completeness, a proxy of genome assembly quality, was a covariate. We followed this analysis with Tukey’s Honestly Significant Difference (HSD) test for pairwise comparisons among species (function glht, library multcomp, (Hothorn et al., 2016, 2012)). To account for the effect of phylogenetic relationships on the extent of differentiation between isolates, we used a simulation-based phylogenetic ANOVA (Garland et al. 1993) using the function phylANOVA (library phytools, (Revell, 2012; Revell and Revell, 2014)).

We calculated the correlation between different types of TE content using the function cor (library stats, (R Core Team, 2018)); we used the function corrplot (library corrplot, (Wei et al., 2017)) to plot the correlation matrix. To visualize TEs across species of Histoplasma, we used principal component analysis (PCA). We used the function prcomp (library “stats”; (R Core Team, 2018)) to calculate the PCA loadings. To visualize the PCA, we used the function fviz_pca_ind (library “factoextra”; (Kassambara and Mundt, 2017)). The 95% confidence interval for the TE content distribution of each species was plotted using the option ellipse.type with a multinomial distribution.

iii). Phylogenetic Signal and Evolutionary Models

Next, we investigated whether TE content across the Histoplasma genus followed a phylogenetic conservatism pattern (i.e., the tendency for closely related species to share similar mean trait values) or if, on the contrary, they had diverged more than expected from their phylogenetic relationships. We used the phylogenetic tree (57 isolates of Histoplasma and one isolate of B. parvus; described immediately above) and the TE content inferred by EDTA-RM, to measure phylogenetic signal in TE content treated as a trait using two different, and complementary metrics: Pagel’s λ (Pagel, 1999) and Blomberg’s K (Blomberg et al., 2003). The former measures the extent to which related species resemble each other more than expected by chance, and ranges from 0 (traits evolve independently of phylogeny) to 1 (traits are fully influenced by phylogeny); the latter also indicates how closely related species resemble each other, but with an explicit comparison to a Brownian motion null model of evolution. A λ value close to 1 suggests trait variation aligns with phylogeny, while values near 0 imply weak phylogenetic signal, with traits evolving independently of relatedness. Blomberg’s K provides a subtly different and complementary interpretation. When K = 1, the observed phenotypic values reflect a Brownian motion evolutionary process. A value of K < 1 indicates that species are less similar than expected, and is considered evidence against phylogenetic conservatism (Cooper et al., 2011; Losos, 2008; Wiens et al., 2010). By contrast, a value of K > 1 suggests greater trait similarity than would be predicted by a Brownian process. We used the function phylosig (library “phytools” v.1.0–3 (function, (Revell, 2012; Revell and Revell, 2014)) to calculate phylogenetic signal for genome size, RNA TE, and DNA TE content. We also calculated the phylogenetic signal for each order and superfamily of TEs described above.

Finally, we investigated the pattern of TE evolution across the Histoplasma phylogeny. To do this, we fit five alternative models of trait evolution using the function fitcontinuous (library “geiger”; (Harmon et al., 2015, 2008; Pennell et al., 2014)): white noise, Brownian motion (BM), Ornstein–Uhlenbeck (OU) process, Early branching (EB), and Kappa. Descriptions of all five models are provided in Table 2. These models represented a range of evolutionary scenarios, from no phylogenetic signal (i.e., white noise) to models with varying evolutionary modes and tempos. We fit each of these models to genome size, RNA TE content, and DNA TE content, and qualitatively compared the best-fit model among these three traits. To assess model performance we compared the Akaike Information Criterion (AIC) values for each model, and estimated their relative support using Akaike weights (wAIC). These weights represent the relative likelihood of each evolutionary model normalized relative to each other.

TABLE 2. Trait-evolution models suggest that genome size in Histoplasma evolves following a OU model.

BM: Brownian motion, OU: Ornstein–Uhlenbeck, EB: Early burst

Model Model description σ2 Z0 Additional parameters AIC wAIC
White noise Trait values are drawn from a normal distribution, assuming no covariance among species. 22854813386382.375000 27286685.569 1952.906 < 0.0001
BM The correlation among trait values is proportional to the shared ancestry between pairs of tips (Felsenstein, 1973). 1036381135067571.875 26104147.054 1917.545636 0.252
OU Trait evolution is best described by a random walk with a central tendency, where the strength of attraction is governed by the parameter α, ranging from ~0 to 2.72 (Butler and King, 2004). 1088460801721879.500 26384062.407 α = 2.718 1915.808 0.537
EB The rate of evolution is modeled as an exponential function of time: rt = r0 × e(a × t), with r0 representing the initial rate, a is the parameter governing rate change, and t time. 1371384672604576.7500 26130875.02 a = −2.199956 1918.650 0.129
Kappa Character divergence is linked to speciation events in the phylogeny (Pagel, 1999). κ ranges between ~0 and 1 1036381307234625.625 26104147.054 κ = 1.000 1919.546 0.083

Correlation between genome length and TE content

TE content has been proposed to be correlated and even causal of enlarged genomes across taxa. We thus evaluated whether DNA and RNA TE content were correlated with genome length. We used two complementary approaches. First, we fit ordinary least squares (OLS) regressions using the function lm (library ‘stats’, (R Core Team, 2018)) with genome length as the response and the content of DNA or RNA TEs as the explanatory variable. Second, we fit phylogenetic generalized least squares (PGLS) using the function pgls(library ‘caper, (Freckleton et al., 2002) REF); genome length was the response, and TE content, either DNA or RNA, was the explanatory variable. We used the phylogenetic tree described above (Phylogenetic signal and models of evolution of genome length and TE content: Phylogenetic tree ).

RESULTS

Histoplasma genome assemblies

We assembled the genomes of 57 Histoplasma isolates using previously published chromosomal assemblies (Voorhies et al., 2022) as reference genomes. The depth of coverage ranged from 4x-170x coverage across all 57 isolates. We calculated three assembly statistics for each of the 57 isolates: L50, N50, and BUSCO gene completeness (Figure 1, Table S2). In general, we detected differences among species in all summary assembly statistics (Table S4). Genome completeness ranged from 50.3% (505, H. ohiense) to 97.6% (WGS_S11 and 343_p_18_S1, India) and significantly differed among species (One way ANOVA: F9,48 = 19.16, P = 4.51 × 10−13). Of the eight phylogenetic species, India showed more complete genomes than the other species (96.4%, Figure 1A), while mz5-like showed the lowest genome completeness (63.75%). (The assembly of the isolate B05821 contained 60.60% of the BUSCO genes.) Similarly, genome contiguity–as measured by N50–differs between species, with the more complete genomes (India, Africa) being more contiguous than species with lower BUSCO completeness (Figure 1B). Tables S5 and S6 show all the pairwise comparisons between species. This level of assembly quality variance indicates that genome assembly might be a source of variation when studying genome features, and for that reason we use N50 as a covariate in the forthcoming linear regressions to study genome size (See below).

FIGURE 1. Metrics of genome assembly in eight phylogenetic species of Histoplasma (ten lineages).

FIGURE 1.

A. Percentage of BUSCO genes that are complete in the assembly. B. N50 as a metric of genome contiguity.

We used these assemblies to generate a phylogenetic tree from concatenated markers with the 57 samples included in the study (Figure 2A). The resulting tree recapitulates the phylogenetic relationships between isolates that have been previously described for Histoplasma, and all the eight previously described species appear as monophyletic groups (Figure 2A, (Jofre et al., 2022; Li et al., 2025; Mapengo et al., 2025; Sepúlveda et al., 2017)).

FIGURE 2. Maximum likelihood tree and distribution of genome size in Histoplasma. evolution and TEs content.

FIGURE 2.

A. A maximum likelihood phylogeny showing the relationships between Histoplasma isolates. B. Genome size estimates for each included isolate. C. Boxplot for genome size for each Histoplasma species. D. Heatmap showing the significance of pairwise comparisons. H. miss: H. mississippiense, H. caps: H. capsulatum sensu stricto, H ohie: H. ohiense, Ama_II: Amazon-II, H. sura: H. suramericanum.

The differences in genome size among Histoplasma species are explained by phylogenetic structure

Previous studies suggested variation in the genome size of Histoplasma (Carr and Shearer, 1998). We followed up on this observation using the genome assemblies described above. Genome size ranged between 19.67Mb (FGPOE2043, an mz5-like isolate) and 47.94Mb; FGLIN2055, an Amazon-III isolate). Note that the variance within species is considerable (Figures 2A2B), but the biggest differences occur across species boundaries (Figure 2B). The inferred genome size of the short-read assemblies is correlated to the inferred genome size previously inferred with long-read assemblies (N = 5, ρ = 0.964, P = 0.008), indicating that short-read assemblies can broadly recapitulate the genome size variance in the clade. An ANOVA that included BUSCO completeness as a covariate suggested strong differences between species (F9.47 = 9.714, P = 3.356 × 10−8; ANOVAs using N50 or L50 showed a similar result). The species with the largest mean genome size was Africa (mean = 32.99Mb; sd = 5.64Mb; Figures 2B2C), and the species with the smallest mean genome size was H. suramericanum (mean = 22.4Mb, sd = 0.24Mb). The differences between species persisted even after we removed the three outliers that show genome sizes larger than 33.5Mb (F1,44=13.503, P= 1.241× 10−9), but not phylogenetic corrections (PhyloANOVA, F = 3.364, P = 0.993). Table S7 shows all the pairwise comparisons between species. In general, these results support the hypothesis of interspecific differences in genome size between Histoplasma species.

Next, we studied the evolution of genome size in Histoplasma. Metrics of phylogenetic signal showed evidence of conservatism. Pagel’s λ for genome size (N=58, λ > 0.999) is significantly larger than 0 (P = 7.987 × 10−15, Figure 3A) and close to the maximum (1.00), thus indicating evidence of conservatism. Blomberg’s K (K = 0.152) was significantly lower than 1 (P < 0.001) and significantly larger than 0 (P = 1 × 10−5, Figure 3B), suggesting that genome size among Histoplasma isolates resembled each other less than expected under Brownian motion evolution. In light of these results, we explored which evolutionary model best explained the evolution of genome size in Histoplasma. Among the five evolutionary models we tested, the model with the best support was OU, which suggests that genome evolution is best explained by an Ornstein–Uhlenbeck process (Table 2). This result suggests that genome size is either evolving toward one or more evolutionary optimum, or stabilizing around existing optima.

FIGURE 3. Phylogenetic signal of genome size.

FIGURE 3.

Pagel’s λ and Blomberg’s K are two metrics of phylogenetic signal of genome size along the Histoplasma phylogeny. A. Pagel’s λ value. Estimated likelihood of Pagel’s λ values. The dotted line marks the maximum likelihood estimate (MLE) of λ. B. Blomberg’s K. The histogram shows the distribution of simulated values of Blomberg’s K; the blue arrow shows the observed value of Blomberg’s K.

Method comparison

We used four different approaches to detect TEs in Histoplasma. Table 3 shows the TE genomic percentages that were inferred by each of the four methods. The two EDTA-based methods and dnaPipeTE EDTA lib show variable levels of concordance, but overall they all report the presence of DNA, RNA TEs, and MITEs. dnaPipeTE EDTA lib identifies a higher proportion of RNA TEs than the EDTA methods. dnaPipeTE RM lib detected a much lower genomic percentage of both DNA and RNA TEs than any of the other three methods and also found no MITEs. All methods detect complete TE copies (Table 4, Tables S810).

TABLE 3. Genomic percentages of TEs and low complexity motifs inferred by four computational pipelines.

Class Average genomic percent EDTA-RM Average genomic percent EDTA-only Average genomic percent dnaPipeTE EDTA lib Average genomic percent dnaPipeTE RM lib
DNA TE 1.219% 1.637% 1.231% 0.071%
Low complexity 0.318% 0.0% 0.200% 0.149%
MITE 0.019% 0.018% 0.021% 0.0%
RNA TE 8.713% 8.649% 13.728% 1.147%
Satellite 3.0 × 10−5 % 0.0% 2.310 × 10−4% 0.015%
Simple repeat 1.313% 0.0% 1.066% 1.307%
Unknown 3.851% 3.517% 6.123% 0.018%

TABLE 4. TE completion by class for EDTA-RM.

Note that the categories are nested. Tables S8S10 show similar analyses with other methods to detect TEs.

Class Superfamily per_100 per_gt_95 per_gt_90 per_gt_50
DNA TE DNA/DTA 1.6 1.7 1.7 12.5
DNA TE DNA/DTB 21.7 33.3 44.2 66.7
DNA TE DNA/DTC 3.8 8.8 9.9 26.7
DNA TE DNA/DTH 3.2 4.3 4.9 20.5
DNA TE DNA/DTM 2.1 3.8 4.1 15.8
DNA TE DNA/DTT 7.6 12.2 14.2 41.1
DNA TE DNA/Helitron 0.6 0.9 1.3 5.6
RNA TE LINE/unknown 4.8 7.7 8.9 35
RNA TE LTR/Copia 18.6 22 24 48.3
RNA TE LTR/TRIM 0 0 0 0
RNA TE LTR/Ty3 10.6 13.3 15.6 34.3
RNA TE LTR/unknown 10.9 14.8 16.4 39.8

The results from dnaPipeTE are contingent on the specified TE library (dnaPipeTE EDTA lib vs dnaPipeTE RM lib; Table 3), and dnaPipeTE infers a lower genomic percentage of all TE subclasses when we use the default RM library. EDTA-only and EDTA-RM inferred similar genomic proportions of DNA and RNA TEs across isolates, and the inferred proportions across individuals are highly correlated between the two methods (Figure 4, DNA TEs: R2 =0.985, P < 1 × 10−10, RNA TEs: R2 = 1, P < 11 × 10−10). This is an expected result given the fact the two pipelines share the same initial analytical steps. The correlation is lower, but still significant, between the proportions inferred by the two EDTA-based pipelines, and dnaPipeTE (all P < 1 × 10−10; Figure 4BC). The correlation between the inferred genomic percentage of DNA and RNA TEs when estimated at the individual isolate level using the two dnaPipeTE-based methods was low for DNA TEs (slope = 2.95×10−3, R2 = 6.9×10−3, P = 0.539), and high for RNA TEs (slope = 0.078, R2 = 0.918, P < 1 × 10−30; Figure 4). In both cases, the slope of the regression was much lower than the expected value of 1. The correlations between the inferred proportions by the EDTA-based methods and dnaPipeTE RM library are much lower (Figure 4).

FIGURE 4. Correlations between methods.

FIGURE 4.

EDTA-based methods show strong levels of correlation across samples. The twelve panels show the six possible pairwise comparisons among the four methods to detect TEs for the two classes of TEs (DNA and RNA TEs).

Since dnaPipeTE is designed for instances of low coverage, which do not apply to the Histoplasma samples in our study, and because dnaPipeTE seems to strongly depend on sampling parameters, we proceeded with our study using the inferences from EDTA-RM but report the results of all the methods where relevant.

TE content patterns

We characterized the relative abundance of TEs in Histoplasma by partitioning the TE content in each Histoplasma genome into TE classes. While the results of the EDTA-based approaches are generally consistent, EDTA-RM identifies a higher genomic percentage of RNA TEs (8.79% vs 8.72%) but also a lower percentage of DNA TEs (1.23% vs 1.66%) than EDTA-only.

Figure 5 shows the DNA and RNA TE genome proportions across phylogenetic species. H. suramericanum (median = 2.530%, 1 isolate) and Amazon-II (median = 2.177%, 3 isolates) had the highest mean DNA TE content, and India had the lowest (median = 0.5%, 15 isolates; Figure 5A). A two-way ANOVA with BUSCO completeness as a covariate indicated differences between species in DNA TE content (F9,47 = 2.866, P = 8.847 × 10−3). Pairwise comparisons indicated Africa had higher DNA TE content than H. mississippiense and India; Figure 5B). The heterogeneity remained after removing outliers (F8,44=2.347, P=0.034), but not after correcting for phylogeny (F = 2.298, P = 0.997).

FIGURE 5. RNA and DNA TE content.

FIGURE 5.

A. DNA TE content across species of Histoplasma. B. Heatmap showing the results from all pairwise comparisons among species in DNA TE content (P-value from a Tukey HSD test). C. RNA TE content across species of Histoplasma. D. Heatmap showing the results from all pairwise comparisons among species in RNA TE content (P-value from a Tukey HSD test).

RNA TEs were generally — but not universally — more abundant than DNA TEs across Histoplasma species (Figure 5C). RNA TEs also show differences among species in two tests: a two-way ANOVA (F8,44 = 53.510, P < 1 × 10−10), and an ANOVA that excluded outliers (F1,48 = 20.190, P = 1.745 × 10−13), but not in a phylogenetically informed ANOVA (F = 54.640, P = 0.077). B05821 (median = 45.43%, 1 isolate) and Africa (median = 22.51%, 6 isolates) had the highest RNA TE content, and India had the lowest (median = 0.344%, 15 isolates; Figure 5C). Pairwise comparisons indicated that B05821 and Africa had higher RNA TE content than other Histoplasma species; Figure 5D). Table S11 shows the results for linear models using EDTA and dnaPipeTE, which are qualitatively similar.

We also studied whether Histoplasma species differed in their TE content at the level of TE order and superfamilies. A previous set of chromosomal assemblies revealed differences in TE content between isolates from different species (Voorhies et al., 2022). We expanded these results and used our genome assemblies to identify TEs and compared them across the Histoplasma phylogenetic species. The only element detected by EDTA-RM and not by EDTA-only was LTR/TRIM which was detected at low levels in eleven out of sixteen isolates from the India phylogenetic species (Figure 6). No element was detected by EDTA-only but not by EDTA-RM. Figures S1S3 show the abundances of the inferred TEs by other methods. Of note, dnaPipeTE RM lib detected evidence for other TEs not detected by the methods that used the EDTA library (e.g., Mavericks, SINE) at low frequency (Figure S3). All methods inferred that the most common TE order in the Histoplasma genome are LTRs, consistent with previous chromosomal assemblies of five Histoplasma isolates (Voorhies et al., 2022). We found evidence for two LTR superfamilies (Ty3 and Copia) and for uncharacterized LTRs, with Ty3 being the most common of the two classified superfamilies. (dnaPipeTE RM lib finds indications of other LTRs at low abundance, Figures S3.) All isolates harbored both Ty3 and Copia, with just a few exceptions (e.g., HC1070058, Belem3).

FIGURE 6. DNA and RNA orders and superfamily.

FIGURE 6.

The color in the heatmap shows the genome proportion of TE types (orders and superfamilies) as inferred by EDTA-RM in each of the 57 isolates of Histoplasma and the outgroup, B. parvus. Columns show TE orders and superfamilies; rows show isolates. Figures S1S3 shows similar analyses with other tools to infer TE proportions.

We find evidence for LINEs representing up to 3.1% of the genome (e.g., CI_10, H. ohiense), but LINEs are also absent in several isolates from the Indian phylogenetic species per all TE-detection methods (Figure 6; Figures S1S3). Within DNA TE subclass 1, we found evidence for five orders of TIRs (DNA/DTA, DNA/DTC, DNA/DTH, DNA/DTM, DNA/DTT; Figure 6) but not for Cryptons. Within DNA TE subclass 2, we found evidence for Helitrons, which in some cases can represent close to 3% of the genome (e.g., H. ohiense isolates G217B and G222B and Africa isolate ES90). Tables S12 and S13 show the results of linear models comparing the content of each superfamily across species while controlling for genome assembly quality. DNA/DTC, DNA/DTH, LINE, LTR/Ty3, LTR/unknown, MITE/DTA, MITE/DTH, and MITE/DTT all differed between species in both EDTA-based methods. This significance did not persist for any of the TE orders after accounting for phylogenetic relationships.

Pairs of TE orders and superfamilies showed high levels of positive (e.g., DNA/DTC and MITE/DTT: ρ =0.61) and negative correlations (e.g., LTR.TRIM and LTR.unknown: ρ = −0.390; Figure 7A). We used a Principal Component Analysis (PCA) based on TE abundance to determine if Histoplasma species differ in their TE load. The majority of the variance (57.33%) is explained by the first five PCs. PC1 explained 18.0% of the variance; PC2 explained 11.4% (Figure 7B). The largest contributor to this abundance differentiation in PC1 were LTRs, while the largest contributors for PC2 were MITE/DTT and DNA/DTC (Figure 7C, Figure S4). The combination of these two PCs leads to a modest differentiation between some species pairs (Figure 7D). For example, India and Africa show little overlap, but H. capsulatum ss shows overlap with all of the other species. PC3 and PC4 together explain ~18.2% of the variance, and unlike the PC1/PC2 combination, all the species show highly overlapping distributions (Figure S5). The largest contributors to these PCs were MITE/DTA and DNA/DTH (Figure S5CD). Collectively these results indicate differences in TE abundance across individual isolates and in some instances modest differentiation across species of Histoplasma. PCAs of results from the other TE-detection methods also showed no large-scale differentiation in TE abundances among Histoplasma species (EDTA: Figures S6S8, dnaPipeTE-RM lib: Figures S9S12, dnaPipeTE-EDTA lib: Figures S13S14) which in turns suggest that regardless of the approach to identify TEs, TE abundance is not a diagnostic trait for any of the species

FIGURE 7. Multivariate analyses of TE content in Histoplasma inferred with EDTA-RM.

FIGURE 7.

A. Pairwise correlations between the TE orders and superfamilies. B-D. Principal component analyses to study the relative contribution of different TEs frequency to genomic differences among Histoplasma isolates. B. Scree plot showing the contribution of the different PCs. C. PCA loadings. D. Biplot of PC1 and PC2 showing the contributions of each TE. Figure S5 shows similar information for PC3 and PC4.

Phylogenetic signal of TE content

We also measured the phylogenetic signal of DNA and RNA TE content. Strong phylogenetic signal would indicate that closely related isolates, and species, are more similar in their TE proportions than more distantly related isolates and species. Low values of phylogenetic signal would suggest that variation in Histoplasma TE content reflects distinct expansions and/or contractions along the phylogeny. The Pagel’s λ of TE content when using EDTA-RM to estimate TE proportion was close to one in both cases (Pagel’s λDNA-TEs = 0.972, P = 2.603 × 10−10; Pagel’s λRNA-TEs = 0.995, P = 2.018 × 10−25) indicating that, for both classes, TE-content variation is structured by the phylogenetic relationships of the isolates (Figure 7A,C).

The results from Blomberg’s K differed between the two classes of TEs (Figure 7B,D). In the case of DNA TEs, K was low (K = 0.089) but still significantly higher than zero (P = 6.60 × 10−4); in the case of RNA TEs, K was higher (0.636) but also significantly different from 1 (P = 1 × 10−5). Other TE detection methods revealed a qualitatively similar result (Table S14). Thus, while Pagel’s λ confirmed the influence of phylogenetic history on the evolution of Histoplasma TE content and the necessity of accounting for it in our statistical tests (Pagel, 1999), Blomberg’s K suggested the possibility of distinct evolutionary histories between TE classes (Blomberg et al., 2003).

Therefore, we next identified the best-fitting evolutionary model for the two classes of TEs. In the case of DNA TEs, the overwhelmingly best fit corresponded to a Kappa model in which DNA TE total content follows a punctuational model of trait evolution (Table 5). EDTA-only and dnaPipeTE-RM showed similar results (Table S15). Collectively, these results indicate that different TEs classes have experienced different evolutionary trajectories in Histoplasma.

TABLE 5. Trait-evolution models for DNA and RNA TEs content in Histoplasma.

Abbreviations and descriptions of each model are listed in the legend of Table 1.

TE type Model σ2 Z0 Additional parameters AIC wAIC
DNA White noise 1.602 1.213 NA 195.945 8.378 × 10−4
DNA BM 129.167 1.350 NA 194.168 0.002
DNA OU 131.592 1.272 α= 2.718 191.110 0.010
DNA EB 129.146 1.350 a = −1.00 × 10−6 196.168 6.690 × 10−4
DNA Kappa 6.420 1.058 κ = 0.472873 181.577 0.986
RNA White noise 85.414 8.668 NA 426.553 1.549 × 10−25
RNA BM 1016.915 12.449 NA 313.847 0.461
RNA OU 1090.897 12.296 α=1.953 314.962 0.236
RNA EB 1016.923 12.449 a = −1.00 × 10−6 315.847 0.152
RNA Kappa 1016.915 12.449 κ = 1.000 315.847 0.152

Next, we studied whether each individual type of TE showed a phylogenetic signal (Figure 9). For this portion of the study, we focused on the results obtained with EDTA-RM and EDTA-only (Figure S15) due to the sampling effects observed in methods based on dnapipeTE. Pagel’s λ estimates for the seven out of eight orders of DNA (all but DNA.DTB), all the five RNA TE categories, and four MITES (all but MITE.DTM) were significantly larger than zero. Blomberg’s K for individual TE order and subfamilies was significant in about half the evaluated TEs. Phylogenetic signal estimates were essentially identical between EDTA-only and EDTA-RM results (Table S16).

FIGURE 9. Pagel’s λ (A) and Blomberg’s K (B) for the TE orders and superfamilies detected by EDTA-RM.

FIGURE 9.

Figure S15 shows the results for EDTA-only.

Genome size and TEs in Histoplasma

Finally, we tested whether there was a relationship between TE content and genome size in Histoplasma. We used four different methods to infer TE content, and three types of linear regressions, accounting for differences in genome assembly and phylogenetic history. Neither DNA TE nor RNA TE content was positively correlated with genome size (PGLS, Figure 10). None of the regressions showed a significant positive effect of TE content in genome size after correcting for multiple comparisons (Table S17).

FIGURE 10. Correlations between genomic TE percentages and genome length.

FIGURE 10.

Right panels: DNA TEs. Left Panels: RNA TEs.

DISCUSSION

Transposable elements (TEs) are structural variants that can play important roles in evolution. Besides inducing structural changes in the genome, TEs can be mutagenic (Chao et al., 1983; Lambert et al., 1988), give rise to regulatory regions (Feschotte, 2008; Jordan et al., 2003; Sundaram et al., 2014), and facilitate gene duplication (Cerbin and Jiang, 2018; Krasileva, 2019). TEs can also generate adaptive variants (Casacuberta and González, 2013; Lin et al., 2024; Nakamoto et al., 2023; Singh et al., 2021; Stapley et al., 2015) and promote recombination events (Huang et al., 2025; Kent et al., 2017; Pettrich et al., 2024; Rizzon et al., 2002), and they have been implicated in the emergence of inversions (Araya et al., 2025; Cáceres et al., 1999). Further understanding the importance of TEs to evolution requires surveys of TE content across taxa. In this report, we characterize and describe TE content across species of the genus Histoplasma, a human pathogen of critical public health importance.

First, we studied the extent of variance in genome size and established the mode of evolution of this trait. Pulse field electrophoresis previously revealed differences in chromosome size between isolates of Histoplasma (Canteros et al., 2005; Steele et al., 1989), but this predated the description of cryptic phylogenetic species in Histoplasma (Kasuga et al., 2003; Sepúlveda et al., 2017). Thus, the question of genome size differences along species boundaries remained unclear. Our sampling revealed genome size differences among species of Histoplasma, as well as the existence of size variance within each species. The strong phylogenetic signal for genome size, which reflects similarity among closely related isolates and species, is similar to patterns observed in palms (λ = 0.933, (Schley et al., 2022)), diatoms (λ =0.998, (Roberts et al., 2024)), crustaceans and insects (both λ > 0.98 (Alfsnes et al., 2017)). Also similar to our findings in Histoplasma, palm genome size evolution is best supported by an Ornstein–Uhlenbeck model which indicates evolution towards trait optima across the tree (Schley et al., 2022). Since eukaryote genome size shows a proportional model of evolution and larger genomes show faster rates of change (Oliver et al., 2007), determining the range of genome size in pathogenic fungi might also reveal variation in rates of genome evolution.

We also quantified the abundance of DNA and RNA TEs in short-read assemblies of Histoplasma. Previous measurements of the relative abundance of TEs in a set of five chromosomal assemblies from four species found differences between isolates (Voorhies et al., 2022). However, since these studies only used one or two isolates per species, these differences could not be ascribed to species boundaries. Here we expanded these analyses using genome assemblies derived from short reads for 57 Histoplasma isolates. While the general quality of these assemblies was predictably lower than those generated with long-reads, they were sufficient to detect TEs. Our findings revealed differences in TE content between species when we pool TE types (i.e., DNA TE and RNA TE as a whole) and among the individual TE varieties. Some TE types are mostly restricted to a single species (e.g., DNA/DTA, LTR/TRIM, MITE/DTH) which might indicate TE invasions since speciation. These results indicate that these TEs might have crossed species boundaries and that these TEs might have been introduced to Histoplasma from a different genus (or unsampled species of Histoplasma) through horizontal gene transfer. The dynamics of TE transmission and inheritance, especially across species boundaries, is an active area of research in multiple taxa, and more dense sampling of Histoplasma and other related fungi might reveal the directionality and timing of transfer. Nonetheless, establishing the source of these invasions is not always straightforward (e.g., P-elements–DTB– in Drosophila, (Engels, 1992; O’Brochta and Handler, 1988)).

Comparative approaches have the potential to reveal the extent to which TEs shape genome evolution and whether TE-mediated changes are lineage specific. Our findings indicate that TE content and genome length are not correlated within the genus Histoplasma. In contrast, TE amplifications are correlated with genome size in basidiomycetes (Castanera et al., 2017), and obligate biotrophs have been hypothesized to have more TEs, and larger genomes, than necrotrophs (Amselem et al., 2015a). Moreover, TEs have been implicated in the evolution of genome structure in human pathogens by mediating gene movement (McDonald et al., 2019), and they have been implicated in gene family expansion among both plant pathogens (Faino et al., 2016) and animal pathogens (Wacker et al., 2023) (reviewed in (Dong et al., 2015; Torres et al., 2020). Of note is that while the proportion of TEs in the Histoplasma genome is fairly consistent, it is generally higher than in other species of fungal pathogens. For example, the genome of Candida albicans harbors 0.1% DNA TEs and 0.7% RNA TEs. In Cryptococcus, the proportion of DNA TEs is 0.8% and that of RNA TEs is 4.4%, and these proportions vary with serotype (not including the proportion of unknown TEs) (reviewed in (Maxwell, 2020)). On the other hand, other fungal pathogens from the Onygenales order have higher TE content. For example, the Coccidioides genome is composed of 17–19% TEs (Kirkland et al., 2018), and Blastomyces is about 50% (McTaggart et al., 2016; Muñoz et al., 2015), with some variation among species within each genus. This variation in fungal pathogen TE content must be caveated, however, with the fact that different studies have used different tools to estimate the proportions of TEs in the genomes of different pathogens. Therefore, whether the apparent genome expansions of genera like Coccidioides and Blastomyces reflect biological reality or methodological differences requires further study.

An additional implication of our study is that although all computational pipelines detected the presence of TEs, the reported proportions varied among them, with dnaPipeTE RM lib being the most divergent from the other approaches. There is precedence for the importance of both the TE library and the analytical tool in TE assessments. For example, in Arabidopsis genomes, dnaPipeTE detects a lower proportion of TEs than LTRHarvest (Berthelier et al., 2018). Similarly, in trypanosomatids, dnaPipeTE identifies fewer putative TEs than RepeatModeler (Tullume-Vergara et al., 2025). Iterative and periodic benchmarking efforts (Gozashti and Hoekstra, 2024; Hoen et al., 2015; Ou et al., 2019; Vendrell-Mir et al., 2019) are essential for the success of comparative approaches aimed at understanding the roles of TEs across different taxa.

Life history traits, genome size, and TE content seem to be correlated across fungal species (Muszewska et al., 2019) cf. (Fijarczyk et al., 2022). TEs have also been implicated in differences in virulence (e.g., (Fouché et al., 2020), reviewed in (Möller and Stukenbrock, 2017). For example, stressful conditions such as starvation and host infection stress induce the expression of DNA/Ty3 and MITEs in Zymoseptoria tritici. In this pathogen, the expression and epigenetic repression of a major virulence factor, Avr3D1, is affected by TE insertions (DNA/DTM, Crypton, and others) (Fouché et al., 2020). Similarly, TEs have been identified in avirulence factors AvrPib and AvrPizt of Magnaporthe grisea (Bao et al., 2017), which suggests that associations between TEs and virulence might generalize across mechanisms of virulence among fungal pathogens. In addition, TEs can facilitate rapid responses to changing environments among experimentally evolved fungal populations (Gusa et al., 2023, 2020; López Díaz et al., 2025), and TEs have been implicated in drug resistance in a variety of species (e.g., Cryptococcius neoformans: (Priest et al., 2022); Monilinia: (Durak and Özkılınç, 2025). Fine-scale trait mapping across species will reveal whether TEs are a substantial source of variation for adaptation in the mutational repertoire of fungi.

Two additional caveats warrant discussion. First, we used short read assemblies that can bias the detection of TEs. For example, our results indicate that, although short-read assemblies can provide a proxy for genome size, they consistently underestimate it. More importantly, short read sequencing underestimates the proportion of TEs when compared with long-read sequencing (Bao et al., 2017) Second, comparative phylogenetic methods benefit from sampling completeness (Wiens, 2006). Our report is missing several reported phylogenetic species of Histoplasma (e.g., RJ, LamB (Almeida-Silva et al., 2021; Li et al., 2025)) and lineages that seem to be genetically isolated but whose genetic variation has not been catalogued (Kasuga et al., 2003; Quan et al., 2025). Further, intraclonal variation in TEs has not been assessed in fungi, but comparative methods that incorporate uncertainty at the tips and variation within isolates (Beaulieu and O’Meara, 2025) can be incorporated to address the role of variation within isolates.

Overall, these limitations warrant the continued study of TEs in Histoplasma and in other fungal pathogens. Future studies should address the fitness effects of TE movement, modes of inheritance within clades, and evidence of horizontal gene transfer across clades.While we provide a series of examples of insertions, the dynamics of TE insertions along the genome and in the evolutionary history of Histoplasma will require more highly contiguous assemblies than those currently available. Similarly, assessment of the precise rates of insertions, deletions, and the extent of selection (as hypothesized in (Petrov, 2002)) that might determine genome size in Histoplasma and other fungal pathogens will require estimates of mutational rates, and thus denser sampling within each species.

In conclusion, our results indicate substantial variance in TE content and mode of evolution in Histoplasma. The evolution of these elements has been correlated with the evolution of genome size in this genus of pathogenic fungi. Our results contribute to a growing body of observations that TE content variation is structured by biogeography and species boundaries. New technical developments and analytical tools are ripe for dissecting the mode of evolution of TEs, and other complex mutations, in human fungal pathogens.

Supplementary Material

Supplementary Material

FIGURE 8. Phylogenetic signal of total DNA and total RNA TE content.

FIGURE 8.

A. Pagel’s λ DNA TE. B. Blomberg’s K DNA TE. C. Pagel’s λ RNA TE. D. Blomberg’s K RNA TE. Top panels show the estimated likelihood of Pagel’s λ values. The dotted line marks the maximum likelihood estimate (MLE) of λ. Bottom panels show a histogram of the distribution of simulated values of Blomberg’s K and the observed value of Blomberg’s K (blue arrow). Table S14 shows the phylogenetic signal for other methods to detect TEs.

ACKNOWLEDGEMENTS

We would like to thank the Matute lab for comments and feedback on early versions of the manuscript.

FUNDING SOURCES

This work was supported by the National Institute of General Medical Sciences (R35GM148244). JAR was supported by the National Institute of Environmental Health Sciences under award F32ES035271. PK is supported by the National Institute of Allergy and Infectious Disease (2T32AI052080).

DATA AVAILABILITY

The github repositories along with the saved Zenodo repositories can be found for each portion of the analysis below. Each step was split into different github repositories for clarity and brevity. https://github.com/tania-k/Histoplasma_Assembly doi: 10.5281/zenodo.12747066 https://github.com/tania-k/Histo_Annotations doi: 10.5281/zenodo.12747074 https://github.com/tania-k/Histo_TE_pipeline doi: 10.5281/zenodo.12747068 https://github.com/dturissini/Histoplasma_TE_manuscript

REFERENCES

  1. Abraham LN, Oggenfuss U, Croll D, 2024. Population-level transposable element expression dynamics influence trait evolution in a fungal crop pathogen. mBio 15, e02840–23. 10.1128/mbio.02840-23 [DOI] [PMC free article] [PubMed] [Google Scholar]
  2. Adenis AA, Aznar C, Couppié P, 2014a. Histoplasmosis in HIV-Infected Patients: A Review of New Developments and Remaining Gaps. Curr Trop Med Rep. 10.1007/s40475-014-0017-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  3. Adenis AA, Aznar C, Couppié P, 2014b. Histoplasmosis in HIV-Infected Patients: A Review of New Developments and Remaining Gaps. Curr Trop Med Rep. 10.1007/s40475-014-0017-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  4. Al Mheiri FG, Joseph M, Joseph S, Alqassim M, Kinne J, Wernery U, 2024. An overview of various stages and morphology of Histoplasma capsulatum var. farciminosum in the horse. Vet Res Commun 48, 3483–3487. 10.1007/s11259-024-10483-0 [DOI] [PubMed] [Google Scholar]
  5. Al-Ani FK, 1999. Epizootic lymphangitis in horses: a review of the literature. Revue scientifique et technique (International Office of Epizootics) 18, 691–699. [PubMed] [Google Scholar]
  6. Alfsnes K, Leinaas HP, Hessen DO, 2017. Genome size in arthropods; different roles of phylogeny, habitat and life history in insects and crustaceans. Ecology and Evolution 7, 5939–5947. 10.1002/ece3.3163 [DOI] [PMC free article] [PubMed] [Google Scholar]
  7. Allard A, Décarie D, Grenier J-L, Lacombe M-C, Levac F, 2014. Histoplasmosis Outbreak Associated with the Renovation of an Old House—Quebec, Canada, 2013. MMWR: Morbidity & Mortality Weekly Report 62. [PMC free article] [PubMed] [Google Scholar]
  8. Almeida-Silva F, de Melo Teixeira M, Matute DR, de Faria Ferreira M, Barker BM, Almeida-Paes R, Guimarães AJ, Zancopé-Oliveira RM, 2021. Genomic Diversity Analysis Reveals a Strong Population Structure in Histoplasma capsulatum LAmA (Histoplasma suramericanum). J Fungi (Basel) 7, 865. 10.3390/jof7100865 [DOI] [PMC free article] [PubMed] [Google Scholar]
  9. Alonge M, Lebeigle L, Kirsche M, Jenike K, Ou S, Aganezov S, Wang X, Lippman ZB, Schatz MC, Soyk S, 2022. Automated assembly scaffolding using RagTag elevates a new tomato system for high-throughput genome editing. Genome Biol 23, 258. 10.1186/s13059-022-02823-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
  10. Ameni G, 2006. Epidemiology of equine histoplasmosis (epizootic lymphangitis) in carthorses in Ethiopia. The Veterinary Journal 172, 160–165. [DOI] [PubMed] [Google Scholar]
  11. Amselem J, Lebrun M-H, Quesneville H, 2015. Whole genome comparative analysis of transposable elements provides new insight into mechanisms of their inactivation in fungal genomes. BMC Genomics 16, 141. 10.1186/s12864-015-1347-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  12. Araya RA, Reinar WB, Tørresen OK, Goubert C, Daughton TJ, Hoff SNK, Baalsrud HT, Brieuc MSO, Komisarczuk A, Jentoft S, 2025. Chromosomal inversions mediated by tandem insertions of transposable elements. bioRxiv 2025–01. [DOI] [PMC free article] [PubMed] [Google Scholar]
  13. Arkhipova IR, 2006. Distribution and phylogeny of Penelope-like elements in eukaryotes. Systematic biology 55, 875–885. [DOI] [PubMed] [Google Scholar]
  14. Azar MM, Loyd JL, Relich RF, Wheat LJ, Hage CA, 2020. Current Concepts in the Epidemiology, Diagnosis, and Management of Histoplasmosis Syndromes. Semin Respir Crit Care Med 41, 013–030. 10.1055/s-0039-1698429 [DOI] [PubMed] [Google Scholar]
  15. Bankevich A, Nurk S, Antipov D, Gurevich AA, Dvorkin M, Kulikov AS, Lesin VM, Nikolenko SI, Pham S, Prjibelski AD, Pyshkin AV, Sirotkin AV, Vyahhi N, Tesler G, Alekseyev MA, Pevzner PA, 2012. SPAdes: A New Genome Assembly Algorithm and Its Applications to Single-Cell Sequencing. Journal of Computational Biology 19, 455–477. 10.1089/cmb.2012.0021 [DOI] [PMC free article] [PubMed] [Google Scholar]
  16. Bao J, Chen M, Zhong Z, Tang W, Lin L, Zhang X, Jiang H, Zhang D, Miao C, Tang H, 2017. PacBio sequencing reveals transposable elements as a key contributor to genomic plasticity and virulence variation in Magnaporthe oryzae. Molecular plant 10, 1465–1468. [DOI] [PubMed] [Google Scholar]
  17. Bao W, Kojima KK, Kohany O, 2015. Repbase Update, a database of repetitive elements in eukaryotic genomes. Mobile DNA 6, 11. 10.1186/s13100-015-0041-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  18. Bateman A, Coin L, Durbin R, Finn RD, Hollich V, Griffiths-Jones S, Khanna A, Marshall M, Moxon S, Sonnhammer EL, 2004. The Pfam protein families database. Nucleic acids research 32, D138–D141. [DOI] [PMC free article] [PubMed] [Google Scholar]
  19. Beaulieu JM, O’Meara BC, 2025. Navigating “tip fog”: Embracing uncertainty in tip measurements. Evolution qpaf067. [DOI] [PubMed] [Google Scholar]
  20. Benedict K, Mody RK, 2016. Epidemiology of histoplasmosis outbreaks, United States, 1938–2013. Emerging infectious diseases 22, 370. [DOI] [PMC free article] [PubMed] [Google Scholar]
  21. Bernardes JPRA, Tenorio BG, Lucas JM Jr, Paternina CEM, Silva RKM, Erazo FAH, Silva-Pereira I, Alves L.G. de B., Arenas PHR, Bueno-Rocha ID, 2025. Mapping Histoplasma in Bats and Cave Ecosystems: Evidence from Midwestern Brazil. bioRxiv 2025–03. [Google Scholar]
  22. Berthelier J, Casse N, Daccord N, Jamilloux V, Saint-Jean B, Carrier G, 2018. A transposable element annotation pipeline and expression analysis reveal potentially active elements in the microalga Tisochrysis lutea. BMC Genomics 19, 378. 10.1186/s12864-018-4763-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  23. Beyhan S, Sil A, 2019. Sensing the heat and the host: virulence determinants of Histoplasma capsulatum. Virulence 10, 793–800. [DOI] [PMC free article] [PubMed] [Google Scholar]
  24. Biedler J, Tu Z, 2003. Non-LTR retrotransposons in the African malaria mosquito, Anopheles gambiae: unprecedented diversity and evidence of recent activity. Molecular biology and evolution 20, 1811–1825. [DOI] [PubMed] [Google Scholar]
  25. Bleykasten-Grosshans C, Friedrich A, Schacherer J, 2013. Genome-wide analysis of intraspecific transposon diversity in yeast. BMC Genomics 14, 399. 10.1186/1471-2164-14-399 [DOI] [PMC free article] [PubMed] [Google Scholar]
  26. Bleykasten-Grosshans C, Neuvéglise C, 2011. Transposable elements in yeasts. Comptes rendus biologies 334, 679–686. [DOI] [PubMed] [Google Scholar]
  27. Blin K, Shaw S, Kloosterman AM, Charlop-Powers Z, Van Wezel GP, Medema MH, Weber T, 2021. antiSMASH 6.0: improving cluster detection and comparison capabilities. Nucleic acids research 49, W29–W35. [DOI] [PMC free article] [PubMed] [Google Scholar]
  28. Blomberg SP, Garland T, Ives AR, 2003. Testing for phylogenetic signal in comparative data: behavioral traits are more labile. Evolution 57, 717–745. 10.1111/j.0014-3820.2003.tb00285.x [DOI] [PubMed] [Google Scholar]
  29. Böhne A, Brunet F, Galiana-Arnoux D, Schultheis C, Volff J-N, 2008. Transposable elements as drivers of genomic and biological diversity in vertebrates. Chromosome research 16, 203–215. [DOI] [PubMed] [Google Scholar]
  30. Bongomin F, Kwizera R, Denning DW, 2019. Getting histoplasmosis on the map of international recommendations for patients with advanced HIV disease. Journal of Fungi 5, 80. [DOI] [PMC free article] [PubMed] [Google Scholar]
  31. Brömel C, Sykes JE, 2005. Epidemiology, diagnosis, and treatment of blastomycosis in dogs and cats. Clin Tech Small Anim Pract 20, 233–239. 10.1053/j.ctsap.2005.07.004 [DOI] [PubMed] [Google Scholar]
  32. Bruuna T, Lomsadze A, Borodovsky M, 2020. GeneMark-EP+: eukaryotic gene prediction with self-training in the space of genes and proteins. NAR genomics and bioinformatics 2, lqaa026. [DOI] [PMC free article] [PubMed] [Google Scholar]
  33. Bushnell B, 2014. BBMap: a fast, accurate, splice-aware aligner. [Google Scholar]
  34. Butler MA, King AA, 2004. Phylogenetic Comparative Analysis: A Modeling Approach for Adaptive Evolution. The American Naturalist 164, 683–695. 10.1086/426002 [DOI] [PubMed] [Google Scholar]
  35. Ca M, Pv B,R,R, Nair LR, 2013. Disseminated histoplasmosis. IJOO 47, 639–642. 10.4103/0019-5413.121601 [DOI] [PMC free article] [PubMed] [Google Scholar]
  36. Caballero Van Dyke MC, Teixeira MM, Barker BM, 2019. Fantastic yeasts and where to find them: the hidden diversity of dimorphic fungal pathogens. Curr Opin Microbiol 52, 55–63. 10.1016/j.mib.2019.05.002 [DOI] [PMC free article] [PubMed] [Google Scholar]
  37. Cáceres M, Ranz JM, Barbadilla A, Long M, Ruiz A, 1999. Generation of a Widespread Drosophila Inversion by a Transposable Element. Science 285, 415–418. 10.1126/science.285.5426.415 [DOI] [PubMed] [Google Scholar]
  38. Canteros CE, Zuiani MF, Ritacco V, Perrotta DE, Reyes-Montes MR, Granados J, Zuniga G, Taylor ML, Davel G, 2005. Electrophoresis karyotype and chromosome-length polymorphism of Histoplasma capsulatum clinical isolates from Latin America. FEMS Immunology & Medical Microbiology 45, 423–428. [DOI] [PubMed] [Google Scholar]
  39. Carleton KL, Conte MA, Malinsky M, Nandamuri SP, Sandkam BA, Meier JI, Mwaiko S, Seehausen O, Kocher TD, 2020. Movement of transposable elements contributes to cichlid diversity. Molecular Ecology 29, 4956–4969. 10.1111/mec.15685 [DOI] [PubMed] [Google Scholar]
  40. Carr J, Shearer G, 1998. Genome Size, Complexity, and Ploidy of the Pathogenic Fungus Histoplasma capsulatum. J Bacteriol 180, 6697–6703. 10.1128/JB.180.24.6697-6703.1998 [DOI] [PMC free article] [PubMed] [Google Scholar]
  41. Casacuberta E, González J, 2013. The impact of transposable elements in environmental adaptation. Molecular Ecology 22, 1503–1517. 10.1111/mec.12170 [DOI] [PubMed] [Google Scholar]
  42. Castanera R, Borgognone A, Pisabarro AG, Ramírez L, 2017. Biology, dynamics, and applications of transposable elements in basidiomycete fungi. Applied microbiology and biotechnology 101, 1337–1350. [DOI] [PubMed] [Google Scholar]
  43. Cerbin S, Jiang N, 2018. Duplication of host genes by transposable elements. Current Opinion in Genetics & Development 49, 63–69. [DOI] [PubMed] [Google Scholar]
  44. Chan PP, Lin BY, Mak AJ, Lowe TM, 2021. tRNAscan-SE 2.0: improved detection and functional classification of transfer RNA genes. Nucleic acids research 49, 9077–9096. [DOI] [PMC free article] [PubMed] [Google Scholar]
  45. Chang J, Duan G, Li W, Yau TO, Liu C, Cui J, Xue H, Bu W, Hu Y, Gao S, 2023. The first discovery of Tc1 transposons in yeast. Frontiers in Microbiology 14, 1141495. [DOI] [PMC free article] [PubMed] [Google Scholar]
  46. Chao L, Vargas C, Spear BB, Cox EC, 1983. Transposable elements as mutator genes in evolution. Nature 303, 633–635. [DOI] [PubMed] [Google Scholar]
  47. Clinkenbeard KD, Cowell RL, Tyler RD, 1988. Disseminated histoplasmosis in dogs: 12 cases (1981–1986). Journal of the American Veterinary Medical Association 193, 1443–1447. [PubMed] [Google Scholar]
  48. Cooper N, Freckleton RP, Jetz W, 2011. Phylogenetic conservatism of environmental niches in mammals. Proc. R. Soc. B. 278, 2384–2391. 10.1098/rspb.2010.2207 [DOI] [PMC free article] [PubMed] [Google Scholar]
  49. de Perio MA, Benedict K, Williams SL, Niemeier-Walsh C, Green BJ, Coffey C, Di Giuseppe M, Toda M, Park J-H, Bailey RL, 2021. Occupational histoplasmosis: epidemiology and prevention measures. Journal of Fungi 7, 510. [DOI] [PMC free article] [PubMed] [Google Scholar]
  50. Dong S, Raffaele S, Kamoun S, 2015. The two-speed genomes of filamentous pathogens: waltz with plants. Current opinion in genetics & development 35, 57–65. [DOI] [PubMed] [Google Scholar]
  51. Drongitis D, Aniello F, Fucci L, Donizetti A, 2019. Roles of transposable elements in the different layers of gene expression regulation. International Journal of Molecular Sciences 20, 5755. [DOI] [PMC free article] [PubMed] [Google Scholar]
  52. Dubin CA, Voorhies M, Sil A, Teixeira MM, Barker BM, Brem RB, 2022. Genome Organization and Copy-Number Variation Reveal Clues to Virulence Evolution in Coccidioides posadasii. Journal of Fungi 8, 1235. [DOI] [PMC free article] [PubMed] [Google Scholar]
  53. Durak MR, Özkılınç H, 2025. Transposable elements in genomic architecture of Monilinia fungal phytopathogens and TE-driven DMI-resistance adaptation. Mobile DNA 16, 8. 10.1186/s13100-025-00343-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  54. Elliott TA, Gregory TR, 2015. Do larger genomes contain more diverse transposable elements? BMC Evol Biol 15, 69. 10.1186/s12862-015-0339-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  55. Emmons CW, Morlan HB, Hill EL, 1949. Histoplasmosis in rats and skunks in Georgia. Public Health Reports (1896–1970) 1423–1431. [PubMed] [Google Scholar]
  56. Engels WR, 1992. The origin of P elements in Drosophila melanogaster. BioEssays 14, 681–686. 10.1002/bies.950141007 [DOI] [PubMed] [Google Scholar]
  57. Faino L, Seidl MF, Shi-Kunne X, Pauper M, van den Berg GC, Wittenberg AH, Thomma BP, 2016. Transposons passively and actively contribute to evolution of the two-speed genome of a fungal pathogen. Genome research 26, 1091–1100. [DOI] [PMC free article] [PubMed] [Google Scholar]
  58. Felsenstein J, 1973. Maximum likelihood and minimum-steps methods for estimating evolutionary trees from data on discrete characters. Systematic Biology 22, 240–249. [Google Scholar]
  59. Feschotte C, 2008. Transposable elements and the evolution of regulatory networks. Nature Reviews Genetics 9, 397–405. [DOI] [PMC free article] [PubMed] [Google Scholar]
  60. Fijarczyk A, Hessenauer P, Hamelin RC, Landry CR, 2022. Lifestyles shape genome size and gene content in fungal pathogens. Biorxiv 2022–08. [DOI] [PMC free article] [PubMed] [Google Scholar]
  61. Flynn JM, Hubley R, Goubert C, Rosen J, Clark AG, Feschotte C, Smit AF, 2020. RepeatModeler2 for automated genomic discovery of transposable element families. Proc. Natl. Acad. Sci. U.S.A. 117, 9451–9457. 10.1073/pnas.1921046117 [DOI] [PMC free article] [PubMed] [Google Scholar]
  62. Fouché S, Badet T, Oggenfuss U, Plissonneau C, Francisco CS, Croll D, 2020. Stress-driven transposable element de-repression dynamics and virulence evolution in a fungal pathogen. Molecular biology and evolution 37, 221–239. [DOI] [PubMed] [Google Scholar]
  63. Fouché S, Oggenfuss U, Chanclud E, Croll D, 2022. A devil’s bargain with transposable elements in plant pathogens. Trends in Genetics 38, 222–230. [DOI] [PubMed] [Google Scholar]
  64. Freckleton RP, Harvey PH, Pagel M, 2002. Phylogenetic Analysis and Comparative Data: A Test and Review of Evidence. The American Naturalist 160, 712–726. 10.1086/343873 [DOI] [PubMed] [Google Scholar]
  65. Gebrie A, 2023. Transposable elements as essential elements in the control of gene expression. Mobile DNA 14, 9. 10.1186/s13100-023-00297-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  66. Goodwin TJD, Butler MI, Poulter RTM, 2003. Cryptons: a group of tyrosine-recombinase-encoding DNA transposons from pathogenic fungi. Microbiology 149, 3099–3109. 10.1099/mic.0.26529-0 [DOI] [PubMed] [Google Scholar]
  67. Goubert C, Modolo L, Vieira C, ValienteMoro C, Mavingui P, Boulesteix M, 2015. De novo assembly and annotation of the Asian tiger mosquito (Aedes albopictus) repeatome with dnaPipeTE from raw genomic reads and comparative analysis with the yellow fever mosquito (Aedes aegypti). Genome biology and evolution 7, 1192–1205. [DOI] [PMC free article] [PubMed] [Google Scholar]
  68. Gozashti L, Hoekstra HE, 2024. Accounting for diverse transposable element landscapes is key to developing and evaluating accurate de novo annotation strategies. Genome Biol 25, 4. 10.1186/s13059-023-03118-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  69. Grabherr MG, Haas BJ, Yassour M, Levin JZ, Thompson DA, Amit I, Adiconis X, Fan L, Raychowdhury R, Zeng Q, 2011. Trinity: reconstructing a full-length transcriptome without a genome from RNA-Seq data. Nature biotechnology 29, 644. [DOI] [PMC free article] [PubMed] [Google Scholar]
  70. Gugnani HC, Denning DW, 2023. Infection of bats with Histoplasma species. Medical mycology 61, myad080. [DOI] [PMC free article] [PubMed] [Google Scholar]
  71. Gugnani HC, Muotoe-Okafor FA, Kaufman L, Dupont B, 1994. A natural focus of Histoplasma capsulatum var. duboisii is a bat cave. Mycopathologia 127, 151–157. [DOI] [PubMed] [Google Scholar]
  72. Guillot J, Guérin C, Chermette R, 2018. Histoplasmosis in Animals, in: Seyedmousavi S, De Hoog GS, Guillot J, Verweij PE (Eds.), Emerging and Epizootic Fungal Infections in Animals. Springer International Publishing, Cham, pp. 115–128. 10.1007/978-3-319-72093-7_5 [DOI] [Google Scholar]
  73. Gusa A, Williams JD, Cho J-E, Averette AF, Sun S, Shouse EM, Heitman J, Alspaugh JA, Jinks-Robertson S, 2020. Transposon mobilization in the human fungal pathogen Cryptococcus is mutagenic during infection and promotes drug resistance in vitro. Proc. Natl. Acad. Sci. U.S.A. 117, 9973–9980. 10.1073/pnas.2001451117 [DOI] [PMC free article] [PubMed] [Google Scholar]
  74. Gusa A, Yadav V, Roth C, Williams JD, Shouse EM, Magwene P, Heitman J, Jinks-Robertson S, 2023. Genome-wide analysis of heat stress-stimulated transposon mobility in the human fungal pathogen Cryptococcus deneoformans. Proc. Natl. Acad. Sci. U.S.A. 120, e2209831120. 10.1073/pnas.2209831120 [DOI] [PMC free article] [PubMed] [Google Scholar]
  75. Haas BJ, Salzberg SL, Zhu W, Pertea M, Allen JE, Orvis J, White O, Buell CR, Wortman JR, 2008. Automated eukaryotic gene structure annotation using EVidenceModeler and the Program to Assemble Spliced Alignments. Genome Biol 9, R7. 10.1186/gb-2008-9-1-r7 [DOI] [PMC free article] [PubMed] [Google Scholar]
  76. Harmon L, Weir J, Brock C, Glor R, Challenger W, Hunt G, FitzJohn R, Pennell M, Slater G, Brown J, 2015. Package ‘geiger.’ R package version 2, 1–74. [Google Scholar]
  77. Harmon LJ, Weir JT, Brock CD, Glor RE, Challenger W, 2008. GEIGER: investigating evolutionary radiations. Bioinformatics 24, 129–131. [DOI] [PubMed] [Google Scholar]
  78. Havecker ER, Gao X, Voytas DF, 2004. The diversity of LTR retrotransposons. Genome Biol 5, 225. 10.1186/gb-2004-5-6-225 [DOI] [PMC free article] [PubMed] [Google Scholar]
  79. Hirsch CD, Springer NM, 2017. Transposable element influences on gene expression in plants. Biochimica et Biophysica Acta (BBA)-Gene Regulatory Mechanisms 1860, 157–165. [DOI] [PubMed] [Google Scholar]
  80. Hoen DR, Hickey G, Bourque G, Casacuberta J, Cordaux R, Feschotte C, Fiston-Lavier A-S, Hua-Van A, Hubley R, Kapusta A, Lerat E, Maumus F, Pollock DD, Quesneville H, Smit A, Wheeler TJ, Bureau TE, Blanchette M, 2015. A call for benchmarking transposable element annotation methods. Mobile DNA 6, 13. 10.1186/s13100-015-0044-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  81. Höft MA, Duvenage L, Hoving JC, n.d. Key thermally dimorphic fungal pathogens: shaping host immunity. Open Biol 12, 210219. 10.1098/rsob.210219 [DOI] [PMC free article] [PubMed] [Google Scholar]
  82. Hothorn T, Bretz F, Westfall P, Heiberger RM, 2012. Multcomp: simultaneous inference for general linear hypotheses. URL http://CRAN.R-project.org/package=multcomp, R package version 1–2. [Google Scholar]
  83. Hothorn T, Bretz F, Westfall P, Heiberger RM, Schuetzenmeister A, Scheibe S, Hothorn MT, 2016. Package ‘multcomp.’ Simultaneous inference in general parametric models. Project for Statistical Computing, Vienna, Austria. [Google Scholar]
  84. Huang Y, Gao ZY, Ly K, Lin L, Lambooij J-P, King EG, Janssen A, Wei KH-C, Lee YCG, 2025. Polymorphic transposable elements contribute to variation in recombination landscapes. Proc. Natl. Acad. Sci. U.S.A. 122, e2427312122. 10.1073/pnas.2427312122 [DOI] [PMC free article] [PubMed] [Google Scholar]
  85. Huerta-Cepas J, Szklarczyk D, Heller D, Hernández-Plaza A, Forslund SK, Cook H, Mende DR, Letunic I, Rattei T, Jensen LJ, 2019. eggNOG 5.0: a hierarchical, functionally and phylogenetically annotated orthology resource based on 5090 organisms and 2502 viruses. Nucleic acids research 47, D309–D314. [DOI] [PMC free article] [PubMed] [Google Scholar]
  86. Jing X, Liu X, Yuan H, Dai Y, Zheng Y, Zhao L, Ma L, Huang Y, 2024. Evolutionary dynamics of genome size and transposable elements in crickets (ENSIFERA: GRYLLIDEA). Systematic Entomology 49, 549–564. 10.1111/syen.12629 [DOI] [Google Scholar]
  87. Jofre GI, Singh A, Mavengere H, Sundar G, D’Agostino E, Chowdhary A, Matute DR, 2022. An Indian lineage of Histoplasma with strong signatures of differentiation and selection. Fungal Genet Biol 158, 103654. 10.1016/j.fgb.2021.103654 [DOI] [PMC free article] [PubMed] [Google Scholar]
  88. Johnson M, Zaretskaya I, Raytselis Y, Merezhuk Y, McGinnis S, Madden TL, 2008. NCBI BLAST: a better web interface. Nucleic acids research 36, W5–W9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  89. Jordan IK, Rogozin IB, Glazko GV, Koonin EV, 2003. Origin of a substantial fraction of human regulatory sequences from transposable elements. TRENDS in Genetics 19, 68–72. [DOI] [PubMed] [Google Scholar]
  90. Kapitonov VV, Jurka J, 2008. A universal classification of eukaryotic transposable elements implemented in Repbase. Nature Reviews Genetics 9, 411–412. [DOI] [PubMed] [Google Scholar]
  91. Kassambara A, Mundt F, 2017. Package ‘factoextra.’ Extract and visualize the results of multivariate data analyses 76, 10–18637. [Google Scholar]
  92. Kasuga T, White TJ, Koenig G, McEwen J, Restrepo A, Castañeda E, Da Silva Lacaz C, Heins-Vaccari EM, De Freitas RS, Zancopé-Oliveira RM, Qin Z, Negroni R, Carter DA, Mikami Y, Tamura M, Taylor ML, Miller GF, Poonwan N, Taylor JW, 2003. Phylogeography of the fungal pathogen Histoplasma capsulatum. Mol Ecol 12, 3383–3401. 10.1046/j.1365-294x.2003.01995.x [DOI] [PubMed] [Google Scholar]
  93. Kent TV, Uzunović J, Wright SI, 2017. Coevolution between transposable elements and recombination. Phil. Trans. R. Soc. B 372, 20160458. 10.1098/rstb.2016.0458 [DOI] [PMC free article] [PubMed] [Google Scholar]
  94. Kidwell MG, 2002. Transposable elements and the evolution of genome size in eukaryotes. Genetica 115, 49–63. [DOI] [PubMed] [Google Scholar]
  95. Kim JM, Vanguri S, Boeke JD, Gabriel A, Voytas DF, 1998. Transposable elements and genome organization: a comprehensive survey of retrotransposons revealed by the complete Saccharomyces cerevisiae genome sequence. Genome research 8, 464–478. [DOI] [PubMed] [Google Scholar]
  96. Kirkland TN, Muszewska A, Stajich JE, 2018. Analysis of transposable elements in Coccidioides species. Journal of Fungi 4, 13. [DOI] [PMC free article] [PubMed] [Google Scholar]
  97. Klein BS, Tebbets B, 2007. Dimorphism and virulence in fungi. Curr Opin Microbiol 10, 314–319. 10.1016/j.mib.2007.04.002 [DOI] [PMC free article] [PubMed] [Google Scholar]
  98. Korf I, 2004. Gene finding in novel genomes. BMC Bioinformatics 5, 59. 10.1186/1471-2105-5-59 [DOI] [PMC free article] [PubMed] [Google Scholar]
  99. Krasileva KV, 2019. The role of transposable elements and DNA damage repair mechanisms in gene duplications and gene fusions in plant genomes. Current Opinion in Plant Biology 48, 18–25. [DOI] [PubMed] [Google Scholar]
  100. Lambert ME, McDonald JF, Weinstein IB, 1988. Eukaryotic transposable elements as mutagenic agents. Cold Spring Harbor Lab. of Quantitative Biology, ME; (USA: ). [Google Scholar]
  101. Li H, 2018. Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics 34, 3094–3100. [DOI] [PMC free article] [PubMed] [Google Scholar]
  102. Li T, Teixeira M. de M., Jofre-Rodriguez G, Blanchet D, Mc Donald S, Alvarado P, Marques da Silva SH, Sepulveda VE, Zeb Q, Vreden S, 2025. High genetic diversity of Histoplasma in the Amazon basin, 2006–2017. medRxiv 2025–04. [DOI] [PMC free article] [PubMed] [Google Scholar]
  103. Lin Lianyu, Sun T, Guo J, Lin Lili, Chen M, Wang Zhe, Bao J, Norvienyeku J, Zhang D, Han Y, Lu G, Rensing C, Zheng H, Zhong Z, Wang Zonghua, 2024. Transposable elements impact the population divergence of rice blast fungus Magnaporthe oryzae. mBio 15, e00086–24. 10.1128/mbio.00086-24 [DOI] [PMC free article] [PubMed] [Google Scholar]
  104. Linder KA, Kauffman CA, 2019. Histoplasmosis: Epidemiology, Diagnosis, and Clinical Manifestations. Curr Fungal Infect Rep 13, 120–128. 10.1007/s12281-019-00341-x [DOI] [Google Scholar]
  105. López Díaz C, Ayhan DH, Rodríguez López A, Gómez Gil L, Ma L-J, Di Pietro A, 2025. Transposons and accessory genes drive adaptation in a clonally evolving fungal pathogen. Nature Communications 16, 6982. [DOI] [PMC free article] [PubMed] [Google Scholar]
  106. Loreto EL, Melo E.S. de, Wallau GL, Gomes TM, 2023. The good, the bad and the ugly of transposable elements annotation tools. Genetics and Molecular Biology 46, e20230138. [DOI] [PMC free article] [PubMed] [Google Scholar]
  107. Losos JB, 2008. Phylogenetic niche conservatism, phylogenetic signal and the relationship between phylogenetic relatedness and ecological similarity among species. Ecology Letters 11, 995–1003. 10.1111/j.1461-0248.2008.01229.x [DOI] [PubMed] [Google Scholar]
  108. Lottenberg R, Waldman RH, Ajello L, Hoff GL, Bigler W, Zellner SR, 1979. Pulmonary histoplasmosis associated with exploration of a bat cave. American journal of epidemiology 110, 156–161. [DOI] [PubMed] [Google Scholar]
  109. Lowe TM, Chan PP, 2016. tRNAscan-SE On-line: integrating search and context for analysis of transfer RNA genes. Nucleic acids research 44, W54–W57. [DOI] [PMC free article] [PubMed] [Google Scholar]
  110. Ly T, de Melo Teixeira M, Jofre GI, Blanchet D, MacDonald S, Alvarado P, da Silva SHM, Sepúlveda VE, Zeb Q, Vreden S, 2025. High Genetic Diversity of Histoplasma in the Amazon Basin, 2006–2017. Emerging Infectious Diseases 31, 1169. [DOI] [PMC free article] [PubMed] [Google Scholar]
  111. Lyon GM, Bravo AV, Espino A, Lindsley MD, Gutierrez RE, Rodriguez I, Corella ANA, Carrillo F, McNeil MM, Warnock DW, 2004. Histoplasmosis associated with exploring a bat-inhabited cave in Costa Rica, 1998–1999. The American journal of tropical medicine and hygiene 70, 438–442. [PubMed] [Google Scholar]
  112. Mackey AI, Fraunfelter V, Shaltz S, McCormick J, Schroeder C, Perfect JR, Feschotte C, Magwene PM, Gusa A, 2025. Temperature and genetic background drive mobilization of diverse transposable elements in a critical human fungal pathogen. bioRxiv. [DOI] [PMC free article] [PubMed] [Google Scholar]
  113. Majoros WH, Pertea M, Salzberg SL, 2004. TigrScan and GlimmerHMM: two open source ab initio eukaryotic gene-finders. Bioinformatics 20, 2878–2879. [DOI] [PubMed] [Google Scholar]
  114. Manni M, Berkeley MR, Seppey M, Simão FA, Zdobnov EM, 2021. BUSCO update: novel and streamlined workflows along with broader and deeper phylogenetic coverage for scoring of eukaryotic, prokaryotic, and viral genomes. Molecular biology and evolution 38, 4647–4654. [DOI] [PMC free article] [PubMed] [Google Scholar]
  115. Mapengo RE, Maphanga TG, Jofre GI, Rader JA, Turissini DA, Birkhead M, Kwabia SA, Sepúlveda VE, Buitrago MJ, Teixeira MDM, Barker BM, Alanio A, Sturny-Leclère A, Garcia-Hermoso D, Govender NP, Matute DR, 2025. Genomic epidemiology of Histoplasma in Africa. mBio. 10.1128/mbio.00564-25 [DOI] [PMC free article] [PubMed] [Google Scholar]
  116. Maxwell PH, 2020. Diverse transposable element landscapes in pathogenic and nonpathogenic yeast models: the value of a comparative perspective. Mobile DNA 11, 16. 10.1186/s13100-020-00215-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  117. McDonald MC, Taranto AP, Hill E, Schwessinger B, Liu Z, Simpfendorfer S, Milgate A, Solomon PS, 2019. Transposon-Mediated Horizontal Transfer of the Host-Specific Virulence Protein ToxA between Three Fungal Wheat Pathogens. mBio 10, e01515–19. 10.1128/mBio.01515-19 [DOI] [PMC free article] [PubMed] [Google Scholar]
  118. McTaggart LR, Brown EM, Richardson SE, 2016. Phylogeographic analysis of Blastomyces dermatitidis and Blastomyces gilchristii reveals an association with North American freshwater drainage basins. PLoS One 11, e0159396. [DOI] [PMC free article] [PubMed] [Google Scholar]
  119. Möller M, Stukenbrock EH, 2017. Evolution and genome architecture in fungal plant pathogens. Nature Reviews Microbiology 15, 756–771. [DOI] [PubMed] [Google Scholar]
  120. Morgan J, Cano MV, Feikin DR, Phelan M, Monroy OV, Morales PK, Carpenter J, Weltman A, Spitzer PG, Liu HH, 2003. A large outbreak of histoplasmosis among American travelers associated with a hotel in Acapulco, Mexico, spring 2001. The American journal of tropical medicine and hygiene 69, 663–669. [PubMed] [Google Scholar]
  121. Mulder N, Apweiler R, 2007. InterPro and InterProScan: tools for protein sequence classification and comparison. Methods Mol Biol 396, 59–70. 10.1007/978-1-59745-515-2_5 [DOI] [PubMed] [Google Scholar]
  122. Mulec J, Simčič S, Kotar T, Kofol R, Stopinšek S, 2020. Survey of Histoplasma capsulatum in bat guano and status of histoplasmosis in Slovenia, Central Europe. International journal of speleology 49, 1. [Google Scholar]
  123. Muñoz JF, Gauthier GM, Desjardins CA, Gallo JE, Holder J, Sullivan TD, Marty AJ, Carmen JC, Chen Z, Ding L, 2015. The dynamic genome and transcriptome of the human fungal pathogen Blastomyces and close relative Emmonsia. PLoS genetics 11, e1005493. [DOI] [PMC free article] [PubMed] [Google Scholar]
  124. Muszewska A, Hoffman-Sommer M, Grynberg M, 2011. LTR retrotransposons in fungi. PloS one 6, e29425. [DOI] [PMC free article] [PubMed] [Google Scholar]
  125. Muszewska A, Steczkiewicz K, Stepniewska-Dziubinska M, Ginalski K, 2019. Transposable elements contribute to fungal genes and impact fungal lifestyle. Scientific Reports 9, 4307. [DOI] [PMC free article] [PubMed] [Google Scholar]
  126. Myint T, Leedy N, Villacorta Cari E, Wheat LJ, 2020. HIV-Associated Histoplasmosis: Current Perspectives. HIV Volume 12, 113–125. 10.2147/HIV.S185631 [DOI] [PMC free article] [PubMed] [Google Scholar]
  127. Nacher M, Couppié P, Epelboin L, Djossou F, Demar M, Adenis A, 2020. Disseminated Histoplasmosis: Fighting a neglected killer of patients with advanced HIV disease in Latin America. PLoS pathogens 16, e1008449. [DOI] [PMC free article] [PubMed] [Google Scholar]
  128. Nakamoto AA, Joubert PM, Krasileva KV, 2023. Intraspecific variation of transposable elements reveals differences in the evolutionary history of fungal phytopathogen pathotypes. Genome Biology and Evolution 15, evad206. [DOI] [PMC free article] [PubMed] [Google Scholar]
  129. Neafsey DE, Barker BM, Sharpton TJ, Stajich JE, Park DJ, Whiston E, Hung C-Y, McMahan C, White J, Sykes S, 2010. Population genomic sequencing of Coccidioides fungi reveals recent hybridization and transposon control. Genome research 20, 938–946. [DOI] [PMC free article] [PubMed] [Google Scholar]
  130. Nekrutenko A, Li W-H, 2001. Transposable elements are found in a large number of human protein-coding genes. TRENDS in Genetics 17, 619–621. [DOI] [PubMed] [Google Scholar]
  131. Nemecek JC, Wüthrich M, Klein BS, 2006. Global Control of Dimorphism and Virulence in Fungi. Science 312, 583–588. 10.1126/science.1124105 [DOI] [PubMed] [Google Scholar]
  132. Norkaew T, Ohno H, Sriburee P, Tanabe K, Tharavichitkul P, Takarn P, Puengchan T, Bumrungsri S, Miyazaki Y, 2013. Detection of environmental sources of Histoplasma capsulatum in Chiang Mai, Thailand, by nested PCR. Mycopathologia 176, 395–402. [DOI] [PubMed] [Google Scholar]
  133. O’Brochta DA, Handler AM, 1988. Mobility of P elements in drosophilids and nondrosophilids. Proc. Natl. Acad. Sci. U.S.A. 85, 6052–6056. 10.1073/pnas.85.16.6052 [DOI] [PMC free article] [PubMed] [Google Scholar]
  134. Oggenfuss U, Todd RT, Soisangwan N, Kemp B, Guyer A, Beach A, Selmecki A, 2025. Candida albicans isolates contain frequent heterozygous structural variants and transposable elements within genes and centromeres. Genome research 35, 824–838. [DOI] [PMC free article] [PubMed] [Google Scholar]
  135. Oliver KR, Greene WK, 2009. Transposable elements: powerful facilitators of evolution. BioEssays 31, 703–714. 10.1002/bies.200800219 [DOI] [PubMed] [Google Scholar]
  136. Oliver MJ, Petrov D, Ackerly D, Falkowski P, Schofield OM, 2007. The mode and tempo of genome size evolution in eukaryotes. Genome research 17, 594–601. [DOI] [PMC free article] [PubMed] [Google Scholar]
  137. Organization WH, 2022. WHO fungal priority pathogens list to guide research, development and public health action. World Health Organization. [Google Scholar]
  138. Ou S, Su W, Liao Y, Chougule K, Agda JRA, Hellinga AJ, Lugo CSB, Elliott TA, Ware D, Peterson T, Jiang N, Hirsch CN, Hufford MB, 2019. Benchmarking transposable element annotation methods for creation of a streamlined, comprehensive pipeline. Genome Biol 20, 275. 10.1186/s13059-019-1905-y [DOI] [PMC free article] [PubMed] [Google Scholar]
  139. Pagel M, 1999. Inferring the historical patterns of biological evolution. Nature 401, 877–884. [DOI] [PubMed] [Google Scholar]
  140. Palmer J, Stajich J, 2020. Funannotate Eukaryotic genome annotation. [Google Scholar]
  141. Palmer JM, Stajich JE, 2022. Automatic assembly for the fungi (AAFTF): genome assembly pipeline. Automatic Assembly for the fungi (AAFTF): genome assembly pipeline. [Google Scholar]
  142. Pasqualotto AC, Quieroz-Telles F, 2018. Histoplasmosis dethrones tuberculosis in Latin America. The Lancet Infectious diseases 18, 1058–1060. [DOI] [PubMed] [Google Scholar]
  143. Paysan-Lafosse T, Andreeva A, Blum M, Chuguransky SR, Grego T, Pinto BL, Salazar GA, Bileschi ML, Llinares-López F, Meng-Papaxanthos L, 2025. The Pfam protein families database: embracing AI/ML. Nucleic acids research 53, D523–D534. [DOI] [PMC free article] [PubMed] [Google Scholar]
  144. Pennell MW, Eastman JM, Slater GJ, Brown JW, Uyeda JC, FitzJohn RG, Alfaro ME, Harmon LJ, 2014. geiger v2. 0: an expanded suite of methods for fitting macroevolutionary models to phylogenetic trees. Bioinformatics 30, 2216–2218. [DOI] [PubMed] [Google Scholar]
  145. Petrov DA, 2002. Mutational equilibrium model of genome size evolution. Theoretical population biology 61, 531–544. [DOI] [PubMed] [Google Scholar]
  146. Pettrich LC, King R, Field LM, Waldvogel A-M, 2024. The interplay of recombination landscape, a transposable element and population history in European populations of Chironomus riparius. bioRxiv 2024–04. [DOI] [PMC free article] [PubMed] [Google Scholar]
  147. Potocki L, Kuna E, Filip K, Kasprzyk B, Lewinska A, Wnuk M, 2019. Activation of transposable elements and genetic instability during long-term culture of the human fungal pathogen Candida albicans. Biogerontology 20, 457–474. 10.1007/s10522-019-09809-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  148. Potter SC, Luciani A, Eddy SR, Park Y, Lopez R, Finn RD, 2018. HMMER web server: 2018 update. Nucleic acids research 46, W200–W204. [DOI] [PMC free article] [PubMed] [Google Scholar]
  149. Priest SJ, Yadav V, Roth C, Dahlmann TA, Kück U, Magwene PM, Heitman J, 2022. Uncontrolled transposition following RNAi loss causes hypermutation and antifungal drug resistance in clinical isolates of Cryptococcus neoformans. Nature microbiology 7, 1239–1251. [DOI] [PMC free article] [PubMed] [Google Scholar]
  150. Quan Y, Zhou X, Belmonte-Lopes R, Li N, Wahyuningsih R, Chowdhary A, Hawksworth DL, Stielow JB, Walsh TJ, Zhang S, 2025. Potential predictive value of phylogenetic novelties in clinical fungi, illustrated by Histoplasma. IMA fungus 16, e145658. [DOI] [PMC free article] [PubMed] [Google Scholar]
  151. R Core Team, Rf., 2018. R: A language and environment for statistical computing. [Google Scholar]
  152. Rawlings ND, Waller M, Barrett AJ, Bateman A, 2014. MEROPS: the database of proteolytic enzymes, their substrates and inhibitors. Nucleic acids research 42, D503–D509. [DOI] [PMC free article] [PubMed] [Google Scholar]
  153. Revell LJ, 2012. phytools: an R package for phylogenetic comparative biology (and other things). Methods in ecology and evolution 217–223. [Google Scholar]
  154. Revell LJ, Revell MLJ, 2014. Package ‘phytools.’ Website: https://cran.r-project.org/web/packages/phytools. [Google Scholar]
  155. Rezabek GB, Donahue JM, Giles RC, Petrites-Murphy MB, Poonacha KB, Rooney JR, Smith BJ, Swerczek TW, Tramontin RR, 1993. Histoplasmosis in horses. Journal of comparative pathology 109, 47–55. [DOI] [PubMed] [Google Scholar]
  156. Rizzon C, Marais G, Gouy M, Biémont C, 2002. Recombination rate and the distribution of transposable elements in the Drosophila melanogaster genome. Genome research 12, 400–407. [DOI] [PMC free article] [PubMed] [Google Scholar]
  157. Roberts WR, Siepielski AM, Alverson AJ, 2024. Diatom abundance in the polar oceans is predicted by genome size. PLoS Biology 22, e3002733. [DOI] [PMC free article] [PubMed] [Google Scholar]
  158. Schäffer AA, Nawrocki EP, Choi Y, Kitts PA, Karsch-Mizrachi I, McVeigh R, 2018. VecScreen_plus_taxonomy: imposing a tax (onomy) increase on vector contamination screening. Bioinformatics 34, 755–759. [DOI] [PMC free article] [PubMed] [Google Scholar]
  159. Schley RJ, Pellicer J, Ge X, Barrett C, Bellot S, Guignard MS, Novák P, Suda J, Fraser D, Baker WJ, Dodsworth S, Macas J, Leitch AR, Leitch IJ, 2022. The ecology of palm genomes: repeat-associated genome size expansion is constrained by aridity. New Phytologist 236, 433–446. 10.1111/nph.18323 [DOI] [PMC free article] [PubMed] [Google Scholar]
  160. Schrader L, Kim JW, Ence D, Zimin A, Klein A, Wyschetzki K, Weichselgartner T, Kemena C, Stökl J, Schultner E, 2014. Transposable element islands facilitate adaptation to novel environments in an invasive species. Nature communications 5, 5495. [DOI] [PMC free article] [PubMed] [Google Scholar]
  161. Schrader L, Schmitz J, 2019. The impact of transposable elements in adaptive evolution. Molecular Ecology 28, 1537–1549. 10.1111/mec.14794 [DOI] [PubMed] [Google Scholar]
  162. Sepúlveda VE, Goldman WE, Matute DR, 2024. Genotypic diversity, virulence, and molecular genetic tools in Histoplasma. Microbiol Mol Biol Rev 88, e00076–23. 10.1128/mmbr.00076-23 [DOI] [PMC free article] [PubMed] [Google Scholar]
  163. Sepúlveda VE, Márquez R, Turissini DA, Goldman WE, Matute DR, 2017. Genome sequences reveal cryptic speciation in the human pathogen Histoplasma capsulatum. MBio 8, e01339–17. [DOI] [PMC free article] [PubMed] [Google Scholar]
  164. Serrato-Capuchina A, Matute DR, 2018. The role of transposable elements in speciation. Genes 9, 254. [DOI] [PMC free article] [PubMed] [Google Scholar]
  165. Sharma SP, Zuo T, Peterson T, 2021. Transposon-induced inversions activate gene expression in the maize pericarp. Genetics 218, iyab062. [DOI] [PMC free article] [PubMed] [Google Scholar]
  166. Sil A, 2019. Molecular regulation of Histoplasma dimorphism. Current opinion in microbiology 52, 151–157. [DOI] [PMC free article] [PubMed] [Google Scholar]
  167. Singh NK, Badet T, Abraham L, Croll D, 2021. Rapid sequence evolution driven by transposable elements at a virulence locus in a fungal wheat pathogen. BMC Genomics 22, 393. 10.1186/s12864-021-07691-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  168. Slater GSC, Birney E, 2005. Automated generation of heuristics for biological sequence comparison. BMC Bioinformatics 6, 31. 10.1186/1471-2105-6-31 [DOI] [PMC free article] [PubMed] [Google Scholar]
  169. Smit AF, 2004. Repeat-masker open-3.0. http://www.repeatmasker.org. [Google Scholar]
  170. Stajich JE, Lovett B, Lee E, Macias AM, Hajek AE, de Bivort BL, Kasson MT, Licht HHDF, Elya C, 2024. Signatures of transposon-mediated genome inflation, host specialization, and photoentrainment in Entomophthora muscae and allied entomophthoralean fungi. Elife 12, RP92863. [DOI] [PMC free article] [PubMed] [Google Scholar]
  171. Stanke M, Morgenstern B, 2005. AUGUSTUS: a web server for gene prediction in eukaryotes that allows user-defined constraints. Nucleic Acids Res 33, W465–467. 10.1093/nar/gki458 [DOI] [PMC free article] [PubMed] [Google Scholar]
  172. Stapley J, Santure AW, Dennis SR, 2015. Transposable elements as agents of rapid adaptation may explain the genetic paradox of invasive species. Molecular Ecology 24, 2241–2252. 10.1111/mec.13089 [DOI] [PubMed] [Google Scholar]
  173. Steele PE, Carle GF, Kobayashi GS, Medoff G, 1989. Electrophoretic Analysis of Histoplasma capsulatum Chromosomal DNA. Molecular and Cellular Biology 9, 983–987. 10.1128/mcb.9.3.983-987.1989 [DOI] [PMC free article] [PubMed] [Google Scholar]
  174. Sudnik P, Passarelli P, Branche A, Giampoli E, Louie T, 2024. Histoplasmosis Associated With Bat Guano Exposure in Cannabis Growers: 2 Cases, in: Open Forum Infectious Diseases. Oxford University Press US, p. ofae711. [DOI] [PMC free article] [PubMed] [Google Scholar]
  175. Sun C, Shepard DB, Chong RA, López Arriaza J, Hall K, Castoe TA, Feschotte C, Pollock DD, Mueller RL, 2012. LTR retrotransposons contribute to genomic gigantism in plethodontid salamanders. Genome biology and evolution 4, 168–183. [DOI] [PMC free article] [PubMed] [Google Scholar]
  176. Sundaram V, Cheng Y, Ma Z, Li D, Xing X, Edge P, Snyder MP, Wang T, 2014. Widespread contribution of transposable elements to the innovation of gene regulatory networks. Genome research 24, 1963–1976. [DOI] [PMC free article] [PubMed] [Google Scholar]
  177. Taylor RL, Shacklette MH, Kelley HB, 1962. Isolation of Histoplasma capsulatum and Microsporum gypseum from soil and bat guano in Panama and the Canal Zone. The American Journal of Tropical Medicine and Hygiene 11, 790–795. [DOI] [PubMed] [Google Scholar]
  178. Torres DE, Oggenfuss U, Croll D, Seidl MF, 2020. Genome evolution in fungal plant pathogens: looking beyond the two-speed genome model. Fungal Biology Reviews 34, 136–143. [Google Scholar]
  179. Torres DE, Thomma BP, Seidl MF, 2021. Transposable elements contribute to genome dynamics and gene expression variation in the fungal plant pathogen Verticillium dahliae. Genome Biology and Evolution 13, evab135. [DOI] [PMC free article] [PubMed] [Google Scholar]
  180. Tullume-Vergara PO, Ludwig A, Yurchenko V, Coelho AC, Coser EM, Krieger MA, Teixeira MM, Shaw JJ, Alves JM, 2025. Comparative analysis of the mobilome yields new insights into its diversity, dynamics and evolution in parasites of the Trypanosomatidae family. Parasitology 1–16. [DOI] [PMC free article] [PubMed] [Google Scholar]
  181. Vendrell-Mir P, Barteri F, Merenciano M, González J, Casacuberta JM, Castanera R, 2019. A benchmark of transposon insertion detection tools using real data. Mobile DNA 10, 53. 10.1186/s13100-019-0197-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  182. Vite-Garín T, Estrada-Bárcenas DA, Gernandt DS, Reyes-Montes M. del R., Sahaza, Canteros CE, Ramírez JA, Rodríguez-Arellanes G, Serra-Damasceno L, Zancopé-Oliveira RM, 2021. Histoplasma capsulatum isolated from Tadarida brasiliensis bats captured in Mexico form a sister group to North American class 2 clade. Journal of Fungi 7, 529. [DOI] [PMC free article] [PubMed] [Google Scholar]
  183. Volff J, 2006. Turning junk into gold: domestication of transposable elements and the creation of new genes in eukaryotes. BioEssays 28, 913–922. 10.1002/bies.20452 [DOI] [PubMed] [Google Scholar]
  184. Voorhies M, Cohen S, Shea TP, Petrus S, Muñoz JF, Poplawski S, Goldman WE, Michael TP, Cuomo CA, Sil A, 2022. Chromosome-level genome assembly of a human fungal pathogen reveals synteny among geographically distinct species. MBio 13, e02574–21. [DOI] [PMC free article] [PubMed] [Google Scholar]
  185. Wacker T, Helmstetter N, Wilson D, Fisher MC, Studholme DJ, Farrer RA, 2023. Two-speed genome evolution drives pathogenicity in fungal pathogens of animals. Proc. Natl. Acad. Sci. U.S.A. 120, e2212633120. 10.1073/pnas.2212633120 [DOI] [PMC free article] [PubMed] [Google Scholar]
  186. Wei T, Simko V, Levy M, Xie Y, Jin Y, Zemla J, 2017. Package ‘corrplot.’ Statistician 56, e24. [Google Scholar]
  187. Wicker T, Sabot F, Hua-Van A, Bennetzen JL, Capy P, Chalhoub B, Flavell A, Leroy P, Morgante M, Panaud O, 2007. A unified classification system for eukaryotic transposable elements. Nature reviews genetics 8, 973–982. [DOI] [PubMed] [Google Scholar]
  188. Wiens JJ, 2006. Missing data and the design of phylogenetic analyses. Journal of biomedical informatics 39, 34–42. [DOI] [PubMed] [Google Scholar]
  189. Wiens JJ, Ackerly DD, Allen AP, Anacker BL, Buckley LB, Cornell HV, Damschen EI, Jonathan Davies T, Grytnes J, Harrison SP, Hawkins BA, Holt RD, McCain CM, Stephens PR, 2010. Niche conservatism as an emerging principle in ecology and conservation biology. Ecology Letters 13, 1310–1324. 10.1111/j.1461-0248.2010.01515.x [DOI] [PubMed] [Google Scholar]
  190. Wilmes D, Mayer U, Wohlsein P, Suntz M, Gerkrath J, Schulze C, Holst I, von Bomhard W, Rickerts V, 2022. Animal Histoplasmosis in Europe: Review of the literature and molecular typing of the etiological agents. Journal of Fungi 8, 833. [DOI] [PMC free article] [PubMed] [Google Scholar]
  191. Zhang H, Yohe T, Huang L, Entwistle S, Wu P, Yang Z, Busk PK, Xu Y, Yin Y, 2018. dbCAN2: a meta server for automated carbohydrate-active enzyme annotation. Nucleic acids research 46, W95–W101. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Material

Data Availability Statement

The github repositories along with the saved Zenodo repositories can be found for each portion of the analysis below. Each step was split into different github repositories for clarity and brevity. https://github.com/tania-k/Histoplasma_Assembly doi: 10.5281/zenodo.12747066 https://github.com/tania-k/Histo_Annotations doi: 10.5281/zenodo.12747074 https://github.com/tania-k/Histo_TE_pipeline doi: 10.5281/zenodo.12747068 https://github.com/dturissini/Histoplasma_TE_manuscript

RESOURCES