Skip to main content
Scientific Reports logoLink to Scientific Reports
. 2025 Nov 21;15:41249. doi: 10.1038/s41598-025-25108-y

Enhanced blood parasite species identification using V4–V9 18S rDNA barcoding by universal primers on a nanopore platform

Tatsuki Sugi 1,✉,#, Patrick Reteng 1,#, Alex K Gaithuma 1, Naoko Kawai 1, Boniface Namangala 2, Mable M Mutengo 3, Nobuhisa Ishiguro 4, Kyoko Hayashida 1,2,6, Junya Yamagishi 1,5,6,7
PMCID: PMC12639079  PMID: 41271826

Abstract

Microscopic examination is commonly used for blood parasite detection in resource-limited settings due to its low cost and simplicity. However, it requires microscopy experts and has poor species-level identification. This study proposes a targeted next-generation sequencing (NGS) approach using a portable nanopore platform to enable accurate and sensitive parasite detection in such settings. To improve species identification on the error-prone nanopore sequencer, we designed a DNA barcoding strategy targeting the 18S rDNA V4–V9 region, which outperformed the commonly used V9 region. To enrich parasite DNA and reduce host contamination, two blocking primers were developed: a C3 spacer-modified oligo competing with the universal reverse primer and a peptide nucleic acid (PNA) oligo that inhibits polymerase elongation. These were combined to selectively reduce the amplification of host’s DNA from blood samples. The developed targeted NGS test successfully detected Trypanosoma brucei rhodesiense, Plasmodium falciparum, and Babesia bovis in human blood samples spiked with as few as 1, 4, and 4 parasites per microliter, respectively. Validation study using field cattle blood samples revealed that this test could detect multiple Theileria species co-infections in the same cattle. The established parasite targeted NGS test using a portable nanopore platform enables comprehensive parasite detection with high sensitivity and accurate species identification.

Supplementary Information

The online version contains supplementary material available at 10.1038/s41598-025-25108-y.

Keywords: Portable nanopore sequencing, Parasite targeted NGS test, DNA barcoding, 18S rDNA, Host blocking primer

Subject terms: Next-generation sequencing, Targeted resequencing, Malaria, Parasitic infection

Introduction

Blood parasites of medical and veterinary importance are found in several taxonomical lineages within eukaryotic organisms. Protozoan parasites from the phylum Apicomplexa include the known etiologies of malaria (order Haemosporida) and piroplasmosis (order Piroplasmida), and those from phylum Euglenozoa include the causative etiologies of trypanosomiasis and leishmaniasis (order Trypanosomatida). Helminth parasites also range throughout two taxonomic lineages, including parasites from the phylum Nematoda causing angiostrongyliasis, onchocerciasis, loiasis, and wuchereriasis (order Rhabditida) and another phylum Platyhelminthes causing schistosomiasis (order Strigeidida)1.

For parasitic diseases, timely diagnosis is important for both medical treatment and management of diseases. Microscopic analysis is an affordable diagnostic method to broadly detect haemoparasites rapidly without any sophisticated laboratory setups. For example, the World Health Organization (WHO) guidelines for malaria diagnosis2 positioned the microscopic test as one of the first-line diagnostic tests. The microscopic test could capture even unrecognized parasites. For example, Singh et al. reported a wide endemic of the monkey malaria parasite Plasmodium knowlesi among human malaria patients in Malaysia3. Before the report, the parasite was not considered a human malaria parasite and was misdiagnosed as P. malariae3. This case highlights a weak point of the microscopic test for parasite species identification and highlights the power of the test in detecting parasites comprehensively.

Besides microscopic examination, blood parasites can also be detected by identifying parasite nucleic acids through nucleic acid amplification tests (NAATs) or by detecting parasite antigens or antibodies with immunological tests, in general. Conventional NAATs, including polymerase chain reaction (PCR), loop-mediated isothermal amplification (LAMP), and realtime PCR, use specific primers to target parasites and thereby achieve high sensitivity. Immunological tests, particularly rapid diagnostic tests (RDTs), offer quick and cost-effective detection; however, they rely on specific antibodies to recognize parasite antigens or on specific antigens to capture antibodies against the parasites. For example, in malaria detection, both NAATs and RDTs have demonstrated high sensitivity and have been proposed as alternative methods in settings where expert microscopists are difficult to secure4,5. However, these molecular tests can only detect the targeted parasites. Thus, prior knowledge of the target parasite is required. Recently, unrecognized parasites have been reported to be associated with human diseases. One example of a novel parasite in human diseases is the Colpodella-like parasite6. This type of unexpected novel parasite cannot be detected by molecular testing targeting specific pathogens of interest. Therefore, test methods that complement microscopic tests while maintaining comprehensiveness are required.

To achieve comprehensive pathogen detection from patient samples in laboratory settings where RDTs or microscopic tests are the first choice, the application of metagenomic next-generation sequencing (mNGS) analysis using a portable nanopore sequencing platform has been expanded in virology7,8 and microbiology9,10. Because the portable nanopore sequencing platform has lower sequencing outputs than high-capacity NGS platforms like Illumina, those reports enriched pathogen-specific sequences (such as 16S rDNA for bacteria) before NGS analysis to reduce the required amount of sequence data to obtain pathogen sequences, the strategy called as a targeted NGS. To apply this approach to parasite detection, an 18S rDNA barcode11 is a promising tool applicable to eukaryotic pathogens. When clinical specimens like fecal samples with low host cell amounts were investigated, researchers successfully applied 18S rDNA barcoding for parasite detection12,13. However, if the clinical specimens contain a lot of host cells, like whole blood samples, the 18S rDNA barcode approach will suffer from overwhelming host 18S rDNA, which will also be amplified by the universal primer for 18S rDNA barcodes.

To overcome these hosts’ overwhelming problems, Flaherty et al. used restriction enzymes to specifically cut the host 18S rDNA barcodes to detect parasites14. Another workaround involves sequence-dependent amplification suppression using blocking primers15,16. Blocking primers are designed with sequences specific to the target DNA, ensuring precise binding to the template15,16. They carry a 3′-terminal modification, such as a C3 spacer, to halt polymerase extension15. Alternatively, peptide nucleic acid (PNA) can be used to block extension at the binding site16. This combination of sequence-specific binding and polymerase inhibition suppresses amplification of the target sequence during PCR. Lager et al. used a combination of universal primers to amplify eukaryotic organisms and a blocking primer to suppress the mammalian 18S rDNA barcodes to achieve enrichment of eukaryotic pathogen barcodes17. The approaches described above used short 18S rDNA amplicons, which can be analyzed using highly accurate Illumina NGS platforms. This DNA barcode length will cause problems with parasite species identification using error-prone portable nanopore sequencing platforms.

To enhance species-level bacterial identification using the portable nanopore sequencing platform, nearly full length 16S rDNA barcodes were used18. For the identification of parasites in blood samples using the portable nanopore sequencing platform, Huggins et al. successfully used > 1 kb 18S rDNA barcodes for species-level parasite identification, although the target parasites were restricted to apicomplexan parasites19. To date, no parasite targeted NGS test for blood samples using universal primers for wide taxa of eukaryotic organisms with the portable nanopore sequencing platform has been reported.

In this report, we designed a parasite targeted NGS test using a portable nanopore sequencing device. To obtain species-level resolutions, we selected universal primers to obtain > 1 kb 18S rDNA. To avoid overwhelming host 18S rDNA, we designed novel blocking primers to suppress mammalian 18S rDNA amplification. By using three blood parasites, Trypanosoma, Plasmodium, and Babesia, as model blood parasites, the performance of the established tests was validated with parasite-spiked human blood samples. At last, established tests were applied to the field cattle specimens.

Results

Primer design

To cover blood parasites from several taxonomic lineages and to get enough species-level resolutions, universal primers covering a high ratio of eukaryotic organisms with a > 1 kb 18S rDNA barcode were selected (Table 1; Fig. 1A). Target regions spanned from variable area 4 (V4) to V9, based on the naming scheme according to the previous publication20. To simulate the classification of parasite species with error-prone portable NGS sequence data, one thousand error-containing sequences for V9 and V4-9 were obtained by introducing random mutations with various error ratios per base from the reference 18S rDNA sequence of P. falciparum, knowlesi, malariae, ovale and vivax. When these sequences were classified by blastn with modified parameters (-task blastn for somewhat similar sequences), up to 1.7% o top hit were misassigned to another species with V9 regions depending on the error rate (Table 2). When these sequences were classified by ribosomal database project (RDP) naive Bayesian classifier method, proportion of sequences without classification above the bootstrap values threshold (> 50%) increased in error rate dependent manner with V9 region (Fig. 1B).

Table 1.

Primer sequences.

Name(a) Sequence (5’-3’)
F566 CAGCAGCCGCGGTAATTCC
1776R TACRGMWACCTTGTTACGAC
3SpC3_Hs1829R CGACTTTTACTTCCTCTAGATAGTCIIIIIIGACCGTCTTCTCAGCGCTCCG-3SpC3(b)
PNA_Hs733F(c) CCCCGCCCCTTGCCTC

(a) Numbers in primer name are derived from the 5’ first base position of primer annealing site in S. cerevisiae 18S rRNA sequence (NR_132213.1: RDN18-1).

(b) 3SpC3 : C3-spacer modification at 3’ end of the primer.

(c) A number of primer name starting from “Hs” is derived from the 5’ first base position of primer annealing site in Human 18S rRNA sequence (NR145820: RNA18SN1).

Underlined sequences of 1776R and 3SpC3_Hs1829R bind to the same location in 18S rDNA.

Fig. 1.

Fig. 1

18S rDNA barcode and primers in this study. (A) Schematic illustration of the primers used in this study. Universal Primers (F566 and 1776R, Table 1) to amplify V4-9 region are shown in white arrows. Blocking primers (PNA_Hs733F and 3SpC3_Hs1829R) are shown in black arrows. Variable areas of 18S rDNA gene are shown in white boxes. Universal primers (1391f and EukBr) targeting V9 region are shown in grey arrows. Size of the sequence and primer arrows are not scaled to the actual lengths. 3SpC3_Hs1829R and 1776R have actual overlapping bases for four bases. (B) RDP naive bayesian classifier method based classification results of the simulated error containing 18S rDNA barcodes. Proportions of the assigned classification in species levels are shown. NA means that taxonomic assignment were not beyond the bootstrap value threshold (> 50%).

Table 2.

BLASTn based miss-classified sequence number out of thousand simulated error containing 18S rDNA barcodes.

error rate
0.05 0.1
18S rDNA barcode regions V4-9 V9 V4-9 V9 misassigned species
P. falciparum 0 4 0 10

P. sp. DRC-Itaito

P. gabonim

P. sp. gorilla clade G2

P. knowlesi 0 4 0 9

P. vivax

Hepatocystis sp. MFRC11

P. cynomologi

P. falciparum

P. ovale 0 0 0 2

P. sp DRC-Itaito

P. vivax

P. vivax 0 3 0 17

P. fragile

P. cynomolgi

P. coatneyi

P. inui

P. knowlesi

P. malariae 0 0 0 0 NA

The parameter adjustment of blast search was critical to get similar sequence by blastn with error prone sequence data. With default parameter settings of balstn (-task megablast), more than 50% of the sequences were classified as “no hit” when V9 sequences were analyzed, depending on the error rates even with the large database of NCBI nt database (Supplemental Fig. 1).

F566 primer targets the fourth conserved area before the V4 (Figs. 1A and 2A), 1776R primer targets the 10th conserved area after V9 (Figs. 1A and 2B), and both primers covered the wide range of eukaryotic organisms including the representative eukaryotic pathogens from kingdom Fungi, phylum Nematoda, Platyhelminth, Apicomplexa and Euglenozoa (Fig. 2A and B). Of note, six nucleotides (nts) from the 5’ terminus of F566 had one mismatch with the 18S rDNA of Trypanosoma cruzi, T. brucei, and Leishmania donovani.

Fig. 2.

Fig. 2

Multiple Sequence alignment of primer annealing site of 18S rDNA barcodes. Regions for Universal primer F566 (A), 1776R and blocking primer 3SpC3_Hs1829R (B), and PNA_Hs733F (C) are shown. Bases in human 18S rDNA are shown in color background and bases different from human sequence in the corresponding regions are shown in color background. Black boxes show universal primer annealing sites and red boxes show blocking primers annealing sites. Reverse primers are shown in reverse complement from the actual primer sequences. Reverse complement of Inosine bases for 3SpC3_Hs1829R are shown as N for visualization.

Next, non-biased ribosomal RNA small subunit (SSU) sequences from all domains (bacteria, archaea, and eukaryotes) were analyzed with the primer annealing sites. SSU sequences were retrieved from Silva database21. The number of organisms having SSU sequence(s) with primer annealing sites was counted. Universal primers (F566 and 1776R, Table 1) had annealing sites with fewer than three total mismatches in over 60% ad less than 1% of SSU entries from eukaryotic or non-eukaryotic organisms, respectively (Supplemental Fig. 2A). This primer covers a reasonable number of organisms deposited to the database in the following taxonomic lineages for blood parasites, order Haemosporida and Piroplasmida in phylum Apicomplexa, order Trypanosomatida in phylum Euglenozoa, order Rhabditida in phylum Nematoda, and order Strigeidida in phylum Platyhelminth (Supplemental Fig. 2B).

Blocking primers to suppress the overwhelming host DNA amplification by universal primer

To suppress overwhelming host DNA amplification by the universal primer for pan-eukaryotic organisms, we designed two blocking primers (Tables 1 and Fig. 1A). 3SpC3_Hs1829R was designed to overlap with universal reverse primer 1776R, to have C3 spacer modification at 3’ terminal to block the extension of the polymerase by competitive clumping manner, and to have host mammalian 18S rDNA specific sequence at 3’ end with middle six inosine babbles, according to the previously validated blocking primer design scheme22. Sequence alignment shows that mammalian host sequences have the blocking primer annealing site with fewer than two mismatches, but the other eukaryotic organisms have more than ten mismatches (Fig. 2B). In silico search of this blocking primer sequence in the SSU sequences in the SILVA database also showed that this blocking primer is selective to mammalian hosts but not in the other eukaryotic organisms for known blood pathogens (Supplemental Fig. 2C). Another blocking primer, PNA_Hs733F, was designed to have perfect matches to mammalian sequences in the V4 region (Figs. 1A and 2C) and was selected to have the least appearance of the blocking primer sequences in non-mammalian sequences throughout the SSU sequence between universal primers F566 and 1776R (Supplemental Fig. 1D). PNA_Hs733F was synthesized to have a peptide backbone to suppress the polymerase extension by an elongation arrest manner. To observe the effect of the designed blocking primers, DNA extracted from human whole blood spiked with T. brucei rhodesiense parasites was amplified with universal primers with various concentrations of blocking primers (Fig. 3). Both of 3SpC3_Hs1829R (Fig. 3A) and PNA_Hs733F (Fig. 3B) suppressed the amplification of human 18S rDNA in a concentration-dependent manner, whereas the amplification of parasite 18S rDNA was not affected. The combination of two blocking primers further suppressed human 18S rDNA amplification (Supplemental Fig. 3). These qualitative results were further validated with the quantitative analysis including the nanopore portable NGS sequencing analysis in the following experiments.

Fig. 3.

Fig. 3

Concentration dependent effects of blocking primers. Various concentrations of blocking primers were added while amplifying 18S rDNA barcode from the T. brucei rhodesiense spiked human blood samples (105 parasites/mL) for 3SpC3_Hs1829R (A) and PNA_Hs733F (B). Bands above 1500 bp are Trypanosoma derived 18S rDNA and ~ 1200 bp amplicons are Human derived 18S rDNA. When 3SpC3_Hs1829R was added to the PCR reaction, primer dimers were observed below 250 bp.

Sensitivity of the designed parasite targeted NGS test using universal primers

Next, we combined DNA barcode amplification using universal primers and blocking primers with actual NGS analysis using a portable nanopore sequencing device to establish a parasite targeted NGS test. For the proof of concept, various numbers of parasites were spiked in the healthy human whole blood, and the extracted DNA was used for the PCR amplification of 18S rDNA barcodes. PCR without blocking primers using the universal primer amplifies human 18S rDNA, detected as dominant bands around 1.2 kb throughout the various numbers of T. brucei rhodesiense (mock control and from 10, 102, 103, 104 parasites/mL) spiked samples (Fig. 4A, lanes 2–6). When the blocking primers were added to the PCR reaction, Trypanosoma-derived 18S rDNA amplicons (~ 1.5 kb) were visible up to the spiked samples with 103 parasites/mL (Fig. 4A, lane 11). NGS analysis of those amplicons by the portable nanopore sequencing device also validated that Trypanosoma sequences were detectable from the samples with 103 and 104 parasites/mL (Fig. 4A, lanes 11 and 12, respectively). With this experiment, we observed a significant proportion of the sequences obtained by the portable nanopore sequencing was found to be nonspecifically amplified genomic regions from host genome (host non-18S) (Fig. 4A). Thus, to decrease the possibility of amplifying those DNA non-specifically with many PCR cycles, we decreased the PCR cycles from 40 to 35 in following experiments. When P. falciparum and B. bovis was spiked in the healthy human whole blood (mock control and from 4 × 103 to 1.5 × 106parasites/mL, lanes 1 to 6) PCR with blocking primers successfully amplified detectable 18S rDNA barcdoes (Fig. 4B and C, respectively). As common in the other eukaryotic organisms, B. bovis (Fig. 4C) has a matching size of 18S rDNA with human 18S rDNA (~ 1.2 kb), thus the DNA electrophoresis itself cannot ascertain the origin of the amplicon. NGS analysis of those amplicons by portable nanopore sequencer validated that parasites derived sequences were dominant with the samples from 2 × 104 and 4 × 103 parasites/mL, (for P. falciparum [lane 3] and B. bovis [lane2], respectively). Of note, some of the reads in the PCR amplicon corresponding to the lane 6 in B. bovis are classified as Trichosporon insectorum, suggesting that the environmental contamination during the PCR steps might happen.

Fig. 4.

Fig. 4

Selective enrichment of parasites 18S rDNA from parasite spiked samples. (A) PCR products from human blood samples spiked with various number of T. brucei rhodesiense are shown. Lanes 1 and 7, no template control. Lanes 2–6 and 8–12, parasite spiked whole blood samples (0, 10, 100, 1000 and 10000 parasites/mL, respectively). Forty cycles of PCR reaction were used for this experiment. (B, C) PCR products from blood samples spiked with various number of ring stage P. falciparum (B) or B. bovis (C) are shown. Lane 1: no template control. Lanes 2–6: parasite spiked whole blood samples (0, 4000, 20000, 105, 5 × 105, 1.5 × 106 parasites/mL, respectively). Thirty-five cycles of PCR reactions were used to obtain these results. (A, B, and C) Bottom panels show the portable nanopore sequencing results of PCR amplicons. Up to five thousand reads were analyzed. host 18S rDNA were first removed by the blast search against eukaryotic 18S rDNA database. All the sequences belonging to phylum Chordata were considered to be host derived barcodes and are shown as red box: “host 18S”. Non-host 18S sequences were further clustered to give accurate estimation of sequence of origin and classified by the blast search against ncbi nt database. All the PCR reactions were performed with 2 µM 3SpC3_Hs1829R blocking primers and 5 µM PNA_Hs733F blocking primers, except panel A, lanes 1–6, which are the control reaction without any blocking primers.

A validation of the test with field cattle blood samples

Finally, we tested the established parasite targeted NGS test with field samples. For this purpose, the blood DNA from the cattle from the area with a high prevalence of parasite infection was used. We analyzed three cattle samples, of which two were already known to be positive for Theileria mutans and T. velifera, and one sample was negative for those Theileria parasites by the conventional PCR test23.

18S rDNA barcodes were amplified with universal primers, with various concentrations of blocking primers, and amplicons were sequenced to see the organisms inside (Fig. 5). To speed up the analytical steps, we first removed host sequences by blasting them to the 18S rDNA database, and after that, the sequence-clusters were made to obtain the accurate 18S rDNA barcode (Fig. 5). Without any blocking primers, host 18S rDNA sequences accounted for more than 90% of the amplicons from all three samples tested (Fig. 5). The addition of 2 µM 3SpC3_Hs1829R suppressed the amplification of host 18S rDNA, and the combination with PNA_Hs733F (1.25 µM and 5 µM final for labels PNA1 and PNA5, respectively) further suppressed the host 18S rDNA amplification in a concentration-dependent manner (Fig. 5, Supplemental Fig. 4). As a result, the sequence reads assigned to the order of piroplasmida increased from < 1.1% to > 37.5% (Supplemental Fig. 4). After the clustering of the error-containing sequence, we could estimate the accurate 18S rDNA barcode [Accession number LC878354 and LC878355], identical to the sequence deposited to the NCBI genbank (KU206307: T. velifera, and KU206320: T. mutans) from all of the three samples (Fig. 5). One bp deletion in T. mutans sequence from BA48 + PNA5 condition, and 1 bp insertion in T. velifera sequence from NN96 + condition were observed (Supplemental Table 1). Considering the other clustered sequences from the same sample with different PCR conditions supported the identical sequence to deposited 18S rDNA barcodes, those consensus sequences could reflect sequencing error of portable nanopore sequencing device. From amplicon of BA38 by universal primers for 18S, we detected 16S rDNA barcode for Anaplasma marginale [Accession number: LC878356] and another genomic DNA sequence [Accession number: LC878357] (supplemental Table 1), 16S rDNA barcode was 100% identical to the reported 16S rRNA sequence of A. marginale KU686794.1, and genomic DNA sequence [LC878357] had 95.48% identity to the whole genome sequence of A. marginale strain Florida CP001079.1. The presence of this pathogen was further validated by the conventional PCR targeting gltA gene [Accession number: LC878358] to be 99.74% identical to the reported A. marginale gltA sequence OQ185251.1. A. marginale is a causative bacterial agent for Anaplasmosis and the universal primers we used had several mismatches to A. marginale SSU (F566; two mismatches in one and three nts from 3’ terminal, and no mismatches in 1776R with A. marginale SSU sequence: KU686778.1).

Fig. 5.

Fig. 5

Identification of multiple pathogens co-infection from cattle blood samples. An analytical pipeline is shown in the left panel. After the standard quality filtering of the portable nanopore sequencing data, host 18S rDNA were first removed by the blast search against eukaryotic 18S rDNA database. All the sequences belonging to phylum Chordata were considered to be host derived barcodes and are shown as white box: “host”. Non-host 18S sequences were further clustered to give accurate estimation of sequence of origin and classified by the blast search against ncbi nt database. Sequences which were not clustered were classified as “not Clustered” and are shown in black. When less than 80 sequences were obtained as non-host 18S reads, those reads number were considered not enough to get accurate clustering, thus were classified as “not Clustered”.

Discussion

In the present report, we designed a parasite targeted NGS test using a portable nanopore sequencing device. The proposed strategy combined DNA barcoding methods with a portable nanopore sequencing device with the aid of a blocking primer to suppress host DNA amplification. The concept of the test was validated with laboratory samples with a known number of parasites in the healthy donor blood and then further validated with actual field cattle blood samples with a practical number of parasites in the blood. Here we showed that the proposed test had enough sensitivity and performance to identify the parasite in the mammalian blood at species levels.

As a low-cost and comprehensive parasite test, the microscopic test is widely available for the parasites that can be identified morphologically in the blood samples. Regarding Plasmodium and Trypanosoma, which we used as model parasites for detection, sensitivity of the microscopical test is ranging from 50 to 100 parasites / µL (5 × 104 – 105/mL) for malaria24,25, and 500–5000 parasites / mL for African Trypanosomiasis26. The presented parasite targeted NGS test detected four thousand parasites / mL in the blood (4 parasites / µL) for malaria and babesiosis, and one thousand parasites /mL (1 parasite /µL) for trypanosomiasis. These values are equivalent to or better than the limit of detection by the microscopy expert and can be used as an alternative method to complement the microscopical examination in resource-scarce laboratory settings. Furthermore, better sensitivity than microscopic test will help to investigate the situation of the submicroscopic infection of malaria which is known to occur all over the world27. Another aspect of this test that can complement the microscopy test is the species identification performance. Microscopic tests have the weak point of insufficient identification of the parasite species, especially in the co-endemic area28. With > 1 kb of DNA barcode region to be used, our parasite targeted NGS test approach could classify the parasite with species levels, as shown in simulated Plasmodium sequences and as validated in the analysis on field cattle samples.

Several reports proposed the comprehensive parasite detection test from host-rich samples using NGS14,17. These tests are based on the short DNA barcode amplicon, which requires precise sequencing quality by Illumina. Considering that the large-scale samples will be characterized at once in cross-sectional research, the Illumina high-throughput NGS is of choice for this kind of study. However, in the resource-limited laboratory settings, the portable nanopore sequencing platform is much easier to implement and for small number of samples like the case of diagnosis usage, small scale NGS is cost efficient. Semi-comprehensive parasites test for blood samples using the portable nanopore sequencing device are also reported by Huggins et al., recently19,29,30. That strategy focused on the known pathogenic blood pathogens by combining multiple PCR tests for each specific pathogen of interest. PCR for the specific target pathogen can solve the problem of overwhelming host-derived DNA barcodes. If there is a priori knowledge of the suspected pathogens, those strategies should be of choice. In our present report, by using the blocking primer, we suppressed the amplification of host DNA rather than amplifying the target pathogen DNA barcode specifically. Thus, our strategy is not restricted by any prior target pathogens.

In this proposed pathogen DNA barcoding with host depletion strategy, we used two different mechanisms of blocking primers. 3SpC3_Hs1829R was designed as a competitive clamping blocking primer with overlapping bases with universal primer 1776R. This could be possible because the V9 region, close to 1776R, had mammalian-specific sequences. As described elsewhere16, the blocking effect was greater in 3SpC3_Hs1829R than in PNA-Hs733F, which used the elongation arrest blocking. The combination of these two primers targeting the sense strand/antisense strand synthesis worked in an additive manner to gain the extensive host suppression effects. Especially in the field cattle samples BA48, the previous attempts of screening Piroplasmida DNA by reverse line blot (RLB)-PCR had a negative result23. Reflecting the negative result from the conventional PCR assay, application of only 3SpC3_Hs1829R blocking primer to the BA48 sample resulted in few reads of Piroplasmida 18S rDNA sequences in the amplicon of BA48 (supplemental Fig. 4) and not resulted in the enough number of sequences to retrieve sequence cluster to be assigned (Fig. 5). Even in this low parasite burden sample, the combination of two blocking primers could suppress the host 18S rDNA amplification, so that we were able to detect two Theileria co-infections from the same sample.

There are several limitations in our present study. Although we designed the universal primers to cover a substantial proportion of the eukaryotic DNA barcodes deposited in the SILVA database, some of the sequences have mismatches in the primer binding sites. Of note, the parasites in our validation study, T. brucei rhodesiense in order Trypanosomatida, have one mismatch base in the primer binding sites. Nevertheless, we obtained good sensitivity for T. brucei rhodesiense spiked samples. Additionally, we observed the presence of bacterial Anaplasma 16S rDNA barcode in sample BA38, to find A. marginale infection, which is known to infect in the red blood cells of cattle to cause anaplasmosis. These results suggest that the PCR condition we used are permissive to several primer mismatches to broaden the coverage of the pathogens, whereas the host suppression by the blocking primers was sufficient to allow those non-optimal amplifications over compete with the host 18S rDNA amplification. Collectively, our universal 18S rDNA primers and PCR conditions may have the potential to amplify related 16S rDNA from certain bacteria including Anaplasma. Therefore, the presence of non-host 18S rDNA cannot be confirmed by amplicon size alone; NGS analysis is required to identify the parasite.

Another limitation is that our validation study was only conducted with several parasite species P. falciparum, T. brucei, and B. bovis in the human blood as clean laboratory setups, and Theileria spp. in cattle blood as field sample setups which we could readily find co-infections of several parasites. Although the universal primers and blocking primers were designed to amplify a broad range of parasites and to suppress amplification of mammalian host sequences, the performance of detection needed to be validated for each parasite of interest under actual infection conditions. For example, in the present study, detection of the human parasite was tested only under laboratory conditions. Therefore, future studies using clinical or field samples were required to confirm the sensitivity of the proposed test and to determine whether the approach described here was applicable in clinical settings. Different host species may affect the efficacy of blocking primers, due to the difference in the sequences of the primer annealing sites. Although sequence analysis ensures that almost all the mammalian host animals can be targeted by the blocking primers proposed, the prior test should be conducted depending on the target host animals for the efficiency of the blocking primers.

When the proposed test is compared with the other widely used parasite NAATs, there are several advantages and disadvantages. For example, in the test for malaria, simple and cost affordable NAATs (including PCR, LAMP, realtime PCR etc.) can detect less than 1 parasite / µL31, outperforms our proposed current test methods in sensitivity. On the other hand, the conventional non-comprehensive type NAATs uses the specific primers for the targeted parasite(s), thus non-suspected parasites (not targeted by specific primers) will be missed. Although, our proposed test utilizes the portable NGS, which enables relatively low-cost NGS analysis, however, it needs additional preparation times and cost for NGS in addition to the PCR step.

Despite some described limitations mentioned above remaining to be solved, as validated by the proof-of-concept work, we propose that the established parasite targeted NGS test with a portable nanopore sequencing device will aid the blood parasites survey.

Methods

Sample collections and ethics approval

Blood samples from healthy human donors were obtained with written informed consent from all subjects according to an approved protocol, which was reviewed by the Hokkaido University ethical committee (ref no SEI023-0114/Z-2023-1). We used three archived cattle DNA samples extracted from whole blood, which had already been screened for the presence of piroplasma infection previously23. BA38 and NN96 were piroplasma-positive by RLB-PCR screening, followed by the species identification by Illumina MiSeq to be positive for both T. velifera and T. mutans. Another BA48 was a piroplasma-negative in the previous RLB-PCR screening23. The bloods for those samples were collected from cattle in Itezhi-Tezhi district in Zambia between April and May 2019 under the protocol approved by ERES Converge IRB, Lusaka, Zambia (ref. No. 2019-Feb-081) as described23. All methods were performed in accordance with the relevant guidelines and regulations. This study is reported in accordance with ARRIVE guidelines.

DNA extraction

Two hundred µL of whole blood was used for the isolation of genomic DNA (gDNA) with the QuickGene DNA Whole-Blood Kit S (Kurabo, Osaka, Japan) following the manufacturer’s instructions. In vitro cultivated parasites were enumerated by hemocytometer (parasites for blood stream form of T. brucei rhodesiense, and total red blood cells for P. falciparum and B. bovis) and Giemsa staining to estimate parasitemia (for P. falciparum and B. bovis). For the P. falciparum, the ring stage of the parasites was used to spike the whole blood to mimic the parasites in the peripheral blood of a patient.

PCR

18S rDNA barcodes were amplified from the extracted gDNA. PCR reactions contain 5 U Kapa taq Extra (KAPA Biosystems, MA, USA), final concentration of 0.3 µM each universal primers (F566 and 1776R) and blocking primers, final concentration of 2 µM 3SpC3_Hs1829R and 5 µM PNA_Hs733F, and 1.5 µL DNA template (< 30 ng total DNA amount) in AmpDirect PCR buffer (Shimadzu Corporation, Kyoto, Japan) in total reaction volume of 15 µl (primer sequences are listed in Table 1). PCR were performed with following thermal cycles, initial denature 95 °C 3 min, amplification cycles [98 °C 20 s-75 °C 15 s (for PNA_Hs733F), 63 °C 15 s (for 3SpC3_Hs1829R), 55 °C 15 s (for universal primers), 72 °C 1 min 30 s] for 35 cycles, and final extension of 72 °C 1 min 30 s, and reactions were stored at < 10 °C before the further analysis. When control reactions without blocking primers were set up, the same thermal cycles were used. For the validation of A. marginale detection, gltA gene was amplified with conventional PCR using primers (F1b and HG1085R) as described in elsewhere32.

Portable NGS sequencing

NGS libraries were prepared from the PCR amplicons using the library preparation kit SQK-NBD114.96 or SQK-NBD112.96 according to the manufacturer’s instructions. To analyze the Plasmodium and Babesia spiked samples, index sequences for each sample were added by an additional twelve cycles of PCR with primers with an 8 bp index nucleotide outside of universal primers using the same PCR conditions mentioned above. If the index sequences were added by PCR, SQK-LSK114 was also used to prepare the sequencing library for nanopore sequencing. Sequencing was conducted on flongle flow cells using a MinION Mk1B device (all the kit and devices were purchased from Oxford Nanopore Technologies).

Sequence analysis

Fast5 files produced with SQK-NBD112.96 were basecalled by guppy (ver 6.5.7 using basecall model: dna_r10.4.1_e8.2_400bps_sup) and POD5 files produced with SQK-NBD114.96 or SQK-LSK114 were basecalled by dorado (ONT basecalling software version 7.2.13 with model: dna_r10.4.1_e8.2_400bps_5khz_sup) to acquire fastq files for the downstream analysis. The fastq sequence files were analyzed as follows. First, the adapter sequences were trimmed off from the end of the sequences by porechop33. Full-length amplicons with universal primer sequences at both ends and with a length range from 800 bp to 1800 bp were extracted using cutadapt34 and seqkit35. To remove host 18S rDNA, NCBI-blastn was used to search for the most similar sequence in the global alignment using the locally built 18S rRNA database built by crabs36. The same blast results were used to estimate the proportion of the spiked parasites and host 18S rDNA. To detect pathogens in an agnostic manner from field samples, sequences after filtering host 18S rDNA were clustered by amplicon_sorter37. Obtained consensus sequence for each sequence cluster was blasted against ncbi nt database.

To estimate the probability of the miss assignment of the DNA barcodes with sequence errors, V4-9 (between primers F566 and 1776R) or V9 (between primers 1391f and EukBr)11 were excised from 18S rDNA sequences for P. falciparum [XR_002966654], P. knowlesi [XR_005506393], P. ovale [AB182491], P. malariae [XR_003751948], and P. vivax [XR_003001217]. Sequencing errors of base substitution were introduced randomly with error ratios of 1, 0.5, and 0.1% per base to make thousand simulated sequences. Mutated sequences were classified by blastn v2.13.0 using SILVA SSU nr99 database with permissive parameter (-task blastn) or searched against nt databse with default parameter (-task megablast), and taxonomical classification was assigned based on the top hit. Thresholds values of (qcov_hsp_perc 70 and evalue 1e-10) were used to get significant blast hits. Sequences without a blast hit were classified as no hit. To infer the probability of the taxon assignment, RDP naive bayesian classifier method38 implemented in dada239 were performed using SILVA SSU nr 99 v132 train set (formatted for dada240. Species identification was performed using default parameter set and assigned taxon with bootstrap values over 50% were analyzed. When subspecies level classification were assigned, those classifications were aggregated to species level classification manually.

Supplementary Information

Below is the link to the electronic supplementary material.

Supplementary Material 1 (1.3MB, docx)
Supplementary Material 2 (12.9KB, xlsx)

Acknowledgements

This study was supported by following grants, Japan Society for the Promotion of Science Core-to-Core program (JPJSCCB20230007), Heiwa Nakajima Foundation, Ohyama Health Foundation, Japan Agency for Medical Research and Development (JP24ym0126801), Japan Program for Infectious Diseases Research and Infrastructure: JIDRI (JP20wm0125008). The funder had no role in the design of the study, data collection, analysis, decision to publish, or preparation of the manuscript.

Author contributions

TS wrote the main manuscript text. TS, PR, NK performed the experiment to acquire data, TS, PR, AG, BN, MM, NI, KH, JY designed the work, interpret data. All authors reviewed the manuscript.

Data availability

All generated and analyzed data during this study are included in this published article and its Supplementary Information files. The pathogen DNA sequences were deposited to DDBJ repository under the accession IDs from LC878354 to LC878358. Blast database built with 18S rDNA barcodes were made with the script available in https://github.com/Tasu/18SmNGS. Additional data is available from the corresponding author, Tatsuki Sugi, upon reasonable request.

Declarations

Competing interests

The author(s) declare the following potential conflict of interest: TS, KH, and JY are inventors on a patent application related to the blocking primers described in this manuscript. The patent application has been filed in Japan under application number JP2025-005095.

Footnotes

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Tatsuki Sugi and Patrick Reteng contributed equally to this work.

References

  • 1.Bogitsh, B., Carter, C. & Oeltmann, T. Human Parasitology 3rd edn (Academic Press, 2005).
  • 2.WHO. WHO guideline for malaria diagnosis. https://app.magicapp.org/#/guideline/LwRMXj/section/L0v9rE
  • 3.Singh, B. et al. A large focus of naturally acquired plasmodium Knowlesi infections in human beings. Lancet363, 1017–1024 (2004). [DOI] [PubMed] [Google Scholar]
  • 4.Lin, K. et al. Evaluation of malaria standard microscopy and rapid diagnostic tests for Screening - Southern Tanzania, 2018–2019. China CDC Wkly.4, 605–608 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Fançony, C., Sebastião, Y. V., Pires, J. E., Gamboa, D. & Nery, S. V. Performance of microscopy and RDTs in the context of a malaria prevalence survey in angola: a comparison using PCR as the gold standard. Malar. J.12, 284 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Yuan, C. L. et al. Colpodella spp.-like parasite infection in woman, China. Emerg. Infect. Dis.18, 125–127 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Reteng, P. et al. A targeted approach with nanopore sequencing for the universal detection and identification of flaviviruses. Sci. Rep.11, 19031 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Xavier, J. et al. A multiplex nanopore sequencing approach for the detection of multiple arboviral species. Viruses16, 23 (2023). [DOI] [PMC free article] [PubMed]
  • 9.Lao, H. Y. et al. The clinical utility of nanopore 16S rRNA gene sequencing for direct bacterial identification in normally sterile body fluids. Front. Microbiol.14, 1324494 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Nakagawa, S. et al. Rapid sequencing-based diagnosis of infectious bacterial species from meningitis patients in Zambia. Clin. Transl Immunol.8, e01087 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Thompson, L. R. et al. A communal catalogue reveals earth’s multiscale microbial diversity. Nature551, 457–463 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Popovic, A. et al. Design and application of a novel two-amplicon approach for defining eukaryotic microbiota. Microbiome6, 228 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Kounosu, A., Murase, K., Yoshida, A., Maruyama, H. & Kikuchi, T. Improved 18S and 28S rDNA primer sets for NGS-based parasite detection. Sci. Rep.9, 15789 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Flaherty, B. R. et al. Restriction enzyme digestion of host DNA enhances universal detection of parasitic pathogens in blood via targeted amplicon deep sequencing. Microbiome6, 164 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Vestheim, H. & Jarman, S. N. Blocking primers to enhance PCR amplification of rare sequences in mixed samples - a case study on prey DNA in Antarctic Krill stomachs. Front. Zool.5, 12 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.von Wintzingerode, F., Landt, O., Ehrlich, A. & Göbel, U. B. Peptide nucleic acid-mediated PCR clamping as a useful supplement in the determination of microbial diversity. Appl. Environ. Microbiol.66, 549–557 (2000). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Lager, S. et al. Detecting eukaryotic microbiota with single-cell sensitivity in human tissue. Microbiome6, 151 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Curry, K. D. et al. Emu: species-level microbial community profiling of full-length 16S rRNA Oxford nanopore sequencing data. Nat. Methods. 19, 845–853 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Huggins, L. G., Colella, V., Young, N. D. & Traub, R. J. Metabarcoding using nanopore long-read sequencing for the unbiased characterization of apicomplexan haemoparasites. Mol. Ecol. Resour.24, e13878 (2024). [DOI] [PubMed] [Google Scholar]
  • 20.Huysmans, E. & De Wachter, R. Compilation of small ribosomal subunit RNA sequences. Nucleic Acids Res.14 Suppl, r73–118 (1986). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Quast, C. et al. The SILVA ribosomal RNA gene database project: improved data processing and web-based tools. Nucleic Acids Res.41, D590–D596 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Belda, E. et al. Preferential suppression of Anopheles Gambiae host sequences allows detection of the mosquito eukaryotic Microbiome. Sci. Rep.7, 3241 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Squarre, D. et al. Investigation of the Piroplasm diversity Circulating in wildlife and cattle of the greater Kafue ecosystem, Zambia. Parasit. Vectors. 13, 599 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Kobayashi, T. et al. Malaria diagnosis across the international centers of excellence for malaria research: Platforms, performance, and standardization. Am. J. Trop. Med. Hyg.93, 99–109 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Joanny, F., Löhr, S. J. Z., Thomas, E., Lell, B., Mordmüller, B. & & Limit of blank and limit of detection of plasmodium falciparum Thick blood smear microscopy in a routine setting in central Africa. Malar. J.13, 234 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Chappuis, F., Loutan, L., Simarro, P., Lejon, V. & Büscher, P. Options for field diagnosis of human African trypanosomiasis. Clin. Microbiol. Rev.18, 133–146 (2005). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Whittaker, C. et al. Global patterns of submicroscopic plasmodium falciparum malaria infection: insights from a systematic review and meta-analysis of population surveys. Lancet Microbe. 2, e366–e374 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Barber, B. E., William, T., Grigg, M. J., Yeo, T. W. & Anstey, N. M. Limitations of microscopy to differentiate plasmodium species in a region co-endemic for plasmodium falciparum, plasmodium Vivax and plasmodium Knowlesi. Malar. J.12, 8 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Huggins, L. G. et al. Assessment of a metabarcoding approach for the characterisation of vector-borne bacteria in canines from Bangkok, Thailand. Parasit. Vectors. 12, 394 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Huggins, L. G. et al. A novel metabarcoding diagnostic tool to explore protozoan haemoparasite diversity in mammals: a proof-of-concept study using canines from the tropics. Sci. Rep.9, 12644 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Patiño, L. H. et al. Validation of real-time PCR assays for detecting plasmodium and Babesia DNA species in blood samples. Acta Trop.258, 107350 (2024). [DOI] [PubMed] [Google Scholar]
  • 32.Inokuma, H., Brouqui, P., Drancourt, M. & Raoult, D. Citrate synthase gene sequence: a new tool for phylogenetic analysis and identification of Ehrlichia. J. Clin. Microbiol.39, 3031–3039 (2001). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Wick, R. R., Judd, L. M., Gorrie, C. L. & Holt, K. E. Completing bacterial genome assemblies with multiplex minion sequencing. Microb. Genom. 3, e000132 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Martin, M. Cutadapt removes adapter sequences from high-throughput sequencing reads. EMBnet J.17, 10 (2011). [Google Scholar]
  • 35.Shen, W., Le, S., Li, Y., Hu, F. & SeqKit A Cross-Platform and ultrafast toolkit for FASTA/Q file manipulation. PLoS One. 11, e0163962 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Jeunen, G. J. et al. crabs-A software program to generate curated reference databases for metabarcoding sequencing data. Mol. Ecol. Resour.23, 725–738 (2023). [DOI] [PubMed] [Google Scholar]
  • 37.Vierstraete, A. R. & Braeckman, B. P. Amplicon_sorter: A tool for reference-free amplicon sorting based on sequence similarity and for Building consensus sequences. Ecol. Evol.12, e8603 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Wang, Q., Garrity, G. M., Tiedje, J. M. & Cole, J. R. Naive bayesian classifier for rapid assignment of rRNA sequences into the new bacterial taxonomy. Appl. Environ. Microbiol.73, 5261–5267 (2007). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Callahan, B. J. et al. DADA2: High-resolution sample inference from illumina amplicon data. Nat. Methods. 13, 581–583 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Morien, E. & Parfrey, L. W. SILVA v128 and v132 dada2 formatted 18s ‘train sets’ (1.0). Zenodo10.5281/zenodo.1447330 (2018).

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Material 1 (1.3MB, docx)
Supplementary Material 2 (12.9KB, xlsx)

Data Availability Statement

All generated and analyzed data during this study are included in this published article and its Supplementary Information files. The pathogen DNA sequences were deposited to DDBJ repository under the accession IDs from LC878354 to LC878358. Blast database built with 18S rDNA barcodes were made with the script available in https://github.com/Tasu/18SmNGS. Additional data is available from the corresponding author, Tatsuki Sugi, upon reasonable request.


Articles from Scientific Reports are provided here courtesy of Nature Publishing Group

RESOURCES