Abstract
In this study, we designed a gene panel based on Nanopore long-read sequencing using adaptive sampling, targeting n = 564 genes associated with Parkinson’s disease (PD) and repeat expansion disorders. We investigated its diagnostic utility in n = 18 patients with (1) pathogenic variants in LRRK2, PRKN, SNCA, and RAB32 (n = 7); (2) idiopathic PD or FTD negative for known genetic causes (n = 3); (3) repeat expansions in ATXN1, ATXN2, ATXN3, C9orf72, DAB1, FGF14, HTT, and TAF1-SVA (n = 8). Three individuals were multiplexed per flow cell, achieving a mean coverage of 24X (SD = ± 9X) per sample. Ultimately, 17/20 (85%; Clopper–Pearson 95% CI: 62–97%) expected pathogenic variants were identified; i.e., an SNCA triplication and two repeat expansions in C9orf72 and TAF1-SVA were missed. Additional variants in STXBP2, AP4S1, RARS2, ALS2, and CNBP were identified in five patients. Adaptive sampling is a versatile genetic diagnostic tool in neurodegenerative disorders, while enabling cost-effective multiplexing. Further validation and improvements in bioinformatic analysis are needed.
Subject terms: Diseases, Genetics, Neurology, Neuroscience
Introduction
Parkinson’s disease (PD) is the second-largest and fastest-growing neurodegenerative disorder, with global patient numbers expected to increase by 112% from 2021 to 2050, primarily due to an aging population1–3. The diagnostic criteria for PD are currently still based on motor features caused by neurodegeneration4–6. The diagnostic process can be challenging due to the heterogeneity of signs and symptoms and is supported by pathologic presynaptic dopaminergic imaging. At present, the α-synuclein seed amplification assay (α-syn SAA)7 is not yet approved for the clinical diagnosis of PD, but it is used in research and is advancing efforts in biomarker development. Genetic testing is increasingly used to aid diagnosis and provide insights into treatment responsiveness and familial risk8,9. Genetic variants associated with PD can be identified in up to 15% of patients; however, variable penetrance and variants of unknown significance (VUS) complicate testing5,10,11.
The clinical features and family history often inform the choice of genetic test12,13, with sequential assays being performed if the genetic cause cannot be determined, e.g., exome sequencing, followed by multiplex ligation-dependent probe amplification (MLPA) to detect structural variants (SVs)14. This can be costly and time‑consuming. Exploring new methods for genetic screening is warranted, as costs and access to genetic testing vary by region and clinical setting5,15.
Novel sequencing techniques have significantly advanced genetic research, from Sanger sequencing to Next Generation Sequencing (NGS) using short reads, and now to long-read sequencing (LRS)16,17. LRS has enabled the discovery of multiple repeat expansion disorder loci18,19 and can sequence challenging regions, including large SVs and repetitive regions, such as short tandem repeats (STRs)20–22. Oxford Nanopore Technologies (ONT) specifically offers targeted LRS called adaptive sampling, which is based on mapping actively sequenced reads to predefined regions of interest. Only reads aligning to those regions are completely sequenced. In contrast, off-target reads are ejected from the Nanopore, resulting in targeted enrichment and deeper coverage of predefined targets, e.g., disease-relevant genes. Adaptive sampling has been previously used for gene panels with the potential to improve diagnostic rates for repeat expansion disorders18,23,24, and has shown promise in detecting diverse pathogenic variant types25. While the use of adaptive sampling for the detection of single-nucleotide variants (SNVs), indels, and repeat expansions has been investigated, especially in ataxias18,23, other movement disorders with pathogenic variants, i.e., large structural variants, remain understudied. Here, we designed a novel gene panel comprising 564 loci related to PD, repeat expansion disorders, and other neurodegenerative and movement disorders, and performed adaptive sampling to jointly interrogate phased SNVs, indels, large SVs, and repeat expansions. We investigate the technical feasibility of our approach and its potential diagnostic utility (Fig. 1).
Fig. 1. Overview of the workflow to evaluate the panel’s diagnostic utility on samples with known variants.

A gene panel for neurodegenerative movement disorders (NMDs) was created, consisting of 564 loci linked to Parkinson’s disease and repeat expansion disorders. DNA was isolated from n = 18 individuals with known genetic variants and sequenced using ONT long-read sequencing to test this panel’s diagnostic utility. Adaptive sampling was applied during the sequencing to enrich the gene loci of the panel. The human variation workflow from EPI2ME Labs was used to align reads and call variants. The visualization of the aligned reads was performed by the Integrative Genomics Viewer (IGV). The Noise Cancelling Repeat Finder (NCRF) was used to analyze the repeat expansion loci not included in the EPI2ME human variation workflow. Additional pathogenic variants detected by the EPI2ME human variation workflow were validated using PCR-based Sanger sequencing for SNVs and indels or agarose gel electrophoresis for SVs and repeat expansions. ONT: Oxford Nanopore Technologies, SNV: Single-nucleotide variant. This figure was created in Created in BioRender. Fienemann, A. (2026) https://BioRender.com/hwptm5k.
Results
Read quality and coverage of the gene panel
The average number of mapped on-target reads per sample was 420,408 (SD = ± 244,251) (Table 1, Supplementary Table 2). Across all samples, a mean Phred quality score of 19.7 (SD = ± 2.2) was achieved. The mean on-target read N50 of all samples was 21.5 kb (SD = ± 11 kb). Across all samples, a mean on-target coverage of 24X (SD = ± 9X) was achieved. The coverage of all 563 loci of the gene panel (excluding the mtDNA) differed both between samples (Fig. 2A) and across the individual gene loci (Fig. 2B). In particular, the genes affected by known pathogenic variants (Table 2) showed differences in the mean coverage depths (Fig. 2C). SNCA and LRRK2 were among the genes with a mean coverage of ≥25X (Fig. 2B, C), while GBA1 and TAF1 showed a mean coverage of approximately 15X (Fig. 2B, D). The remaining selected gene targets exhibited coverage ranging from 23–28X (Fig. 2B). For the mtDNA, a mean coverage of 1256X was obtained.
Table 1.
Read statistics from the ONT long-read adaptive sampling runs
| Total mapped on-target reads | On-target coverage (X) | Quality (Q-score) | Read length (kb) | On-target Read N50 (kb) | |
|---|---|---|---|---|---|
| Mean | 418,989 | 24.0 | 19.6 | 5.37 | 21.5 |
| Median | 382,251 | 21.7 | 19.4 | 4.72 | 18.6 |
| SD | 251,251 | 9.0 | 2.2 | 3.49 | 11.3 |
| IQR | 380,679 | 12.5 | 3.4 | 4.32 | 20.1 |
Descriptive statistics were calculated based on all 18 samples of the panel.
SD standard deviation, IQR interquartile range, kb kilobase, Q-score read quality (Phred), X depth of coverage.
Fig. 2. Sequencing coverage across genes for the targeted long-read sequencing panel.

Average coverage depth of gene targets obtained for A each sample, B the samples obtained for selected gene targets on the targeted long-read sequencing panel, C the gene targets with an average coverage of ≥25X, and D the gene targets with an average coverage of <15X. The mtDNA was excluded from the coverage analyses.
Table 2.
Patient demographics and types of prior genetic testing for the gene panel for Parkinson’s disease and repeat expansion disorders
| Sample ID | Sex | AAO | AAE | Known pathogenic variant | Prior genetic testing |
|---|---|---|---|---|---|
| L-15936 | female | NA | 44 |
GBA1 c.1223C > T LRRK2 c.6055G > A (p.G2019S) |
Sanger sequencing |
| iPS-L-3244 | female | 39 | 45 | PRKN c.823C > T, PRKN ex1del | OGM, LRS |
| SFC817 | male | 64 | 75 | PRKN c.971del, PRKN ex7del | OGM, LRS |
| L-13754 | female | 29 | 71 | PRKN ex2-3del, PRKN c.101_102del | MLPA |
| iPS-L-3034 | male | 38 | 53 | PRKN ex2del, PRKN ex3-5del | OGM, LRS |
| L-26618 | male | 53 | 60 | RAB32 c.213C > G | NGS gene panel, WGS |
| SFC831 | female | 35-40 | 55 | SNCA triplication | OGM, LRS |
| L-20456 | male | 33 | 46 | ATXN1 repeat expansion | PCR-based LRS |
| L-20468 | male | 41 | 44 | ATXN2 repeat expansion | PCR-based LRS |
| L-10180 | female | 46 | 52 | ATXN3 repeat expansion | PCR-based LRS |
| L-20677 | male | 65 | 67 | HTT repeat expansion | NA |
| L-1689 | female | 15 | 80 | DAB1 repeat expansion | PCR-based LRS |
| L-15764 | male | 63 | 70 | FGF14 repeat expansion | PCR-based LRS |
| L-20799 | male | 50 | 60 | C9orf72 repeat expansion | Repeat-primed PCR |
| L-14372 | male | 41 | 46 | TAF1-SVA repeat expansion | PCR-based LRS |
| D298 | female | 58 | 59 | None (idiopathic FTD) | WES, Repeat-primed PCR |
| L-4474 | female | 47 | 62 | None (idiopathic PD) | OGM, LRS |
| L-8319 | male | NA | 31 | None (idiopathic PD) | NA |
AAO Age At Onset, AAE Age At Examination, FTD Frontotemporal Dementia, LRS Long-read Sequencing, MLPA Multiplex Ligation-dependent Probe Amplification, NA not applicable or not available, NGS Next Generation Sequencing, OGM Optical Genome Mapping, PCR Polymerase Chain Reaction, PD Parkinson’s Disease, WES Whole Exome Sequencing, WGS Whole Genome Sequencing.
Detection of the known pathogenic variants with EPI2ME
The human variation workflow from the ONT bioinformatics software collection, EPI2ME, was applied as a screening tool to validate the performance of the gene panel using pathogenic variants detected by prior genetic testing in eighteen individuals from established patient cohorts (Table 2). This includes SNVs, indels, SVs, and repeat expansions.
Regarding the detection of SNVs and indels, the small variant caller Clair3 in the EPI2ME workflow correctly detected all six pathogenic small variants: the c.1223C > T (p.T408M) GBA1 variant and the c.6055G > A (p.G2019S) LRRK2 variant in L-15936, the three PRKN variants c.823C > T (p.R275W), c.971del (p.V324fs), and c.101_102del (p.Q34fs) in iPS-L-3244, SFC817, and L-13754, and the c.213C > G (p.S71R) RAB32 variant in L-26618 (Table 3). The small variant caller BCFtools, applied after the initial analysis, also confirmed all six variants. Regarding the detection of SVs, the SV caller Sniffles2, included in the EPI2ME workflow, correctly detected two large PRKN deletions (≥200 kb) in SFC817 and iPS-L-3034 with default parameters (–long-del-length: 50000, –long-del-coverage: 0.66). However, it failed to identify three additional PRKN deletions in iPS-L-3244, L-13754, and iPS-L-3034, as well as a 1.7 Mb triplication encompassing the entire SNCA gene in SFC831 (Table 3). For SFC817, the read-based phasing software Whatshap confirmed compound heterozygosity of the two PRKN variants (c.971del and ex7del). Using adjusted parameters to account for the detection of SVs >100 kb, both CuteSV and SVIM, two SV callers applied after the initial analysis, showed performance equal to Sniffles2, with CuteSV additionally detecting the two missed PRKN deletions in L-13754 (320 kb) and iPS-L-3034 (112 kb). Regarding the detection of repeat expansions, the tandem repeat variant caller Straglr, included in the EPI2ME workflow, correctly identified four pathogenic repeat expansions: an ATXN1 (CAG)45 repeat expansion in L-20456 (Fig. 3A), an ATXN2 (CAG)34 repeat expansion in L-20468 (Fig. 3B), an ATXN3 (CAG)70 repeat expansion in L-10180 (Fig. 3C), and an HTT (CAG)45 repeat expansion in L-20677 (Fig. 3D). The expected repeat number (RN) based on prior genetic testing was 44 for ATXN1, 34 for ATXN2, 73 for ATXN3, and 43 for HTT.
Table 3.
Overview of the detection results for the known pathogenic variants sequenced with ONT long-read adaptive sampling
| Single-nucleotide variants and Indels | ||||
|---|---|---|---|---|
| Sample ID | Known pathogenic variant | EPI2ME | NCRF | IGV |
| L-15936 | GBA1 c.1223C > T, p.T408M | Detected | - | - |
| L-15936 | LRRK2 c.6055G > A, p.G2019S | Detected | - | - |
| iPS-L-3244 | PRKN c.823C > T, p.R275W | Detected | - | - |
| SFC817 | PRKN c.971del, p.V324Afs*111 | Detected | - | - |
| L-13754 | PRKN c.101_102del, p.Q34Rfs*5 | Detected | - | - |
| L-26618 | RAB32 c.213C > G, p.S71R | Detected | - | - |
| Structural variants | ||||
| Sample ID | Known pathogenic variant | EPI2ME | NCRF | IGV |
| SFC817 | PRKN ex7del | Detected | - | - |
| iPS-L-3244 | PRKN ex1del | Not detected | - | Observed |
| L-13754 | PRKN ex2-3del | Not detected | - | Observed |
| iPS-L-3034 | PRKN ex2del | Not detected | - | Observed |
| iPS-L-3034 | PRKN ex3-5del | Detected | - | - |
| SFC831 | SNCA triplication | Not detected | - | Not observed |
| Repeat expansions | ||||
| Sample ID | Known pathogenic variant | EPI2ME | NCRF | IGV |
| L-20456 | ATXN1 repeat expansion | Detected | - | - |
| L-20468 | ATXN2 repeat expansion | Detected | - | - |
| L-10180 | ATXN3 repeat expansion | Detected | - | - |
| L-20677 | HTT repeat expansion | Detected | - | - |
| L-1689 | DAB1 repeat expansion | Not detected | Detected | - |
| L-15764 | FGF14 repeat expansion | Not detected | Detected | - |
| L-20799 | C9orf72 repeat expansion | Inconclusive | Not detected | Inconclusive |
| L-14372 | TAF1-SVA repeat expansion | Not detected | Not detected | Inconclusive |
D298, L-4474, and L-8319 are not listed as these individuals were idiopathic and had no known pathogenic variants.
EPI2ME Nanopore human variation workflow, NCRF Noise Cancelling Repeat Finder, IGV Integrative Genomics viewer, “-” The method was not necessary to employ after the detection of the variant by either EPI2ME or NCRF.
Fig. 3. Histograms of the repeat expansions detected using EPI2ME.

A. ATXN1 (CAG)45 repeat expansion in L-20456. B. ATXN2 (CAG)34 repeat expansion in L-20468. C. ATXN3 (CAG)70 repeat expansion in L-10180. D. HTT (CAG)45 repeat expansion in L-20677. For each repeat expansion, the top panel shows the histogram with the number of reads for each repeat length based on the number of repetitions of the repeat motif. The bottom panel shows the repeat length of the short and expanded allele in base pairs (bp) with the intermediate (black line) and pathogenic threshold (red line).
Currently, EPI2ME uses an input file of known repeat expansion loci, repeat motifs, pathogenic repeat lengths, and associated repeat expansion disorders, which does not yet include the genes FGF14, DAB1, and TAF1. Thus, Straglr could not detect the pathogenic repeat expansions in DAB1, FGF14, and the TAF1-SVA found in L-1689, L-15764, and L-14372, respectively (Table 3). The standalone version of Straglr, using a custom input BED file, detected a pathogenic DAB1 (ATTTC)123 repeat expansion for L-1689 based on two reads (original RN = 116) and a pathogenic FGF14 (GAA)350 repeat expansion for L-15764 based on 31 reads (original RN = 330). The TAF1-SVA repeat expansion was not detected.
Prior analysis of L-20799 using Repeat-primed PCR revealed an expansion in the C9orf72 locus, but no exact RN was determined. Out of 41 reads covering C9orf72 in L-20799, the workflow version of Straglr identified six expanded reads with an RN between 482 and 1660, but called the expansion ultimately as benign.
Detection of the known pathogenic variants with Noise Cancelling Repeat Finder
Noise Cancelling Repeat Finder (NCRF), a short tandem repeat alignment tool, confirmed the repeat expansions in ATXN1 (RN = 42), ATXN2 (RN = 34), ATXN3 (RN = 70), and HTT (RN = 41) detected by Straglr with a slight deviation of 3-4 repeat units. It also detected the pathogenic DAB1 (ATTTC)105 repeat expansion for L-1689 (Supplementary Fig. 1) and the pathogenic FGF14 (GAA)349 repeat expansion for L-15764 (Supplementary Fig. 2), with a bigger deviation of 18 repeat units for DAB1 compared to the Straglr result. While NCRF detected 13 reads with high repeat lengths (2000–10,000 bp) in C9orf72, no exact repeat number could be determined due to the high variability (Supplementary Fig. 3). The expected repeat expansion in TAF1-SVA could not be detected (Table 3).
Evaluation of the variants in the Integrative Genomics Viewer
To thoroughly assess the detection limit of our adaptive sampling panel, the Integrative Genomics Viewer (IGV) was used to visually inspect six known pathogenic variants that EPI2ME and NCRF did not detect. The presence of a 20 kb deletion spanning exon 1 of PRKN was identified in the region chr6:162,710,190-162,729,371 for iPS-L-3244 (Supplementary Fig. 4). A total of 20 reads of haplotype 1 (red) mapped to the deleted region. A single read of haplotype 2 (blue) carried the left breakpoint of the deletion as highlighted by a stretch of mismatched bases (Supplementary Fig. 4A). The PRKN SNV c.823C > T was located on the other allele, i.e., haplotype 1 (red) (Supplementary Fig. 4B).
A 320 kb deletion spanning exons 2–3 of PRKN was identified in the region chr6:162,232,182-162,555,779 for L-13754 (Supplementary Fig. 5). Compared to the surrounding coverage of approximately 38X, a drop to 21X was observed in the deleted region, with mostly reads of haplotype 1 (red) remaining (Supplementary Fig. 5A). The small PRKN deletion c.101_102del was located in the deleted stretch on haplotype 1 (red) (Supplementary Fig. 5B).
Both a 250 kb deletion in the region chr6:162,029,818-162,280,411 spanning exon 3–5 of PRKN, and a 112 kb deletion in the region chr6:162,336,433-162,448,673 spanning exon 2 of PRKN, were visually confirmed in iPS-L-3034 (Supplementary Fig. 6). Compared to the surrounding coverage of 26X, the deleted regions showed a coverage of 14X and 10X, respectively. Inspection of the breakpoint regions revealed that the 250 kb deletion is located on haplotype 1 (red) (Supplementary Fig. 6A, B), while the other deletion is present on haplotype 2 (blue) (Supplementary Fig. 6C). Thus, all PRKN variants were confirmed to be compound heterozygous.
The presence of the 1.7 Mb triplication spanning the whole SNCA gene in SFC831 was previously validated using both long-read whole-genome sequencing and optical genome mapping26. While the variant could not be visually identified in the IGV reads, there was a 2.5-fold increase in SNCA coverage as compared to the mean coverage of all other gene loci in SFC831 (Fig. 4). In line with the EPI2ME results by Straglr, six out of the 41 reads covering C9orf72 in L-20799 showed expansions ranging from 2877 to 9944 bp (Supplementary Fig. 7). At the position of the TAF1 insertion in L-14372, three out of the seven on-target reads carried an insertion of approximately 2.7 kb (Supplementary Fig. 8). The low number of reads with a repeat expansion and the difference in repeat length make a reliable and accurate detection difficult. Therefore, both the C9orf72 and TAF1-SVA repeat expansions were categorized as inconclusive. Seventeen out of the 20 known pathogenic variants across 15 individuals were clearly identified using the gene panel, while three variants (one SNCA triplication and two large repeat expansions in C9orf72 and TAF1-SVA) were not accurately detected. No pathogenic variants were detected in the three individuals included as negative controls.
Fig. 4. SNCA coverage in SFC831.

The coverage of all 563 gene loci (mtDNA not considered) for the sample SFC831 is depicted in a box plot. A much deeper coverage (29.5X) was obtained for SNCA compared to all other genes, indicating the presence of the expected triplication spanning SNCA.
Detection of additional pathogenic variants
Filtering of the EPI2ME output yielded 14 additional candidate variants not previously reported for these individuals. While nine variants could not be validated (Supplementary Table 4), four potentially pathogenic SNVs and SVs and one repeat expansion of uncertain significance (Table 4) were validated using either Sanger sequencing (Supplementary Fig. 9) or PCR and Agarose gel electrophoresis (Supplementary Fig. 10). For L-13754, who carried compound heterozygous PRKN deletions, Clair3 detected a nonsynonymous SNV in STXBP2 (c.1247-1G > C). The deep learning-based splice variant prediction tool, SpliceAI, predicted a splice acceptor loss with a Delta score of 0.99, and both the CADD score of 32 and the ACMG score of 16 categorized the variant as pathogenic. The STXBP2 gene associated with familial hemophagocytic lymphohistiocytosis (fHLH)27,28. The heterozygous SNV was present in the Sanger Chromatogram of the STXBP2 PCR product (Supplementary Fig. 9A).
Table 4.
The EPI2ME human variation workflow identified four potentially pathogenic variants and one variant of unknown significance that had not been previously reported for these individuals
| Known pathogenic variants | ||||
|---|---|---|---|---|
| Sample ID | Gene | GRCh38 position | Variant | Associated phenotype |
| L-13754 | PRKN | chr6:162,443,378 | c.101_102del | Early-Onset Parkinson’s Disease |
| PRKN | chr6:162,229,153-162,547,507 | 320 kb deletion (exon 2-exon 3) | Early-Onset Parkinson’s Disease | |
| L-1689 | DAB1 | chr1:57,367,043 | repeat expansion (ATTTC)105 | Spinocerebellar Ataxia 37 |
| SFC817 | PRKN | chr6:161,548,965 | c.971del | Early-Onset Parkinson’s Disease |
| PRKN | chr6:161,590,316-161,794,731 | 204 kb deletion (exon 7) | Early-Onset Parkinson’s Disease | |
| iPS-L-3034 | PRKN | chr6:162,335,133-162,449,511 | 112 kb deletion (exon 2) | Early-Onset Parkinson’s Disease |
| PRKN | chr6:162,028,488-162,281,482 | 250 kb deletion (exon 3-exon 5) | Early-Onset Parkinson’s Disease | |
| L-15764 | FGF14 | chr13:102,161,570 | repeat expansion (GAA)349 | Spinocerebellar Ataxia 27B |
| Additionally detected genetic variants | ||||
| Sample ID | Gene | GRCh38 position | Variant | Associated phenotype |
| L-13754 | STXBP2 | chr19:7,645,196 | c.1247-1G > C | Familial Hemophagocytic Lymphohistiocytosis 5 |
| L-1689 | AP4S1 | chr14:31,066,239 | c.43C > T | Spastic Paraplegia 52 |
| SFC817 | RARS2 | chr6:87,564,181 | c.160_161del | Pontocerebellar Hypoplasia |
| iPS-L-3034 | ALS2 | chr2:201,703,123- 201,704,615 | 1.4 kb deletion (exon 32-intron 33) | Amyotrophic Lateral Sclerosis 2 |
| L-15764 | CNBP | chr3:129,172,624 | repeat expansion (CCTG)137 | Myotonic Dystrophy 2 |
These variants were validated using Sanger sequencing or PCR.
In L-1689, who carried a DAB1 repeat expansion, a nonsynonymous SNV was identified in AP4S1 (c.43C > T). A high CADD score of 39 and an ACMG score of 14 classified the variant as pathogenic, which is known to be associated with spastic paraplegia 52 (SPG52)29. The heterozygous SNV was present in the Sanger Chromatogram of the AP4S1 PCR product (Supplementary Fig. 9B).
A small indel affecting RARS2 (c.160_161delAA) was detected in SFC817, who carried compound heterozygous PRKN variants. As a frameshift indel, it was predicted to result in a complete loss of function. Pathogenic variants in RARS2 are associated with pontocerebellar hypoplasia type 6 (PCH6)30. The heterozygous indel was present in the Sanger chromatogram of the RARS2 PCR product (Supplementary Fig. 9C).
Sniffles2 called a 1.4 kb deletion spanning exon 32 to intron 33 of ALS2 in iPS-L-3034, who carried compound heterozygous PRKN deletions. The ALS2 deletion had a CADD score of 11.16 and an ACMG score of 0.9. Biallelic variants, including deletions, in ALS2 are associated with juvenile amyotrophic lateral sclerosis 231. Two discrete bands at the height of approximately 2.1 kb and 700 bp were identified for the ALS2 PCR product (Supplementary Fig. 10A). The presence of the deletion was also validated in the native sample prior to iPSC reprogramming.
A repeat expansion of 137 repeat units (RU) with the motif CCTG was detected by Stranglr for one CNBP allele in L-15764, who carried an FGF14 repeat expansion. NCRF also detected the variant with an RN of 140.8 (Supplementary Fig. 11). Repeat expansions over a threshold of 75 repeats with the TCTG and CCTG motifs are known to cause myotonic dystrophy 2 (DM2)32,33. Three discrete bands at the height of approximately 1.5 kb, 900 bp, and 700 bp were identified for the CNBP PCR product (Supplementary Fig. 10B). Thus, five patients with pathogenic variants associated with early-onset PD or spinocerebellar ataxia (SCA) also carried variants in genes linked to other neurodegenerative diseases or conditions, such as immune deficiency, but without corresponding clinical phenotypes, indicating that these represent incidental findings requiring cautious interpretation rather than established additional diagnoses. Interestingly, 13 patients also carried benign (AAAAG)n expansions in RFC1, with repeat numbers ranging from 57 to 144 (Supplementary Table 5), below the pathogenic threshold of 250 for RFC134.
Discussion
In this study, we developed and evaluated a long-read adaptive sampling gene panel consisting of 564 gene loci related to PD, repeat expansion disorders, and other neurodegenerative movement disorders and assessed its technical feasibility and potential diagnostic utility in 18 individuals with prior genetic testing, including 15 positive controls and 3 negative disease controls.
By focusing on relevant genes, we achieved a mean coverage of >20X, comparable to whole-genome sequencing (WGS) for a single sample on the PromethION, with three samples multiplexed per flow cell. Our targeted long-read sequencing approach thus allows for cost reduction and lower data storage and transfer requirements. This is sufficient for the robust detection of SNVs, small Indels, and SVs35. An adjustment would be to reduce the multiplexing to 2 samples to achieve a mean coverage of >30X, thereby further optimizing the detection of repeat expansions36. The coverage over the different gene loci showed high variance, which could be attributed in part to some genes being located in hard-to-map regions. The genes VARS1, VARS2, TRIM40, and NEU1 are all located within or near the major histocompatibility complex (MHC) on chromosome 6, a highly polymorphic and complex region that can hinder accurate mapping37. Similarly, HLA-DRB5 is located in the highly polymorphic HLA class II region38. Twenty out of the 31 genes with a mean coverage <15X are also X-linked (e.g., TAF1) and thus less abundant in the ten male individuals of the cohort, while two genes on chr5, OCLN and SLC6A3, are known challenging medically relevant genes (CMRGs)39 that are hard to characterize using NGS. Software frameworks like BOSS-RUNS40, based on readfish41, would allow for dynamic adjustments during adaptive sampling to, for example, favor hard-to-map regions and achieve a more continuous coverage over all gene loci. However, in this case, as large copy number variants were assessed, the underlying coverage information needed to be preserved40,42. To this end, we selected the built-in adaptive sampling solution of the ONT operating software, MinKNOW, to benefit from active support and future performance improvements.
The mean on-target read N50 of 21.5 kb was reported to be suitable for the detection of SVs43, but it also showed high variance (SD = ± 11.3 kb), which may be attributed to samples originating from different countries and varying storage durations and conditions. In particular, L-20456, L-20468, and L-14372 showed high fragmentation (on-target read N50 < 6 kb). The iPSC samples, on the other hand, generally showed high on-target read N50 values (>30 kb) as they were freshly generated and processed for this study, comparable to more recently collected in-house samples like L-15764 (Supplementary Table 2).
Using the EPI2ME human variation workflow, 12 of the 20 expected pathogenic variants were directly called (60%; Clopper-Pearson 95% CI: 36%–81%), while the presence of five additional variants was verified using either NCRF, CuteSV, or manual inspection in IGV. This includes three initially missed PRKN SVs: two were not called due to SV caller constraints, and one was missed due to the limited buffer sequence of 50 kb. The other two were repeat expansions that were likely not directly called due to limited read length and coverage, and to a limitation of the GRCh38 reference regarding the TAF1-SVA insertion. In total, 17/20 (85%; Clopper–Pearson 95% CI: 62%–97%) expected pathogenic variants were identified (15% missed; Clopper-Pearson 95% CI: 3%–38%).
All six known SNVs and small indels investigated in this study play a role in either early-onset PD (PRKN)44, late-onset PD (LRRK2, RAB32)45,46, or are associated with increased PD risk (GBA1 p.T408M)47. Their detection was especially successful, as all expected variants were identified. The base-calling accuracy of ONT has improved significantly over the years, with Q-scores exceeding 20 now being routinely achievable using R10 flow cell chemistry and super-accurate base calling25,48. This advancement enabled the robust analysis of small variants using Clair3 on our targeted long-read data. Nevertheless, further improvements are still warranted, as the accuracy in SNV detection within complex regions remains limited compared to NGS49. All of the six known pathogenic SVs investigated in this study, i.e., the PRKN deletions and SNCA triplication, are known to cause early-onset PD44,50. The current generation of SV-callers generally shows comparable performance51, yet all have struggled to detect large (>50 kb) SVs in the past52. Recent improvements in Sniffles2 version 2.5 enhanced the detection of large (>50 kb) variants53. Using default parameters, Sniffles2 correctly identified two PRKN deletions of 200 and 250 kb in SFC817 and iPS-L-3034. However, using a manually adjusted size parameter, SVIM showed the same performance, whereas CuteSV additionally detected the missed PRKN deletions of 112 and 320 kb. This highlights that there is still no universal gold standard for SV calling of long-read data that does not require manual adjustment, which complicates automation. The universally missed 20 kb PRKN deletion in iPS-L-3244 spanned exon 1 and was located on the boundary of the sequenced region of interest, potentially preventing its detection. While flanking sequences of 50 kb were previously successful in similar studies24,54, we recommend to increase the flanking sequence for regions with a higher likelihood of SV formation, e.g., PRKN, located in a common fragile site55,56, or in general for genes known to be challenging to characterize due to their position in complex or highly repetitive regions (Supplementary Table 1). EPI2ME also includes the read depth based variant caller Spectre, which is optimized for CNVs ≥100 kb and could likely identify the large PRKN deletions in WGS data. However, it is not designed for targeted sequencing data and was therefore not applicable. While haplotagged reads enabled the detection and phasing of all PRKN variants in IGV, there remains potential to optimize the detection of large SVs, as manual inspection (~30 min per sample) might not be feasible in every analysis setting. Similarly, the SNCA triplication in SFC831 was not detected by any SV-caller, as the original variant spanned 1.7 Mb and increased the coverage of the entire SNCA locus. This emphasizes the need for specialized tools to detect copy number changes in targeted sequencing data57. The data generated here could serve as a benchmark for advancing these approaches.
The eight known pathogenic repeat expansions in this study cause different types of spinocerebellar ataxia (ATXN1, ATXN2, ATXN3, DAB1, FGF14)58–62, Huntington’s disease (HTT)63, amyotrophic lateral sclerosis (ALS) or FTD (C9orf72)64,65, and X-linked dystonia-parkinsonism (XDP-SVA)66. EPI2ME and NCRF showed concordance for the RN of four repeat expansions, while the pathogenic repeat expansions in DAB1 and FGF14 were only detected by the standalone version of Straglr and NCRF. The deviation of 18 repeat units observed in the DAB1 repeat expansion is likely due to different custom references. We expect a more standardized workflow to emerge as new findings (i.e., new loci and repeat motifs) are added to the reference list of Straglr in the workflow. The repeat expansion in C9orf72 was inconclusive. Straglr detected six reads carrying a repeat expansion in C9orf72, which did not suffice to categorize it as pathogenic. NCRF also detected expanded reads; however, the varying repeat lengths prevented the determination of a reliable repeat number. The very long GC-rich repeats of the C9orf72 repeat expansion, which can also form G-quadruplex structures67, require deep coverage to be characterized beyond the repeat number68–70. A recent study on the utility of adaptive sampling in the diagnosis of cerebellar ataxia also presents the read depth as a limitation of adaptive sampling and suggests a coverage of greater than 33X for the interpretation of repeat expansions in a diagnostic setting36. While we achieved a coverage of 43X over the C9orf72 locus, it was likely still insufficient, as the obtained read N50 of 11 kb (Supplementary Table 3) was low. This case emphasizes the current limitations in studying very large repeat expansions. There is a trade-off between improved enrichment from a smaller targeted genome fraction and pore blocking from increased strand rejection42. However, in a diagnostic-focused setting, it may be warranted to reduce the number of target regions (Supplementary Table 1) to further optimize coverage, while remaining in the advised 1%–5% fraction of the genome, e.g., by removing the GWAS loci that are primarily relevant for research. Similar problems also arose in detecting the repeat expansion in TAF1-SVA, which was a special case because it is located within a SINE-VNTR-Alu (SVA) retrotransposon insertion in the gene TAF166,71. This SVA retrotransposon is absent in the GRCh38 human reference genome, which likely led to reduced coverage as some reads of interest were falsely discarded as off-target during adaptive sampling. Additionally, the TAF1 gene is located on the X chromosome, resulting in lower coverage in men. Therefore, we obtained only three on-target reads with the SVA retrotransposon in the TAF1 gene, which was insufficient for NCRF to accurately determine the RN.
Using our adaptive sampling panel, we identified four additional potentially pathogenic variants and one variant of uncertain significance across all three tested types (SNVs, SVs, and repeat expansions), demonstrating the broad utility compared to more specialized diagnostic tests. The three SNVs in STXBP2, AP4S1, and RARS2, as well as the 1.4 kb deletion in ALS2, were heterozygous, and no additional SNVs or SVs on the other allele were identified. All four patients showed no signs of comorbidity consistent with either fHLH, SPG52, PCH6, or ALS2, which is in line with our expectations, as these conditions exhibit autosomal recessive inheritance. DM2 has an autosomal dominant inheritance pattern, with a (CCTG)137 repeat expansion in CNBP being considered pathogenic53,54. However, no myotonic signs were observed, and all clinical signs are consistent with the initial diagnosis of autosomal dominant spinocerebellar ataxia 27 (SCA27), which is caused by the validated repeat expansion in FGF1458,72. A further complication is the presence of unspecific bands in the Agarose gel of the CNBP PCR validation. In addition to the expected band of the PCR product carrying the CNBP repeat expansion at approximately 700 bp, two additional bands are present at 900 bp and 1.5 kb. Furthermore, the median age of symptom onset for DM2 is 48 years73, and the patient was last examined at 73 years of age, underscoring the need for cautious interpretation and genetic counseling when incidental findings lack a clear clinical correlate. Further validation of the CNBP repeat expansion using an orthogonal method would be advised.
Our ONT adaptive sampling gene panel demonstrated strong detection capacities across all tested variant types, including phased large SVs and long repeat expansions, highlighting its potential as a cost-effective approach to leverage the advantages of long-read sequencing for clinical diagnosis in the future. A study focused on the diagnosis of spastic-ataxia disorders further emphasizes the method’s technical potential, as another ONT adaptive sampling panel detected pathogenic repeat expansions in 9/23 (39%) individuals with prior negative genetic test results23. Recent studies also present other specialized applications for adaptive sampling, e.g., to specifically validate CNVs74 or complex rearrangements57 in the clinic.
A possible use case of our neurology-focused panel in the context of genetic testing would be for an experienced clinician to recommend customizations based on specific clinical indicators and to apply the adaptive sampling panel as a first-line method, i.e., as a combined approach (capturing phased SNVs, large SVs and repeat expansions) for neurological disorders, or as a second-line method after a negative or inconclusive NGS sequencing result (Supplementary Fig. 12).
Under the research conditions described here, running the panel for three individuals multiplexed on a single flow cell costs the same as an ONT WGS run for a single individual (approximately 330–830 € per patient sample). Depending on the target query (e.g., large repeat expansions), multiplexing only two individuals would significantly increase coverage while remaining more cost-efficient than ONT WGS (approximately 500–1250 € per patient sample). It is still warranted to validate both pathogenic and inconclusive variants using methods such as PCR or Sanger sequencing, increasing total costs to approximately 500–1000 € per patient sample as a first-line test or to 600-1700 € as a second-line test (Supplementary Fig. 12). The costs for each genetic test (per sample with ~20-30X Coverage) were estimated based on international studies75–77, and on Chapter 11.4 of the EBM (German Uniform Evaluation Standard) catalog for the second quarter of 2026, provided by the National Association of Statutory Health Insurance Physicians (NASHIP)78. The long-read sequencing costs are based on self-estimates for research purposes. In clinical practice, however, this would currently be an underestimation, as additional factors (e.g., genetic counseling) contribute to the total cost.
A significant advantage is that the panel can be flexibly adapted by simply changing the text in the provided BED file. As newly discovered genes, risk loci79,80, and repeat motifs implicated in neurodegenerative movement disorder emerge, the panel should be updated and customized. The wide variety of targeted genes included in our study accounts for combined or atypical phenotypes81 and addresses limitations of a solely phenotype-driven analysis, as often many different genes can cause a given clinical phenotype82. At the same time, the panel can serve as a more accessible entry for non-clinicians to perform a combined assessment of SNVs, SVs, and repeat expansions.
There are limitations to address. This study demonstrates the technical feasibility of our adaptive sampling gene panel; however, a larger cohort would strengthen this potential, while a separate cohort of patients without prior knowledge of pathogenic variants would be needed to make broader claims about its performance in a genetic diagnostic setting. For four patients, genomic DNA was isolated from iPSC lines, as no blood was available. This allowed us to compare results from both biomaterials and was a practical approach to investigate the variants in this technical proof, but it would not be suitable for clinical diagnostics, as somatic variants might be introduced during cell reprogramming. Roadblocks, such as the need for extensive supporting infrastructure e.g., extraction and shearing of high-molecular-weight DNA83, limited bioinformatic resources (e.g., technology-matched variant catalogs and gold-standard variant callers)84, and a high concentration of DNA36, which currently limit the acceptance of long-read sequencing in clinical testing, also generally apply to our panel. As our data indicate, it is also currently advisable to manually inspect potential pathogenic variants, further increasing the need for specialized expertise and time. Finally, the obtained mean coverage of approximately 20X was sufficient for detecting large SVs but was likely too low to allow in-depth characterization of large repeat expansions. Most long-read software tools, especially variant callers, are currently optimized for WGS data, which may affect their performance when applied indiscriminately to adaptive sampling data, as highlighted for the CNV caller Spectre. We believe that as more specialized tools and settings for targeted long-read sequencing data become available, the advantages of performing a single comprehensive test will outweigh current concerns. To this end, the standalone EPI2ME software would likely be the solution of choice for improving automation compared to the manual command-line approach presented here. Further validation in larger and more heterogeneous patient cohorts will provide stronger support for the panel’s clinical potential.
Methods
Gene panel design
A novel gene panel focused on PD and repeat expansion disorders was designed using genes selected from established NGS-based genetic diagnostic panels for neurodegeneration offered by the Center for Genomics and Transcriptomics (CeGaT)85, the International Parkinson and Movement Disorder Society (MDS) Taskforce86, 90 PD genome-wide association study (GWAS) loci87, and known genes associated with repeat expansion disorders. In total, the panel consists of 564 target regions generated from the position of each gene locus ± 50 kb and the mitochondrial genome (Supplementary Table 1). The size distribution of the target regions and the total numbers for each chromosome are summarized in the supplementary materials (Supplementary Figs. 13 and 14).
Patient demographics
Eighteen individuals from established patient cohorts with prior genetic testing were sequenced using adaptive sampling from ONT. Fifteen individuals were selected as positive controls based on their known pathogenic variants, and two individuals with idiopathic PD, as well as one individual with idiopathic frontotemporal dementia (FTD), were included as negative controls. Individuals carrying known pathogenic SVs or repeat expansions were also selected for investigation, as those are more difficult to identify with short-read sequencing. Patient demographics are provided in Table 2. The RAB32 c.213C > G single-nucleotide variant (SNV) in L-2661888, four large SVs spanning PRKN and SNCA in iPS-L-3244, iPS-L-3034, and SFC83126, as well as the DAB1 repeat expansion in L-168989 and the FGF14 repeat expansion in L-1576458,90,91 were previously described.
Generation of induced pluripotent stem cells (iPSC)
Four skin fibroblast-derived induced pluripotent stem cell (iPSC) lines were examined in this study, as no whole blood was available for the patients. All iPSC lines were generated by overexpressing OCT4, SOX2, KLF4, and cMYC using Sendai virus to infect the fibroblast cultures, according to the manufacturer’s protocol (CytoTune Reprogramming Kit; Thermo Fisher Scientific). iPSC lines SFC831, iPS-L-3244, and iPS-L-3034 have been reported previously26, whereas iPSC line SFC817 has been newly generated. The clearance of CytoTune Sendai virus-delivered reprogramming genes has been confirmed for all lines (data not shown). The iPSCs were cultured in mTeSR1 medium (StemCell Technologies) on Matrigel-coated plates (BD Bioscience).
ONT long-read sequencing
Genomic DNA was isolated from whole blood or iPSCs using automated DNA precipitation with the FlexiGene DNA Whole Blood Kit (Autogen), spin column purification with the High Pure Viral Nucleic Acid Kit (Roche), manual DNA precipitation, or with the Bionano Prep SP-G2 Blood & Cell Culture DNA Isolation Kit (Bionano) as previously described26.
Controlled shearing of the DNA was performed using the Megaruptor 3 from Diagenode, followed by size selection with the BluePippin from Sage Science, aiming at fragments of ≥30 kb92. The DNA quantity was monitored after each step using the Qubit fluorometer, and the fragment size distribution was assessed using either the Femto Pulse System or the Tapestation 4200 from Agilent Technologies, before and after shearing and size selection.
Library preparation was conducted according to an adapted version of the “1D Native barcoding genomic DNA Protocol” using the Native 96 Barcoding Kit (EXP-NBD114-96). This included ligation of the Barcodes to the ends of the DNA fragments and ligation and cleaning of the sequencing adaptors. Three samples were sequenced per R10.4.1 flow cell (FLO-PRO114M) on the PromethION from Oxford Nanopore Technologies (ONT) for 72 h, respectively (Supplementary Fig. 15).
Adaptive sampling (using the built-in Read until API) was activated in the device’s operating software, MinKNOW (version 24.06.15) to specifically enrich the regions of all gene loci in the panel, which were provided in a BED file containing the GRCh38 coordinates (Supplementary Table 1). Flanking regions of ±50 kb were included for each target gene locus to avoid coverage loss at the edges. All regions together correspond to a total length of 113,562,275 bp or 3.44% of the GRCh38 reference sequence. A fraction between 1% and 5% is recommended for optimal enrichment with adaptive sampling42.
Bioinformatic analysis
The POD5 files were demultiplexed, and Dorado base-calling (version 0.5.1) was performed using the super accuracy (SUP) model (dna_r10.4.1_e8.2_400bps_sup@v4.3.0). Samtools (version 1.15)93 enabled the conversion and handling of FASTQ and BAM files. Unmapped BAM files served as input for the human variation workflow from the ONT bioinformatics software collection EPI2ME (version 2.7.0). The human variation workflow uses Nextflow (version 24.04.4)94 to integrate a collection of bioinformatic tools and establish reproducible and efficient conditions for the detection of genetic variants. This includes alignment to the GRCh38 genome reference using Minimap2 (version 2.24)95, SNV and indel calling using Clair3 (version 1.0.8)96, SV calling using Sniffles2 (version 2.6.2)97, and repeat expansion genotyping using Straglr (version 1.4.5)98. WhatsHap (version 2.0)99 performed phasing of the reads and identified variants. Sniffles2 was applied using default parameters (–long-del-length: 50000, –long-del-coverage: 0.66). The detection of pathogenic repeat expansions by Straglr relies on an input BED file of known repeat expansion loci, which currently excludes the genes FGF14, DAB1, and TAF1 in the workflow. After the initial analysis, the standalone version of Straglr was applied to the GRCh38-aligned CRAM file using a custom BED file containing FGF14, DAB1, and TAF1.
The data were visualized with R (version 4.4.0)100 and the packages tidyverse (v2.0.0)101, dplyr (v1.1.4)102, ggplot2 (v3.5.2)103, and data.table (v1.17.2)104. SnpEff105 and ClinVar106 provided annotations for the SNVs and indels. The SVs were annotated outside of EPI2ME using SVAFotate107 and AnnotSV (version 3.4.2)108.
Subsequent filtering of the variants was performed in R, focusing on rare variants (total genome and exome population allele frequency ≤1% in gnomAD 4.1.0) or those classified as “pathogenic” and “likely pathogenic” according to the ACMG guidelines109. The number of filtered candidates was further reduced by removing variants detected in multiple samples and STRs with benign motifs or repeat lengths well below the pathogenic threshold.
A workflow based on the short tandem repeat alignment tool, Noise Cancelling Repeat Finder (NCRF) (version 1.01.02)110, was used to detect repeat expansions that were not included or missed by the EPI2ME human variation workflow. Minimap2 (version 2.22) was applied to align the unmapped BAM files to the corresponding reference sequence of the gene, which was modified to include an expanded tract of the repeat motif. The SAM/BAM files were again handled with Samtools. The identification of repeat motifs was performed by NCRF, with a minimum of 10 detected repeat units. A maximum noise of 80% was applied as a filter, as previously tested for FGF14 repeat expansions90. The resulting output was analyzed in R to calculate the number of repeats of the pathogenic allele. The detailed commands used in this study are provided on GitHub (https://github.com/aFienemann/Adaptive-sampling-gene-panel-for-PD-and-REDs). After all annotation and filtering steps, only the barcodes without context for the expected variants were provided to reduce confirmation bias during the analysis. The regions of all known pathogenic variants not detected using Sniffles2, Straglr, or NCRF were visualized using the Integrative Genomics Viewer (IGV)111. For this, the CRAM file containing the aligned reads produced by the EPI2ME human variation workflow was loaded into IGV using GRCh38 as the reference. SNVs and indels missed by the EPI2ME human variation workflow were additionally investigated by BCFtools (version 1.9)112, while missed SVs were further examined using both CuteSV (version 2.1.3)113 and SVIM (version 2.0)114. Parameters of both CuteSV (–max_size 350000) and SVIM (–max_sv_size 350000) were increased from the default value of 100,000 to allow for the detection of SVs up to 350 kb instead of 100 kb. As the cohort size was small, the Clopper-Pearson exact method was applied to obtain confidence intervals for the detection rate. Known variants that could be identified in the data but were not sufficiently characterized (e.g., repeat number) by our approach without prior knowledge were termed “inconclusive”. Using our filtering strategy, we detected additional potentially pathogenic variants besides the known pathogenic ones. These variants were validated using PCR-based Sanger sequencing for SNVs and indels (Supplementary Table 3) or agarose gel electrophoresis for SVs and repeat expansions.
European Nucleotide Archive (ENA)
Mapped BAM files of all 15 positive controls with known pathogenic variants used in this study were deposited as a resource in the European Nucleotide Archive (ENA) at EMBL-EBI under accession number PRJEB105200 (https://www.ebi.ac.uk/ena/browser/view/PRJEB105200).
Ethics approval and consent to participate
Written informed consent was obtained from all individuals and the study was conducted according to the guidelines of the Declaration of Helsinki, and approved by the Ethics Committees of the University of Lübeck, Germany (protocol code 16-039, date of approval 27 September 2019) and the P2N supervisory board, Kiel University, Germany (protocol code 2021-037, date of approval 16 September 2021).
Supplementary information
Acknowledgements
This research was funded by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) (TR 1714/4-1 and TR 1714/8-1) and a Heisenberg Grant (TR 1714/7-1) to J.T. A.N.B. and T.G.D. gratefully acknowledge the use of the services and facilities of the Koç University Research Center for Translational Medicine (KUTTAM). G.S. is supported by the Wolf Chair for Neurodevelopmental Psychiatry, a joint Hospital-University Named Chair between the University of Toronto, UHN, and the UHN Foundation. G.S. received honoraria for educative activities by the International Movement Disorder Society and Merz Pharmaceuticals. TerK would like to acknowledge support from the Global Parkinson’s Genetics Program (GP2) funded by the Aligning Science Across Parkinson’s initiative and implemented by The Michael J. Fox Foundation.
Author contributions
A.F. contributed to the conceptualization of the study, data curation, formal analysis, methodology, investigation, software, visualization, writing of the original draft, and final approval of the manuscript. J.C.P. was involved in the conceptualization of the study, data curation, formal analysis, methodology, investigation, software, visualization, writing of the original draft, and final approval of the manuscript. J.L. was involved in conceptualization, data curation, methodology, software, supervision, and final approval of the manuscript. C.M. and S.S. participated in the methodology, investigation, and final approval of the manuscript. T.L. and Car.G. contributed to the methodology, data curation, software, and final approval of the manuscript. A.Z., E.S., The.K., Chr.G., T.G.D., A.N.B., R.D.G.J., R.L.R., G.S., C.C.E.D., M.M., M.B., Ter.K., A.B. and N.B. provided patient samples, clinical information of patients, and read and approved the final manuscript. P.S. participated in the methodology and investigation, provided iPSC lines, and approved the final manuscript. C.K. contributed to the conceptualization, project administration, funding acquisition, provision of resources, and final approval of the manuscript. J.T. was involved in conceptualization, funding acquisition, project administration, resource provision, supervision, and final approval of the manuscript.
Funding
Open Access funding enabled and organized by Projekt DEAL.
Data availability
Mapped BAM files of all 15 positive controls with known pathogenic variants used in this study were deposited as a resource in the European Nucleotide Archive (ENA) at EMBL-EBI under accession number PRJEB105200 (https://www.ebi.ac.uk/ena/browser/view/PRJEB105200). The remaining datasets used and/or analyzed during the current study are available from the corresponding author on reasonable request.
Code availability
The detailed commands used in this study are provided on GitHub (github.com/aFienemann/Adaptive-sampling-gene-panel-for-PD-and-REDs).
Declarations
Competing interests
The authors declare the following competing interests: RDGJ has received travel grants and speakers’ honoraria from Royal Care Super Specialty Hospital, Darya-Varya Laboratoria, and the Philippine offices of AbbVie, Exeltis, HI-Eisai, Innogen, Medichem, Natrapharm, Sun, Torrent, and Vexxa. M.B. received honoraria from Bial Germany. N.B. received honoraria from AbbVie, Esteve, Ipsen, Merz, Takeda, Teva, and Zambon. J.T. has received travel funding to speak on behalf of O.N.T. C.K. has served as a medical advisor to Centogene, Takeda, and Biogen and received speakers’ honoraria from Bial, as well as royalties from Oxford University Press and Springer Nature. The other authors do not have a competing interest.
Footnotes
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
These authors contributed equally: André Fienemann, Julia C. Prietzsche.
Supplementary information
The online version contains supplementary material available at https://doi.org/10.1038/s41531-026-01585-4.
References
- 1.Su, D. et al. Projections for prevalence of Parkinson’s disease and its driving factors in 195 countries and territories to 2050: modelling study of Global Burden of Disease Study 2021. BMJ388, e080952 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Collaborators GBDNSD Global, regional, and national burden of disorders affecting the nervous system, 1990-2021: a systematic analysis for the Global Burden of Disease Study 2021. Lancet Neurol.23, 344–381 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Ou, Z. et al. Global trends in the incidence, prevalence, and years lived with disability of Parkinson’s disease in 204 countries/territories from 1990 to 2019. Front. Public Health9, 776847 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Trevisan, L. et al. Genetics in Parkinson’s disease, state-of-the-art and future perspectives. Br. Med. Bull.149, 60–71 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Pal, G. et al. Genetic testing in Parkinson’s disease. Mov. Disord.38, 1384–1396 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Postuma, R. B. et al. MDS clinical diagnostic criteria for Parkinson’s disease. Mov. Disord.30, 1591–1601 (2015). [DOI] [PubMed] [Google Scholar]
- 7.Siderowf, A. et al. Assessment of heterogeneity among participants in the Parkinson’s progression markers initiative cohort using alpha-synuclein seed amplification: a cross-sectional study. Lancet Neurol.22, 407–417 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Hopfner, F., Hoglinger, G., German Parkinson’s Guidelines, G. & Trenkwalder, C. Definition and diagnosis of Parkinson’s disease: guideline “Parkinson’s disease” of the German Society of Neurology. J. Neurol.271, 7102–7119 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Tolosa, E., Garrido, A., Scholz, S. W. & Poewe, W. Challenges in the diagnosis of Parkinson’s disease. Lancet Neurol.20, 385–397 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Westenberger, A. et al. Relevance of genetic testing in the gene-targeted trial era: the Rostock Parkinson’s disease study. Brain147, 2652–2667 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Hoglinger, G. U. et al. A biological classification of Parkinson’s disease: the SynNeurGe research diagnostic criteria. Lancet Neurol.23, 191–204 (2024). [DOI] [PubMed] [Google Scholar]
- 12.Gasser, T. Genetic testing for Parkinson’s disease in clinical practice. J. Neural Transm.130, 777–782 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Cook, L., Schulze, J., Naito, A. & Alcalay, R. N. The role of genetic testing for Parkinson’s disease. Curr. Neurol. Neurosci. Rep.21, 17 (2021). [DOI] [PubMed] [Google Scholar]
- 14.Cook, L. et al. The commercial genetic testing landscape for Parkinson’s disease. Park. Relat. Disord.92, 107–111 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Crook, A., Jacobs, C., Newton-John, T., O’Shea, R. & McEwen, A. Genetic counseling and testing practices for late-onset neurodegenerative disease: a systematic review. J. Neurol.269, 676–692 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Mantere, T., Kersten, S. & Hoischen, A. Long-read sequencing emerging in medical genetics. Front. Genet.10, 426 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Shendure, J. et al. DNA sequencing at 40: past, present and future. Nature550, 345–353 (2017). [DOI] [PubMed] [Google Scholar]
- 18.de Boer E. N. et al. Nanopore long-read sequencing as a first-tier diagnostic test to detect repeat expansions in neurological disorders. Int. J. Mol. Sci.26, 2850 (2025). [DOI] [PMC free article] [PubMed]
- 19.Depienne, C. & Mandel, J. L. 30 years of repeat expansion disorders: what have we learned and what are the remaining challenges? Am. J. Hum. Genet.108, 764–785 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Wirth, T., Kumar, K. R. & Zech, M. Long-read sequencing: the third generation of diagnostic testing for dystonia. Mov. Disord.40, 1009–1019 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Owusu, R. & Savarese, M. Long-read sequencing improves diagnostic rate in neuromuscular disorders. Acta Myol.42, 123–128 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.van Dijk, E. L., Jaszczyszyn, Y., Naquin, D. & Thermes, C. The third revolution in sequencing technology. Trends Genet.34, 666–681 (2018). [DOI] [PubMed] [Google Scholar]
- 23.Rudaks, L. I. et al. Targeted long-read sequencing as a single assay improves the diagnosis of spastic-ataxia disorders. Ann. Clin. Transl. Neurol.12, 832–841 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Stevanovski, I. et al. Comprehensive genetic diagnosis of tandem repeat expansion disorders with programmable targeted nanopore sequencing. Sci. Adv.8, eabm5386 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Chevrier, S., Richard, C., Mille, M., Bertrand, D. & Boidot, R. Nanopore adaptive sampling accurately detects nucleotide variants and improves the characterization of large-scale rearrangement for the diagnosis of cancer predisposition. Clin. Transl. Med.15, e70138 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Trinh, J. et al. Optical genome mapping of structural variants in Parkinson’s disease-related induced pluripotent stem cells. BMC Genom.25, 980 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Dantas, V. M. et al. Germline compound heterozygous variants identified in the STXBP2 gene leading to a familial hemophagocytic lymphohistiocytosis type 5: a case report. Front. Pediatr.9, 633996 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Tang, X., Guo, X., Li, Q. & Huang, Z. Familial hemophagocytic lymphohistiocytosis type 5 in a Chinese Tibetan patient caused by a novel compound heterozygous mutation in STXBP2. SciMed Cent.98, e17674 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Tessa, A. et al. Identification of mutations in AP4S1/SPG52 through next generation sequencing in three families. Eur. J. Neurol.23, 1580–1587 (2016). [DOI] [PubMed] [Google Scholar]
- 30.Lax, N. Z. et al. Neuropathologic characterization of pontocerebellar hypoplasia type 6 associated with cardiomyopathy and hydrops fetalis and severe multisystem respiratory chain deficiency due to novel RARS2 mutations. J. Neuropathol. Exp. Neurol.74, 688–703 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Hadano, S. et al. A gene encoding a putative GTPase regulator is mutated in familial amyotrophic lateral sclerosis 2. Nat. Genet.29, 166–173 (2001). [DOI] [PubMed] [Google Scholar]
- 32.Wendlandt, M. et al. Updated structure of CNBP repeat expansions in patients with myotonic dystrophy type 2 and its implication for standard diagnostics. Neurol. Genet.11, e200220 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Liquori, C. L. et al. Myotonic dystrophy type 2 caused by a CCTG expansion in intron 1 of ZNF9. Science293, 864–867 (2001). [DOI] [PubMed] [Google Scholar]
- 34.Curro, R. et al. Role of the repeat expansion size in predicting age of onset and severity in RFC1 disease. Brain147, 1887–1898 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Harvey, W. T. et al. Whole-genome long-read sequencing downsampling and its effect on variant-calling precision and recall. Genome Res.33, 2029–2040 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Rafehi, H. et al. A prospective trial comparing programmable targeted long-read sequencing and short-read genome sequencing for genetic diagnosis of cerebellar ataxia. Genome Res35, 769–785 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Mosbruger T. L. et al. Haplotypic resolution of the challenging genomic regions of MHC and KIR using a combination of targeted sequencing and a novel assembly pipeline. Nucleic Acids Res.53, gkaf441 (2025). [DOI] [PMC free article] [PubMed]
- 38.Klasberg, S., Surendranath, V., Lange, V. & Schofl, G. Bioinformatics strategies, challenges, and opportunities for next generation sequencing-based HLA genotyping. Transfus. Med Hemother.46, 312–325 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Ji, Y., Zhao, J., Gong, J., Sedlazeck, F. J. & Fan, S. Unveiling novel genetic variants in 370 challenging medically relevant genes using the long read sequencing data of 41 samples from 19 global populations. Mol. Genet. Genom.299, 65 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Weilguny, L. et al. Dynamic, adaptive sampling during nanopore sequencing using Bayesian experimental design. Nat. Biotechnol.41, 1018–1025 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Payne, A. et al. Readfish enables targeted nanopore sequencing of gigabase-sized genomes. Nat. Biotechnol.39, 442–450 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Iyer, S. V., Goodwin, S. & McCombie, W. R. Leveraging the power of long reads for targeted sequencing. Genome Res.34, 1701–1718 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.De Coster, W., Strazisar, M. & De Rijk, P. Critical length in long-read resequencing. NAR Genom. Bioinform.2, lqz027 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Kitada, T. et al. Mutations in the parkin gene cause autosomal recessive juvenile parkinsonism. Nature392, 605–608 (1998). [DOI] [PubMed] [Google Scholar]
- 45.Gustavsson, E. K. et al. RAB32 Ser71Arg in autosomal dominant Parkinson’s disease: linkage, association, and functional analyses. Lancet Neurol.23, 603–614 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.Zimprich, A. et al. Mutations in LRRK2 cause autosomal-dominant parkinsonism with pleomorphic pathology. Neuron44, 601–607 (2004). [DOI] [PubMed] [Google Scholar]
- 47.Greuel, A. et al. GBA variants in Parkinson’s disease: clinical, metabolomic, and multimodal neuroimaging phenotypes. Mov. Disord.35, 2201–2210 (2020). [DOI] [PubMed] [Google Scholar]
- 48.Moustakli, E. et al. Long-read sequencing and structural variant detection: unlocking the hidden genome in rare genetic disorders. Diagnostics15, 1803 (2025). [DOI] [PMC free article] [PubMed]
- 49.Santos, R. et al. Investigating the performance of Oxford Nanopore long-read sequencing with respect to Illumina microarrays and short-read sequencing. Int. J. Mol. Sci. 26, 4492 (2025). [DOI] [PMC free article] [PubMed]
- 50.Singleton, A. B. et al. alpha-Synuclein locus triplication causes Parkinson’s disease. Science302, 841 (2003). [DOI] [PubMed] [Google Scholar]
- 51.Nardone G. G. et al. A hitchhiker guide to structural variant calling: a comprehensive benchmark through different sequencing technologies. Biomedicines13, 1949 (2025). [DOI] [PMC free article] [PubMed]
- 52.Cuenca-Guardiola, J. et al. Improvement of large copy number variant detection by whole genome nanopore sequencing. J. Adv. Res.50, 145–158 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53.Fienemann, A. et al. Complementarity of long-reads and optical mapping in Parkinson’s disease for structural variants. Ann. Clin. Transl. Neurol. 13, 1467–1481 (2026). [DOI] [PMC free article] [PubMed]
- 54.Nakamichi, K. et al. Targeted long-read sequencing enriches disease-relevant genomic regions of interest to provide complete Mendelian disease diagnostics. JCI Insight. 9, e183902 (2024). [DOI] [PMC free article] [PubMed]
- 55.Munk, S. H. N., Voutsinos, V. & Oestergaard, V. H. Large intronic deletion of the fragile site gene PRKN dramatically lowers its fragility without impacting gene expression. Front Genet.12, 695172 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 56.Denison, S. R. et al. Alterations in the common fragile site gene Parkin in ovarian and other cancers. Oncogene22, 8370–8378 (2003). [DOI] [PubMed] [Google Scholar]
- 57.Paivandy, A. et al. Flexible and rapid validation of structural variation using adaptive sampling. Eur J Hum Genet. 34, 649–657 (2026). [DOI] [PMC free article] [PubMed]
- 58.Rafehi, H. et al. An intronic GAA repeat expansion in FGF14 causes the autosomal-dominant adult-onset ataxia SCA50/ATX-FGF14. Am. J. Hum. Genet.110, 105–119 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 59.Seixas, A. I. et al. A pentanucleotide ATTTC repeat insertion in the non-coding region of DAB1, mapping to SCA37, causes spinocerebellar ataxia. Am. J. Hum. Genet.101, 87–103 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 60.Imbert, G. et al. Cloning of the gene for spinocerebellar ataxia 2 reveals a locus with high sensitivity to expanded CAG/glutamine repeats. Nat. Genet.14, 285–291 (1996). [DOI] [PubMed] [Google Scholar]
- 61.Kawaguchi, Y. et al. CAG expansions in a novel gene for Machado-Joseph disease at chromosome 14q32.1. Nat. Genet.8, 221–228 (1994). [DOI] [PubMed] [Google Scholar]
- 62.Orr, H. T. et al. Expansion of an unstable trinucleotide CAG repeat in spinocerebellar ataxia type 1. Nat. Genet.4, 221–226 (1993). [DOI] [PubMed] [Google Scholar]
- 63.A novel gene containing a trinucleotide repeat that is expanded and unstable on Huntington’s disease chromosomes. The Huntington’s Disease Collaborative Research Group. Cell72, 971–983 (1993). [DOI] [PubMed]
- 64.Renton, A. E. et al. A hexanucleotide repeat expansion in C9ORF72 is the cause of chromosome 9p21-linked ALS-FTD. Neuron72, 257–268 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 65.DeJesus-Hernandez, M. et al. Expanded GGGGCC hexanucleotide repeat in noncoding region of C9ORF72 causes chromosome 9p-linked FTD and ALS. Neuron72, 245–256 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 66.Bragg, D. C. et al. Disease onset in X-linked dystonia-parkinsonism correlates with expansion of a hexameric repeat within an SVA retrotransposon in TAF1. Proc. Natl. Acad. Sci. USA.114, E11020–E11028 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 67.Fratta, P. et al. C9orf72 hexanucleotide repeat associated with amyotrophic lateral sclerosis and frontotemporal dementia forms RNA G-quadruplexes. Sci. Rep.2, 1016 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 68.Udine, E. et al. Targeted long-read sequencing to quantify methylation of the C9orf72 repeat expansion. Mol. Neurodegener.19, 99 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 69.DeJesus-Hernandez, M. et al. Long-read targeted sequencing uncovers clinicopathological associations for C9orf72-linked diseases. Brain144, 1082–1088 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 70.Ebbert, M. T. W. et al. Long-read sequencing across the C9orf72 ‘GGGGCC’ repeat expansion: implications for clinical use and genetic discovery efforts in human disease. Mol. Neurodegener.13, 46 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 71.Westenberger, A. et al. A hexanucleotide repeat modifies expressivity of X-linked dystonia parkinsonism. Ann. Neurol.85, 812–822 (2019). [DOI] [PubMed] [Google Scholar]
- 72.Pellerin, D. et al. Deep intronic FGF14 GAA repeat expansion in late-onset cerebellar ataxia. N. Engl. J. Med.388, 128–141 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 73.Meola, G. & Cardani, R. Myotonic dystrophy type 2: an update on clinical aspects, genetic and pathomolecular mechanism. J. Neuromuscul. Dis.2, S59–S71 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 74.Greer, S. U. et al. Implementation of Nanopore sequencing as a pragmatic workflow for copy number variant confirmation in the clinic. J. Transl. Med.21, 378 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 75.Ehman, M. et al. The cost and cost trajectory of genome sequencing and bioinformatics analysis for Indigenous children with suspected rare diseases. Genet. Med.27, 101568 (2025). [DOI] [PubMed] [Google Scholar]
- 76.Schwarze, K. et al. The complete costs of genome sequencing: a microcosting study in cancer and rare diseases from a single center in the United Kingdom. Genet. Med.22, 85–94 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 77.Marino, P. et al. Cost of cancer diagnosis using next-generation sequencing targeted gene panels in routine practice: a nationwide French study. Eur. J. Hum. Genet.26, 314–323 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 78.(NASHIP) NAoSHIP. Chapter 11.4 of the EBM (German Uniform Evaluation Standard) 2026. Available from: https://www.kbv.de/documents/praxis/abrechnung/ebm/2026-2-ebm.pdf.
- 79.Leonard, H. L. & Global Parkinson’s Genetics, P. Novel Parkinson’s Disease Genetic Risk Factors Within and Across European Populations. medRxiv. 10.1101/2025.03.14.24319455 (2025). [DOI]
- 80.Kim, J. J. et al. Multi-ancestry genome-wide association meta-analysis of Parkinson’s disease. Nat. Genet.56, 27–36 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 81.Jia, F., Fellner, A. & Kumar K. R. Monogenic Parkinson’s disease: genotype, phenotype, pathophysiology, and genetic testing. Genes13, 471 (2022). [DOI] [PMC free article] [PubMed]
- 82.Dratch, L. et al. Genetic testing in adults with neurologic disorders: indications, approach, and clinical impacts. J. Neurol.271, 733–747 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 83.Mahmoud, M., Agustinho, D. P. & Sedlazeck, F. J. A Hitchhiker’s Guide to long-read genomic analysis. Genome Res.35, 545–558 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 84.Cogan, G., Daida, K., Blauwendraat, C., Billingsley, K. & Brice, A. Exploration of neurodegenerative diseases using long-read sequencing and optical genome mapping technologies. Mov. Disord.40, 996–1008 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 85.CeGaT GmbH. National Library of Medicine 2025. Available from: https://www.ncbi.nlm.nih.gov/gtr/all/tests/?term=Neurodegenerative+Diseases. [DOI] [PubMed]
- 86.Lange, L. M. et al. Nomenclature of genetic movement disorders: recommendations of the International Parkinson and Movement Disorder Society Task Force - An update. Mov. Disord.37, 905–935 (2022). [DOI] [PubMed] [Google Scholar]
- 87.Nalls, M. A. et al. Identification of novel risk loci, causal insights, and heritable risk for Parkinson’s disease: a meta-analysis of genome-wide association studies. Lancet Neurol.18, 1091–1102 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 88.Kleinz, T. et al. RAB32-Linked Parkinson’s disease: deep phenotyping, MDSgene literature review, and application of SynNeurGe criteria. Mov Disord. 40, 2746–2769 (2025). [DOI] [PMC free article] [PubMed]
- 89.Rosenbohm, A. et al. Familial cerebellar ataxia and amyotrophic lateral sclerosis/frontotemporal dementia with DAB1 and C9ORF72 repeat expansions: an 18-Year study. Mov. Disord.37, 2427–2439 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 90.Lass, J. et al. FGF14 repeat length and mosaic interruptions: modifiers of spinocerebellar ataxia 27B? Brain148, 4072–4083 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 91.Borsche, M. et al. Bilateral vestibulopathy in RFC1-positive CANVAS is distinctly different compared to FGF14-linked spinocerebellar ataxia 27B. J. Neurol.271, 1023–1027 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 92.Wenghöfer, A. et al. On behalf of the CARD Long-read Team 2024. Processing frozen archival human DNA samples for large-scale SQK-LSK114 Oxford Nanopore long-read DNA sequencing SOP v1. protocols.io 10.17504/protocols.io.5jyl82morl2w/v1 (2024). [DOI]
- 93.Li, H. et al. The Sequence Alignment/Map format and SAMtools. Bioinformatics25, 2078–2079 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 94.Di Tommaso, P. et al. Nextflow enables reproducible computational workflows. Nat. Biotechnol.35, 316–319 (2017). [DOI] [PubMed] [Google Scholar]
- 95.Li, H. Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics34, 3094–3100 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 96.Zheng, Z. et al. Symphonizing pileup and full-alignment for deep learning-based long-read variant calling. Nat. Comput Sci.2, 797–803 (2022). [DOI] [PubMed] [Google Scholar]
- 97.Smolka, M. et al. Detection of mosaic and population-level structural variants with Sniffles2. Nat. Biotechnol.42, 1616 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 98.Chiu, R., Rajan-Babu, I. S., Friedman, J. M. & Birol, I. Straglr: discovering and genotyping tandem repeat expansions using whole genome long-read sequences. Genome Biol.22, 224 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 99.Martin, M., Ebert, P. & Marschall, T. Read-based phasing and analysis of phased variants with WhatsHap. Methods Mol. Biol.2590, 127–138 (2023). [DOI] [PubMed] [Google Scholar]
- 100.R Core Team. R: A Language and Environment for Statistical Computing. R Foundation for Statistical Computing, Vienna, Austria. https://www.r-project.org/ (2025).
- 101.Wickham, H. et al. Welcome to the Tidyverse. J. Open Source Softw.4, 1686 (2019). [Google Scholar]
- 102.Wickham, H. et al. dplyr: A Grammar of Data Manipulation. R package version 1.2.1. https://dplyr.tidyverse.org/ (2026).
- 103.Wickham, H. ggplot2: Elegant Graphics for Data Analysis. Springer-Verlag New York. ISBN 978-3-319-24277-4. https://ggplot2.tidyverse.org/ (2016).
- 104.Barrett, T. et al. data.table: Extension of ‘data.frame‘. R package version 1.18.99. https://r-datatable.com (2026).
- 105.Cingolani, P. et al. A program for annotating and predicting the effects of single nucleotide polymorphisms, SnpEff: SNPs in the genome of Drosophila melanogaster strain w1118; iso-2; iso-3. Fly6, 80–92 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 106.Landrum, M. J. et al. ClinVar: improving access to variant interpretations and supporting evidence. Nucleic Acids Res.46, D1062–D1067 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 107.Nicholas, T. J., Cormier, M. J. & Quinlan, A. R. Annotation of structural variants with reported allele frequencies and related metrics from multiple datasets using SVAFotate. BMC Bioinform.23, 490 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 108.Geoffroy, V. et al. AnnotSV: an integrated tool for structural variations annotation. Bioinformatics34, 3572–3574 (2018). [DOI] [PubMed] [Google Scholar]
- 109.Richards, S. et al. Standards and guidelines for the interpretation of sequence variants: a joint consensus recommendation of the American College of Medical Genetics and Genomics and the Association for Molecular Pathology. Genet. Med.17, 405–424 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 110.Harris, R. S., Cechova, M. & Makova, K. D. Noise-cancelling repeat finder: uncovering tandem repeats in error-prone long-read sequencing data. Bioinformatics35, 4809–4811 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 111.Robinson, J. T. et al. Integrative genomics viewer. Nat. Biotechnol.29, 24–26 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 112.Li, H. A statistical framework for SNP calling, mutation discovery, association mapping and population genetical parameter estimation from sequencing data. Bioinformatics27, 2987–2993 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 113.Jiang, T. et al. Long-read-based human genomic structural variation detection with cuteSV. Genome Biol.21, 189 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 114.Heller, D. & Vingron, M. SVIM: structural variant identification using mapped long reads. Bioinformatics35, 2907–2915 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
Mapped BAM files of all 15 positive controls with known pathogenic variants used in this study were deposited as a resource in the European Nucleotide Archive (ENA) at EMBL-EBI under accession number PRJEB105200 (https://www.ebi.ac.uk/ena/browser/view/PRJEB105200). The remaining datasets used and/or analyzed during the current study are available from the corresponding author on reasonable request.
The detailed commands used in this study are provided on GitHub (github.com/aFienemann/Adaptive-sampling-gene-panel-for-PD-and-REDs).
