Skip to main content
Scientific Reports logoLink to Scientific Reports
. 2026 Aug 28;16:27095. doi: 10.1038/s41598-026-61605-4

Learning precise segmentation of neurofibrillary tangles from rapid manual point annotations

Sina Ghandian 1,2,3,4, Liane Albarghouthi 1,2,3,4, Kiana Nava 5, Shivam R Rai Sharma 6,7, Lise Minaud 1,2,3,4, Laurel Beckett 8, Naomi Saito 8, Charles DeCarli 9, Robert A Rissman 10, Andrew F Teich 11, Lee-Way Jin 5, Brittany N Dugger 5,✉, Michael J Keiser 1,2,3,4,✉
PMCID: PMC13524952  PMID: 42665582

Abstract

Accumulation of abnormal tau protein into neurofibrillary tangles (NFTs) is a pathologic hallmark of Alzheimer disease (AD). Accurate detection of NFTs in tissue samples can reveal relationships with clinical, demographic, and genetic features through deep phenotyping. However, expert manual analysis is time-consuming, subject to observer variability, and cannot handle the data amounts generated by modern imaging. We present a scalable, open-source, deep-learning approach to quantify NFT burden in digital whole slide images (WSIs) of post-mortem human brain tissue. To achieve this, we developed a method to generate detailed NFT boundaries directly from single-point-per-NFT annotations. We then trained a semantic segmentation model on 45 annotated 2400 μm by 1200 μm regions of interest (ROIs) selected from 15 unique temporal cortex WSIs of AD cases from three institutions (University of California (UC)-Davis, UC-San Diego, and Columbia University). Segmenting NFTs at the single-pixel level, the model achieved an area under the receiver operating characteristic of 0.832 and an F1 of 0.527 (196-fold over random) on a held-out test set of 664 NFTs from 20 ROIs (7 WSIs). We compared this to deep object detection, which achieved comparable but coarser-grained performance that was 60% faster. The segmentation and object detection models correlated well with expert semi-quantitative scores at the whole-slide level (Spearman’s rho ρ = 0.654 (p = 6.50e-5) and ρ = 0.513 (p = 3.18e-3), respectively). We openly release this multi-institution deep-learning pipeline to provide detailed NFT spatial distribution and morphology analysis capability at a scale otherwise infeasible by manual assessment.

Supplementary Information

The online version contains supplementary material available at https://doi.org/10.1038/s41598-026-61605-4.

Keywords: Digital pathology, Neurofibrillary tangles, NFT, Tau, Deep learning, Semantic segmentation, Object detection, Open source, Alzheimer disease, Neuropathology

Subject terms: Alzheimer's disease, Machine learning, Software, Image processing, Immunohistochemistry, Neurodegeneration, Alzheimer's disease, Neurodegeneration, Alzheimer's disease, Prion diseases, Alzheimer's disease, Alzheimer's disease, Prion diseases, Neurodegeneration, Data processing

Introduction

Accurate identification and quantification of neuropathological hallmarks such as neurofibrillary tangles (NFTs) may be crucial for advancing our knowledge of Alzheimer disease progression and developing effective interventions1. Convolutional Neural Networks (CNNs) and their variants have demonstrated remarkable capabilities for image recognition and segmentation2 tasks in the medical domain3. In neuropathology, deep learning on digitized whole slide images (WSIs) of brain tissue can automate detecting and quantifying distinct pathological features such as amyloid beta plaques4–6. This includes recognizing and quantifying the NFTs central to AD diagnosis and staging7–10. However, several challenges persist. Variability in staining techniques, tissue preparation, and imaging conditions across laboratories hinders the generalization of deep learning models11,12. Additionally, limited expert annotator bandwidth creates a scarcity of large, well-annotated datasets for neuropathologies. Previous studies addressing NFT and tau pathology quantification have required substantial annotation investments, employing approaches ranging from semi-automated color segmentation (Wurts et al.7 Dice: 0.611) to extensive manual pixel-level outlining (Signaevsky et al.8, F1: 0.81; Ingrassia et al.16, Dice: 0.77–0.83). While manual outlining produces high-quality ground truth, it demands hours per whole slide image, creating a practical bottleneck that prevents scaling to multi-institutional cohorts, which hampers consistency and reproducibility. This has encouraged research in label-efficient modeling strategies such as self-supervision, weak supervision, and multiple-instance learning13–15. The multi-class nature of tau pathology complicates this further. Studies have variably focused on either NFT subtypes — pre-NFTs, mature NFTs, and ghost tangles8—or labeled other lesions such as neuritic plaques16–18 or tufted astrocytes19,20. This heterogeneity in task definition makes cross-study comparisons difficult and complicates efforts to build consensus on best practices. Addressing these challenges is essential to deploy deep learning in neuropathology as consistent and reproducible analyses.

Building on our deep-learning-based neuropathological image analysis research4,5,11,12,21, we introduce a unique open-source robust algorithm for automated detection, segmentation, and quantification of mature NFTs in the temporal lobe of AD brain tissue WSIs. Crucially, we develop a framework for automatically converting point-annotated NFTs to detailed ground-truth segmentation masks to maximize annotator bandwidth and harness active learning approaches more effectively. We leverage a carefully curated dataset from multiple Alzheimer’s Disease Research Centers (ADRCs) and employ a straightforward and reproducible modeling architecture to segment NFTs. The objective is to provide researchers with a freely accessible, efficient, and reliable tool and framework to enhance NFT burden quantification to ultimately advance our understanding of Alzheimer disease and other neurodegenerative diseases. Quantitative data on NFT burden can aid in more robust correlations to clinical, demographic, and other data collected.

The model strongly correlates to expert-assigned WSI semi-quantitative scores on a 24-case hold-out set. These “CERAD-like” scores follow a scale for NFTs similar to that from the original CERAD criteria for neuritic plaques5,22. We present the methodology, including dataset curation, deep learning model architecture, evaluation metrics, and object detection benchmarks. We also discuss the approach’s clinical implications and potential benefits in the broader neuropathology research context of harnessing computational methods for more accurate and consistent analysis. Finally, we are among the first groups to publicly release our complete semantic segmentation and object detection pipelines, including code, datasets, and pre-trained models, enabling independent validation and reducing barriers to entry for other researchers.

Methods

Dataset curation

Cohort selection

We obtained de-identified autopsy brain tissue samples devoid of personal identifiers and compliant with HIPAA regulations consistent with previous practices4,5,11,12,21,23. This study only utilized human post-mortem tissues and is not considered human subject research, as only living subjects are defined as Human Subjects under federal law (45 CFR 46, Protection of Human Subjects). For participants during life, ethical approval for this study was granted by the Columbia Institutional Review Board, the University of California San Diego Institutional Review Board, and the University of California Davis Institutional Review Board and was performed in accordance with the Declaration of Helsinki and other state/federal and institutional guidelines. At death, autopsies were performed after legal consent for autopsy was provided by appropriate family members. As the dataset originates from ADRCs, we consistently collected select data using standardized forms from the National Alzheimer’s Coordinating Center to ensure data integrity and consistency across cases24. The annotated dataset comprises a subset of 22 cases from 295 cases collected from three distinct Alzheimer’s Disease Research Centers (ADRCs): the University of California Davis ADRC, the Columbia ADRC, and the University of California San Diego ADRC, following published case-specific inclusion/exclusion criteria23. Both the training (n = 22 WSIs) and validation batches (n = 24 WSIs) were randomly sampled from the larger dataset of cases across the 3 ADRCs included in the study, whose source details have been published23. In this study, we used batches assigned by permuted block randomization within ADRC, gender, and ethnicity strata23. Table 1 reports this demographic information. These cases came from a diverse pool of research subjects recruited from various sources, including the practices of participating neurologists and community-based recruitment. The source publication delineates additional recruitment strategy information23. All cases met the pathological criteria for Alzheimer disease (AD), meeting NIA Reagan or NIA-AA intermediate/high criteria25,26.

Table 1.

Demographics grouped by dataset split.

Grouped by Dataset Split
Missing Overall Test Set Training Set Validation Batch P-Value
n 46 7 15 24
Sex, n (%) Female 1 26 (57.8) 3 (42.9) 9 (60.0) 14 (60.9) 0.684
Male 19 (42.2) 4 (57.1) 6 (40.0) 9 (39.1)
Age at Death, Yrs median [Q1,Q3] 1 83.0 [77.0,88.0] 83.0 [78.0,86.0] 84.0 [78.0,90.0] 82.0 [78.0,86.5] 0.574
Hispanic ethnicity, n (%) Yes 1 13 (28.9) 1 (14.3) 4 (26.7) 8 (34.8) 0.562
No 32 (71.1) 6 (85.7) 11 (73.3) 15 (65.2)
Hispanic origins, n (%) Cuban 33 1 (7.7) 1 (12.5) 0.785
Dominican 2 (15.4) 1 (25.0) 1 (12.5)
Mexican 4 (30.8) 1 (25.0) 3 (37.5)
Puerto Rican 3 (23.1) 1 (25.0) 2 (25.0)
Unknown 3 (23.1) 1 (100.0) 1 (25.0) 1 (12.5)
Race, n (%) African American 1 1 (2.2) 1 (6.7) 0.496
Other 4 (8.9) 2 (13.3) 2 (8.7)
Unknown 3 (6.7) 2 (13.3) 1 (4.3)
White 37 (82.2) 7 (100.0) 10 (66.7) 20 (87.0)
Years of Education, median [Q1,Q3] 1 15.0 [12.0,16.0] 14.0 [12.5,15.5] 16.0 [10.5,16.0] 14.0 [12.0,16.0] 0.902
CERAD-like* NFT Score, n (%) 0 1 1 (2.2) 1 (14.3) 0.044
1 8 (17.8) 1 (14.3) 5 (33.3) 2 (8.7)
2 16 (35.6) 4 (57.1) 5 (33.3) 7 (30.4)
3 20 (44.4) 1 (14.3) 5 (33.3) 14 (60.9)
PMI, median [Q1,Q3] 35 6.0 [2.8,21.9] 24.0 [22.4,36.0] 6.0 [3.1,14.5] 5.2 [0.0,6.0] 0.057
Primary Clinical Diagnosis, n (%) AD 2 20 (45.5) 4 (57.1) 4 (30.8) 12 (50.0) 0.145
CERAD Definite AD 2 (4.5) 1 (14.3) 1 (7.7)
CERAD Probable AD 1 (2.3) 1 (7.7)
Definite AD 11 (25.0) 5 (38.5) 6 (25.0)
LB variant AD 9 (20.5) 1 (14.3) 2 (15.4) 6 (25.0)
Probable AD 1 (2.3) 1 (14.3)
ADRC Origin, n (%) Columbia 2 21 (47.7) 2 (28.6) 7 (53.8) 12 (50.0) 0.624
UC Davis 9 (20.5) 1 (14.3) 3 (23.1) 5 (20.81)
UC San Diego 14 (31.8) 4 (57.1) 3 (23.1) 7 (29.2)
Braak NFT Stage, n (%) III 16 1 (3.3) 1 (14.3) 0.280
IV 4 (13.3) 1 (14.3) 2 (25.0) 1 (6.7)
V 6 (20.0) 2 (28.6) 4 (26.7)
VI 19 (63.3) 3 (42.9) 6 (75.0) 10 (66.7)
CERAD, n (%) Sparse 13 2 (6.1) 1 (10.0) 1 (5.9) 0.496
Moderate 4 (12.1) 2 (33.3) 1 (10.0) 1 (5.9)
Frequent 20 (60.6) 4 (66.7) 5 (50.0) 11 (64.7)
Not assessed 7 (21.2) 3 (30.0) 4 (23.5)
Thal, n (%) 4 32 1 (7.1) 1 (25.0) 0.559
5 10 (71.4) 2 (50.0) 3 (75.0) 5 (83.3)
Not assessed 3 (21.4) 1 (25.0) 1 (25.0) 1 (16.7)

Demographics, neuropathologic variables, and clinical diagnoses. P-values were calculated between datasets using appropriate statistical tests; continuous and normally distributed variables were tested using one-way ANOVA, continuous and non-normal variables were tested using Kruskal-Wallis, and categorical variables were tested via Chi-squared test. PMI: Post-mortem interval; LB: Lewy body. One external-validation case lacked demographic data and is counted under Missing for all demographic variables.

Histology and slide-level assessments

This study used 5–7 μm formalin-fixed paraffin-embedded (FFPE) sections from the temporal cortex. These sections arose from designated anatomical regions available at each ADRC. Each ADRC prepared its FFPE sections, mounted the slides, and shipped unstained slides to the University of California Davis (UCD) for staining to minimize batch effects. As previously published23, we performed all antibody staining procedures under laboratory best practice standards, meeting Federal, State of California, and UC Davis guidelines and regulations. We used appropriate positive and negative controls for each antibody in each run.

We stained temporal cortex slides with the AT8 antibody (1:1000, ThermoFisher Scientific Cat# MN1020, RRID: AB_223647). All slide sections were digitized, capturing whole slide images (WSI) using a Zeiss Axio Scan Z.1 microscope at 40x magnification, creating images with a 0.11 μm/pixel resolution saved in the proprietary Carl Zeiss (.CZI) format.

An expert (BND) performed semi-quantitative histopathological assessments of NFTs on each WSI, blinded to demographic, clinical, and genetic information. The assessments followed semi-quantitative protocols outlined by the Consortium to Establish a Registry for Alzheimer’s Disease (CERAD)22. This protocol consisted of denoting the densest 1mm2 area of NFTs per WSI as none (no NFTs present), sparse (0–5 NFTs), moderate (6–20 NFTs), or frequent (greater than 20 NFTs).

Data annotation (annotations version one)

We used Zen Blue 3.2 software for WSI annotation. We visually explored the tissue and identified three regions of interest (ROIs), each measuring 10,680 × 21,236 pixels with a resolution of 0.11 microns per pixel. We designated two gray matter ROIs that spanned the entire cortical region and one ROI along the gray matter-white matter junction (e.g., Fig. 1a). The annotation process referred to the visualizations and descriptors outlined in Moloney et al. 20211 to determine whether a neuron met the criteria for being a neurofibrillary tangle. While Moloney et al. 2021 define criteria for identifying pre-tangles, mature tangles, and ghost tangles, we focused solely on marking mature tangles within a cell with a clearly defined nucleolus. The trained annotator (KN) meticulously scanned the ROI, looking for flame-shaped tangles that conformed to the shape of the respective neuron defined by a visible nucleolus.

Fig. 1.

Fig. 1

Neurofibrillary tangle model annotation and training pipeline. (A) Representative slide with three regions of interest and their point-annotated NFTs. Processing steps are as follows: (B) Isolate rotated ROIs and transform global NFT coordinates to local coordinates. (C) Zoom to point annotation (right), isolate the DAB channel (middle), and apply morphological operations and Otsu thresholding to segment large “blobs” (left). (D) Apply a center bias to remove off-center NFT candidates, retaining a single blob per tile. (E) Stitch tiles into a single large ground truth mask for each ROI, replacing point annotations. (F) Generate class-balanced batches for training via random NFT and tile sampling from the ROI using an input tile size of 1024 × 1024 pixels. (G) Slice WSIs in a structured format with a stride equal to tile size. (H) Feed tiles into UNet with ResNet50 backbone to generate prediction masks. (I) Stitch tile predictions and overlay them onto the WSI to visualize the heatmap.

We annotated neurofibrillary tangles that exhibited defined boundaries with smooth curves, contained a nucleolus, had AT8 staining filling the cell, were located in gray matter, and typically featured 1–2 protrusions. The NFT was marked at the nucleolus (Fig. 1b) if and only if it met all criteria; otherwise, it was left unmarked. Identifying the nucleolus with a cross marking ensured a consistent size and relative location, which was crucial for subsequent training. We used a clear-cut framework following a uniform NFT definition to enhance consistency in the automated detection model (Supp. Figure 1). A total of 1476 NFTs were annotated across the 74 ROIs.

Datatype conversion

We converted the Carl Zeiss proprietary format images into the open-source Zarr file format to facilitate data analysis pipelines27. We loaded the highest resolution view of each WSI into an intermediary numpy28 array in memory and then saved it to disk in Zarr. We set the storage chunk size to 5000 pixels in X and Y dimensions. When images had multiple scenes, we separated them into distinct files with the same base name and a suffix indicating the corresponding scene (e.g., 1-343-Temporal_AT8_s1). The Carl Zeiss format’s native compression method (JPEG XR) saves disk space at the cost of read and write speed. We used a lossless compressor (Blosc-zstd, clevel = 5, bit shuffling enabled) that achieves significantly higher read and write speeds during WSI-level segmentation at the cost of disk space usage. Other users can change this compressor to suit their needs. We provide WSIs in the data repository in Zarr using the JPEG XR compression algorithm to reduce disk space and transfer bandwidth. Users will benefit from recompressing the files with the Blosc-zstd compressor for faster processing.

Rotated ROI correction

Annotated region of interest (ROI) orientations within WSIs were not guaranteed to align with slide or image edges. This introduced complexity when loading the ROI, as arrays require slicing along fixed columns and rows parallel to boundaries. Ensuring the isolation of the ROI was crucial to prevent the inadvertent introduction of non-annotated NFTs into the dataset while preserving the total count of manually annotated NFTs. We first sliced the minimum inscribing region around the rotated ROI to address this. Second, we used the skimage library’s transform.warp method to crop the ROI. We stored each corrected ROI as a Zarr file and attached its path to a custom WSIAnnotation Python object, which was saved to disk using the pickle format for easy downstream processing29.

Point annotation to segmentation masks

We procedurally converted NFT point annotations to ground truth NFT pixel-boundary masks to train a semantic segmentation model. Bootstrapping point annotations into masks saves expert annotators time and scales to more annotations, facilitating efficient expert pathologist active-learning iterations. Additionally, this approach unlocks retrospective analysis for old datasets, reducing the annotation starting requirement from bounding boxes or masks to simple point annotations. Semantic segmentation enables detailed morphological analyses, WSI counts, and spatial distribution analyses of NFTs10.

We began this procedure by cropping 400 × 400px tiles around each NFT in the dataset (Figs. 1c and 2a). We padded NFTs at the ROI boundary to ensure consistent centering of the NFT in the cropped tile. Each tile, centered on an NFT, underwent a custom conventional image-processing segmentation pipeline to generate a 400 × 400px binary mask. This pipeline involved color deconvolution using skimage’s color.rgb2hed method to convert the tile to HED color channels, followed by min-max normalization for enhanced diaminobenzidine (DAB) channel extraction30. The tile was then Otsu thresholded and binarized, followed by post-processing with morphological opening and closing operations31, isolating the AT8 immunohistochemically (IHC) stained tissue (Fig. 2b).

Fig. 2.

Fig. 2

NFT point-to-mask pipeline detail. (A) We convert ROIs with human NFT point annotations to detailed segmentation masks. We do so by (B) generating 400 × 400 pixel tiles centered on the label, (C) applying color deconvolution, otsu thresholding, and morphological cleaning operations, and (D) performing blob detection with center-biasing and size filtering to obtain a mask. (E) We then stitch each mask into its corresponding location in the ROI via union operations.

Certain tiles posed challenges, such as those having closely clustered NFTs or background noise due to the high density of phosphorylated tau protein in the surrounding tissue. To specifically segment the target NFT, we used skimage’s measure.label method to label contiguous regions of the IHC mask. Subsequently, we iteratively removed all labeled regions (blobs) whose center of mass was not within 80px of the tile’s center. This “center bias” segmentation procedure favored the NFT most centrally located within the tile. Finally, we identified the largest blob in this region, and we removed all other blobs < 50% of its pixel size (Fig. 2c). We ran this pipeline on all annotations (Fig. 1d, Supp. Figure 2), then plotted the distribution of ground-truth NFT segmentations across all ROIs to check for outlying tiles; we verified that no tiles had mask sizes much smaller than expected (close to 0) (Supp. Figure 3). Finally, we employed union operations and the coordinates of each tile to stitch together all masks corresponding to a single ROI upon an empty numpy array with a shape equal to its corresponding ROI, creating an ROI mask with each NFT segmented rather than point-annotated (Figs. 1e and 2d). We detail all steps in the pipeline further in Supp. Figure 4.

We applied this procedure to each WSI in the dataset (ntrain/test = 22, nval-batch = 24). We encapsulated all relevant annotation information for a WSI into a WSIAnnotation object and then saved it to disk using Python’s native object serialization library, pickle32. We generated lightweight pickle objects by employing Zarr for both the cropped and ground truth ROIs — we stored independent Zarr arrays and simply referenced their paths in the pickled objects. This strategy eliminated the requirement for time-consuming unpickling of the WSIAnnotation during model training and evaluation. While we successfully parallelized the creation of WSIAnnotation objects at the WSI level, the maximum RAM capacity of a workstation can still be a bottleneck. Each worker loaded a distinct WSI; considering each worker’s large memory footprint, this constraint was a significant factor.

Model training and evaluation

We conducted 10-fold cross-validation to train and evaluate the segmentation model. For each fold, we split the dataset 80/20 into training (n = 15) and hold-out test (n = 7) datasets stratified by WSI. We further split the training dataset into static training (n = 10) and validation (n = 5) sets in the same proportions stratified by WSI (Supp. Table 1).

To load tiles and their segmentation masks from disk efficiently, we designed a custom PyTorch33 Dataset combining Zarr with weighted random sampling. We created a static data frame with rows containing coordinates generated with a stride of 1024 (equivalent to the tile size) and mappings to a single ROI/WSI pair. We assigned it a positive binary label if an NFT point annotation was present within a 1024 × 1024 tile loaded at the given coordinate. We employed a custom PyTorch weighted random sampler with 50/50 class balancing during training. These oversampled tiles contain NFTs across all ROIs in the training set, addressing the sparsity of NFT tiles when randomly sampled. For instance, with a batch size of 32 in our training dataloader, we generated 32 1024 × 1024px tiles from random locations across all training set ROIs, ensuring that, on average, 50% contained NFTs instead of including many empty ground truth tiles. In validation, testing, and inference, we loaded tiles across ROIs non-randomly to guarantee full coverage of the input image.

We deployed Kornia’s geometric and color augmentation procedures to increase our training set size and enhance the model’s robustness on the hold-out test set34. Kornia utilizes GPU processing to generate these augmentations and dynamically considers which transformations should apply to the image and mask, respectively (e.g., masks should not be ColorJittered). Augmentations included horizontal and vertical flips, affine transformations (e.g., rotation, rescaling, shear), color jiggle of saturation, hue, brightness, and contrast, and Gaussian blurring.

We tested several different state-of-the-art model architectures (FPN and PAN) and encoders (resnet34, resnet101, resnext101_32×4d, vgg19, efficientnet-b5) using the `segmentation-models-pytorch` library35. Our internal results, however, indicated model performances did not achieve substantively different F1 values. Consequently, we chose a standard UNet model with a ResNet50 encoder, pre-trained on ImageNet35–37. We ImageNet-normalized the input images to account for this pretraining. Due to the imbalanced nature of the ground truth masks (i.e., NFTs take up much less than 50% of a tile), we desired a loss function that allowed for weighing false negative classifications more heavily. We chose Tversky loss with alpha and beta parameters defined by hyperparameter optimization via Weights and Biases’ (W&B) sweeps38,39. We monitored the training and validation losses, F1 scores, and positive IOUs using the W&B integration with Pytorch Lightning40.

We conducted training across two NVIDIA RTX 3090 GPUs using Pytorch Lightning’s distributed data-parallel scheme. We selected hyperparameters using W&B’s Bayesian optimizer and Hyperband early-stopping procedures41. Sweep parameters included learning rate, weight decay, optimizers, momentum, and alpha parameters for the Tversky loss function (Supp. Figure 5a). We include the sweep parameter ranges in the GitHub repo in the file titled “wandb_sweep_config.yaml”; the file titled “best_hyperparameters_UNet.yaml” contains the set of hyperparameters that achieves the lowest validation loss in fold0. We observed that the model either did not converge to a solution for most hyperparameter combinations, or converged to comparable results and thus fixed the hyperparameters to a working set and reported their performance accordingly. We tested on a single GPU with non-overlapping tiles to avoid potentially duplicated samples. Key metrics for assessing the best model performance included positive IOU, F1 score, and mIOU. We chose the epoch with the lowest validation loss for the final evaluation of the test set.

Agreement maps

We visualized model performance across varying fields of view to better understand the model’s selection criteria for NFTs. We thresholded predictions at 0.5 to convert pixel values from continuous to binary class labels that correspond to whether an NFT is present (value of 1) or absent (value of 0). We then constructed ‘agreement maps’ by directly comparing pixel-wise values between the ground truth and predictions (Fig. 3). Pixels are ‘true positive’ (TP) if their ground truth and predicted value at a given index are equal to 1 and ‘true negative’ (TN) if the ground truth and predicted values are 0. We assigned ‘false positive’ (FP) labels where the model incorrectly predicts an NFT that is not present in the ground truth (i.e., ground truth pixel value = 0, predicted pixel value = 1) and ‘false negative’ (FN) labels to the inverse. We used Matplotlib’s LinearSegmentedColormap class to distinguish TP, FP, and FN masks42. All other regions of the ROI, i.e., those without annotated NFTs, were considered TN.

Fig. 3.

Fig. 3

Agreement maps display how NFT predictions compare to human labels. Agreement maps from four different ROIs in order of increasing semi-quantitative severity. As the slide-level severity score increases, the model detects more NFTs. However, False Positives (FPs, yellow) and False Negatives (FNs, magenta) increasingly appear as the severity increases. (A) None, (B) Mild, (C) Moderate, and (D) Severe CERAD-like NFT burden scoring by expert assessment. Notably, True Positive (TP, cyan) predictions display high pixel-wise boundary fidelity, and disagreement between human labels and model predictions more typically applies to whether the entire object qualifies as a mature NFT.

Data re-annotation (annotations version two)

When examining preliminary results from the agreement maps, we observed many FP masks across the ROIs throughout the entire dataset (Figs. 3c and 4a). A closer examination of the FP-labeled raw image areas revealed many plausible predicted NFTs meeting the annotation criteria described earlier (which was carried out by a trained novice) (Fig. 4a,c). To stress-test the correctness and validity of our ground truth data, we initiated a re-annotation experiment conducted by an expert (BND) to rescue any NFTs overlooked in the initial round of annotations.

Fig. 4.

Fig. 4

Expert reannotation rescues putative predicted NFTs. (A) ROI-level agreement map comparing trainee-annotated ground truth NFT labels versus “version 1” model-predicted segmentation masks. (B) Agreement map of the same ROI using expert-reannotated ground truth NFT labels versus “version 2” model-predicted masks. (C) Zoomed view highlighting initial trainee-vs-model agreement maps. (D) Zoomed view highlighting annotated expert-vs-model agreement maps. Multiple newly expert-annotated (“rescued”) NFTs that were previously labeled FPs or TNs become TPs after integrating the reannotated data and retraining the model.

We sliced each ROI into fifteen 4247 × 3560 pixel “super-tiles,” which were large enough for our expert annotator to process efficiently while maintaining sufficient resolution to evaluate NFTs displayed in the super-tiles. We uploaded all super-tiles to the SuperAnnotate43 platform in a randomized order. Each re-annotation study image consisted of two identical side-by-side super-tiles — with the raw super-tile image on the right and the same image overlaid with the agreement map labels on the left (Supp. Figure 6). We chose this approach to accommodate annotator bandwidth limitations. As the focus was on correcting potentially mislabeled FP objects, the expert annotator concentrated on objects with FP masks, using SuperAnnotate’s point annotation tool to rapidly annotate any object meeting the NFT annotation criteria described earlier. No objects/NFTs were removed during the re-annotation experiment.

Following the expert’s examination of the super-tiles in the re-annotation procedure, we downloaded the re-annotation data from SuperAnnotate, which included JSON files for each image with new point annotations. We parsed these files and generated a table containing the WSI and ROI IDs, super-tile coordinates relative to the ROIs they originated from, and the coordinates of the new point annotations. In total, 280 new point annotations (a 19% addition overall) across 17 WSIs were added in the re-annotation experiment, providing higher-quality ground truth labels for training and evaluating the model (Fig. 4b,d). The entire dataset was re-annotated by BND in less than 90 min, facilitated by the platform and the outlined scheme.

Correlating WSI-level model scores with semi-quantitative scoring

We performed rapid inference using structured slices (sliding window approach) to construct WSI-level segmentations with the model. This allows for manual inspection of segmentations across the entire WSI by overlapping the produced segmentation mask and the original WSI (Fig. 5a). Subsequently, we generated a score summarizing the image’s NFT burden. We opted for a count-based score, which tallies the total number of contiguous blobs in a downscaled image using skimage’s measure.label function. The downscaled tissue area then normalizes this count to ensure consistent burden scores. The tissue area is calculated by applying histomicsTK’s63 tissue detection algorithm to isolate a binary mask and then summing the pixel area. Finally, we rescaled scores by the median tissue area of the training set to produce an interpretable result. We refer to this score as an NFTDetector score.

Fig. 5.

Fig. 5

Model-derived NFTDetector scores correlated with expert CERAD-like scores. (A) Representative WSIs spanning the four semi-quantitative stages of increasing NFT burden severity, ordered by CERAD-like scoring. Model NFT detections overlaid in green. (B) Area-normalized NFTDetector scores on test set WSIs correlate with semi-quantitative ordinal categories. (C) Area-normalized NFTDetector scores calculated on the full WSI dataset versus area-normalized counts for the same WSIs derived from the annotator labels. **: 0.001 < p < = 0.01 by Welch’s t-test. We used different assessment datasets in (B) vs. (C) because the external test set lacks human NFT point annotations.

We ran all slides in the training, validation, and testing datasets and a separate hold-out batch of images (not point-annotated, only having slide-level scores) through the model to generate a distribution of NFTDetector scores (Supp. Tables 2, 3). We then assign the WSIs and their NFTDetector scores into groups corresponding to their slide-level CERAD-like scores assigned by an expert pathologist (four grades ranging from None to Severe) and tested for statistical difference via Welch’s t-test using the statannotations44 library (Fig. 5b).

We compared the model’s slide-level semi-quantitative correlative performance with the trainee’s by applying the same score generation logic directly to the ground truth point annotations. By counting the number of point annotations across all ROIs for a given WSI and then normalizing by the total pixel area of the ROI, we generated a human analog to the NFTDetector score. We scaled scores by the median tissue area and min-max normalized to achieve better alignment of the y-axis for plotting purposes only. The null hypotheses aimed to validate whether the trainee’s semi-quantitative score groups were statistically different; this was confirmed via Welch’s t-test, as scores were not guaranteed to be normally distributed (Fig. 5b).

Comparing model predictions with AT8 burden and semi-quantitative scores

We generated simple AT8 burden scores by calculating the proportion of DAB signal present in an ROI. To do this, we applied a similar pipeline to the initial steps of the point-to-mask pipeline. We first converted the image from RGB to HED channels via skimage, then binarized the image using Otsu thresholding (Supp. Figure 7a). We then counted the total number of detected DAB positive pixels and normalized by the total number of pixels in the grayscale image to generate a DAB Proportion score (Supp. Figure 7b). We used this DAB Proportion score and the expert-assigned semi-quantitative CERAD-like scores as correlative variables with the models’ performance metrics (Supp. Figure 7b,c). We constructed boxplots and scatterplots with seaborn and tested for significant differences with statannotations44,45.

Statistical methods

This study operates at any of four levels. From the most to least granular: individual NFT objects (ntest=664), tiles (ntest=4620), ROIs (ntest=20), and WSI images (or decedents, ntest=7, nhold-out-batch=24). Only the object detection analyses (YOLOv8, segmentation bounding boxes) can operate at the level of individual NFTs (“object-level”). We use this level for Fig. 7 or whenever explicitly denoted. We use a tile-level loss function to train the segmentation and object detection models, reported in Supp. Figure 5. We report ROI-level metrics in Fig. 6b because the ground truth only exists at the ROI level. Performance and other characteristics are available in detail at the ROI level (Supp. Tables 4–7). Finally, at the WSI level (1-to-1 with decedent), we report correlations against semi-quantitative scores as in Fig. 5b,c, and 7c (Supp. Table 2). The sample sizes for the training data were as follows: objects (ntrain=792), tiles (ntrain=13860), ROIs (ntrain=45), and WSI images (or decedents, ntrain=15).

Fig. 6.

Fig. 6

Segmentation numerical performance metrics were not always visually intuitive. (A) Progressing in order of decreasing performance (left to right): Four 1024 × 1024 pixel tile-level agreement maps with corresponding numerical performance metrics for segmentation performance at the NFT object level from two different ROIs (rows). The F1, mean, and positive IOU scores are specific to each 1024 × 1024 tile. IOU and F1 scores suffer the most from false negatives, even when the model predicts other NFT pixels correctly. (B) At the ROI level, model performance on the test set, assessed by F1 and mean IOU, improved after reannotation, with substantial qualitative improvement evident by visualization.

Fig. 7.

Fig. 7

Repurposing the NFT dataset for object-detection models yields similar overall performance to the segmentation approach. (A) The precision-recall curve plots the tradeoff between the YOLO model’s object-detection precision and recall across all potential confidence thresholds aggregated across all objects in the test dataset. (B) Plot of aggregated object-wise F1 versus confidence threshold; low (permissive) thresholds yielded the highest F1 scores. (C) WSI-level NFT counts from the YOLO model correlated with CERAD-like semi-quantitative scoring, similar to the segmentation-derived NFTDetector scores in Fig. 5b. **: 0.001 < p < = 0.01 by Welch’s t-test. (D) Example tiles from different ROIs overlaid with ground truth (cyan) and predicted (yellow) bounding boxes. The values highlighted in yellow are confidence scores for the predictions. Reference (B) to determine which NFTs would be detected.

To calculate statistical significance, we proceeded as follows. In Table 1, we tested continuous and normally distributed variables using one-way ANOVA, continuous and non-normal variables using Kruskal-Wallis, and categorical variables via Chi-squared (Supp. Table 8). We used Welch’s t-test to compare continuous vs. categorical variables, as in Figs. 5b and 7c, and Supp. Figures 5b-c, and 9, and Student’s t-test for Supp. Figure 8a. Combining annotated (n = 22) and model (n = 46) scores of 46 decedents, we assessed score differences among Severe, Moderate, and Mild CERAD-like categories and score type (annotator vs. model). We excluded the None CERAD-like category due to insufficient examples (n = 1); however, both the annotator and the model correctly assigned this slide. To meet model assumptions, we used a mixed-effect model and transformed scores using the natural logarithm. For Fig. 6b, we use linear mixed models (Statsmodels.formula.api mixedlm46 to estimate the mean values and 95% confidence intervals for ROI-level metrics (e.g., pIOU, F1, and mIOU) from 7 decedents in the test set with repeated samples from each decedent, typically corresponding to 3 ROIs.

No formal a priori power calculation was performed; the training set size (n = 15 WSIs, 45 ROIs) was determined by the annotation effort feasible within the study and by case availability across the three ADRCs. To characterize the sensitivity of the primary segmentation endpoint post hoc, we computed the minimum detectable effect size (MDES) at the ROI level. We derived an effective standard error (SE ≈ 0.041) from the mixed-model 95% CI for ROI-level F1 [0.45, 0.61], which accounts for within-WSI clustering of repeated ROI measurements. At 80% power and one-tailed α = 0.05, the MDES ranges from 0.11 to 0.12 F1 units across the plausible range of effective degrees of freedom (df = 6 to df = 19, bounding cluster- and observation-level estimates); code implementing these calculations is available in the study’s GitHub repository. For the WSI-level Spearman correlation (n = 31 WSIs; 7 test and 24 external validation), Fisher’s z-transformation gives an MDES of ρ ≥ 0.44 (one-tailed, α = 0.05) or ρ ≥ 0.49 (two-tailed) at 80% power, well below the observed ρ = 0.654 (p = 6.50 × 10− 5).

Metric choices

Object detection and instance segmentation models typically report mean average precision (mAP), while fewer studies report segmentation model performance by mAP. Since we perform binary segmentation, mAP is equivalent to AUPRC for the positive class47. We observed that the model’s confidences are highly skewed towards 0 or 1, such that varying the confidence threshold for the segmentation predictions does not meaningfully alter performance (Supp. Figure 9). For this reason, we report F1 instead (i.e., Dice score for pixel-based calculations), as it is approximately the same as mAP given the above constraints. To determine a random F1 baseline, we used a null random model that achieves a precision equal to the test set ROIs’ positive pixel prevalence of 0.000898 (CI: [0.000205, 0.00159]) and a recall of 0.5. F1 can be further deconstructed into precision and recall at the chosen threshold, displayed within the figures of this study. We also report mean intersection over union (mIOU) — the average of positive and negative IOUs — and another standard segmentation metric. We calculate AUROC using sklearn.metrics.RocCurveDisplay48. We aggregate these metrics across all pixels in the dataset during training (i.e., flattening the predictions) but report the test set metrics at the ROI level.

Object detection: converting masks to bounding boxes for instance detection metrics

To align with Signaevsky et al.8 when comparing model performance, we transformed NFT segmentations into object detections to obtain object-level metrics, as they appear to have done in their study (we could not find precise metric details). While this assessment method may not offer the granularity of pixel-level metrics, it can be more forgiving if an entire NFT prediction does not perfectly align with its ground truth mask.

To implement this, we leveraged histomicsTK’s contour generation functions to generate bounding boxes from stitched ROI prediction masks. We merged nearby bounding boxes within 150 pixels from center to center via a custom algorithm (Supp. Figure 10a). We generated bounding boxes around the ground truth masks by taking the minimum and maximum coordinates of the mask in a cropped 400 × 400px view of each point-annotated NFT, consistent with the tile size used to construct the masks (Supp. Figure 10b). We then matched bounding boxes from the same ROI for greatest IOU. We assigned unmatched and very low IOU prediction boxes (i.e., IOU < 0.001) as false positives and unmatched ground truth boxes as false negatives (Supp. Table 9a). We also calculate performance at IOU = 0.5 to facilitate comparison with studies using this metric (Supp. Table 9b). We used these values to calculate our dataset splits’ precision, recall, and F1 scores.

Object detection: training an object detection model from bootstrapped bounding boxes

We generated ROI-level ground truth bounding boxes as described above (i.e., min/max coordinates of ground truth masks cropped around each NFT, performed at the ROI level). We then leveraged these ROI-level bounding boxes to create tile-level bounding boxes to train a YOLOv8 model49,50. We assigned an ROI-level bounding box to a single tile if it contained the center of the bounding box. We saved non-overlapping 1024 × 1024px tiles from the ROIs as static PNG files due to a limitation of the Ultralytics library, which we used to train the YOLOv8 model49,50. We shifted bounding box coordinates according to their assigned tile’s coordinates, clipped them to fit within the tile, converted them to YOLO format ([label, xmin, ymin, xmax, ymax]), and then saved them into a text file with the same name as its corresponding tile PNG.

We followed the protocol to generate a YOLOv8 configuration file in YAML format. In this file, we defined the paths to the train, validation, test directories, and the names of our labels (e.g., 0: NFT). Training a YOLOv8 model was straightforward after constructing the dataset with the correct input format (see Supp. Figure 5b for loss curves). We observed low variability in the segmentation model’s performance across folds (Supp. Table 10), so we report YOLOv8 results on a single fold (fold0). To maximize performance, we optimized hyperparameters with the default search parameters suggested by Ultralytics and their built-in hyperparameter search evolution algorithm via model.tune() for 30 iterations. The best hyperparameter settings are included in the Github repo file, “best_hyperparameters_yolo.yaml,” and tuning results can be found in Supp. Figure 11.

Use of LLMs

We used ChatGPT (https://chat.openai.com/, February 2024) to provide editing feedback for the manuscript and code. We also used it to convert tables from raw text into Markdown and to generate a first draft of the Abstract from the study results and field context, which we further edited. During manuscript editing, we subjected specific sections of the text to editing review by ChatGPT to enhance clarity and conciseness, manually incorporating suggestions as appropriate.

Results

We automatically generated pixel-wise masks of NFTs from single-point annotations

We developed an automated point-to-mask image processing pipeline to reproducibly generate pixel-level masks (outlines) around NFTs solely from manual point annotations. The motivation was that manually outlining these objects at a sufficient scale to train a deep learning model is exceedingly time-consuming, nuanced, and prone to high degrees of interobserver variation in determining the precise boundary of an NFT. The resulting point-to-mask pipeline generated high-quality ground truth segmentations with few visually erroneous examples despite leveraging solely conventional image-processing libraries (see Methods). We sampled 10 ROIs across semi-quantitative categories to assess the generated segmentation masks’ accuracy, inspecting 50 ground truth NFT segmentations per ROI. We also empirically validated results by generating a distribution of NFT ground-truth masked pixel areas and then visually inspecting examples of the largest and smallest masks in the dataset.

Using this automated pipeline, we processed 1476 objects across the 74 ROIs. The resulting masks, on average, comprised 5.5% of a tile (or 8800 pixels in area at 0.11 microns per pixel), with a minimum of 0.25% (~ 400 pixels), a maximum of 32% (51,000 pixels), and a standard deviation of 3% (4940 pixels). By this method, we proceeded through an iterative process to refine parameters of the point-to-mask pipeline, such as the center bias and minimum size variables, to ensure NFT masks were sufficiently large and to minimize the chance that they are over-segmented as sometimes seen in high-background regions (Supp. Figure 3). In the originally calculated masks, many objects would be missed due to an over-stringent center bias parameter. We corrected this by increasing the “center bias” parameter from 80 pixels to 120 pixels. The center bias procedure ensured that the detected blob’s center of mass was within a predefined number of pixels from the center of the point annotation. This automatically fixed the ~ 7% of the tiles missed in the first iteration. NFTs that were typically affected by this procedure were those with longer tails and contiguous staining from nucleus to axon. After applying the filtering pipeline, we inspected the histogram of mask sizes and found the proportion of over-segmented tiles to be acceptable (< 1% of tiles). Rather than manually adjust these masks, we included them to maintain a fully automated pipeline and to determine whether the trained model would be robust to the noisier ground truth labels.

We compared how guided expert reannotation affected segmentation model quality

As there is no single metric to score how well a model identifies (segments) an object at the pixel level, we used standard segmentation metrics of mIOU and F1. As it may not be clear from numerical scores alone what constitutes mediocre, acceptable, versus exceptional performance, we frequently inspected representative examples visually (Fig. 6a). For the segmentation model, we conducted 10-fold cross-validation and reported pixel-level performance metrics. We observed consistent performance across folds (F1 = 0.51, CI: [0.49, 0.53]) so we selected a single fold (fold0) for downstream comparison (Supp. Table 10).

After training a “version 1” model, we conducted a comprehensive reannotation procedure in which an expert manually reviewed each false positive prediction to see whether they met the study’s definition of an NFT (i.e., clear nucleolus present). In doing so, we increased the number of labeled NFTs in our dataset by 280, a roughly 19% gain. Furthermore, this indicated the trainee had nearly a 20% false negative rate relative to an expert in identifying NFTs meeting the study’s labeling criteria. In the “version 1” model before expert reannotation, the F1 at the ROI level was 0.49 (CI: [0.44, 0.54]). Following retraining on the re-annotated dataset, the final (“version 2”) model performance improved by about 8% at the ROI level to a precision of 0.53 (CI: [0.37, 0.7]), recall of 0.60 (CI: [0.49, 0.71]), F1 of 0.53 (CI: [0.45, 0.61]), mIOU of 0.68 (CI: [0.65, 0.72]) and AUROC of 0.832 (CI: [0.781, 0.883]) on a test set of seven held-out whole slide images (WSIs) (Fig. 6b). Additionally, the model’s area-normalized WSI score increased its correlation with semi-quantitative scores on the hold-out test sets (Spearman’s rho, ρ, increased from 0.624 to 0.654) (Supp. Figure 12).

Segmentation and repurposed instance-detection interpretations were consistent

Compared with other studies, we also repurposed the segmentation model predictions as though the masks only predicted objects – i.e., independently of whether the predicted NFT pixel boundaries were precisely accurate. We calculated approximate bounding boxes from the ground truth and prediction masks (see Methods) to do this. This allowed us to assess the NFT segmentation model as a coarse-grained object detector model instead. We scored NFTs as true positive (TP) if the intersection over union (IOU) between the prediction and ground truth bounding boxes exceeded 0, as reported previously in the field51. False positive (FP) NFTs were predictions that did not overlap with ground truth bounding boxes, while FNs were ground truth bounding boxes that did not overlap with predictions. Intriguingly, this more categorical and coarse-grained object-level performance nonetheless mirrors pixel-level performance, yielding a test set F1 of 0.53 for objects with IOU > 0.001 (CI: [0.45, 0.62]; Supp. Table 9a). Requiring a higher IOU threshold of 0.5 slightly reduces F1 to 0.45 (CI: [0.38, 0.52]; Supp. Table 9b), although it would be a future direction to explore whether the deep learning prediction surpasses the original annotation mask fidelity, perhaps in blinded comparative expert review studies.

A dedicated YOLOv8 model did not improve NFT instance detection

Whereas the above approach reinterpreted the NFT segmentation model post hoc as an NFT object detector, we posited that an established object-detection architecture trained from scratch on the bootstrapped bounding box dataset might perform better at this task. Accordingly, we trained a standalone YOLOv8 model with the sole task of object detection using the ground truth bounding boxes constructed for the object detection metric baseline. The YOLOv8 model performed comparably but no better than the repurposed segmentation model, achieving an ROI-level F1 of 0.53 (CI: [0.46, 0.60], Fig. 7b) and an aggregate object-level mAP50 (equivalent to AUPRC for the NFT class) of 0.485 (Fig. 7a). The model also displayed similar effectiveness in categorizing slides into “CERAD-like” NFT semi-quantitative grades (ρ = 0.513, Fig. 7c). Visual inspection revealed that FPs were predominantly NFT-like objects failing to meet our particular annotation criteria. At the same time, FNs likely resulted from this specificity (Fig. 7d).

Performance did not vary by AT8 stain burden or CERAD-like category

We noticed that the point-to-mask pipeline generated larger masks from dark AT8 staining around cropped NFTs, so we posited that some ground truth masks could be lower quality in tissue regions with higher overall AT8 stain burden. To assess this, we analyzed whether the model performed differently in images with higher AT8 stain burden, using semi-quantitative scores and quantified diaminobenzidine (DAB) signals as correlative variables. We found the model exhibited consistent performance (mIOU) across images despite varying AT8 staining intensities (R2 = 0.109, Supp. Figure 7c) and semi-quantitative grades (ρ = 0.026; Supp. Figure 7c). Furthermore, the model effectively differentiated slide-level semi-quantitative grades assigned by an expert, outperforming grades assigned using DAB signal alone, particularly between the Moderate and Severe CERAD-like categories (Supp. Figure 7b).

Models identify NFT objects down to the pixel on one GPU an order of magnitude faster than humans can roughly point-annotate them

The conversion from the .czi file type to the .zarr file type averaged 7.7 min per WSI and required enough random-access memory (RAM) to load the entire uncompressed WSI (50.2 ± 21.4 GB). A user with RAM limits could use the aicspylibczi library instead to load regions of .czi files into memory rather than the entire image and build the Zarr files in successive chunks.

Generating a WSI-level segmentation takes approximately 32 min per WSI per GPU on a dataset with WSIs that average roughly 160,000 by 220,000px. This was 2.5x faster than reported by Wurts et al. 20207, who segmented WSIs of size 120,000 × 120,000 pixels at 20 min using four Volta V100 GPUs, which are comparable to the NVIDIA RTX 3090s we used. Further reduction in time could be achieved by generating a tissue mask via a lower-resolution image and running inference solely on the masked region. Generating WSI segmentations for the entire cohort (48 WSIs) took 12.8 h across two NVIDIA RTX 3090s. The object detection model’s inference speed was approximately 60% faster, taking about 20 min per WSI per GPU; the entire cohort of 48 WSIs took 8.1 h with 2 NVIDIA RTX 3090s.

To calculate the NFTDetector score, we loaded each segmented WSI at 1/64 of the original resolution and post-processed it through the blob and tissue detection pipelines (Methods). On average, this took 4 min and 1.5 min per WSI, respectively. The YOLOv8 model’s WSI-level score required a post-processing non-maximum suppression procedure to eliminate overlapping bounding boxes, averaging 5.2 s per WSI and totaling 4 min across the dataset.

The version-1 annotation files the trainee (KN) generated did not have timestamps per point annotation. However, through personal communications, we determined that the average ROI that contained NFTs took approximately 30 min to annotate. We estimate that generating the initial point annotations for this dataset took 33 h (n = 1476 NFT annotations, or 44.7 annotations/hour). For the version-2 reannotation study, we directly parsed the image metadata from SuperAnnotate to assess the time it took for our expert annotator to add new point annotations to the dataset. Reannotating the entire dataset, which consisted of 975 “super-tiles,” took 82 min (11.9 super-tiles/minute). This reannotation process, which ultimately rescued 280 new annotations (or 204.9 annotations/hour), was completed in one sitting.

As the WSIs in the study had a mean of 1028 ± 957 detected NFTs, the model’s 1927.5 annotations/hour rate can automatically process entire slides at a ratio of 9.4x-43x faster than human annotators per single GPU.

Slide-level neuropathologic burden scores from the NFT model correlated with separate expert semiquantitative assessments

While pixel-level identification of NFTs enables deep neuropathological phenotyping, many studies rely instead on WSI-level CERAD-like semiquantitative assessments of neuropathology burden. To calculate an automated per-slide single score for NFTs, we generated area-normalized scores from WSI-wide segmentations by counting the number of detected NFTs and normalizing the count by the tissue area (Fig. 5a). We then compared these calculated scores versus previously expert-assigned semi-quantitative scores and subjected them to statistical analysis using Spearman’s rho, ρ = 0.654 (Fig. 5b). The model accurately discriminated semi-quantitative categories, particularly in distinguishing Mild versus Moderate categories (Fig. 5b). As all cases had an intermediate/high AD neuropathologic diagnosis, few cases had none or very few NFTs; despite this, we qualitatively observed that the model discriminated between the None and Mild categories (i.e., very few detections in slides assigned None or Mild (Fig. 5c)).

For comparison, we also constructed an area-normalized WSI-score benchmark directly from the NFT point annotations provided by the novice using the same methodology as the model’s WSI-level scores. We used mixed effect models to characterize the difference in scoring between the model and annotator at each CERAD level while accounting appropriately for the strong similarity of samples from the same decedent. We found that both the model and the annotator scored the Mild-or-Less category significantly lower than the Moderate category, with a mean 38% lower for the novice (95% CI (0.15, 0.98)) and a somewhat greater drop for the model (24% vs. 38%, but not significantly different from the novice). Neither model nor novice showed statistically significant differences between samples from the Frequent and Moderate CERAD categories after accounting for repeated measurements from decedents. The model and novice did not score significantly differently from each other within the Moderate or Mild-or-Less categories; the estimated scores for the model were 63% higher in the Frequent category than for the novice, a near-significant difference (p = 0.061, 95% CI 2% less to 2.7-fold greater.)

Lastly, we performed a correlation analysis between demographic features, semi-quantitative scores, and WSI-level predictions. Categorical differences were not significant, except for the relationship between maximal educational attainment and whether the individual identified as Hispanic (p < 0.01, Supp. Figure 8a). We observed strong positive correlations between the annotator’s tissue area-normalized score and both models’ tissue area-normalized scores per WSI (ρCNNv2 = 0.704, ρYOLOv8 = 0.773) and notable negative correlations between both models’ area-normalized scores and the age at death (ρCNNv2= -0.375, ρYOLOv8 = -0.405). We also observed a notable positive correlation between the annotator’s tissue area-normalized score and years of education but did not see the same results when comparing against either model’s scores (Supp. Figure 8b).

Discussion

We present a scalable deep-learning pipeline to detect the location and precise boundaries of mature neurofibrillary tangles in immunohistochemically stained whole slide images despite only using rapid single-point human annotations as training data. The models performed comparably to expert semi-quantitative slide-level scoring despite not being explicitly trained at this task and equivalently to the trainee annotations it learned on (Fig. 5b). The models also outperform both the trainee and the expert in NFT detection speed with pixel-level specificity and visually include only tangle-like objects in their predictions (i.e., other tau lesions such as tufted astrocytes, astrocytic plaques, or neuritic plaques are not predicted as mature tangles). Further, they demonstrate similar performance even as distinguishing pathologies from background signal becomes more difficult; more specifically, model performance remains strong as the AT8 burden in the images increases. We achieve this performance with a framework that is easily extended to other datasets and enables rapid integration and conversion of point annotations into either ground truth masks or bounding boxes. This pipeline exhibits stability in a cohort sampled across three different ADRCs and is likely to perform similarly on slides sourced from other institutions. Additionally, we publish the data, annotations, model weights, and easily extensible code along with the study if others wish to fine-tune the model or modify the pipeline to suit their dataset.

Studies using deep learning methods to detect NFTs predominantly omit segmentation models due to the massive annotator investment required to collect manual ground truth training masks at the pixel level9,10,52 despite the reported research advantages of object- and pixel-based NFT counts over typical IHC positive pixel count methods53. Indeed, segmentation models unlock deeper neuropathological phenotyping than the standard pathological analyses that compress pathological information into a single category or semi-quantitative grade — examples include nuclei and tissue segmentation as tools to improve cancer grading schemes such as Gleason and C-Path scores54–56. Object detection models address some concerns, but the predicted bounding boxes they output do not directly capture morphological patterns in the detected NFTs. Although substantial morphological differences are well-documented between pretangles and mature tangles, NFTs may have undiscovered morphological nuances, particularly those characterized by distinct fibril structures, which are only unveiled through cryo-EM imaging1,57. Our study releases open-source NFT segmentation models trained solely from point annotations, requiring less annotator investment, delivering comparable object detection performance to published models, and exhibiting strong correlation with expert-assigned whole-slide image semi-quantitative grades6,8,10,19,58. In a future direction, we imagine this tool may enable retrospective analysis of existing datasets that contain only point or bounding box annotations, transforming them into suitable training input for high-resolution segmentation models.

Unlike other studies (see Supp. Table 11, summarizing study task, methods, performance, and the accessibility of the code, dataset, and model weights, links), we narrowed the focus to cells containing a nucleolus and mature tangles instead of annotating pre-tangles, ghost tangles, or cells that may appear neuronal but for which a nucleolus was not apparent. This likely decreased numerical model performance, as the decision boundary between mature tangles and other tau tangle categories is more complex than deciding between a tangle phenotype versus other objects or backgrounds in the slide. In addition, all annotations in this study are point annotations curated by a single trainee, then refined in an iteration by a single expert. Wurts et al.7 use a semi-automated color segmentation approach to generate “noisily labeled data”. However, their dataset spanned 27 WSIs from 1 patient, limiting population generalizability; they were also restricted to labeling all tau-positive inclusions as a single class, unlike our study. By contrast, Signaevsky et al.8 perform NFT segmentation using manual expert pixel-level annotations but report object-level F1 in their test set at 0.8; however, their methods for calculating object-level F1, such as IOU threshold, are not reported. Unlike our mature-tangle annotation criteria, they annotate NFT subtypes: pre-tangles, mature tangles, and ghost tangles. The dataset included 22 cases of tauopathy sampled from the hippocampal formation and dorsolateral prefrontal cortex stained with AT8; ground truth mask annotations were manually generated across 178 ROIs by three expert neuropathologists (ntrain= 14 WSIs). Maňoušková et al.51 conduct NFT and NP segmentation, also relying on manual pixel-level outline annotations by expert pathologists. They report high performance (F1 = 0.91), however, only on segmentation performance amongst tiles already suspected to have NFTs rather than at the ROI or slide-level; detection F1 is more modest at 0.76. Further, they do not provide code or data to reproduce their results. Ingrassia et al.16 also rely on manual pixel-level outline annotations by expert pathologists, but they leverage an AI-assisted framework to improve efficiency. They integrate proprietary software (Visiopharm) with a UNet trained on an initial small set of hand-drawn NFT and NP masks to iteratively improve their ground truth label quality. Two expert pathologists then validate the label masks post hoc. Ingrassia et al. acknowledge that the manual method is time-consuming but do not specify how long the reannotation tasks to complete. By contrast, our framework reduces annotation input to a single point per NFT and automatically generates boundary masks using conventional image-processing techniques. We record extensive time quantifications for many parts of our pipeline, and find overall that it achieves a reannotation of the entire training dataset in 45 min.

In terms of model architectures, Zhao et al. performed extensive benchmarking with state-of-the-art models against UNet and saw marginal to no gains in performance18. This was not entirely surprising, as more complex models such as Attention U-Net59 or Swin-Unet60 can require large dataset sizes to achieve similar performance to smaller models. Similarly, we favored scientific parsimony as the larger and more parameterized models, which can be more prone to shortcut learning61 and other complications, did not deliver enhanced performance in our own testing. Consequently, we employed the standard UNet and ResNet50 encoder for parsimony and parity with existing studies performing segmentation in medical imaging.

Moving to object detection, Vizcarra et al.10 assess various brain regions and multiple annotators (nexpert = 5, nnovice= 3), achieving a top object-detection macro F1 performance in the temporal region of 0.44 ± 0.06. This average considered pre-NFT and iNFT performance (F1pre-tangle = 0.20, F1iNFT = 0.71). Since we did not annotate and subcategorize pre-NFTs, they may be lumped into the model’s predictions for mature tangles; this would numerically penalize apparent model performance. Also, Vizcarra et al. note that there are generally more pre-NFTs predicted in the UC Davis cohort than iNFTs compared to Emory. By contrast, in their cohort, the opposite is true (albeit the model performed poorly in the UC Davis cohort overall). They proposed site-specific differences, such as differing tau antibodies. This lends evidence that our model suffers from the lack of pre-NFT labels. Their dataset, curated by five experts and three novices, underwent eight iterations of manual adjustment, removal, and addition of bounding boxes, whereas our study underwent a single iteration, collecting additional point annotations only.

Ramaswamy et al. 2023 report an amyloid plaque object detection average precision (AP) averaged over IOUs thresholds from 0.5 to 0.95 (AP@50:95) of 42.2. Wong et al. 2023 report 0.64 (model) vs. 0.64 (cohort) average precision (AP) for cored plaques and 0.75 vs. 0.51 AP for CAAs at a 0.5 IOU threshold (AP@50). We report an object detection AP@50:95 of 32.2 and AP@50 of 0.49 for mature NFTs programmatically bootstrapped from point annotations. In a third case, Koga et al. 2022 performed tau neuronal inclusion object detection in 2,522 CP13-stained images from 10 cases each of AD, PSP, or CBD across a range of brain regions; they achieved an average precision of 0.827. Koga et al. appear to use a more inclusive definition of tau neuronal inclusion by inspecting a representative lesion19,62. If we were to apply a similarly inclusive definition of NFT and hypothesize that 50% of our false positives were pre-NFTs or mature tangles outside of our annotation criteria, it would have a marked impact on the F1 score (from 0.58 to 0.74), but this is speculation.

We noted several caveats, which more typically arose from data than model architecture limitations. We could not assess whether the model can distinguish “None” versus “Mild” semi-quantitative categories due to the sample size (n = 1 “None” test WSIs), although the model scored the “None” slide perfectly with an NFTDetector score of 0. Performance may vary by factors such as tissue staining variations due to overfixation or differences in scoring methodologies — semi-quantitative scoring analyzes the densest 1mm2 region of pathology, in contrast to the NFTDetector, which gives a WSI-level quantification. The standalone object detection model trained using bounding boxes derived from the ground truth segmentation masks likewise correlated with semi-quantitative grades to the same extent as the segmentation model. This suggested that both model strategies – segmentation and object detection – learned approximately equally well and possibly to their limit from the training data. The point-to-mask pipeline may also need to be tuned slightly for each AT8-stained dataset, although we would expect this to be uncommon, as the code benefits from an autothresholding strategy.

We focused on samples from individuals with pathologically defined AD, specifically from the temporal lobe, which may limit generalizability to other brain regions and disease presentations. To aid in development of more generalizable ML models, additional studies could include additional anatomic regions, neurodegenerative diagnoses (such as other tauopathies), and staining protocols. All slides were processed at UC Davis, minimizing variability in the preparation process such as scanning and staining. While this standardization increases data quality, it may in turn limit the resulting models’ generalizability. For comparison, the recent ADNP-15 dataset18 release included 15 slides collected from 4 different biobanks and scanned with two scanners. However, the staining protocol was the same across all slides and it was unclear whether they were done at the same site. Further, the ADNP-15 study did not disaggregate performance across biobanks, as they lacked enough slides to power such an analysis. This highlights a crucial need to assess generalizability across a range of preanalytical variables, which we encourage through the full release of our dataset and pipeline11.

Critically, our pixel-level predictions achieve a strong WSI-level correlation with expert semi-quantitative scoring (Spearman’s ρ = 0.654, p < 0.0001), supporting clinical utility. Likewise, when we convert segmentation masks to object-level bounding boxes for comparison with prior work, we achieve comparable F1 while still maintaining pixel-level morphological information. Notably, we are among the first to publicly release an expert-annotated dataset, trained model weights, and codebase for NFT detection and segmentation to support reproducibility and broad real-world use (https://github.com/keiserlab/tangle-tracer). To our knowledge, only Vizcarra et al. have released their dataset, codebase, and model weights, although they performed NFT object detection alone10. We hope that by releasing our annotated batch 1 and batch 2 WSIs (n = 46 WSIs: 22 annotated WSIs, 24 validation WSIs, all slide-level scored), this will help enable the field as a whole, comprising some three-fold the number of WSIs in the recent ADNP-1518 dataset. These models can be used as-is, or the reader may wish to fine-tune or retrain them for institution- or goal-specific tasks. Instead of working directly with Python code, researchers and pathologists can use the platform without programming via its command line interface to convert new point annotations to segmentation ground truth masks (see the README.md in the code repository for details). We hope to integrate this pipeline and model with other pathology tools, such as the HistomicsUI63 and active learning frameworks. Future studies may benefit from performing deeper morphological and correlative analyses between NFT features and clinical or genetic features using methods such as spatial transcriptomics.

Conclusion

We introduce a robust and scalable deep learning approach to generate and detect ground truth masks for mature neurofibrillary tangles (NFTs). The segmentation model precisely detects and delineates mature tangles, enabling deeper morphological neuropathological analyses that may ultimately relate to subtle clinical and genetic factors. The model generates assessments comparable to expert semi-quantitative scoring when aggregated across an entire slide into a single score. Where speed is a concern, object detection models trained on bounding boxes calculated from the same data achieve similar but more rapid detection performance while giving up pixel-accurate NFT boundaries. Unlike other segmentation pipelines, the platform scales readily to new institutions because it “bootstraps” high-resolution NFT boundary training data automatically from simple and comparatively rapid human NFT single-point annotations. This tool reveals detailed maps of NFT distribution and morphology within tissue samples. We hope the open-source annotated dataset, trained model weights, and codebase we release here will facilitate diverse research directions, new uses, and collaborations in digital neuropathology.

Supplementary Information

Below is the link to the electronic supplementary material.

Supplementary Material 1 (20.4MB, pdf)
Supplementary Material 2 (165.1KB, xlsx)

Acknowledgements

The authors thank the families and participants of the University of California Davis, University of California San Diego, and Columbia University Alzheimer’s Disease Research Centers (ADRC) for their generous donations, as well as ADRC staff and faculty for their contributions. This project was made possible by grant DAF2018-191905 (https://doi.org/10.37921/550142lkcjzw) from the Chan Zuckerberg Initiative DAF, an advised fund of the Silicon Valley Community Foundation (funder https://doi.org/10.13039/100014989) (M.J.K.) and grants from the National Institute on Aging (NIA) of the National Institutes of Health (NIH) under Award Numbers R01AG062517 (B.N.D.), P30AG072972 (C.D.), P30AG062429 (C.D.), P50AG008702 (A.F.T., Neuropathology Core), and P30AG066462 (A.F.T., Neuropathology Core). The views and opinions expressed in this article are those of the authors and do not necessarily reflect the official policy or position of any public health agency or the US government. The authors would also like to thank Mikio Tada and Irene Wang for their contributions throughout the project, as well as David Gutman and JC Vizcarra for sharing ideas and data to strengthen this study.

Abbreviations

ML

Machine learning

DL

Deep learning

NFT

Neurofibrillary tangle

WSI

Whole slide image

Pre-NFT

Pre-tangle stage of NFT

AD

Alzheimer’s Disease

PMI

Post-mortem interval

Aβ

Amyloid beta

DAB

Diaminobenzidine

IHC

Immuno-histochemical

NIA-AA

National Institute on Aging - Alzheimer’s Association

CERAD

Consortium to Establish a Registry for Alzheimer’s Disease

YOLO

You Only Look Once model

ADRC

Alzheimer’s Disease Research Center

ROI

Region of interest

GPU

Graphical processing unit

IOU

Intersection over union

CLI

Command line interface

TP/FP

True positive/false positive

ρ

Spearman’s rho

Author contributions

Conceptualization: BND, MJK; Methodology: SG; Software: SG, LA; Validation: BND, LA, SG, SRRS, MJK; Formal analysis: SG, LA, LB, NS, MJK; Investigation: SG, BND, KN, LWJ, RR, AFT, CD; Resources: BND, MJK, CD, RR, AFT, LWJ; Data Curation: SG, BND, LM, KN; Writing - Original Draft: SG; Writing - Review & Editing: SG, BND, MJK; Visualization: SG, LA; Supervision: BND, MJK; Project administration: BND, MJK; Funding acquisition: BND, MJK. All authors reviewed the manuscript.

Data availability

The images (rotated and cropped regions of interest and their segmentation masks) used to train and evaluate all machine learning models and data created and used in this project, including results, are available at https://www.ebi.ac.uk/biostudies/bioimages/studies/S-BIAD1165. The raw WSIs are available for download in Zarr format (must be uncompressed with imagecodecs’ Jpegxr codec; see repo below for example). The Python scripts, Jupyter notebooks, and additional code information are available at https://github.com/keiserlab/tangle-tracer. The README.md file contains detailed information on how to execute the pipeline and recreate results.

Declarations

Competing interests

Brittany N. Dugger reports a relationship with External Advisory Board, University of Southern California Alzheimer’s Disease Research Center that includes: consulting or advisory. The other authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

Ethics and approval and consent to participate

This study only utilized human post-mortem tissues and is not considered human subject research, as only living subjects are defined as Human Subjects under federal law (45 CFR 46, Protection of Human Subjects). For participants during life, ethical approval for this study was granted by the Columbia Institutional Review Board, the University of California San Diego Institutional Review Board, and the University of California Davis Institutional Review Board and was performed in accordance with the Declaration of Helsinki and other state/federal and institutional guidelines. At death, autopsies were performed for individuals of which written informed consent for autopsy was provided by the individual and/or appropriate family members.

Footnotes

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Contributor Information

Brittany N. Dugger, Email: bndugger@health.ucdavis.edu

Michael J. Keiser, Email: keiser@keiserlab.org

References

  • 1.Moloney, C. M., Lowe, V. J. & Murray, M. E. Visualization of neurofibrillary tangle maturity in Alzheimer’s disease: A clinicopathologic perspective for biomarker research. Alzheimers Dement.17, 1554–1574 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Ma, J. et al. Segment Anything in Medical Images. arXiv [eess.IV] (2023).
  • 3.Ronneberger, O., Fischer, P. & Brox, T. U-Net: Convolutional Networks for Biomedical Image Segmentation. arXiv [cs.CV] (2015).
  • 4.Tang, Z. et al. Interpretable classification of Alzheimer’s disease pathologies with a convolutional neural network pipeline. Nat. Commun.10, 2173 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Wong, D. R. et al. Learning fast and fine-grained detection of amyloid neuropathologies from coarse-grained expert labels. Commun. Biol.6, 668 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Ramaswamy, V. G. et al. A Scalable High Throughput Fully Automated Pipeline for the Quantification of Amyloid Pathology in Alzheimer’s Disease using Deep Learning Algorithms. bioRxiv 2023.05.19.541376 (2023). 10.1101/2023.05.19.541376 [DOI]
  • 7.Wurts, A., Oakley, D. H., Hyman, B. T. & Samsi, S. Segmentation of tau stained Alzheimers brain tissue using convolutional neural networks. Annu. Int. Conf. IEEE Eng. Med. Biol. Soc.2020, 1420–1423 (2020). [DOI] [PubMed] [Google Scholar]
  • 8.Signaevsky, M. et al. Artificial intelligence in neuropathology: deep learning-based assessment of tauopathy. Lab. Invest.99, 1019–1029 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Ingrassia, L. et al. Automated segmentation by deep learning of neuritic plaques and neurofibrillary tangles in brain sections of Alzheimer’s Disease Patients. bioRxiv2023.10.31.564976 10.1101/2023.10.31.564976 (2023). [DOI]
  • 10.Vizcarra, J. C. et al. Toward a generalizable machine learning workflow for disease staging with focus on neurofibrillary tangles. Acta Neuropathol. Commun. (2023). [DOI] [PMC free article] [PubMed]
  • 11.Oliveira, L. C. et al. Preanalytic variable effects on segmentation and quantification machine learning algorithms for amyloid-β analyses on digitized human brain slides. J. Neuropathol. Exp. Neurol.82, 212–220 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Vizcarra, J. C. et al. Validation of machine learning models to detect amyloid pathologies across institutions. Acta Neuropathol. Commun.8, 59 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Kim, M. et al. Diagnosis of Alzheimer’s Disease and Tauopathies on Whole Slide Histopathology Images Using a Weakly Supervised Deep Learning Algorithm. (2023). https://www.researchsquare.com/article/rs-2459626/v2 [DOI] [PubMed]
  • 14.Lu, M. Y. et al. Data-efficient and weakly supervised computational pathology on whole-slide images. Nat. Biomed. Eng.5, 555–570 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Koohbanani, N. A., Unnikrishnan, B., Khurram, S. A., Krishnaswamy, P. & Rajpoot, N. Self-Path: Self-Supervision for Classification of Pathology Images With Limited Annotations. IEEE Trans. Med. Imaging. 40, 2845–2856 (2021). [DOI] [PubMed] [Google Scholar]
  • 16.Ingrassia, L. et al. Automated deep learning segmentation of neuritic plaques and neurofibrillary tangles in Alzheimer disease brain sections using a proprietary software. J. Neuropathol. Exp. Neurol.83, 752–762 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Jimenez, G. et al. Visual deep learning-based explanation for neuritic plaques segmentation in Alzheimer’s Disease using weakly annotated whole slide histopathological images. arXiv [eess.IV] (2023).
  • 18.Zhao, C. et al. ADNP-15: An open-source histopathological dataset for neuritic plaque segmentation in human brain whole slide images with frequency domain image enhancement for stain normalization. IRBM46, 100913 (2025). [Google Scholar]
  • 19.Koga, S., Ikeda, A. & Dickson, D. W. Deep learning-based model for diagnosing Alzheimer’s disease and tauopathies. Neuropathol. Appl. Neurobiol.48, e12759 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Pansuwan, T. et al. Accurate digital quantification of tau pathology in progressive supranuclear palsy. Acta Neuropathol. Commun.11, 178 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Wong, D. R. et al. Deep learning from multiple experts improves identification of amyloid neuropathologies. Acta Neuropathol. Commun.10, 66 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Mirra, S. S. et al. The Consortium to Establish a Registry for Alzheimer’s Disease (CERAD). Part II. Standardization of the neuropathologic assessment of Alzheimer’s disease. Neurology41, 479–486 (1991). [DOI] [PubMed] [Google Scholar]
  • 23.Scalco, R. et al. The neuropathological landscape of Hispanic and non-Hispanic White decedents with Alzheimer disease. Acta Neuropathol. Commun.11, 105 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Besser, L. M. et al. The Revised National Alzheimer’s Coordinating Center’s Neuropathology Form-Available Data and New Analyses. J. Neuropathol. Exp. Neurol.77, 717–726 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Montine, T. J. et al. National Institute on Aging-Alzheimer’s Association guidelines for the neuropathologic assessment of Alzheimer’s disease: a practical approach. Acta Neuropathol.123, 1–11 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Consensus recommendations for the postmortem diagnosis of Alzheimer’s disease. The National Institute on Aging, and Reagan Institute Working Group on Diagnostic Criteria for the Neuropathological Assessment of Alzheimer’s Disease. Neurobiol. Aging. 18, S1–2 (1997). [PubMed] [Google Scholar]
  • 27.Miles, A. et al. Zarr-Developers/zarr-Python: v2.16.1Zenodo (2023). 10.5281/ZENODO.3773449 [DOI]
  • 28.Harris, C. R. et al. Array programming with NumPy. Nature585, 357–362 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.van der Walt, S. et al. scikit-image: image processing in Python. PeerJ2, e453 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Ruifrok, A. C. & Johnston, D. A. Quantification of histochemical staining by color deconvolution. Anal. Quant. Cytol. Histol.23, 291–299 (2001). [PubMed] [Google Scholar]
  • 31.Otsu, N. A. Threshold Selection Method from Gray-Level Histograms. IEEE Trans. Syst. Man. Cybern. 9, 62–66 (1979). [Google Scholar]
  • 32.Van Rossum, G. The python library reference, release 3.8. 2. Python Software Foundation.
  • 33.Paszke, A. et al. PyTorch: An Imperative Style, High-Performance Deep Learning Library. arXiv [cs.LG] (2019).
  • 34.Riba, E., Mishkin, D., Ponsa, D., Rublee, E. & Bradski, G. Kornia: an Open Source Differentiable Computer Vision Library for PyTorch. in IEEE Winter Conference on Applications of Computer Vision (WACV) 3674–3683 (IEEE, 2020). 10.1109/wacv45572.2020.9093363 [DOI]
  • 35.Iakubovskii, P. Segmentation Models Pytorch. GitHub Repository (2019). https://github.com/qubvel/segmentation_models.pytorch
  • 36.Deng, J. et al. ImageNet: A large-scale hierarchical image database. in IEEE Conference on Computer Vision and Pattern Recognition 248–255 (IEEE, 2009). 10.1109/CVPR.2009.5206848 [DOI]
  • 37.He, K., Zhang, X., Ren, S. & Sun, J. Deep Residual Learning for Image Recognition. arXiv [cs.CV] (2015).
  • 38.Salehi, S. S. M., Erdogmus, D. & Gholipour, A. Tversky loss function for image segmentation using 3D fully convolutional deep networks. arXiv [cs.CV] (2017).
  • 39.Biewald, L. Experiment Tracking with Weights and Biases (2020). https://www.wandb.com/
  • 40.Falcon, W. et al. PyTorchLightning/pytorch-Lightning: 0.7.6 ReleaseZenodo (2020). 10.5281/ZENODO.3828935 [DOI]
  • 41.Falkner, S., Klein, A. & Hutter, F. BOHB: Robust and Efficient Hyperparameter Optimization at Scale (2018). 10.48550/arXiv.1807.01774 [DOI]
  • 42.Hunter, J. D. & Matplotlib A 2D Graphics Environment. Comput Sci. Eng9, 90–95 (2007).
  • 43.Annotation Automation Platform for Computer Vision. SuperAnnotatehttps://www.superannotate.com/
  • 44.Charlier, F. et al. Trevismd/statannotations: v0.5Zenodo (2022). 10.5281/ZENODO.7213391 [DOI]
  • 45.Waskom, M., Botvinnik, O. & Hobson, P. Seaborn V0. 6.0 (2015).
  • 46.Perktold, J. et al. Statsmodels/statsmodels: Release 0.14.2Zenodo (2024). 10.5281/ZENODO.10984387 [DOI]
  • 47.Padilla, R., Netto, S. L. & da Silva, E. A. B. A Survey on Performance Metrics for Object-Detection Algorithms. in International Conference on Systems, Signals and Image Processing (IWSSIP) 237–242 (IEEE, 2020). 10.1109/IWSSIP48289.2020.9145130 [DOI]
  • 48.Gaël, V. & Édouard, D. Scikit-learn: Machine Learning in Python. J. Mach. Learn. Res. 10.5555/1953048.2078195 (2011). [DOI]
  • 49.Jocher, G., Chaurasia, A. & Qiu, J. YOLO by Ultralytics. (2023).
  • 50.Jocher, G. et al. ultralytics/yolov5: v3.0. Zenodo (2020). 10.5281/zenodo.3983579 [DOI]
  • 51.Maňoušková, K. et al. SPIE,. Tau protein discrete aggregates in Alzheimer’s disease: neuritic plaques and tangles detection and segmentation using computational histopathology. in Medical Imaging 2022: Digital and Computational Pathology vol. 12039 33–39 (2022).
  • 52.Zhang, L. et al. Quantitative Assessment of Hippocampal Tau Pathology in AD and PART. J. Mol. Neurosci.70, 1808–1811 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Marx, G. A. et al. Artificial intelligence-derived neurofibrillary tangle burden is associated with antemortem cognitive impairment. Acta Neuropathol. Commun.10, 157 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Janowczyk, A. & Madabhushi, A. Deep learning for digital pathology image analysis: A comprehensive tutorial with selected use cases. J. Pathol. Inf.7, 29 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Humphrey, P. A. Gleason grading and prognostic factors in carcinoma of the prostate. Mod. Pathol.17, 292–306 (2004). [DOI] [PubMed] [Google Scholar]
  • 56.Beck, A. H. et al. Systematic analysis of breast cancer morphology uncovers stromal features associated with survival. Sci. Transl Med.3, 108–113 (2011). [DOI] [PubMed] [Google Scholar]
  • 57.Limorenko, G. & Lashuel, H. A. Revisiting the grammar of Tau aggregation and pathology formation: how new insights from brain pathology are shaping how we study and target Tauopathies. Chem. Soc. Rev.51, 513–565 (2022). [DOI] [PubMed] [Google Scholar]
  • 58.Koga, S., Ghayal, N. B. & Dickson, D. W. Deep Learning-Based Image Classification in Differentiating Tufted Astrocytes, Astrocytic Plaques, and Neuritic Plaques. J. Neuropathol. Exp. Neurol.80, 306–312 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59.Oktay, O. et al. Attention U-Net: Learning where to look for the pancreas. arXiv [cs.CV] (2018).
  • 60.Hu, C. et al. Swin-Unet: Unet-like Pure Transformer for Medical Image Segmentation. arXiv [eess.IV] (2021).
  • 61.Geirhos, R. et al. Shortcut learning in deep neural networks. Nat. Mach. Intell.2, 665–673 (2020). [Google Scholar]
  • 62.DeTure, M. A. & Dickson, D. W. The neuropathological diagnosis of Alzheimer’s disease. Mol. Neurodegener. 14, 32 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 63.Gutman, D. A. et al. The Digital Slide Archive: A Software Platform for Management, Integration, and Analysis of Histology for Cancer Research. Cancer Res.77, e75–e78 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

The images (rotated and cropped regions of interest and their segmentation masks) used to train and evaluate all machine learning models and data created and used in this project, including results, are available at https://www.ebi.ac.uk/biostudies/bioimages/studies/S-BIAD1165. The raw WSIs are available for download in Zarr format (must be uncompressed with imagecodecs’ Jpegxr codec; see repo below for example). The Python scripts, Jupyter notebooks, and additional code information are available at https://github.com/keiserlab/tangle-tracer. The README.md file contains detailed information on how to execute the pipeline and recreate results.


Articles from Scientific Reports are provided here courtesy of Nature Publishing Group

RESOURCES