Skip to main content
Wiley Open Access Collection logoLink to Wiley Open Access Collection
. 2026 Sep 23;134(10):e70149. doi: 10.1002/cncy.70149

Artificial intelligence–assisted digital thyroid FNA cytology: Improved agreement and sensitivity for higher‐risk Bethesda categories with enhanced screening efficiency

Swati Satturwar 1, Zaibo Li 1, Chi‐Shun Yang 2, Yi‐Jyun Lin 3, Wei‐Lei Yang 4, Ming‐Yu Lin 4, Cheng‐Hung Yeh 4, Shih‐Wen Hsu 4, Yi‐Siou Liu 4, Guowei Shao 4, Tien‐Jen Liu 4, Chih‐Jung Chen 2,5,6,✉, Barbara A Crothers 4,✉
PMCID: PMC13601818  PMID: 42779199

Short abstract

Artificial intelligence–assisted digital review of thyroid FNA cytology improved agreement with consensus diagnoses in higher‐risk Bethesda categories and increased sensitivity for detecting The Bethesda System III+ cases compared with conventional microscopy. Both single‐ and seven‐layer Z‐stack whole‐slide imaging substantially reduced review time, supporting artificial intelligence–assisted digital cytology as a feasible adjunct for improving diagnostic consistency and workflow efficiency.

Keywords: artificial intelligence, deep learning, digital cytology, The Bethesda System for Reporting Thyroid Cytopathology, thyroid fine‐needle aspiration cytology, whole‐slide image, Z‐stack scanning

Abstract

Background

Accurate cytologic classification of thyroid nodules is essential for clinical management, but interobserver variability and indeterminate interpretations remain persistent challenges. The clinical feasibility of AIxTHY, a disease‐specific deep‐learning algorithm integrated into a digital cytology platform, was evaluated for assisting thyroid fine‐needle aspiration (FNA) diagnosis using whole‐slide imaging (WSI).

Methods

Two cytopathologists and one cytologist independently reviewed 100 archival ThinPrep FNA slides using three modalities: microscopy, artificial intelligence (AI)–assisted single‐layer WSI (S‐WSI), and AI‐assisted seven‐layer Z‐stack WSI (7‐WSI), with 2‐week washout intervals. Reviewers assigned The Bethesda System (TBS) categories, and review times were recorded. Diagnoses were compared with expert cytologic consensus as ground truth.

Results

AI‐assisted review modalities improved overall agreement with consensus diagnoses, particularly for higher‐risk categories (TBS III+: atypia of undetermined significance, follicular neoplasm, and malignant), and reduced downgrading compared to microscopy. For binary risk stratification (TBS III+ vs TBS II), AI assistance increased sensitivity from 60% to 78% to 79% and accuracy from 61% to approximately 71%, whereas specificity decreased because of increased false‐positive classifications, particularly among indeterminate cases. AI‐assisted review significantly reduced slide review time by 40% to 65% across diagnostic categories (p < 0.01), with the greatest efficiency gains in positive cases. Compared with S‐WSI, 7‐WSI did not further improve diagnostic concordance or binary classification performance but provided additional reductions in review time.

Conclusions

AIxTHY is a feasible adjunct for improving detection and workflow efficiency in digital thyroid cytopathology, although further refinement is needed to improve specificity in indeterminate lesions.

INTRODUCTION

Thyroid lesions are predominately diagnosed through fine‐needle aspiration (FNA) cytology combined with ultrasound imaging. Inaccurate classification of these lesions can result in delayed diagnoses or unnecessary surgical interventions. The Bethesda System for Reporting Thyroid Cytopathology (TBSRTC) 1 provides standardized terminology to enhance diagnostic consistency; however, challenges persist in the interpretation of “atypia of undetermined significance” (AUS; TBS III), “follicular neoplasm” (FN; TBS IV), and “suspicious for malignancy” (SFM; TBS V) categories. 2 The diagnostic rates of these categories are highly variable 3 yet carry significant risks of malignancy, 4 , 5 , 6 underscoring the need to reduce indeterminant diagnoses. For AUS and follicular lesion of undetermined significance (FLUS), the risks of malignancy in the literature varies from 6% to 72.9%. 5 , 7 , 8 , 9 Inter‐observer variability among pathologists 10 , 11 further complicates clinical decision‐making, often necessitating molecular testing to better assess malignancy risk.

Recent advances in artificial intelligence (AI) have accelerated the development of specialized computational models capable of assisting thyroid cytology interpretation by enabling objective, standardized, and reproducible assessment of diagnostic features. 12 , 13 Despite these advances, the integration of AI‐assisted thyroid cytology into routine clinical workflows remains limited. In particular, quantitative analysis of cytomorphological features at the whole‐slide imaging (WSI) level has not been systematically implemented or comprehensively validated in real‐world practice.

Furthermore, the acquisition of high‐quality, well‐focused WSI in cytology presents unique technical challenges from the inherently three‐dimensional architecture of cytologic specimens. Unlike histologic sections, which are typically 4‐ to 6‐μm thick, liquid‐based cytology preparations may reach thicknesses of up to 30 μm. 14 This increased thickness complicates the identification of optimal focal planes and often necessitates multilayer (Z‐stack) scanning to adequately capture diagnostically relevant cellular details. Additionally, variability among cytologic preparation methods, including conventional smears, cytospin preparations, and liquid‐based cytology, contributes to heterogeneity in slide thickness and cellular distribution, further increasing technical complexity and variability during digital slide acquisition. 15

To address those challenges, we developed AIxTHY (AIxMed, Inc. Santa Clara, CA), a thyroid‐specific AI algorithm integrated into a digital cytopathology platform. In our initial model, AIxTHY applies the diagnostic framework of The Bethesda System (TBS) to perform quantitative detection and characterization of cytomorphologic features associated with malignancy at the WSI level. By systematically identifying and highlighting potentially malignant cells, the platform is designed to standardize interpretation, reduce interobserver variability, and improve both diagnostic accuracy and workflow efficiency in thyroid FNA cytology. In parallel, we implemented a multilayer‐capable WSI scanner that generates both single‐layer images at the optimal focal plane and Z‐stack images from the same cytology slide. These image sets were evaluated within the AI‐assisted workflow to determine whether enhanced focus depth improves cytomorphologic assessment. We hypothesized that AIxTHY‐assisted evaluation would (1) improve the positive predictive value (PPV) of higher‐risk TBSRTC categories, (2) enhance cytomorphologic assessment through Z‐stack imaging by providing better visualization of diagnostically relevant cellular features, and (3) increase overall reviewer screening efficiency without compromising diagnostic performance.

To evaluate the clinical applicability of this platform, we designed a three‐arm comparative study assessing its performance in routine thyroid FNA interpretation. Diagnostic accuracy and workflow efficiency were analyzed across three modalities: microscopy, AI‐assisted single‐layer WSI, and AI‐assisted seven‐layer WSI review. By directly comparing these approaches, we sought to determine whether AI‐supported digital cytology could achieve diagnostic performance comparable to conventional microscopy while enhancing interpretative efficiency and reducing review time.

MATERIALS AND METHODS

For the AI algorithm development, a total of 47 archived and subsequently deidentified thyroid FNA cytology slides were selected to establish the training and cross‐validation (CV) datasets. These comprised two conventional smears, 41 ThinPrep non‐GYN (ThinPrep; Hologic, Marlborough, MA), and 4 BD CytoRich non‐GYN (SurePath; Becton, Dickinson and Company, Franklin Lakes, NJ) preparations. All slides were reviewed by a cytologist to confirm adequate staining and to identify any artifacts such as cracks, bubbles, or debris that might fail WSI. Slides were scanned using the Aperio AT2 scanner (Leica Biosystems, Wetzlar, Germany) at 40× magnification. Before scanning, a cytologist manually selected regions of interest and placed focal points on representative cellular areas to optimize image sharpness and coverage, particularly for liquid‐based preparations such as SurePath slides, where three‐dimensional cellular aggregates require careful focus optimization. These focal points guided the scanner autofocus function to generate a single focal plane in accordance with established protocols. 16 WSIs were generated in proprietary SVS (ScanScope Virtual Slide; Aperio) format. Following digitization, each WSI underwent quality control review by an experienced cytologist to confirm adequate focus, color balance, and overall image fidelity before inclusion in AI model development and evaluation.

Of 47 WSIs used for algorithm development, 37 were assigned to the training dataset (including two smears, 33 ThinPreps and two SurePaths) and 10 to the CV dataset (including eight ThinPreps and two SurePaths), with patient‐level independence maintained in accordance with Food and Drug Administration guidance. For annotation, a senior cytologist selected 2048 × 2048‐pixel regions from each WSI and labeled cellular and background elements, including follicular cells, Hürthle cells, histiocytes, lymphocytes, and colloid, according to TBSRTC criteria (Figure 1). A second cytologist independently reviewed the annotations. Regions with discrepant annotation or cell‐type classification were jointly reexamined by both cytologists using the digital viewer (Figure 2), and a consensus annotation was established before inclusion in the training or CV dataset. This yielded 2425 annotated regions, including 1937 for training and 488 for CV.

FIGURE 1.

FIGURE 1

Example images of AIxTHY identifying target cells by TBS category and cell type. (A) Benign follicular‐cell cluster (FC). (B) Microfollicular proliferation typical of follicular neoplasm (FN). (C) Papillary thyroid carcinoma (PTC). (D) Hürthle‐cell (oncocytic) group. (E) Histiocyte containing granular cytoplasm. (F) Lymphocytes aggregate. (G) Thick colloid fragment. AIxTHY indicates a disease‐specific deep‐learning algorithm integrated into a digital cytology platform; TBS, The Bethesda System.

FIGURE 2.

FIGURE 2

AIxTHY graphical user interface for AI‐assisted thyroid cytology review. (A) Dashboard cell gallery view. Candidate “hot‐spot” cellular clusters detected by the AI algorithm are displayed as clickable tiles (thumbnails) and organized by predicted cytomorphologic categories. Selection of a tile navigates the reviewer to the corresponding whole‐slide image region, where AI‐annotated features (e.g., nuclear enlargement, membrane irregularity, chromatin characteristics) are highlighted for detailed evaluation. (B) Reporting interface. After reviewing the AI‐annotated findings, the user records the final diagnostic interpretation (positive, negative, or nondiagnostic) and assigns the definitive TBS category prior to case submission. AI indicates artificial intelligence; AIxTHY, a disease‐specific deep‐learning algorithm integrated into a digital cytology platform; TBS, The Bethesda System.

The AIxTHY deep‐learning model includes segmentation (isolates cells), morphology (extracts features), and classification (assigns cell types of TBSRTC) submodels, and its outputs are displayed in an interactive digital cytology viewer for expert review (Figure 2). The viewer presents AI‐identified candidate cells and cell groups as clickable tiles linked to their corresponding WSI locations, allowing users to rapidly navigate, tag, and annotate diagnostically relevant findings before assigning the final Bethesda category. This user‐friendly interface is designed to facilitate adoption by supporting a screening workflow similar to AI‐assisted gynecologic cytology, in which algorithmic prescreening prioritizes abnormal cells while maintaining cytopathologist oversight. Active learning refined the model iteratively, with cytologists annotating suboptimal predictions. The final model achieved approximately 67% overall sensitivity, 90% PPV, and 85% category‐level accuracy on the CV dataset. The predefined performance thresholds, at least 85% recall and 60% precision, were met for the target category, and the results were reviewed and confirmed by a senior cytopathologist.

To construct the independent test dataset, 100 archived and deidentified thyroid FNA cases obtained at Taichung Veterans General Hospital between 2020 and 2022 were retrospectively retrieved from the institutional pathology database. Cases were stratified by TBSRTC category to ensure representation across the diagnostic spectrum, followed by random sampling within each category. Eligible cases required histologic follow‐up within 3 to 6 months and an available diagnostic‐quality Pap‐stained ThinPrep slide. Although histologic follow‐up was available for all included cases, the prespecified primary reference standard for this observer‐performance study was the consensus Bethesda diagnosis established independently by two board‐certified senior cytopathologists according to TBSRTC criteria. Histologic findings were used for a separate secondary concordance analysis and did not define the primary ground truth. In routine clinical practice, the original cytologic diagnosis may have incorporated additional cytologic preparations, cell‐block material, ancillary molecular studies, or other clinical information. For standardized AI‐assisted evaluation, however, only one representative ThinPrep slide per case was selected before AI review based on diagnostic adequacy, preservation quality, and representation of the cytomorphologic features supporting the consensus diagnosis. No paired conventional smears, additional liquid‐based slides, or cell‐block preparations were included in the AI analysis.

The study protocol was approved by the Institutional Review Board of Taichung Veterans General Hospital (No. CE23518B‐1). The final consensus diagnoses, serving as the ground truth, were established by two board‐certified senior cytopathologists according to TBSRTC criteria, 1 including five nondiagnostic (NonDx; TBS I), 35 benign (TBS II), 15 AUS (TBS III), 15 FN (TBS IV), no SFM (TBS V), and 30 malignant (TBS VI) cases. All selected slides were scanned using the 3DHISTECH E1000 digital scanner (3DHISTECH Kft., Budapest, Hungary) at 40× objective magnification in two acquisition modes: single‐layer imaging at the optimal autofocus plane (S‐WSI) and seven‐layer Z‐stack imaging at 1‐μm intervals (7‐WSI). All slides were successfully digitized and passed post‐scanning quality control review before algorithmic inference was performed using the AIxTHY platform (Figure 2).

To evaluate the diagnostic performance and clinical utility of the AIxTHY platform, we conducted a three‐arm comparative study involving microscopy (arm 1), AIxTHY‐assisted S‐WSI review (Arm 2), and AIxTHY‐assisted 7‐WSI review (arm 3) (Figure 3). Three independent reviewers, including two board‐certified senior cytopathologists and one experienced cytologist, assessed 100 slides using each modality. In arm 1, reviewers examined the original glass slides by microscope and assigned TBSRTC categories. Subsequently, the same reviewers completed an AIxMed instructor‐led review of thyroid cases spanning all diagnostic categories using the AIxMed Cytology Viewer on the AIxTHY platform to familiarize them with the system. After a 2‐week washout period to minimize recall bias, reviewers performed arm 2 and recorded their TBSRTC diagnoses. Following an additional 2‐week washout interval, arm 3 was conducted. Each arm generated 300 diagnostic assessments, representing 100 cases independently reviewed by three readers, and these assessments were compared with the expert consensus reference standard. Diagnostic performance was evaluated using exact agreement across TBSRTC categories and binary classification for malignancy risk, with TBS II considered benign and TBS III or higher considered positive for malignancy. Workflow efficiency was assessed by recording the time required for each case review, measured from the start of slide evaluation to final diagnostic assignment using a stopwatch, with time recorded in seconds.

FIGURE 3.

FIGURE 3

Study design for evaluating AI‐assisted digital cytopathology in thyroid fine‐needle aspiration (FNA) diagnosis. A set of 100 thyroid FNA cytology (FNAC) slides with consensus reports was digitized on a scanner. Two cytopathologists and one cytologist independently reviewed each case by three sequential modalities separated by 2‐week washout intervals: microscopy examination (arm 1), AI‐assisted S‐WSI review (arm 2), and AI‐assisted 7‐WSI review (arm 3). For every modality, reviewers recorded TBS category and the time required for slide interpretation. Study endpoints included (1) concordance of TBS category with the ground‐truth diagnosis, (2) binary agreement with the expert cytologic consensus reference classification including sensitivity, specificity, positive‐ and negative‐predictive values (PPV/NPV), and accuracy (ACC), and (3) slide‐reporting time. *Seven‐layer‐stacking whole‐slide image.

Diagnostic performance for each reviewer and aggregated results was evaluated using sensitivity, specificity, PPV, negative predictive value (NPV), and accuracy (agreement with consensus diagnoses), with corresponding case numbers and 95% CIs. For the secondary cytology–histology concordance analysis, histologic follow‐up diagnoses were dichotomized as benign or malignant according to prespecified criteria. Definitively benign diagnoses were classified as benign, whereas unequivocally malignant diagnoses were classified as malignant. Borderline or indeterminate diagnoses were assigned according to the final integrated surgical pathology interpretation used for the analysis. Cytologic reviewer interpretations were then compared with the corresponding binary histologic classification across the three review modalities. Screening and evaluation times were summarized using mean, median, and standard deviation. The inter‐observer concordance among the three readers was evaluated using Fleiss' kappa based on binary outcomes. Comparisons of screen time to reporting across diagnostic modalities were performed using the Wilcoxon signed‐rank test. All statistical analyses were two‐sided, with a significance level set at p < 0.05. Data processing and visualization were conducted using R software (version 4.5.0) with the following packages: dplyr, stringr, caret, ggplot2, and ggpubr.

RESULTS

Table 1 summarizes the concordance between reviewers’ diagnoses and cytology consensus across the three diagnostic modalities (N = 300 assessments, respectively). Overall agreement was higher in both AI‐assisted arms compared with microscopy (arm 2: 51.7%; arm 3: 51.0%; arm 1: 44.3%). Category‐specific analysis revealed distinct performance patterns. Microscopy demonstrated the highest agreement in the NonDx and Benign categories (93.3% and 63.8%, respectively), compared with AI‐assisted S‐WSI (73.3% and 58.1%) and 7‐WSI (66.7% and 57.1%). In contrast, AI‐assisted review showed higher concordance in the indeterminate and higher‐risk categories. For AUS, agreement increased from 28.9% with microscopy to 51.1% with both AI‐assisted arms. Similarly, agreement for FN improved from 17.8% with microscopy to 28.9% and 35.6% with AI‐assisted S‐WSI and 7‐WSI. In the Malignant category, concordance increased from 34.4% with microscopy to 52.5% and 48.9% with AI‐assisted S‐WSI and 7‐WSI. Direct comparison between single‐layer and Z‐stack, AI‐assisted review demonstrated modest differences. 7‐WSI showed slightly lower agreement in the NonDx (66.7% vs. 73.3%) and Malignant (48.9% vs. 52.5%) categories but higher agreement in the FN category (35.6% vs. 28.9%) compared to S‐WSI. Notably, AI assistance reduced the proportion of reads classified as NonDx. Three reads in the S‐WSI arm and four reads in the 7‐WSI arm that were designated as NonDx by microscopy were reassigned to definitive TBS categories, contributing to the lower observed NonDx concordance in the AI‐assisted modalities.

TABLE 1.

Agreement rates between reviewers and cytology consensus diagnoses across the Bethesda system diagnostic categories and study modalities.

Agreement to Cytology Consensus Reports (N = 300) Modality
Microscopy

AI‐assisted

S‐WSI Review

AI‐assisted

7‐WSI Review

NonDx (n = 15) 14 (93.3%) 11 (73.3%) 10 (66.7%)
Benign (n = 105) 67 (63.8%) 61 (58.1%) 60 (57.1%)
AUS (n = 45) 13 (28.9%) 23 (51.1%) 23 (51.1%)
FN (n = 45) 8 (17.8%) 13 (28.9%) 16 (35.6%)
Malig (n = 90) 31 (34.4%) 47 (52.5%) 44 (48.9%)
Total 133 (44.3%) 155 (51.7%) 153 (51.0%)

Note: N = total number of reads from three reviewers.

Abbreviations: AUS, atypia of undetermined significance; FN, follicular neoplasm; Malig, malignant; NonDx, nondiagnostic.

Tables 2, 3, 4 present the confusion matrices detailing reviewers’ concordance with consensus diagnoses for each modality. Distinct reclassification patterns were observed across the three approaches. Under microscopy (Table 2), there was a tendency to downgrade a diagnosis. A substantial proportion of consensus AUS (62.2%), FN (33.3%), and Malignant (22.2%) reads were classified as Benign by reviewers. Significantly, 20.0% of consensus Malignant reads were categorized as AUS. These findings indicate a shift toward lower TBS categories relative to the ground truth. In contrast, AI‐assisted S‐WSI review (Table 3) reduced this downgrading pattern. Concordance improved in higher‐risk categories, with 51.1% agreement for AUS, 28.9% for FN, and 52.5% for Malignant reads. Reclassification analysis showed fewer malignant reads assigned to Benign (10.0%) compared with microscopy (22.2%). Moreover, several reads originally interpreted as NonDx by microscopy were reassigned to definitive categories, contributing to redistribution primarily into Benign and AUS categories. Similarly, AI‐assisted 7‐WSI review (Table 4) demonstrated sustained improvement in higher‐risk category concordance (AUS 51.1%, FN 35.6%, Malignant 48.9%). Compared with S‐WSI, 7‐WSI showed a modest increase in FN agreement but slightly lower concordance in the Malignant category. Reclassification patterns were comparable to S‐WSI, with continued reduction in downgrading to Benign and increased assignment to intermediate (AUS) and neoplastic categories. Overall, the confusion matrix analysis indicates that AI assistance mitigated the tendency to downgrade observed with microscopy and facilitated greater alignment with consensus diagnoses in indeterminate and malignant TBS categories.

TABLE 2.

Diagnostic concordance between reviewers and consensus reports across the Bethesda system categories in microscopy modality.

Three Reviewers’ Diagnosis in Microscopy Agreement to Cytology Consensus Reports (N = 300) Total
NonDx Benign AUS FN Malig
NonDx 14 (93.3%) 25 (23.8%) 2 (4.4%) 6 (13.3%) 1 (1.1%) 48
Benign 0 (0.0%) 67 (63.8%) 28 (62.2%) 15 (33.3%) 20 (22.2%) 130
AUS 0 (0.0%) 9 (8.6%) 13 (28.9%) 16 (35.6%) 18 (20.0%) 56
FN 0 (0.0%) 4 (3.8%) 0 (0.0%) 8 (17.8%) 6 (6.7%) 18
SFM 1 (6.7%) 0 (0.0%) 0 (0.0%) 0 (0.0%) 14 (15.6%) 15
Malig 0 (0.0%) 0 (0.0%) 2 (4.4%) 0 (0.0%) 31 (34.4%) 33
Total 15 105 45 45 90 300

Note: N = total number of reads from three reviewers. Shaded cells denote exact agreement between reviewer diagnoses and cytologic ground truth. The sum of these concordant cases represents the raw agreement rate for each respective modality.

Abbreviations: AUS, atypia of undetermined significance; FN, follicular neoplasm; Malig, malignant; NonDx, nondiagnostic; SFM, suspicious for malignancy.

TABLE 3.

Diagnostic concordance between reviewers and consensus reports across the Bethesda system categories in AI‐assisted S‐WSI review modality.

Three Reviewers’ Diagnosis in AI‐assisted S‐WSI Agreement to Cytology Consensus Reports (N = 300) Total
NonDx Benign AUS FN Malig
NonDx 11 (73.3%) 15 (14.3%) 0 (0.0%) 4 (8.9%) 0 (0.0%) 30
Benign 4 (26.7%) 61 (58.1%) 18 (40.0%) 8 (17.8%) 9 (10.0%) 100
AUS 0 (0.0%) 22 (21.0%) 23 (51.1%) 18 (40.0%) 16 (17.8%) 79
FN 0 (0.0%) 6 (5.7%) 2 (4.4%) 13 (28.9%) 4 (4.4%) 25
SFM 0 (0.0%) 1 (1.0%) 2 (4.4%) 2 (4.4%) 14 (15.6%) 19
Malig 0 (0.0%) 0 (0.0%) 0 (0.0%) 0 (0.0%) 47 (52.5%) 47
Total 15 105 45 45 90 300

Note: N = total number of reads from three reviewers. Shaded cells denote exact agreement between reviewer diagnoses and cytologic ground truth. The sum of these concordant cases represents the raw agreement rate for each respective modality.

Abbreviations: AUS, atypia of undetermined significance; FN, follicular neoplasm; Malig, malignant; NonDx, nondiagnostic; SFM, suspicious for malignancy.

TABLE 4.

Diagnostic concordance between reviewers and consensus reports across the Bethesda system categories in AI‐assisted 7‐WSI review modality.

Three Reviewers’ Diagnosis in AI‐assisted 7‐WSI Agreement to Cytology Consensus Reports (N = 300) Total
NonDx Benign AUS FN Malig
NonDx 10 (66.7%) 13 (12.4%) 1 (2.2%) 4 (8.9%) 0 (0.0%) 28
Benign 4 (26.7%) 60 (57.1%) 14 (31.1%) 12 (26.7%) 6 (6.7%) 96
AUS 1 (6.7%) 23 (21.9%) 23 (51.1%) 12 (26.7%) 19 (21.1%) 78
FN 0 (0.0%) 5 (4.8%) 3 (6.7%) 16 (35.6%) 4 (4.4%) 28
SFM 0 (0.0%) 4 (3.8%) 4 (8.9%) 0 (0.0%) 17 (18.9%) 25
Malig 0 (0.0%) 0 (0.0%) 0 (0.0%) 1 (2.2%) 44 (48.9%) 45
Total 15 105 45 45 90 300

Note: N = total number of reads from three reviewers. Shaded cells denote exact agreement between reviewer diagnoses and cytologic ground truth. The sum of these concordant cases represents the raw agreement rate for each respective modality.

Abbreviations: AUS, atypia of undetermined significance; FN, follicular neoplasm; Malig, malignant; NonDx, nondiagnostic; SFM, suspicious for malignancy.

Table 5 summarizes the comparison of slide review time across microscopy (arm 1), AI‐assisted S‐WSI (arm 2) and 7‐WSI (arm 3) stratified by TBS category. Across all categories, both AI‐assisted modalities significantly reduced review time compared with microscopy. For NonDx cases, mean review time decreased from 99.2 seconds with microscopy to 48.7 seconds with S‐WSI (–50.9%) and 47.7 seconds with 7‐WSI (–51.9%) (arm 1 vs 2, p < 0.01; arm 1 vs 3, p < 0.001). For Benign cases, review time was reduced from 153.2 seconds to 88.4 seconds (–42.3%) with S‐WSI and further to 67.4 seconds (–56.0%) with 7‐WSI (all pairwise comparisons p < 0.0001). Similarly, substantial time reductions were observed in indeterminate and higher‐risk categories. In AUS cases, review time decreased from 170.4 seconds with microscopy to 102.6 seconds (–39.8%) with S‐WSI and 69.7 seconds (–59.1%) with 7‐WSI (p < 0.001 for all comparisons). In FN cases, mean review time declined from 185.0 seconds to 96.0 seconds (–48.1%) and 66.2 seconds (–64.2%) for S‐WSI and 7‐WSI, respectively (all p < 0.01). For Malignant cases, review time was reduced from 209.5 seconds with microscopy to 99.2 seconds (–52.6%) and 72.4 seconds (–65.4%) with S‐WSI and 7‐WSI, respectively (all p < 0.01). Direct comparison between the two AI‐assisted modalities demonstrated additional time savings with 7‐WSI in Benign, AUS, FN, and Malignant categories (all statistically significant), whereas review time for NonDx cases was comparable between S‐WSI and 7‐WSI (p > 0.99). In summary, AI assistance resulted in a 40% to 65% reduction in slide review time across TBS categories, with the greatest efficiency gains observed in higher‐risk categories.

TABLE 5.

Comparison of slide reporting time for TBS categorization across three modalities.

Ground Truth TBS Category Modality p Value (Statistical Significance)
Arm 1 Arm 2 Arm 3 Arm 1 vs 2 Arm 1 vs 3 Arm 2 vs 3
NonDx Review Time (seconds) 99.2 48.7 47.7 <0.01 <0.001 >0.99
Standard deviation 30.4 32 31.4
Benign Review Time (seconds) 153.2 88.4 67.4 <0.0001 <0.0001 <0.0001
Standard deviation 73.5 40.7 37.7
AUS Review time (seconds) 170.4 102.6 69.7 <0.0001 <0.0001 <0.001
Standard deviation 68.7 49.8 33.7
FN Review Time (seconds) 185 96 66.2 <0.0001 <0.0001 <0.01
Standard deviation 77.4 45.4 39
Malig Review Time (seconds) 209.5 99.2 72.4 <0.0001 <0.0001 <0.01
Standard deviation 88.9 63.1 51.8

Abbreviations: AUS, atypia of undetermined significance; FN, follicular neoplasm; Malig, malignant; NonDx, nondiagnostic.

For binary classification (benign vs positive for malignancy), 285 diagnostic assessments were analyzed after exclusion of NonDx (TBS I) reads (n = 15). This yielded 180 positive (TBS III–VI) and 105 benign (TBS II) reference diagnoses (Table 6). Microscopy demonstrated a sensitivity of 60.0% and specificity of 63.8%. The PPV was 89.3%, NPV was 51.5%, and overall accuracy was 61.4%. AI‐assisted S‐WSI significantly improved sensitivity to 78.3% and increased NPV to 63.5%. Accuracy also improved to 70.9%; however, specificity and PPV decreased to 58.1% and 82.9%, respectively. Similarly, AI‐assisted 7‐WSI achieved a sensitivity of 79.4% and an NPV of 65.2%, with accuracy increasing to 71.2%. Specificity (57.1%) and PPV (81.7%) were modestly lower compared with microscopy. The results suggest that both AI‐assisted modalities demonstrated substantially higher sensitivity, NPV, and overall accuracy compared with microscopy, with comparable performance between single‐layer and multilayer WSI approaches, albeit at the expense of reduced specificity and PPV.

TABLE 6.

Comparison of diagnostic performance across three modalities.

Modality Microscopy

AI‐assisted

S‐WSI review

AI‐assisted

7‐WSI review

Sensitivity 60.00% 78.30% 79.40%
Read number [108/180] [141/180] [143/180]
95% CI (52.4%–67.2%) (71.6%–84.1%) (72.8%–85.1%)
Specificity 63.80% 58.10% 57.10%
Read number [67/105] [61/105] [60/105]
95% CI (53.9%–73.0%) (48.1%–67.7%) (47.1%– 66.8%)
PPV a 89.30% 82.90% 81.70%
Read number [108/121] [141/170] [143/175]
95% CI (82.3%–94.2%) (76.4%–88.3%) (75.2%–87.1%)
NPV a 51.50% 63.50% 65.20%
Read number [67/130] [61/96] [60/92]
95% CI (42.6%–60.4%) (53.1%–73.1%) (54.6%–74.9%)
Accuracy 61.40% 70.90% 71.20%
Read number [175/285] [202/285] [203/285]
95% CI (55.5%–67.1%) (65.2%–76.1%) (65.6%–76.4%)

Abbreviations: NPV, negative predictive value; PPV, positive predictive value.

a

Exclude diagnosis of TBS I category from calculating PPV and NPV results.

Both AI‐assisted modalities significantly reduced slide reporting time compared with microscopy. In benign cases, the median review time decreased from 149.5 seconds with microscopy (arm 1) to 86.9 seconds with AI‐assisted S‐WSI (arm 2, –41.9%) and 67.6 seconds with AI‐assisted 7‐WSI (arm 3, –54.8%). In positive cases, reductions were more pronounced, with median review time declining from 228.1 seconds in arm 1 to 103.8 seconds in arm 2 (–54.5%) and 69.2 seconds in arm 3 (–69.7%). Overall, the median reporting time for all cases decreased from 174.7 seconds with microscopy to 92.9 seconds (–46.8%) and 68.1 seconds (–61.0%) with arms 2 and 3, respectively (all pairwise comparisons, p < 0.0001). Positive cases required longer review times than benign cases across all modalities; however, AI assistance, particularly with 7‐WSI, conferred the greatest absolute and relative time savings in these higher‐risk cases.

To provide additional context regarding histologic follow‐up, a secondary descriptive analysis was performed using histologic diagnoses dichotomized as benign or malignant according to prespecified criteria (Table S1). Among reads from cases categorized as histologically benign (n = 186), benign cytologic concordance was 52.7% by microscopy, 46.8% by AI‐assisted S‐WSI review, and 44.1% by AI‐assisted 7‐WSI review. Among reads from cases categorized as histologically malignant (n = 114), malignant cytologic concordance increased from 40.4% by microscopy to 55.3% with both AI‐assisted S‐WSI and 7‐WSI review. These findings support the observation that AI‐assisted review increased assignment to higher‐risk cytologic categories among histologically malignant cases, although indeterminate cytologic categories (TBS‐I, III, IV) remained a major source of imperfect cytology–histology alignment. Furthermore, interobserver agreement among the three reviewers was evaluated using Fleiss’ κ based on binary outcomes across the three review modalities. Agreement was fair for microscopy review and AI‐assisted S‐WSI review, with κ values of 0.407 and 0.393, respectively, and increased modestly to moderate agreement for AI‐assisted 7‐WSI review, with a κ value of 0.441. These findings indicate comparable interobserver reproducibility between microscopy and AI‐assisted S‐WSI review, with a modest improvement in reviewer agreement using AI‐assisted 7‐WSI review.

DISCUSSION

The adoption of digital surgical pathology has spurred interest in cytology digitization, driven by advancements in digital imaging that address challenges such as multidimensional imaging, cellular overlap, preparation variability, and large file sizes. Innovations like Z‐stacking, high‐resolution optics, optimized scanning profiles, and compression algorithms have facilitated the integration of AI into cytologic diagnosis. 17 , 18 , 19 Despite these technological advancements, AI in cytology remains nascent, with only one system, the Hologic Genius Digital Diagnostic System (Hologic, Inc., Marlborough, MA), receiving Food and Drug Administration clearance for Pap test interpretation. In nongynecologic cytology, particularly thyroid cytology, AI development has been incremental, with limited studies focusing on this area over the past 6 years. 20 , 21 , 22 , 23 , 24 , 25 , 26 A comprehensive review of machine learning models for thyroid cytology was published by Wong et al. 27 AI holds promise for improving patient outcomes by enhancing diagnostic precision in challenging thyroid cytology categories, such as AUS, FN, and SFM. Up to half of patients with these initial diagnoses will be slated for surgery to further characterize the neoplasm and a large proportion of these will be benign tumors. 28 A more accurate cytologic diagnosis could reduce unnecessary surgeries and follow‐up procedures.

Our proof‐of‐concept pilot study evaluates the feasibility of an initial AI model assistant (AIxTHY) in improving agreement with consensus diagnoses (ground truth) across TBS categories. Inter‐reviewer agreement varied, with the lowest concordance for FN (17.8%–35.6%), followed by AUS (28.9%–51.1%) (Table 1). This aligns with the recognized ambiguity of these TBS categories. 10 , 11 , 29 , 30 , 31 , 32 , 33 , 34 , 35 The cytologic diagnosis of FN depends upon the overall low‐power appearance of cellularity and microfollicle formation, coupled with the absence of colloid, which can be more difficult to assess in a gallery of selected cell groups. Modifications to AIxTHY algorithm training may enhance FN detection.

Microscopy demonstrated modest overall agreement (44.3%), underscoring substantial interobserver variability and highlighting a well‐recognized limitation of this approach in routine thyroid FNA practice. Concordance was highest for NonDx and Benign categories but markedly lower for AUS, FN, and Malignant cases, highlighting diagnostic challenges in indeterminate and higher‐risk lesions. AI assistance improved overall agreement to 51.7% (S‐WSI) and 51.0% (7‐WSI), suggesting improved agreement with the consensus reference diagnosis. The improvement was most evident in clinically significant categories: agreement increased substantially for AUS, FN, and Malignant cases, indicating a tendency toward diagnostic upgrading. In contrast, agreement for Benign and NonDx categories decreased, likely reflecting reclassification of previously under‐called or insufficient cases into more actionable categories. These findings suggest that AI‐assisted review enhances concordance in diagnostically challenging categories while promoting more definitive and higher‐risk classifications.

Notably, AI assistant methods reduced NonDx interpretations to 66.7% and 73.3% (10/15 and 11/15) compared to microscopy and consensus diagnoses (93.3%; 14/15). This improvement likely stems from the capability of AIxTHY to detect and quantify small or infrequent abnormal follicular cell groups, reclassifying hypocellular or obscured cases into benign or abnormal categories. This could reduce the need for repeat FNA procedures. In other words, the specimen may be adequate when cells are viewed in a gallery. TBSRTC requires at least six groups of 10 or more follicular cells for specimen adequacy, 1 excluding cases with mostly lymphocytes or colloid. If a case is hypocellular, or has many obscuring elements, it may be difficult for the reviewer to confirm the presence of sufficient cell groups on microscopy or WSI review. Up to five reads were reclassified into the Benign and AUS categories (Tables 2, 3, 4), potentially saving the cost of unnecessary repeat clinical visits and FNA testing. However, the small sample size (three reviewers, five cases) limits the statistical significance; either more reviewers or more cases might strengthen the significance of this finding. Further investigation into this finding is ongoing.

The confusion matrix analyses (Tables 2, 3, 4) provide insight into diagnostic directionality across modalities. Under microscopy, reviewers demonstrated a clear downgrading tendency relative to consensus diagnoses. A substantial proportion of consensus AUS, FN, and Malignant cases were categorized as Benign or AUS, indicating under‐calling of higher‐risk lesions and reflecting a conservative interpretive bias. With AI‐assisted S‐WSI and 7‐WSI review, this downgrading pattern was reduced. AI assistance reclassified more NonDx cases into Benign and reassigned more Benign cases into AUS compared with microscopy, thereby reducing reliance on indeterminate or insufficient categories. Concordance improved particularly in AUS, FN, and Malignant categories, suggesting that AI‐supported review facilitated recognition of subtle atypical or malignant cytomorphologic features. This shift represents a diagnostic upgrading trend, aligning reviewer interpretations more closely with consensus ground truth. Importantly, the reduction in nondiagnostic agreement with AI does not indicate performance deterioration but rather redistribution into clinically actionable categories. Both single‐layer and multilayer WSI demonstrated comparable reclassification patterns, indicating that the primary driver of change was AI‐assisted highlighting rather than image stacking depth alone. In sum, these findings suggest that AI assistance mitigates the downgrading bias observed with microscopy and promotes more definitive categorization, potentially reducing interobserver variability in diagnostically challenging thyroid FNA cases.

Binary classification analysis supports AIxTHY as a triage‐oriented screening tool (TBS III+ positive; TBS II benign). Compared with microscopy (sensitivity 60.0%, specificity 63.8%), AI‐assisted review markedly increased sensitivity (78.3%–79.4%), NPV, and overall accuracy (70.9%–71.2%), improving detection of higher‐risk cases. This gain occurred with reduced specificity and PPV, consistent with increased false‐positive (particularly AUS) classifications. In a screening context, prioritizing sensitivity to minimize missed malignancies is clinically advantageous. Single‐ and multilayer WSI demonstrated comparable performance, indicating that AI‐driven feature recognition, rather than stacking depth alone, accounted for the improvement. The lower specificity may partly reflect limited user familiarity with the platform, and further model and user optimization may enhance balance between sensitivity and specificity. Overall, AIxTHY functions as an effective detection‐enhancing adjunct within a binary risk stratification framework.

However, because the ground truth was expert cytologic consensus according to TBSRTC, the reported binary performance metrics primarily reflect agreement with the cytologic consensus rather than definitive histologic disease status. To provide additional context, we performed a comparative analysis among cases with histologically benign or malignant results (Table S1). This analysis showed that AI‐assisted review increased malignant cytologic concordance among histologically malignant reads from 40.4% by microscopy to 55.3% with both AI‐assisted modalities, supporting the observation that AI assistance promoted recognition of higher‐risk cytologic features. However, benign cytologic concordance among histologically benign reads was lower with AI‐assisted review than with microscopy, consistent with the reduced specificity observed in the cytologic consensus‐based analysis. In addition, indeterminate Bethesda categories, including NonDx, AUS, and FN, remained a major source of imperfect cytology–histology alignment, reflecting the well‐recognized limitation that thyroid cytology categories do not always correspond directly to binary histologic results. Interobserver agreement analysis further showed fair reproducibility for microscopy and AI‐assisted S‐WSI review, with κ values of 0.407 and 0.393, respectively, and a modest increase to moderate agreement for AI‐assisted 7‐WSI review, with a κ value of 0.441. Taken together, these findings suggest that AI‐assisted review may enhance detection of cytologically higher‐risk cases and modestly improve reviewer consistency with 7‐WSI, but larger studies with uniform histologic follow‐up are needed to determine whether these changes translate into improved prediction of true clinical outcomes.

AI assistance significantly reduced slide review time across all TBS categories and binary risk classifications (Table 5). Both AI‐assisted modalities demonstrated markedly shorter reporting times compared with microscopy, with reductions ranging from approximately 40% to 65% depending on diagnostic category (all p < 0.01). Importantly, seven‐layer AI‐assisted WSI (arm 3) consistently achieved the greatest time savings, regardless of case positivity. Across modalities, positive cases required longer review times than negative cases, reflecting the increased complexity of evaluating cytomorphologic atypia. Notably, the magnitude of time reduction with AI assistance was more pronounced in positive cases, suggesting that algorithmic highlighting of suspicious cell groups streamlines assessment in diagnostically challenging lesions. The seven‐layer stacking approach likely further enhanced efficiency by providing clearer and more optimally focused cellular detail, thereby reducing the need for manual refocusing or prolonged scrutiny of equivocal features. AIxTHY facilitates workflow acceleration by preidentifying and hierarchically organizing candidate cell clusters with suspicious nuclear features, effectively condensing the number of fields requiring detailed review.

Similar to AI‐assisted screening in gynecologic cytology, AIxTHY may improve workflow by directing reviewers to diagnostically relevant cell groups and reducing the time required for manual slide screening. The clickable tile cell gallery, WSI‐linked navigation, and ability to tag or annotate suspicious cells may enhance user friendliness and support adoption in routine thyroid FNA cytology practice, particularly in high‐volume laboratories. The absolute time savings were particularly substantial among cytologists, indicating that AI‐based prescreening may confer a disproportionate efficiency advantage to high‐volume screeners. Collectively, these findings suggest that AI assistance not only offsets the productivity limitations commonly associated with digital transition but may also enhance overall laboratory throughput while maintaining diagnostic rigor.

Although multilayer (Z‐stack) WSI demonstrated a consistent slide review time‐saving advantage, this benefit must be weighed against important operational considerations. In urine cytology, Z‐stacked ThinPrep WSI has been evaluated for high‐grade urothelial carcinoma, and multi–Z‐plane scanning has been shown to improve the capture and evaluation of diagnostically relevant suspicious cells 36 ; however, these benefits are offset by increased scanning time, larger image files, greater storage requirements, and higher computational demands. Similar operational tradeoffs have been reported in digital urine cytology studies comparing multi–Z‐plane scanning with alternative AI‐based scanning approaches. 37 In thyroid FNA cytology, scanning‐parameter studies have shown that increasing the number of focal planes does not necessarily improve diagnostic concordance once file size and storage burden are considered. 38 More recent thyroid FNA WSI and AI studies have incorporated Z‐stacked imaging, but emerging evidence also suggests that routine thyroid FNA adequacy assessment and preliminary Bethesda categorization may be feasible with limited need for Z‐stacking. 39 , 40 These observations are concordant with our findings: seven‐layer WSI shortened reviewer interpretation time, likely by providing more consistently focused cellular detail, but did not materially improve diagnostic concordance or binary classification performance compared with AI‐assisted single‐layer WSI. Therefore, single‐layer AI‐assisted WSI may represent a pragmatic balance for routine implementation, whereas multilayer Z‐stack imaging may be reserved for selected cases with thick cell groups, overlapping clusters, or equivocal cytomorphologic features.

This study has several limitations. First, the sample size was moderate (100 cases with 285 diagnostic reads after exclusion of NonDx cases), resulting in relatively wide confidence intervals, particularly for specificity because of the smaller number of negative cases. In addition, the independent test set consisted entirely of ThinPrep slides from a single institution. Although preparation heterogeneity is an important real‐world challenge in digital cytology and some conventional smears and SurePath preparations were included during algorithm development, performance on conventional smears, cytospin preparations, SurePath slides, and other preparation types were not evaluated in the independent test cohort. Larger, multi‐institutional studies using diverse cytology preparations are needed to provide more precise performance estimates and determine the generalizability of AIxTHY across routine cytopathology workflows. Second, diagnostic consensus in indeterminate categories such as AUS and FN remains inherently challenging, even under conventional microscopy, which may limit the achievable ceiling of agreement. Third, AI analysis was performed using a single representative ThinPrep slide from each case, whereas routine clinical evaluation may incorporate multiple cytologic preparations, cell‐block material, ancillary studies, and subsequent histologic findings. This standardized design minimized variability from additional diagnostic materials and allowed direct comparison of AI‐assisted and human interpretations using the same slide. However, it does not fully reflect routine real‐world diagnostic practice. Fourth, the fixed sequence of review modalities and limited reviewer training represent additional limitations. Although 2‐week washout intervals were used to minimize recall bias, reviewers had evaluated all cases by microscopy and AI‐assisted S‐WSI before performing AI‐assisted 7‐WSI review. This approach prevents “forward contamination,” whereas reviewing AI arms 2 and 3 before microscopy may bias arm 1 by providing cell selection and additional statistical information. The comparable diagnostic concordance between S‐WSI and 7‐WSI may reflect true similarity in modality performance, residual case familiarity, or both. In addition, reviewer familiarity with the digital interface and AI‐assisted workflow was limited to a brief training demonstration, which may have influenced specificity and PPV. Future studies using structured hands‐on training, randomized or counterbalanced review sequences, and larger independent reader cohorts are needed to better assess the incremental diagnostic value of multilayer Z‐stack WSI and further optimize performance, particularly for FN‐patterned lesions. Fifth, while multilayer scanning demonstrated workflow advantages, its increased digitization time, data storage demands, and computational burden may constrain scalability in resource‐limited settings. Finally, because the reference standard in this study was expert cytologic consensus rather than definitive histopathologic outcome, the binary performance measures should be interpreted as indices of agreement with the consensus classification rather than as measures of true disease status. From a strict method‐comparison perspective, positive percent agreement and negative percent agreement would correspond to the quantities conventionally reported as sensitivity and specificity, respectively, when neither method constitutes an independent clinical gold standard. We nevertheless retained the familiar terms sensitivity, specificity, PPV, and NPV to facilitate comparison with prior studies of AI–assisted nongynecologic and thyroid FNA cytology that have used expert consensus or multireader adjudication as the reference standard. 39 , 41 Accordingly, sensitivity and specificity in this study represent agreement with consensus‐positive and consensus‐negative classifications, whereas PPV and NPV represent the probabilities that a reviewer’s positive or negative interpretation, respectively, agrees with the expert cytologic consensus. These measures should therefore not be interpreted as probabilities of histologically confirmed malignancy or benign disease. Fleiss’ κ was additionally reported to characterize interobserver agreement independently of the consensus‐based performance measures.

In summary, AI‐assisted thyroid FNA review improved agreement in higher‐risk categories and increased sensitivity and overall accuracy in binary risk stratification, albeit with reduced specificity because of more false‐positive (particularly AUS) classifications. AI markedly reduced review time, with the greatest efficiency gains in positive cases and among cytologists, supporting its role as a workflow‐enhancing adjunct in digital practice. Although multilayer WSI provided additional time‐saving benefits, it entails longer scanning times and greater data and computational demands. Overall, AI demonstrates feasibility as a detection‐ and efficiency‐enhancing tool, with further optimization needed to improve specificity and performance in indeterminate/FN‐patterned lesions.

AUTHOR CONTRIBUTIONS

Swati Satturwar: Investigation; writing—review and editing; writing—original draft; validation; methodology. Zaibo Li: Investigation; writing—review and editing; writing—original draft; validation; methodology. Chi‐Shun Yang: Validation; data curation. Yi‐Jyun Lin: Investigation; methodology. Wei‐Lei Yang: Project administration; resources; writing—original draft; writing—review and editing; investigation; validation; methodology; data curation; formal analysis. Ming‐Yu Lin: Project administration; resources. Cheng‐Hung Yeh: Data curation; software; visualization. Shih‐Wen Hsu: Data curation; software; visualization. Yi‐Siou Liu: Software; data curation; visualization. Guowei Shao: writing—review and editing; formal analysis. Tien‐Jen Liu: writing—review and editing; funding acquisition; conceptualization; methodology; validation; supervision. Chih‐Jung Chen: writing—review and editing; supervision; funding acquisition; conceptualization; methodology; validation; resources; data curation. Barbara A. Crothers: Writing—review and editing; writing—original draft; supervision; conceptualization; methodology; validation.

CONFLICT OF INTEREST STATEMENT

Swati Satturwar, Zaibo Li, and Yi‐Jyun Lin are medical consultants to AIxMed, Inc. Wei‐Lei Yang, Ming‐Yu Lin, Cheng‐Hung Yeh, Shih‐Wen Hsu, Yi‐Siou Liu, Guowei Shao, Tien‐Jen Liu, and Barbara A Crothers are employees of AIxMed, Inc. The remaining authors declare no competing interests.

Supporting information

Table S1

CNCY-134-0-s001.docx (35.2KB, docx)

ACKNOWLEDGMENTS

This work was supported by a grant from the National Science and Technology Council (NSTC), Taiwan (Grant number 113‐2221‐E‐075A‐006), Ministry of Health and Welfare, Taiwan (Grant number HTS‐114‐115‐A1‐0013), and AIxMed, Inc., United States.

Contributor Information

Chih‐Jung Chen, Email: cjchen1016@gmail.com.

Barbara A. Crothers, Email: barbara.crothers@aixmed.com.

REFERENCES

  • 1. Ali SZ, VanderLaan PA. SpringerLink. The Bethesda System for Reporting Thyroid Cytopathology : Definitions, Criteria, and Explanatory Notes. 3rd 2023. Springer International Publishing; 2023. [Google Scholar]
  • 2. Bongiovanni M, Spitale A, Faquin WC, Mazzucchelli L, Baloch ZW. The Bethesda System for Reporting Thyroid Cytopathology: a meta‐analysis. Acta Cytol. 2012;56(4):333‐339. doi: 10.1159/000339959 [DOI] [PubMed] [Google Scholar]
  • 3. Yaprak Bayrak B, Eruyar AT. Malignancy rates for Bethesda III and IV thyroid nodules: a retrospective study of the correlation between fine‐needle aspiration cytology and histopathology. BMC Endocr Disord. 2020;20(1):48. doi: 10.1186/s12902-020-0530-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4. Hyeon J, Ahn S, Shin JH, Oh YL. The prediction of malignant risk in the category “atypia of undetermined significance/follicular lesion of undetermined significance” of the Bethesda System for Reporting Thyroid Cytopathology using subcategorization and BRAF mutation results. Cancer Cytopathol. 2014;122(5):368‐376. doi: 10.1002/cncy.21396 [DOI] [PubMed] [Google Scholar]
  • 5. Kim SJ, Roh J, Baek JH, et al. Risk of malignancy according to sub‐classification of the atypia of undetermined significance or follicular lesion of undetermined significance (AUS/FLUS) category in the Bethesda system for reporting thyroid cytopathology. Cytopathology. 2017;28(1):65‐73. doi: 10.1111/cyt.12352 [DOI] [PubMed] [Google Scholar]
  • 6. Mahajan S, Srinivasan R, Rajwanshi A, et al. Risk of malignancy and risk of neoplasia in the Bethesda Indeterminate Categories: study on 4,532 thyroid fine‐needle aspirations from a single institution in India. Acta Cytol. 2017;61(2):103‐110. doi: 10.1159/000470825 [DOI] [PubMed] [Google Scholar]
  • 7. Alshalaan AM, Elzain WAD, Alfaifi J, et al. Prevalence of malignancy in thyroid nodules with AUS cytopathology: a retrospective cross‐sectional study. J Fam Med Prim Care. 2024;13(9):3822‐3828. doi: 10.4103/jfmpc.jfmpc_249_24 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8. Hathi K, Rahmeh T, Munro V, Northrup V, Sherazi A, Chin CJ. Rate of malignancy for thyroid nodules with AUS/FLUS cytopathology in a tertiary care center ‐ a retrospective cohort study. J Otolaryngol Head Neck Surg. 2021;50(1):58. doi: 10.1186/s40463-021-00530-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9. Hong SH, Lee H, Cho MS, Lee JE, Sung YA, Hong YS. Malignancy risk and related factors of atypia of undetermined significance/follicular lesion of undetermined significance in thyroid fine needle aspiration. Internet J Endocrinol. 2018;2018:4521984‐4521987. doi: 10.1155/2018/4521984 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10. Padmanabhan V, Marshall CB, Akdas Barkan G, et al. Reproducibility of atypia of undetermined significance/follicular lesion of undetermined significance category using the Bethesda system for reporting thyroid cytology when reviewing slides from different institutions: a study of interobserver variability among cytopathologists. Diagn Cytopathol. 2017;45(5):399‐405. doi: 10.1002/dc.23681 [DOI] [PubMed] [Google Scholar]
  • 11. Bhasin TS, Mannan R, Manjari M, et al. Reproducibility of 'The Bethesda System for reporting Thyroid Cytopathology': a multicenter study with review of the literature. J Clin Diagn Res. 2013;7(6):1051‐1054. doi: 10.7860/JCDR/2013/5754.3087 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12. Slabaugh G, Beltran L, Rizvi H, Deloukas P, Marouli E. Applications of machine and deep learning to thyroid cytology and histopathology: a review. Front Oncol. 2023:13‐2023. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13. Toro‐Tobon D, Loor‐Torres R, Duran M, et al. Artificial intelligence in thyroidology: a narrative review of the current applications, associated challenges, and future directions. Thyroid. 2023;33(8):903‐917. doi: 10.1089/thy.2023.0132 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14. Lee RE, McClintock DS, Laver NM, Yagi Y. Evaluation and optimization for liquid‐based preparation cytology in whole slide imaging. J Pathol Inform. 2011;2(1):46. doi: 10.4103/2153-3539.86285 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15. da Cunha Santos G, Saieg MA. Preanalytic specimen triage: smears, cell blocks, cytospin preparations, transport media, and cytobanking. Cancer Cytopathol. 2017;125(S6):455‐464. doi: 10.1002/cncy.21850 [DOI] [PubMed] [Google Scholar]
  • 16. Liu TJ, Yang WC, Huang SM, et al. Evaluating artificial intelligence‐enhanced digital urine cytology for bladder cancer diagnosis. Cancer Cytopathol. 2024;132(11):686‐695. doi: 10.1002/cncy.22884 [DOI] [PubMed] [Google Scholar]
  • 17. Aeffner F, Zarella MD, Buchbinder N, et al. Introduction to digital image analysis in whole‐slide imaging: a white paper from the Digital Pathology Association. J Pathol Inform. 2019;10(1):9. doi: 10.4103/jpi.jpi_82_18 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18. Jahn SW, Plass M, Moinfar F. Digital pathology: advantages, limitations and emerging perspectives. J Clin Med. 2020;9(11):3697. doi: 10.3390/jcm9113697 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19. Schuffler PJ, Geneslaw L, Yarlagadda DVK, et al. Integrated digital pathology at scale: a solution for clinical diagnostics and cancer research at a large academic medical center. J Am Med Inform Assoc. 2021;28(9):1874‐1884. doi: 10.1093/jamia/ocab085 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20. Gopinath B, Shanthi N. Computer‐aided diagnosis system for classifying benign and malignant thyroid nodules in multi‐stained FNAB cytological images. Australas Phys Eng Sci Med. 2013;36(2):219‐230. doi: 10.1007/s13246-013-0199-8 [DOI] [PubMed] [Google Scholar]
  • 21. Guan Q, Wang Y, Ping B, et al. Deep convolutional neural network VGG‐16 model for differential diagnosing of papillary thyroid carcinomas in cytological images: a pilot study. J Cancer. 2019;10(20):4876‐4882. doi: 10.7150/jca.28769 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22. Fragopoulos C, Pouliakis A, Meristoudis C, et al. Radial basis function artificial neural network for the investigation of thyroid cytological lesions. J Thyroid Res. 2020;2020:5464787‐5464814. doi: 10.1155/2020/5464787 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23. Elliott Range DD, Dov D, Kovalsky SZ, Henao R, Carin L, Cohen J. Application of a machine learning algorithm to predict malignancy in thyroid cytopathology. Cancer Cytopathol. 2020;128(4):287‐295. doi: 10.1002/cncy.22238 [DOI] [PubMed] [Google Scholar]
  • 24. Bohland M, Tharun L, Scherr T, et al. Machine learning methods for automated classification of tumors with papillary thyroid carcinoma‐like nuclei: a quantitative analysis. PLoS One. 2021;16(9):e0257635. doi: 10.1371/journal.pone.0257635 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25. Dov D, Kovalsky SZ, Assaad S, et al. Weakly supervised instance learning for thyroid malignancy prediction from whole slide cytopathology images. Med Image Anal. 2021;67:101814. doi: 10.1016/j.media.2020.101814 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26. Poursina O, Khayyat A, Maleki S, Amin A. Artificial intelligence and whole slide imaging assist in thyroid indeterminate cytology: a systematic review. Acta Cytol. 2025;69(2):161‐170. doi: 10.1159/000543344 [DOI] [PubMed] [Google Scholar]
  • 27. Wong CM, Kezlarian BE, Lin O. Current status of machine learning in thyroid cytopathology. J Pathol Inform. 2023;14:100309. doi: 10.1016/j.jpi.2023.100309 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28. Straccia P, Rossi ED, Bizzarro T, et al. A meta‐analytic review of the Bethesda System for Reporting Thyroid Cytopathology: has the rate of malignancy in indeterminate lesions been underestimated? Cancer Cytopathol. 2015;123(12):713‐722. doi: 10.1002/cncy.21605 [DOI] [PubMed] [Google Scholar]
  • 29. Lewis CM, Chang KP, Pitman M, Faquin WC, Randolph GW. Thyroid fine‐needle aspiration biopsy: variability in reporting. Thyroid. 2009;19(7):717‐723. doi: 10.1089/thy.2008.0425 [DOI] [PubMed] [Google Scholar]
  • 30. Wang CC, Friedman L, Kennedy GC, et al. A large multicenter correlation study of thyroid nodule cytopathology and histopathology. Thyroid. 2011;21(3):243‐251. doi: 10.1089/thy.2010.0243 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31. Seningen JL, Nassar A, Henry MR. Correlation of thyroid nodule fine‐needle aspiration cytology with corresponding histology at Mayo Clinic, 2001‐2007: an institutional experience of 1,945 cases. Diagn Cytopathol. 2012;40(Suppl 1):E27‐E32. doi: 10.1002/dc.21566 [DOI] [PubMed] [Google Scholar]
  • 32. Pathak P, Srivastava R, Singh N, Arora VK, Bhatia A. Implementation of the Bethesda system for reporting thyroid cytopathology: interobserver concordance and reclassification of previously inconclusive aspirates. Diagn Cytopathol. 2014;42(11):944‐949. doi: 10.1002/dc.23162 [DOI] [PubMed] [Google Scholar]
  • 33. Awasthi P, Goel G, Khurana U, Joshi D, Majumdar K, Kapoor N. Reproducibility of “The Bethesda System for Reporting Thyroid Cytopathology:” a retrospective analysis of 107 patients. J Cytol. 2018;35(1):33‐36. doi: 10.4103/joc.joc_215_16 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34. Na HY, Moon JH, Choi JY, et al. Preoperative diagnostic categories of fine needle aspiration cytology for histologically proven thyroid follicular adenoma and carcinoma, and Hurthle cell adenoma and carcinoma: analysis of cause of under‐ or misdiagnoses. PLoS One. 2020;15(11):e0241597. doi: 10.1371/journal.pone.0241597 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35. Panda S, Nayak M, Pattanayak L, Behera PK, Samantaray S, Dash S. Reproducibility of cytomorphological diagnosis and assessment of risk of malignancy of thyroid nodules based on the Bethesda System for Reporting Thyroid Cytopathology: a tertiary cancer center perspective. J Microsc Ultrastruct. 2022;10(4):174‐179. doi: 10.4103/jmau.jmau_88_21 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36. Kim D, Burkhardt R, Alperstein SA, et al. Evaluating the role of Z‐stack to improve the morphologic evaluation of urine cytology whole slide images for high‐grade urothelial carcinoma: results and review of a pilot study. Cancer Cytopathol. 2022;130(8):630‐639. doi: 10.1002/cncy.22595 [DOI] [PubMed] [Google Scholar]
  • 37. Hang JF, Ou YC, Yang WL, et al. Evaluating urine cytology slide digitization efficiency: a comparative study using an artificial intelligence‐based heuristic scanning simulation and multiple Z‐plane scanning. Acta Cytol. 2024;68(4):342‐350. doi: 10.1159/000538985 [DOI] [PubMed] [Google Scholar]
  • 38. Mukherjee MS, Donnelly AD, Lyden ER, et al. Investigation of scanning parameters for thyroid fine needle aspiration cytology specimens: a pilot study. J Pathol Inform. 2015;6(1):43. doi: 10.4103/2153-3539.161610 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39. Lee Y, Alam MR, Park H, et al. Improved diagnostic accuracy of thyroid fine‐needle aspiration cytology with artificial intelligence technology. Thyroid. 2024;34(6):723‐734. doi: 10.1089/thy.2023.0384 [DOI] [PubMed] [Google Scholar]
  • 40. Ahmed MS, Klippel‐Almaraz D, Amin SE, et al. Utility of whole‐slide imaging for rapid evaluation of thyroid FNA: A multireader prospective study. Cancer Cytopathol. 2025;133(9):e70046. doi: 10.1002/cncy.70046 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41. Singh A, Khan AA, Ahluwalia C, Ahuja S, Ranga S. Diagnostic accuracy of the Second Edition of the Paris System for Reporting High‐Grade Urothelial Carcinoma in Urinary Cytology. Acta Cytol. 2024;68(6):525‐531. [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Table S1

CNCY-134-0-s001.docx (35.2KB, docx)

Articles from Cancer Cytopathology are provided here courtesy of Wiley

RESOURCES