Skip to main content
JAMA Network logoLink to JAMA Network
. 2026 Aug 3;83(9):873–884. doi: 10.1001/jamaneurol.2026.2520

Automated Speech Analysis to Identify Clinical, Anatomical, and Pathological Variants of Primary Progressive Aphasia

Jet M J Vonk 1,✉, Giada Antonicelli 1, Siddarth Ramkrishnan 1, Salvatore Spina 1, Howie J Rosen 1, William W Seeley 1, Bruce L Miller 1, Maya L Henry 2, Carly Millanski 2, Maria Luisa Mandelli 1, Zachary Miller 1, Maria Luisa Gorno-Tempini 1
PMCID: PMC13434971  PMID: 42545687

This cross-sectional study investigates if automated analysis of connected speech can distinguish primary progressive aphasia variants and reflect neuroanatomical and neuropathologic substrates.

Key Points

Question

Can automated analysis of 1 to 2 minutes of connected speech yield interpretable profiles that distinguish primary progressive aphasia (PPA) variants and reflect neuroanatomical and neuropathologic substrates?

Findings

In this cross-sectional study of 214 participants, Lasso multinomial modeling generated variant-specific speech profile scores that differentiated nonfluent PPA, logopenic PPA, and semantic PPA with high overall performance and showed expected atrophy associations. In an autopsy-confirmed subset, speech profiles also discriminated common neuropathologic classes.

Meaning

Results suggest that brief automated speech profiling may enable scalable, interpretable support for PPA differential diagnosis and monitoring.

Abstract

Importance

Primary progressive aphasia (PPA) is defined by relatively isolated speech and language symptoms caused by neurodegeneration of language networks; classification of the different clinical, anatomical, and pathological variants relies on time-intensive, expert-dependent assessments that are not widely available. Scalable, interpretable speech-based tools could support diagnosis and monitoring in clinical care and trials.

Objective

To determine whether automated speech analysis of voice recording from a short picture description task can yield clinically interpretable speech and language profiles that (1) distinguish among PPA variants, (2) show variant-specific neuroanatomical correlates, and (3) align with underlying autopsy-confirmed neuropathological diagnoses.

Design, Setting, and Participants

This was a cross-sectional observational study of patients seen between 2001 and 2025 using the participants’ first visit. The setting was a single referral center with external validation in an independent sample from 2 sites. The primary sample included research cohort participants in the following groups: cognitively healthy controls, nonfluent PPA, logopenic PPA, and semantic PPA.

Exposures

Picture description task (1-2 minutes of recorded speech) from which 40 linguistic and acoustic features were automatically extracted.

Main Outcomes and Measures

The main outcomes included variant-specific speech profile scores derived from Lasso multinomial logistic regression; classification performance for clinical variants and most common underlying neuropathology; and voxelwise associations between speech-profile scores and gray matter volume.

Results

A total of 214 participants (mean [SD] age, 65.9 [7.9] years; 118 female [55%]) were included in this analysis (43 in the control group, 50 with nonfluent PPA, 56 with logopenic PPA, and 65 with semantic PPA). Among those with PPA, 64 had postmortem neuropathological data available. Twenty-five features differed between at least 2 PPA variants in 214 patients. Multinomial logistic regression achieved an AUC of 0.90 (95% CI, 0.84-0.97) and generated 3 variant-specific logit scores (speech profiles) using 4 to 8 selected features per variant. External validation in an independent cohort yielded an AUC of 0.90 (95% CI, 0.83-0.97). Profile scores showed associations consistent with established neuroanatomical patterns (n = 195): left superior and middle frontal and premotor cortex in nonfluent PPA, left posterior temporal cortex and angular gyrus in logopenic PPA, and bilateral (left-predominant) anterior temporal lobes in semantic PPA. In an autopsy-confirmed subset with most common underlying pathology (n = 56), speech profile scores discriminated neuropathology with an AUC of 0.90 (95% CI, 0.80-0.96).

Conclusions and Relevance

Results of this cross-sectional study suggest that automated speech analysis of a short audio sample of connected speech yielded interpretable speech profiles that accurately distinguished PPA clinical, anatomical, and neuropathological subtypes. These automated speech profiles may serve as clinical tools to support differential diagnosis and longitudinal monitoring, particularly in settings where specialized speech-language assessment is limited.

Introduction

Primary progressive aphasia (PPA) is a clinical syndrome in which speech and language decline is the earliest and most prominent symptom, with relative preservation of other cognitive domains in the early course.1 PPA comprises 3 consensus-defined variants with distinct linguistic and neuroanatomical profiles as follows: nonfluent or agrammatic PPA, characterized by effortful, agrammatic speech and left frontal-insular atrophy; semantic variant PPA, marked by fluent but semantically impoverished speech and left-predominant anterior temporal degeneration; and logopenic variant PPA, defined by impaired word retrieval and repetition with left temporoparietal atrophy.2,3 Accurate variant classification informs prognosis and underlying pathology and is increasingly relevant for emerging disease-modifying treatments.4 However, diagnosis remains challenging in routine practice, particularly in early or mixed presentations, and access to expert speech-language assessment is limited in many clinical settings.5,6

Connected speech offers a rich window into the multidimensional speech and language deficits in PPA.7,8 Compared with domain-specific tasks (eg, naming, repetition), brief narrative samples capture lexical retrieval, semantic specificity, syntactic complexity, fluency, and prosody within minutes. Manual analyses have shown that connected speech can reliably differentiate PPA variants,8 and automated transcription and speech-analysis tools now enable scalable extraction of linguistic and acoustic features.7,9,10 Beyond reproducing established markers, automated approaches can quantify a broader feature space and detect subtle patterns difficult to assess reliably by hand or have been underemphasized in traditional batteries.

Prior semiautomated and automated studies—often in modest samples—demonstrate that combinations of lexical, syntactic, semantic, and acoustic measures can distinguish PPA variants and controls.7,9,11,12,13,14 Yet individual features overlap across variants, and highly multivariate or end-to-end models may lack clinical interpretability. There remains a need for approaches that treat speech as an integrated multidomain behavior while yielding concise, transparent speech profiles that align with established clinical constructs.

In this study, we derived variant-specific speech profiles from automated analysis of connected speech in a large cohort (>170 individuals with PPA), uniquely integrating detailed linguistic and acoustic features with neuroanatomical and neuropathological data. Using clinically informed and data-driven feature selection, we identified interpretable feature combinations distinguishing nonfluent PPA, semantic PPA, and logopenic PPA. We hypothesized that the nonfluent PPA profile would reflect agrammatism and dysfluency, the semantic PPA profile would reflect reduced semantic specificity, and the logopenic PPA profile would reflect impaired lexical-phonological retrieval and altered timing. We further examined associations with atrophy patterns and neuropathology in an autopsy-confirmed subset. By combining multifeature modeling with interpretability and biological grounding, this work advances phenotypic characterization of PPA and supports clinically usable, speech-based clinical, anatomical, and pathological classification.

Methods

Participants

The primary sample of this cross-sectional study included participants from the University of California at San Francisco (UCSF) Edward and Pearl Fein Memory and Aging Center (MAC) PPA program research cohort seen between 2001 and 2025. All participants or their caregivers provided written informed consent in accordance with the Declaration of Helsinki, and the study was approved by the UCSF institutional review board. Participants with PPA underwent detailed clinical, neuropsychological, and neuroimaging assessments and were diagnosed using established criteria.2,15 All participants underwent research-based genetic testing using a panel of genes associated with autosomal dominant dementia; among them, 1 individual with a diagnosis of nonfluent PPA was identified as carrying a pathogenic variant in the GRN gene. Healthy participants were screened to exclude neurological or psychiatric disorders and assessed with comparable protocols. Exclusion criteria were a diagnosis of mixed PPA or PPA not otherwise specified, nonnative English speaking and/or bilingual status, absence of a picture description speech sample, a Mini-Mental State Examination (MMSE) score less than 10, or less than 50 words produced (to exclude cases too impaired for speech analysis). Prior work supports analyzable features from a minimum of approximately 25 to 50 words, with less than 25 words often excluded and 50 or more words improving reliability.16,17,18,19 We therefore adopted a 50-word minimum to balance robustness and inclusion. For participants with multiple visits, we used the earliest qualifying visit. Participants self-reported the following races and ethnicities: African American or Black, Asian, Hispanic or Latinx, Native American, White, and unknown. Race and ethnicity were reported to characterize the study population and assess the generalizability of the findings. This study is reported in accordance with the Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) reporting guidelines.20

External validation was performed in an independent sample of individuals with PPA combining cases from the UCSF Fein MAC PPA program that were not included in the primary sample and cases from a second site, the University of Texas at Austin (UT Austin) Aphasia Research and Treatment Laboratory. This dual-site external validation sample was constructed to ensure sufficient sample size and broaden inclusion relative to the primary cohort. The UT Austin participants included monolingual native English speakers with PPA enrolled in a speech-language intervention study, with inclusion criteria aligned with the primary cohort and picnic picture description samples acquired pretreatment. The parent study was approved by the UT Austin institutional review board, and all participants provided written informed consent. The UCSF cases included participants who were excluded from the primary sample based on MMSE score, word count, or nonnative English or bilingual status, thereby extending the external validation to include individuals with more severe impairment and more diverse language backgrounds.

Speech Task and Automated Speech and Language Analysis

Speech was elicited with the picnic-scene picture description from the Western Aphasia Battery.21 Participants completed the task individually in a quiet closed-door examination room, with a camcorder on a tripod placed on a table in front of the participant at an approximate distance of 2.5 ft (0.8 m) from the participant’s face. Participants were instructed “Tell me what you see. Try to talk in sentences,” and were allowed to speak without a fixed time limit; samples typically lasted 1 to 2 minutes, with a minority extending beyond 3 minutes. Recordings were processed with our in-house Clinical Linguistic Automated Speech Pipeline (CLASP).17 Acoustic files were diarized and denoised using Audacity and subsequently automatically transcribed using OpenAI’s large Whisper model22 followed by manual quality control in terms of correct wording, disfluencies, punctuation, and grammar. This approach has been validated, demonstrating comparable classification performance of PPA variants using fully manual, semiautomated, and fully automated transcription methods.23 Acoustic and linguistic features were automatically extracted with PRAAT/MyProsody,24 openSMILE,25 spaCy,26 custom Python scripts using lexical databases, and custom large language model–based scripts.

Based on clinical-linguistic expertise and prior literature, we predefined a core set of 40 features, comprising transcript-based linguistic and audio signal–based acoustic measures, reflecting domains previously studied in automated and manual speech analyses of PPA and related neurodegenerative conditions.7,9,11,12,13,14 This set includes measures of overall verbal productivity (total words, sentences), part-of-speech distribution (adjectives + adverbs, adpositions, coordinating conjunctions, demonstratives [eg, this/that/here/there], nouns, pronouns, verbs), lexical selection (content/function words, nouns/pronouns, nouns/verbs), lexical diversity (type-token ratio), psycholinguistic properties of nouns (age of acquisition, ambiguity, concreteness, familiarity, lexical frequency, phonological neighborhood density), emotional-semantic properties (arousal, valence), morphosyntax (finite verb forms), information content (idea density, information content units), discourse or pragmatic markers (epistemic uncertainty, hedging, interjections), speech fluency (articulation rate, loudness peaks per second), pausing (pauses, pause duration), speech continuity (percent speech, segment duration), prosody (fundamental frequency [F0] variability, spectral flux), loudness or intensity (loudness mean and variability), and voice quality (harmonics-to-noise ratio, jitter, shimmer). Detailed descriptions of these features are available in eTable 1 in Supplement 1.

Neuropathological Assessment

Postmortem neuropathological data were available for some patients. Analyses were restricted to cases in which the syndrome was caused by the most common underlying pathology27 as follows: nonfluent PPA with frontotemporal lobar degeneration (FTLD) with accumulation of 4-repeat (4R) tau protein (referred to hereafter as 4R-tau; corticobasal degeneration [CBD], progressive supranuclear palsy [PSP]), semantic PPA with TDP-43 type C pathological form of FTLD (referred to hereafter as TDP-C), and logopenic PPA with Alzheimer disease (AD) pathology. Other combinations were not included in this analysis, including 3 Pick disease in nonfluent PPA, AD with contributing CBD in nonfluent PPA, TDP-43 type A in nonfluent PPA (case with GRN pathogenic variation), TDP-43 type B (TDP-B) with amyotrophic lateral sclerosis–TDP in nonfluent PPA, FTLD-tau/MAPT gene variant in semantic PPA, and TDP-B in semantic PPA. Neuropathological evaluations followed consensus guidelines for FTLD28,29,30 and AD using the National Institute on Aging and the Alzheimer’s Association framework,31 using standard institutional procedures with final pathologic diagnoses assigned by expert consensus.

Neuroimaging Analysis

Magnetic resonance imaging T1-weighted scans, taken within 365 days of the speech recording, were available for the majority of participants and acquired on 1.5-T, 3-T, or 4-T scanners using previously described sequences.3,32,33 All images were processed for gray matter volume using the standard pipeline implemented in the Computational Anatomy Toolbox (CAT12, version 12.9)34 within SPM12 (version 7771),35 following the procedures outlined in Mandelli et al.36 In the full sample, voxelwise general linear models examined associations between gray matter volume and each of the 3 speech profile scores (composite logits for nonfluent PPA, semantic PPA, and logopenic PPA), adjusting for total intracranial volume, variables known to determine gross anatomical differences (ie, age, sex or gender) and a proxy for disease severity (ie, global gray matter volume). Because the multinomial model was trained only on PPA groups, controls did not have profile scores; for the brain-behavior analyses, we therefore assigned all controls a fixed logit of −13.8 on each profile, corresponding to an effectively zero probability of PPA. Resulting t maps were thresholded at peak-level 2-sided P < .05 FWE-corrected with a minimum cluster size of 100 voxels.

Statistical Analysis

Demographic, clinical, and speech variables were summarized with descriptive statistics. For the core features, we inspected distributions and tested group differences to guide feature preselection for regression across controls, nonfluent PPA, semantic PPA, and logopenic PPA using analysis of covariance (ANCOVA) adjusted for age, sex or gender, and education, with false-discovery rate (FDR)–corrected main effect tests (P < .05) and uncorrected post hoc pairwise contrasts thresholded at P < .01 to limit false positives. We pruned 1 feature from any highly correlated pair (Pearson |r|≥0.90). Missing values were minimal and handled with within-feature median imputation.

To derive multifeature speech profiles for the 3 PPA variants, we fit penalized multinomial logistic regression models with an L1 (Lasso) penalty, modeling PPA variant as the outcome and using as predictors the subset of features that differed between at least 1 pair of variants in the ANCOVA post hoc tests. During this feature-selection stage, we applied 5-fold cross-validation to mitigate overfitting, whereby the data were partitioned into 5 folds, with 4 folds used to fit the model and 1 fold held out for validation, rotating across folds. Models used λ-1SE (ie, the largest λ within 1 SE of the minimum deviance). To assess feature-selection stability, this cross-validation procedure was repeated across 30 random seeds, generating different data partitions to reduce dependence on any single data split and improve robustness; features with nonzero coefficients in 90% or more of runs were deemed stable.

We then fit a final unpenalized multinomial logistic regression including only the stable features and evaluated model performance using 5-fold cross-validation with out-of-fold predictions. The resulting class-specific linear predictors (logits) were used as composite speech profile scores for each variant, giving each participant 3 continuous indices of how strongly their speech pattern aligned with each profile. Classification performance was summarized by area under the receiver operating characteristic curve (AUC, with bootstrapped 95% CIs), sensitivity, specificity, positive predicted value (PPV), and a 3-class confusion matrix. We used ANCOVA to test whether the 3 speech profile scores differed across PPA variants, adjusting for age, sex or gender, and education and reporting effect sizes (Cohen d); a sensitivity analysis additionally adjusted for disease severity (Clinical Dementia Rating [CDR]) to assess whether profiles distinguished variants beyond global severity. Additionally, within each PPA variant, disease severity was compared between correctly classified and misclassified cases using Wilcoxon rank sum tests to evaluate whether misclassification was associated with differences in severity.

External validation was performed in an independent sample from 2 sites; sensitivity analyses additionally performed external validation on subsamples per site (UCSF Fein MAC and UT Austin). Preprocessing and feature extraction were conducted as in the primary sample, with missing values imputed using training-derived medians. The final unpenalized multinomial logistic regression model from the primary sample was fixed and applied without refitting. Classification performance was summarized by AUC, sensitivity, specificity, PPV, and confusion matrix.

For neuropathology analyses, we examined whether continuous speech profile logit scores distinguished underlying pathology classes (AD, 4R-tau, TDP-C); in a sensitivity analysis, we split the group with 4R-tau and ran the model on 4 classes (AD, CBD, PSP, TDP-C). We fit 1 multinomial logistic regression model using the 3 speech profile scores as independent variables and neuropathologic diagnosis as the outcome. Model performance was evaluated using multiclass AUC and overall classification accuracy, as well as class-specific sensitivity, specificity, and 1-vs-rest AUCs.

All analyses and visualizations were conducted in R, version 4.4.2 (R Project for Statistical Computing), using a range of packages as listed in the analysis code.37

Results

Participants

A total of 214 participants (mean [SD] age, 65.9 [7.9] years; 118 female [55%]; 96 male [45%]) were included in this analysis. Participant cohorts were as follows: 43 in the cognitively healthy control group, 50 with nonfluent PPA, 56 with logopenic PPA, and 65 with semantic PPA. Participants self-reported the following races and ethnicities: 2 African American or Black (0.9%), 3 Asian (1.4%), 9 Hispanic or Latinx (4.2%), 3 Native American (1.4%), 193 White (90.2%), and 4 unknown (1.9%). Demographic, clinical, and cognitive characteristics of the primary sample are summarized in Table 1, and a flowchart of participant selection is included in eFigure 1 in Supplement 1. Controls were on average older than PPA participants, but age did not differ among PPA variants. Years of education, sex or gender, race and ethnicity, and handedness were comparable across all groups. As expected, PPA variants showed greater impairment than controls on global cognition and dementia severity (MMSE, CDR) and language measures (eg, naming, semantic and letter fluency, repetition, and syntax comprehension). Although participants were allowed to speak without a fixed time limit, samples typically lasted 1 to 2 minutes (mean [SD], 100.5 [54.7] seconds), with a minority extending beyond 2 minutes (18 of 214, of whom 4 had PPA). Demographic characteristics of the external validation sample per site are summarized in eTable 2 in Supplement 1. Among those participants with PPA, 64 had postmortem neuropathological data. External validation was performed in an independent sample of 59 participants (mean [SD] age, 67.6 [6.5]; 31 female [53%]; 28 male [47%]; 19 with nonfluent PPA, 21 with logopenic PPA, and 19 with semantic PPA) from 2 sites. Sensitivity analyses additionally performed external validation on subsamples per site (UCSF Fein MAC, n = 31; 9 nonfluent PPA, 11 logopenic PPA, and 11 semantic PPA and UT Austin, n = 28; 10 nonfluent PPA, 10 logopenic PPA, and 8 semantic PPA). The UCSF Fein MAC external validation subsample had, on average, more impairment indicated on the MMSE than the primary sample, whereas the UT Austin external validation subsample had, on average, less impairment indicated on the MMSE than the primary sample.

Table 1. Demographic, Clinical, and Cognitive Characteristics of the Participant Sample.

Variable Total No. No. (%) P value (all groups)a P value (PPA only)a
Control (n = 43) nfvPPA (n = 50) lvPPA (n = 56) svPPA (n = 65)
Age, mean (SD) [range], y 214 70.21 (6.8) [53.0-87.0] 66.78 (8.5) [51.0-83.0] 63.93 (8.0) [51.0-82.0] 64.03 (6.9) [49.0-88.0] <.001 .14
Sex
Female 214 25 (58) 34 (68) 28 (50) 31 (48) .14 .07
Male 18 (42) 16 (32) 28 (50) 34 (52)
Education, mean (SD) [range], y 208 17.16 (2.0) [12.0-20.0] 15.90 (2.5) [12.0-20.0] 16.35 (2.6) [11.0-20.0] 16.06 (2.4) [11.0-20.0] .06 .50
Race and ethnicity
African American or Black 214 1 (2.3) 0 1 (1.8) 0 .40 .20
Asian 1 (2.3) 1 (2.0) 0 1 (1.5)
Hispanic or Latinx 1 (2.3) 2 (4.0) 1 (1.8) 5 (7.7)
Native American 1 (2.3) 2 (4.0) 0 0
White 39 (91) 45 (90) 51 (91) 58 (89)
Unknown 0 0 3 (5.4) 1 (1.5)
Handedness
Ambidextrous 214 0 0 1 (1.8) 0 .06 .30
Left 33 (77) 47 (94) 47 (84) 61 (94)
Right 10 (23) 3 (6.0) 8 (14) 4 (6.2)
Mini-Mental State Examination, mean (SD) [range] 198 29.29 (1.0) [25.0-30.0] 26.82 (3.0) [12.0-30.0] 22.20 (5.1) [11.0-30.0] 23.09 (5.6) [11.0-30.0] <.001 <.001
Clinical Dementia Rating, mean (SD) [range] 208 0 0.36 (0.3) [0.0-1.0] 0.53 (0.3) [0.0-1.0] 0.71 (0.4) [0.0-2.0] <.001 <.001
Modified trails total time, mean (SD) [range], s 164 0.70 (0.2) [0.3-1.1] 0.29 (0.2) [0.0-0.7] 0.18 (0.2) [0.0-0.9] 0.37 (0.2) [0.0-0.8] <.001 <.001
Backward digit span, mean (SD) [range], No. 175 5.84 (1.2) [3.0-8.0] 3.67 (1.4) [0.0-6.0] 3.00 (0.8) [2.0-5.0] 4.80 (1.3) [2.0-8.0] <.001 <.001
Letter fluency, mean (SD) [range], No. 176 17.66 (4.6) [9.0-27.0] 7.08 (4.2) [2.0-23.0] 7.84 (4.5) [1.0-17.0] 6.94 (4.7) [0.0-19.0] <.001 .60
Animal fluency, mean (SD) [range], No. 194 23.05 (4.5) [16.0-35.0] 13.59 (6.0) [0.0-33.0] 9.87 (5.0) [2.0-24.0] 7.61 (5.0) [0.0-23.0] <.001 <.001
Boston naming test, mean (SD) [range], No. 179 14.70 (0.6) [13.0-15.0] 13.15 (2.1) [6.0-15.0] 9.65 (3.9) [1.0-15.0] 3.98 (3.0) [0.0-13.0] <.001 <.001
WAB auditory word recognition, mean (SD) [range], No. 162 60.00 (0) [60.0-60.0] 59.64 (0.8) [56.0-60.0] 58.51 (2.2) [50.0-60.0] 53.22 (8.6) [25.0-60.0] <.001 <.001
Peabody Picture Vocabulary Test, mean (SD) [range], No. 150 15.80 (0.4) [15.0-16.0] 14.77 (1.4) [11.0-16.0] 13.73 (2.4) [5.0-16.0] 7.57 (4.1) [0.0-15.0] <.001 <.001
Repetition of 5 sentences, mean (SD) [range], No. 169 4.59 (0.6) [3.0-5.0] 2.61 (1.6) [0.0-5.0] 2.13 (1.2) [0.0-5.0] 3.45 (1.4) [0.0-5.0] <.001 <.001
WAB spontaneous speech, mean (SD) [range] 165 20.00 (0.0) [20.0-20.0] 16.27 (2.4) [10.0-20.0] 16.89 (2.1) [12.0-20.0] 17.30 (1.7) [12.0-20.0] .03 .14
Syntax comprehension, mean (SD) [range] 144 4.90 (0.3) [4.0-5.0] 4.03 (1.1) [0.0-5.0] 3.43 (1.2) [0.0-5.0] 4.33 (1.0) [1.0-5.0] <.001 <.001
Motor speech evaluation dysarthria severity, mean (SD) [range] 171 0 1.76 (1.8) [0.0-6.0] 0.02 (0.1) [0.0-1.0] 0 <.001 <.001
Motor speech evaluation apraxia severity, mean (SD) [range] 172 0 2.22 (1.7) [0.0-6.0] 0.08 (0.6) [0.0-4.0] 0.03 (0.3) [0.0-2.0] <.001 <.001

Abbreviations: lvPPA, logopenic primary progressive aphasia; nfvPPA, nonfluent primary progressive aphasia; PPA, primary progressive aphasia; svPPA, semantic primary progressive aphasia; WAB, Western Aphasia Battery.

a

Kruskal-Wallis rank sum test; Pearson χ2 test.

Feature Preselection

Of the 40 core features (missing values, 9 of 40 features; ≤4.7% per feature), collinearity checks reduced the set to 39. We excluded total word count from further analyses, as it was highly correlated with finite verb forms (Pearson r = 0.95) and redundant with number of sentences as an index of verbal productivity. All 39 remaining features showed overall group effects at P < .05 FDR-corrected (eFigure 2 in Supplement 1). Among the PPA variants, 14 features differed between nonfluent PPA and logopenic PPA, 20 between nonfluent PPA and semantic PPA, and 8 between semantic PPA and logopenic PPA, yielding 25 unique features that showed at least 1 pairwise difference between PPA variants (Table 2).

Table 2. Feature Preselection Based on Analysis of Covariance Post Hoc Comparisons at P < .01.

Feature P value Among PPAa
nfvPPA vs lvPPA nfvPPA vs svPPA svPPA vs lvPPA
No. of sentences (count) <.001b <.001b .08 Yes
Adjectives + adverbs/total words (ratio) <.001b <.001b .47 Yes
Adpositions/total words (ratio) .17 .02 .82 No
Coordinating conjunctions/total words (ratio) <.001b <.001b .68 Yes
Demonstratives/total words (ratio) <.001b <.001b <.001b Yes
Nouns/total words (ratio) <.001b <.001b .14 Yes
Pronouns/total words (ratio) <.001b <.001b <.001b Yes
Verbs/total words (ratio) .40 <.001b .009b Yes
Content/function words (ratio) .35 .13 .96 No
Nouns/pronouns (ratio) <.001b <.001b .60 Yes
Nouns/verbs (ratio) <.001b <.001b .08 Yes
Moving-average type/token (ratio) .04 .06 .99 No
Age of acquisition of nouns (mean) .09 <.001b .09 Yes
Ambiguity of nouns (mean) .07 <.001b .11 Yes
Concreteness of nouns (mean) <.001b <.001b .27 Yes
Familiarity of nouns (mean) .15 .42 <.001b Yes
Lexical frequency of nouns (mean) .22 <.001b <.001b Yes
Phonological neighborhood density of nouns (mean) .89 .79 .31 No
Arousal rating of nouns (mean) .69 .23 .86 No
Valence rating of nouns (mean) .12 .09 >.99 No
Finite verb form (count) <.001b <.001b .96 Yes
Idea density (ratio) .26 <.001b .05 Yes
Information content units (ratio) <.001b <.001b .05 Yes
Epistemic uncertainty phrases (count) <.001b <.001b .99 Yes
Hedging phrases (count) <.001b <.001b .90 Yes
Interjections/total words (ratio) .99 .02 .006b Yes
Articulation rate (ratio) .22 <.001b .17 Yes
Loudness peaks per second (ratio) .31 <.001b .03 Yes
No. of pauses (count) .96 .04 .008b Yes
Pause duration (mean) .19 .09 .99 No
Percent speech (ratio) .50 .05 .65 No
Speech segment duration (mean) .11 .02 .95 No
F0 (standard deviation) .95 .72 .96 No
Spectral flux of voiced segments (mean) .15 .68 .004b Yes
Loudness (mean) .14 .93 .02 No
Loudness (SD) .009b .04 .94 Yes
Harmonics to noise (ratio) .40 .63 .97 No
Jitter (mean) .55 .98 .76 No
Shimmer (mean) .13 .23 .99 No

Abbreviations: lvPPA, logopenic primary progressive aphasia; nfvPPA, nonfluent primary progressive aphasia; PPA, primary progressive aphasia; svPPA, semantic primary progressive aphasia.

a

Features for which at least 1 pairwise comparison among PPA groups differed at P < .01.

b

Significant values at P <.01.

Multinomial Speech Profiles for the 3 PPA Variants

Feature selection yielded concise class-specific feature sets. For the nonfluent PPA score, 7 features were selected as follows: less adjectives + adverbs, more coordinating conjunctions, higher information content, higher loudness SD, higher nouns to pronouns ratio, higher mean age of acquisition, and less pronouns relative to the other PPA groups. For the logopenic PPA score, 4 features were selected as follows: more adjectives + adverbs and more hedging phrases (vs nonfluent PPA), and lower noun familiarity and spectral flux (vs both nonfluent PPA and semantic PPA). For the semantic PPA score, 8 features were selected as follows: higher articulation rate, more demonstratives, fewer interjections, more loudness peaks per second, higher mean lexical frequency, fewer pauses, more sentences, and more verbs relative to the other 2 variants.

For each participant, the final multinomial logistic regression model with the selected features generated 3 class-specific logit scores (Figure 1A and B) that quantify how strongly their speech pattern aligned with each variant’s profile, with higher values indicating greater evidence for that subtype relative to the others. The model achieved a mean cross-validated AUC of 0.90 (95% CI, 0.84-0.97), overall accuracy 79%, and macro-averaged sensitivity (true positive rate) and specificity (true negative rate) of 79% and 89%, respectively. Sensitivity varied across variants, with 86% for nonfluent PPA, 70% for logopenic PPA, and 82% for semantic PPA, whereas specificity was consistently high at 91%, 88%, and 90%, respectively. PPV was 80% for nonfluent PPA, 74% for logopenic PPA, and 83% for semantic PPA, indicating that predicted diagnoses were correct in the majority of cases.

Figure 1. Multinomial Speech Profiles for Primary Progressive Aphasia (PPA) Variants.

Data figure with 5 panels: 3D scatter, three confusion matrices, and boxplots. Panel A, title at upper left: 3D scatterplot. A three-dimensional axis frame with light gray gridlines and a black rectangular box. Three axes labeled nfvPPA profile (z), lvPPA profile (z), and svPPA profile (z), each with tick labels spanning approximately minus 2 to plus 2. Points plotted as filled circles in three colors with a legend at upper right of the panel: orange labeled nfvPPA, dark blue labeled lvPPA, and light green labeled svPPA. Orange points cluster mainly toward higher nfvPPA profile values and lower svPPA profile values; dark blue points cluster mainly toward higher lvPPA profile values; light green points extend upward along the svPPA profile axis. Panel B at upper right, title: Confusion matrix-clinical. A 3 by 3 heatmap with horizontal axis labeled Actual diagnosis with column labels nfvPPA, lvPPA, svPPA; vertical axis labeled Predicted diagnosis with row labels svPPA, lvPPA, nfvPPA. Cell values by row: svPPA row 3, 8, 53; lvPPA row 4, 39, 10; nfvPPA row 43, 9, 2. Darker blue shading corresponds to larger numbers. Panel C below B, title: Confusion matrix-external validation. Same axes and labels; cell values by row: svPPA row 1, 2, 16; lvPPA row 2, 16, 2; nfvPPA row 16, 3, 1. Panel D below C, title: Confusion matrix-neuropathological. Horizontal axis labeled Actual pathology with columns 4R-tau, AD, TDP-C; vertical axis labeled Predicted pathology with rows TDP-C, AD, 4R-tau. Cell values by row: TDP-C row 0, 4, 17; AD row 2, 12, 4; 4R-tau row 15, 2, 0. Panel E along the bottom, label at far left: Box plot of each score. Three side-by-side boxplots with vertical axes labeled nfvPPA profile score (logit), lvPPA profile score (logit), and svPPA profile score (logit), each with a horizontal zero line and y-axis range approximately minus 15 to plus 15. Each plot has three colored boxes over x-axis categories nfvPPA, lvPPA, svPPA, using orange, blue, and green fills; whiskers and outlier circles present, including high positive outliers in the nfvPPA profile score plot.

A, Three-dimensional (3D) scatterplot of class-specific logit scores (nonfluent PPA [nfvPPA], logopenic PPA [lvPPA], semantic PPA [svPPA]) showing clustering by clinical diagnosis. B, Three-class confusion matrix summarizing classification performance in all clinical PPA diagnoses in primary sample. C, Three-class confusion matrix summarizing classification performance in clinical PPA diagnoses in external validation sample. D, Three-class confusion matrix summarizing classification performance in most common underlying neuropathology classes in PPA. E, Boxplots of each speech profile score by PPA variant illustrating profile specificity. AD indicates Alzheimer disease; 4R-tau, frontotemporal lobar degeneration with accumulation of 4-repeat tau protein; TDP-C, TDP-43 type C pathological form of frontotemporal lobar degeneration.

Furthermore, ANCOVA showed strong group effects for all 3 speech profiles (score nonfluent PPA: F2,159 = 160.6789; P <.001; score logopenic PPA: F2,159 = 55.3069; P <.001; score semantic PPA: F2,159 = 179.837482; P <.001). Post hoc comparisons indicated that each profile score was highest in its corresponding variant (Figure 1D) as follows: the nonfluent PPA profile score was higher in nonfluent PPA than logopenic PPA (Δ = 7.43, t159 = 9.88, P <.001, Cohen d = 2.02, 95% CI, 1.62-2.43) and semantic PPA (Δ = 12.67, t159 = 17.72, P < .001, Cohen d=3.45, 95% CI, 3.03-3.88); the logopenic PPA profile score was higher in logopenic PPA than nonfluent PPA (Δ = 4.84, t159 = 7.33, P < .001, Cohen d = 1.50, 95% CI, 1.09-1.91) and semantic PPA (Δ = 6.04, t159 = 10.00, P < .001, Cohen d = 1.87, 95% CI, 1.49-2.26); and the semantic PPA profile score was higher in semantic PPA than nonfluent PPA (Δ = 14.46, t159 = 18.34, P < .001, Cohen d = 3.57, 95% CI, 3.16-3.98) and logopenic PPA (Δ = 8.47, t159 = 11.18, P < .001, Cohen d=2.09, 95% CI, 1.72-2.46). In sensitivity analyses additionally adjusting for disease severity (CDR), all group differences remained at P < .001 and effect sizes changed only minimally (Δd = ±0.03-0.18). Within-variant analyses showed no significant differences in disease severity between correctly classified and misclassified cases for nonfluent PPA (mean [SD] CDR, 0.34 [0.30] vs 0.50 [0.32]; P = .23), logopenic PPA (mean [SD], 0.54 [0.27] vs 0.50 [0.20]; P = .63), or semantic PPA (mean [SD], 0.72 [0.43] vs 0.60 [0.22]; P = .60).

External Validation

External validation in an independent cohort (n = 59; nonfluent PPA = 19, logopenic PPA = 21, semantic PPA = 19) demonstrated an AUC of 0.90 (95% CI, 0.83-0.97) and overall accuracy of 81%. Sensitivity varied across variants, with 84% for nonfluent PPA (16 of 19), 76% for logopenic PPA (16 of 21), and 84% for semantic PPA (16 of 19), while specificity was consistently high at 90%, 89%, and 93%, respectively. PPV was 80% for nonfluent PPA, 80% for logopenic PPA, and 84% for semantic PPA. Sensitivity analyses on each subsample by site yielded comparable results in the UCSF Fein MAC subsample (AUC = 0.94; 95% CI, 0.84-1.00; overall accuracy = 87%; sensitivity = 87%; specificity = 93%; PPV = 87%) and UT Austin subsample (AUC = 0.88; 95% CI, 0.76-0.97; overall accuracy = 75%; macro sensitivity = 74%; macro specificity = 87%; macro PPV = 76%).

Anatomy-Speech Profile Associations

MRI scans were available for 195 participants (47 with nonfluent PPA, 53 logopenic PPA, 60 semantic PPA, and 35 controls). Table 3 displays the voxel-based morphometry results for the neural correlates of the 3 PPA speech profiles and Figure 2 provides a visual representation. A higher nonfluent PPA speech profile score was associated with lower brain volume in the left superior and middle frontal gyri and premotor cortex (left precentral gyrus, bilateral supplementary motor area). A higher logopenic PPA speech profile score was associated with lower brain volume in the posterior temporal lobe (left middle and inferior temporal gyri) and angular gyrus. A higher semantic PPA speech profile score was associated with lower brain volume in the bilateral anterior temporal lobes, particularly on the left.

Table 3. Voxel-Based Morphometry: Neural Correlates of the 3 PPA Speech Profiles (Logit Scores).

Speech profile Hemisphere Brain region (AAL atlas) Peak MNI coordinates Peak-level inference Cluster-level inference
x y z t Value P value (FWE) Cluster sizea P value (FWE)
Score nfvPPA Left Superior frontal gyrus, dorsolateral −22 −8 54 6.43 <.001 1265 <.001
Left Precentral gyrus −39 −2 51 5.64 .001
Left Precentral gyrus −30 0 46 5.55 .002
Left Postcentral gyrus −52 −9 33 6.37 <.001 1134 <.001
Left Supplementary motor area −2 4 63 5.66 .001 552 <.001
Supplementary motor area 0 14 62 5.51 .002
Left Supplementary motor area −6 14 51 5.26 .006
Score lvPPA Left Middle temporal gyrus −57 −54 −2 6.39 <.001 2598 <.001
Left Middle temporal gyrus −62 −56 12 6.05 <.001
Left Middle occipital gyrus −36 −72 27 5.83 <.001
Left Middle occipital gyrus −36 −78 0 5.53 .002 152 .001
Left Angular gyrus −40 −69 39 5.26 .006 173 .001
Left Angular gyrus −45 −60 36 5.02 .02
Left Angular gyrus −48 −72 32 4.89 .03
Score svPPA Left Inferior temporal gyrus −30 −4 −40 16.30 <.001 48 408 <.001
Left Fusiform gyrus −34 −12 −34 16.27 <.001
Left Temporal pole −33 8 −26 16.26 <.001
Right Temporal pole 32 9 −34 12.18 <.001 16 192 <.001
Left Anterior cingulate gyrus −6 6 27 5.04 .01 134 .002

Abbreviations: AAL atlas, Automated Anatomical Labeling atlas; FWE, familywise error; IFG, inferior frontal gyrus; lvPPA, logopenic primary progressive aphasia; MNI, Montreal Neurological Institute; nfvPPA, nonfluent primary progressive aphasia; PPA, primary progressive aphasia; svPPA, semantic primary progressive aphasia.

a

Using a cluster-forming voxelwise threshold of P < .001 uncorrected; results displayed with a cluster size of minimum 100 voxels at peak-level P < .001 uncorrected that survived FWE correction at cluster level at P < .05.

Figure 2. Render View of Voxel-Based Morphometry Providing a Visual Representation of Neural Correlates of the 3 Primary Progressive Aphasia (PPA) Speech Profiles.

Three lateral brain renderings with color overlays labeled nfvPPA, lvPPA, svPPA. Three side view cortical surface renderings arranged left to right on a black background, each depicting a single cerebral hemisphere in light gray with sulci and gyri shaded. The left panel title at the top reads nfvPPA; a compact multicolor overlay lies along the superior frontal and perirolandic region near the upper left portion of the hemisphere, with highest intensity in red and orange surrounded by yellow, green, and blue at the margins. Near the lower right of this left panel, white text reads n equals 50. The middle panel title reads lvPPA; a larger multicolor overlay occupies the lateral posterior temporal and inferior parietal region, extending obliquely upward and backward, with a central red and orange area and peripheral yellow, green, and blue. Near the lower right of the middle panel, white text reads n equals 56. The right panel title reads svPPA; an extensive multicolor overlay covers much of the anterior and ventral temporal lobe and adjacent lateral temporal surface, with a broad red and orange region concentrated toward the anteroinferior temporal area and a gradient through yellow and green to blue along more posterior and superior edges. Near the lower right of the right panel, white text reads n equals 65. Centered below the panels is a horizontal color bar transitioning from dark blue at the left through lighter blue, cyan, green, yellow, and orange to red at the right; numbers above the bar read 1, 2, 3, 4, 5, 6 from left to right, and the label beneath reads T-score. In the lower right corner of the overall figure, white text reads F W E-corrected P less than .05.

Brain-behavior associations of nonfluent PPA (nfvPPA), logopenic PPA (lvPPA), and semantic PPA (svPPA) speech profile scores with whole-sample gray matter atrophy with P < .05 familywise error (FWE) correction (produced in Computational Anatomy Toolbox 12 [CAT12] Statistical Parametric Mapping [SPM] [Wellcome Centre for Human Neuroimaging at University College London]).

Pathology-Speech Profiles Associations

In multinomial models using speech profile scores to discriminate among the most common underlying neuropathologic classes in PPA (4R-tau, AD, TDP-C), multiclass AUC was 0.90 (95% CI, 0.80-0.96), and overall classification accuracy was 79% (44 of 56 cases correctly classified). One-vs-rest AUC values were 0.90 (95% CI, 0.78-1.00) for 4R-tau, 0.84 (95% CI, 0.72-0.96) for AD, and 0.96 (95% CI, 0.92-1.00) for TDP-C. For 4R-tau, sensitivity was 0.88 and specificity 0.95; for AD, sensitivity was 0.67 and specificity 0.84; and for TDP-C, sensitivity was 0.81 and specificity 0.86. Misclassifications primarily involved overlap between AD and TDP-C, whereas no 4R-tau cases were classified as TDP-C or vice versa (Figure 1C). Detailed pathology per case is available in eTable 3 in Supplement 1, in which misclassified cases (12 of 56 cases) are highlighted. In a 4-class sensitivity analysis, when CBD and PSP were analyzed separately, overall patterns were similar with most cases correctly classified even within each 4R-tau subtype despite small sample sizes (CBD: 7 of 9; PSP: 5 of 7) (eAppendix and eFigure 3 in Supplement 1).

Discussion

Automated analysis of 1 to 2 minutes of connected speech from a single picture description task yielded concise, interpretable feature subsets that index how closely a speech sample aligns with nonfluent PPA, semantic PPA, or logopenic PPA, with clear neuroanatomical correlates and neuropathological associations in an autopsy-confirmed subset. These findings suggest that complex multidimensional linguistic and acoustic data can be combined into a small set of clinically recognizable markers that support variant classification and align with neuroimaging and neuropathological profiles.

Our multinomial model produced concise feature sets that align with clinical descriptions. Several selected features (eg, pronouns, nouns, verbs, and demonstratives)7 were consistent with prior reports of their relevance to PPA variant classification and were organized into structured, variant-specific profiles. The nonfluent PPA profile combined altered loudness dynamics and lexical selection features with preserved lexical-semantic content features, capturing syntactically reduced, less elaborated, and acoustically uneven speech characteristics of nonfluent PPA with intact conceptual knowledge.8,10 The limited acoustic differences observed in the group with nonfluent PPA may reflect mild apraxia of speech and dysarthria at first visit or differences in sex distribution (despite statistical adjustment). In addition, current Natural Language Processing and signal-processing approaches do not capture phoneme-level distortions or articulatory errors characteristic of motor speech impairment; incorporating phoneme-level and speech-sound distortion measures, as explored in emerging frameworks (eg, Scalable Speech Dysfluency Modeling-Lightweight [SSDM-L]38), represents an important next step. The logopenic PPA profile reflected increased hedging and reduced prosodic variability with relatively preserved semantics, consistent with lexical-phonological retrieval difficulty rather than semantic loss.2,39 The semantic PPA profile showed preserved fluency but impaired lexical-semantic specificity, consistent with fluent yet empty output relying on frequent, vague referential expressions.8,40 Overall discrimination was strong across PPA variants, with high specificity and reliability of predictions across classes. Furthermore, external validation indicated that model performance generalizes well to independent data, with comparable results across site-specific subsamples despite differences in disease severity, suggesting consistent performance across disease stages. These continuous speech profile scores may help visualize diagnostic space, quantifying how typical or atypical a given case is for a variant and help track change over time.

Brain-behavior associations followed canonical atrophy patterns. Nonfluent PPA scores correlated with left frontal and premotor regions, consistent with the frontal-insular speech-motor and syntactic network typically affected in nonfluent PPA.41,42 Logopenic PPA scores correlated with left posterior temporal and angular gyrus regions, in line with the temporoparietal network supporting phonological working memory and sentence repetition that is characteristically impaired in logopenic PPA.36,43 Semantic PPA scores correlated with bilateral, left-predominant anterior temporal atrophy, mirroring the well-established semantic hub degeneration in semantic PPA.44,45 Together, these results support the validity of the speech profiles as not only behavioral but also neuroanatomically grounded markers of the 3 PPA variants.

Speech profile scores also differentiated the most common neuropathologic classes in PPA. Separation was strongest for 4R-tau vs TDP-C, consistent with their typically distinct neuroanatomical signatures in PPA, whereas misclassifications primarily involved AD and TDP-C, which overlap anatomically in the temporal lobes and hence can show partially convergent speech phenotypes. In a 4-class sensitivity analysis separating CBD and PSP, discrimination between the 2 4R-tau subtypes was more limited, as expected given shared molecular pathology and closely overlapping neuroanatomical patterns. This pathological classification should be considered exploratory and constrained by sample composition; speech features may reflect both underlying pathology and clinical phenotype, and copathologies (eg, Lewy body disease in logopenic PPA with AD) may have contributed to observed patterns. With larger and more neuropathologically diverse cohorts, speech-based models may further clarify their utility for distinguishing underlying pathology independent of clinical diagnosis.

Strengths and Limitations

Strengths of the current study include repeated cross-validation to derive stable, sparse feature sets, external validation, feasibility using a brief clinical task, and a large PPA cohort, roughly 2 to 5 times larger than prior studies,7,9,11,12,13,14 supporting model stability. The interpretability of selected features allows them to serve as practical clinical pearls, helping clinicians know what to listen for in each subtype and integrate speech findings with other diagnostic information, even when automated analysis is unavailable. Future work should evaluate the incremental diagnostic value of automated speech features relative to, and in combination with, standard neuropsychological assessments.

This study also has some limitations. Weaknesses include the fact that the primary cohort was a predominantly White race, highly educated, native English-speaking sample from 1 site, which may limit generalizability, although the external validation in an independent multisite and more linguistically diverse cohort provides support for broader applicability. Recording conditions were not standardized in this 25-year clinical cohort and may have introduced variability, although similar conditions across participants make systematic bias unlikely.

Conclusions

Overall, this cross-sectional study suggests proof of concept that automated analysis of 1 to 2 minutes of connected speech samples can yield concise, interpretable, and diagnostically useful speech profiles for PPA. Clinically, these findings support automated speech analysis as a complementary tool for PPA diagnosis and monitoring, particularly in settings where specialized speech-language assessment is limited. The composite scores may also serve as candidate speech biomarkers in clinical trials, pending future work on their reliability and sensitivity to longitudinal change. By treating speech as a multidimensional behavior and integrating linguistic and acoustic information, this approach helps bridge detailed manual connected-speech analyses and opaque black-box models, moving toward tools that clinicians can both trust and understand.

Supplement 1.

eFigure 1. Participant Selection Flowchart

eTable 1. Details of Linguistic and Acoustic Features Extracted for Speech and Language Analysis

eTable 2. Demographic and Clinical Characteristics of the External Validation Sample From the UCSF Fein Memory and Aging Center and the University of Austin Texas

eFigure 2. Box Plots of Performance on Linguistic and Acoustic Features Across Diagnostic Groups

eTable 3. Neuropathological Details per Case With Autopsy

eAppendix. Pathology-Speech Profiles Associations Using 4 Classes

eFigure 3. Confusion Matrix—Neuropathological (4 Classes)

Supplement 2.

Data Sharing Statement.

References

  • 1.Mesulam MM. Primary progressive aphasia. Ann Neurol. 2001;49(4):425-432. doi: 10.1002/ana.91 [DOI] [PubMed] [Google Scholar]
  • 2.Gorno-Tempini ML, Hillis AE, Weintraub S, et al. Classification of primary progressive aphasia and its variants. Neurology. 2011;76(11):1006-1014. doi: 10.1212/WNL.0b013e31821103e6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Gorno-Tempini ML, Dronkers NF, Rankin KP, et al. Cognition and anatomy in 3 variants of primary progressive aphasia. Ann Neurol. 2004;55(3):335-346. doi: 10.1002/ana.10825 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Tippett DC. Classification of primary progressive aphasia: challenges and complexities. F1000Res. 2020;9:F1000 Faculty Rev-64. doi: 10.12688/f1000research.21184.1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Ortiz GG, González-Usigli H, Nava-Escobar ER, et al. Primary progressive aphasias: diagnosis and treatment. Brain Sci. 2025;15(3):245. doi: 10.3390/brainsci15030245 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Belder CRS, Marshall CR, Jiang J, et al. Primary progressive aphasia: 6 questions in search of an answer. J Neurol. 2024;271(2):1028-1046. doi: 10.1007/s00415-023-12030-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Rezaii N, Hochberg D, Quimby M, et al. Artificial intelligence classifies primary progressive aphasia from connected speech. Brain. 2024;147(9):3070-3082. doi: 10.1093/brain/awae196 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Wilson SM, Henry ML, Besbris M, et al. Connected speech production in 3 variants of primary progressive aphasia. Brain. 2010;133(Pt 7):2069-2088. doi: 10.1093/brain/awq129 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Nevler N, Ash S, Irwin DJ, Liberman M, Grossman M. Validated automatic speech biomarkers in primary progressive aphasia. Ann Clin Transl Neurol. 2018;6(1):4-14. doi: 10.1002/acn3.653 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.García AM, Welch AE, Mandelli ML, et al. Automated detection of speech timing alterations in autopsy-confirmed nonfluent/agrammatic variant primary progressive aphasia. Neurology. 2022;99(5):e500-e511. doi: 10.1212/WNL.0000000000200750 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Fraser KC, Meltzer JA, Graham NL, et al. Automated classification of primary progressive aphasia subtypes from narrative speech transcripts. Cortex. 2014;55:43-60. doi: 10.1016/j.cortex.2012.12.006 [DOI] [PubMed] [Google Scholar]
  • 12.Zimmerer VC, Hardy CJD, Eastman J, et al. Automated profiling of spontaneous speech in primary progressive aphasia and behavioral-variant frontotemporal dementia: an approach based on usage-frequency. Cortex. 2020;133:103-119. doi: 10.1016/j.cortex.2020.08.027 [DOI] [PubMed] [Google Scholar]
  • 13.Themistocleous C, Ficek B, Webster K, den Ouden DB, Hillis AE, Tsapkini K. Automatic subtyping of individuals with primary progressive aphasia. J Alzheimers Dis. 2021;79(3):1185-1194. doi: 10.3233/JAD-201101 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Merhbene G, Lecron F, Fortemps P, Dickerson BC, Kurpicz-Briki M, Rezaii N. Detecting primary progressive aphasia (PPA) from text: a benchmarking study. bioRxiv. Preprint posted online February 24, 2025. doi: 10.1101/2025.02.19.639032 [DOI]
  • 15.Neary D, Snowden JS, Gustafson L, et al. Frontotemporal lobar degeneration: a consensus on clinical diagnostic criteria. Neurology. 1998;51(6):1546-1554. doi: 10.1212/WNL.51.6.1546 [DOI] [PubMed] [Google Scholar]
  • 16.Hoemann K, Lee Y, Kuppens P, Gendron M, Boyd RL. Emotional granularity is associated with daily experiential diversity. Affect Sci. 2023;4(2):291-306. doi: 10.1007/s42761-023-00185-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Vonk JMJ, Morin BT, Pillai J, et al. Automated speech analysis to differentiate frontal and right anterior temporal lobe atrophy in frontotemporal dementia. Neurology. 2025;104(9):e213556. doi: 10.1212/WNL.0000000000213556 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Brown W III, Balyan R, Karter AJ, et al. Challenges and solutions to employing natural language processing and machine learning to measure patients’ health literacy and physician writing complexity: the ECLIPPSE study. J Biomed Inform. 2021;113:103658. doi: 10.1016/j.jbi.2020.103658 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Wren Y, Titterington J, White P. How many words make a sample—determining the minimum number of word tokens needed in connected speech samples for child speech assessment. Clin Linguist Phon. 2021;35(8):761-778. doi: 10.1080/02699206.2020.1827458 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.von Elm E, Altman DG, Egger M, Pocock SJ, Gøtzsche PC, Vandenbroucke JP; STROBE Initiative . The Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) statement: guidelines for reporting observational studies. Lancet. 2007;370(9596):1453-1457. doi: 10.1016/S0140-6736(07)61602-X [DOI] [PubMed] [Google Scholar]
  • 21.Kertesz A. Western Aphasia Battery Test Manual. Grune & Stratton; 1982. [Google Scholar]
  • 22.Radford A, Kim JW, Xu T, Brockman G, McLeavey C, Sutskever I. Whisper: robust speech recognition via large-scale weak supervision. arXiv. Preprint posted online December 6, 2022. doi: 10.48550/arXiv.2212.04356 [DOI]
  • 23.Clarke N, Morin B, Bedetti C, et al. Automated transcription in primary progressive aphasia: accuracy and effects on classification. medRxiv. Preprint posted online February 26, 2026. doi: 10.64898/2026.02.24.26346981 [DOI]
  • 24.Sabahi S, Klopp-Tosser A. MyProsody. Accessed December 1, 2022. https://github.com/Shahabks/myprosody
  • 25.Eyben F, Wöllmer M, Schüller B. OpenSMILE: the Munich versatile and fast open-source audio feature extractor. In: Proceedings of the 18th ACM International Conference on Multimedia; 2010;1459-1462. 10.1145/1873951.1874246 [DOI]
  • 26.Honnibal M, Johnson M. An improved nonmonotonic transition system for dependency parsing. In: Marquez L, Callison-Burch C, Su J. Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics; 2015:1373-1378. [Google Scholar]
  • 27.Spinelli EG, Mandelli ML, Miller ZA, et al. Typical and atypical pathology in primary progressive aphasia variants. Ann Neurol. 2017;81(3):430-443. doi: 10.1002/ana.24885 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Mackenzie IR, Neumann M, Baborie A, et al. A harmonized classification system for FTLD-TDP pathology. Acta Neuropathol. 2011;122(1):111-113. doi: 10.1007/s00401-011-0845-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Mackenzie IR, Neumann M, Bigio EH, et al. Nomenclature and nosology for neuropathologic subtypes of frontotemporal lobar degeneration: an update. Acta Neuropathol. 2010;119(1):1-4. doi: 10.1007/s00401-009-0612-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Lee SE, Rabinovici GD, Mayo MC, et al. Clinicopathological correlations in corticobasal degeneration. Ann Neurol. 2011;70(2):327-340. doi: 10.1002/ana.22424 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Montine TJ, Phelps CH, Beach TG, et al. ; National Institute on Aging; Alzheimer’s Association . National Institute on Aging-Alzheimer’s Association guidelines for the neuropathologic assessment of Alzheimer disease: a practical approach. Acta Neuropathol. 2012;123(1):1-11. doi: 10.1007/s00401-011-0910-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Mandelli ML, Caverzasi E, Binney RJ, et al. Frontal white matter tracts sustaining speech production in primary progressive aphasia. J Neurosci. 2014;34(29):9754-9767. doi: 10.1523/JNEUROSCI.3464-13.2014 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Zhang Y, Schuff N, Ching C, et al. Joint assessment of structural, perfusion, and diffusion MRI in Alzheimer disease and frontotemporal dementia. Int J Alzheimers Dis. 2011;2011(1):546871. doi: 10.4061/2011/546871 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Github . Computational Anatomy Toolbox for standardized processing of ENIGMA data. Accessed January 20, 2026. https://neuro-jena.github.io/enigma-cat12/
  • 35.University College London . SPM12. Accessed January 20, 2026. https://www.fil.ion.ucl.ac.uk/spm/software/spm12/
  • 36.Mandelli ML, Lorca-Puls DL, Lukic S, et al. Network anatomy in logopenic variant of primary progressive aphasia. Hum Brain Mapp. 2023;44(11):4390-4406. doi: 10.1002/hbm.26388 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Github . Jet M.J. Vonk repository overview. Accessed May 6, 2026. https://github.com/jmjvonk
  • 38.Vonk JM, Lian J, Cho CJ, et al. AI-based speech error detection to differentiate primary progressive aphasia variants. medRxiv. Preprint posted online February 24, 2026. doi: 10.64898/2026.02.23.26346899 [DOI]
  • 39.Ash S, Evans E, O’Shea J, et al. Differentiating primary progressive aphasias in a brief sample of connected speech. Neurology. 2013;81(4):329-336. doi: 10.1212/WNL.0b013e31829c5d0e [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Rofes A, de Aguiar V, Ficek B, Wendt H, Webster K, Tsapkini K. The role of word properties in performance on fluency tasks in people with primary progressive aphasia. J Alzheimers Dis. 2019;68(4):1521-1534. doi: 10.3233/JAD-180990 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Mandelli ML, Vilaplana E, Brown JA, et al. Healthy brain connectivity predicts atrophy progression in nonfluent variant of primary progressive aphasia. Brain. 2016;139(Pt 10):2778-2791. doi: 10.1093/brain/aww195 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Grossman M. The nonfluent/agrammatic variant of primary progressive aphasia. Lancet Neurol. 2012;11(6):545-555. doi: 10.1016/S1474-4422(12)70099-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Conca F, Esposito V, Giusto G, Cappa SF, Catricalà E. Characterization of the logopenic variant of primary progressive aphasia: a systematic review and meta-analysis. Ageing Res Rev. 2022;82:101760. doi: 10.1016/j.arr.2022.101760 [DOI] [PubMed] [Google Scholar]
  • 44.Collins JA, Montal V, Hochberg D, et al. Focal temporal pole atrophy and network degeneration in semantic variant primary progressive aphasia. Brain. 2017;140(2):457-471. doi: 10.1093/brain/aww313 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Iaccarino L, Crespi C, Della Rosa PA, et al. The semantic variant of primary progressive aphasia: clinical and neuroimaging evidence in single subjects. PLoS One. 2015;10(3):e0120197. doi: 10.1371/journal.pone.0120197 [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplement 1.

eFigure 1. Participant Selection Flowchart

eTable 1. Details of Linguistic and Acoustic Features Extracted for Speech and Language Analysis

eTable 2. Demographic and Clinical Characteristics of the External Validation Sample From the UCSF Fein Memory and Aging Center and the University of Austin Texas

eFigure 2. Box Plots of Performance on Linguistic and Acoustic Features Across Diagnostic Groups

eTable 3. Neuropathological Details per Case With Autopsy

eAppendix. Pathology-Speech Profiles Associations Using 4 Classes

eFigure 3. Confusion Matrix—Neuropathological (4 Classes)

Supplement 2.

Data Sharing Statement.


Articles from JAMA Neurology are provided here courtesy of American Medical Association

RESOURCES