Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2026 Jul 7.
Published before final editing as: Cell Rep. 2026 Mar 30;45(4):117162. doi: 10.1016/j.celrep.2026.117162

Decoding intended speech with an intracortical brain-computer interface in a person with long-standing anarthria and locked-in syndrome

Justin J Jude 1,2,3,10,*, Stephanie Haro 3, Hadar Levi-Aharoni 1,2,3, Hiroaki Hashimoto 1,2, Alexander J Acosta 1, Nicholas S Card 4, Maitreyee Wairagkar 4, David M Brandman 4, Sergey D Stavisky 4, Ziv M Williams 5,6,7, Sydney S Cash 1,2, John D Simeral 3,8,9, Leigh R Hochberg 1,2,3,8,9, Daniel B Rubin 1,2
PMCID: PMC13335209  NIHMSID: NIHMS2170788  PMID: 41920737

SUMMARY

Intracortical brain-computer interfaces (iBCIs) for decoding intended speech have provided individuals with ALS and severe dysarthria an intuitive method for high-throughput communication. These advances have been demonstrated in individuals who are still able to vocalize and move speech articulators. Here, we decoded intended speech from an individual with long-standing anarthria, locked-in syndrome, and ventilator dependence due to advanced symptoms of ALS. We found that phonemes, words, and higher order language units could be decoded well above chance. While sentence decoding accuracy was below that of demonstrations in participants with dysarthria, we attained an extensive characterization of neural signals underlying speech in a person with locked-in syndrome and identify directions for future improvement. These include closed-loop speech imagery training and decoding linguistic (rather than phonemic) units from neural signals in middle precentral gyrus to augment decoding at the sentence level. These results demonstrate that usable speech decoding from motor cortex may be feasible in people with anarthria and ventilator dependence.

In brief

J.J. Jude et al. show that articulatory iBCI speech decoding from people with anarthria and locked-in syndrome is feasible but poses new challenges. Higher level linguistic features and real-time feedback can augment existing phoneme-level neural tuning to improve sentence accuracy.

Graphical abstract

graphic file with name nihms-2170788-f0001.jpg

INTRODUCTION

Loss of communication is one of the most disabling and distressing symptoms of amyotrophic lateral sclerosis (ALS), brainstem stroke, and other neurologic conditions that cause paralysis. The ability to communicate with depth and nuance is one of the most defining characteristics of humans, and for people with ALS, the potential loss of communication is often a determining factor in the decision to continue or withdraw life-sustaining care.1 Commercially available augmentative and assistive devices may enable maintained communication by relying upon residual motor control (e.g., eye gaze), but they are often slow, error-prone, and difficult to use, leading to high rates of abandonment.2,3 For people with ALS, stroke, muscular dystrophy, and other neurological conditions who have lost or are losing dexterous control of their hands and their ability to speak, there is an urgent need to provide intuitive, accurate, and facile tools to restore and maintain communication.

Intracortical brain computer interfaces (iBCIs)4,5 restore lost function by extracting information directly from the cerebral cortex, thereby bypassing the site of primary pathology. Previously, iBCIs have been used to restore communication for people with paralysis by providing point-and-click control of a digital interface,6–10 such as a computer cursor that can be used to select letters from a virtual on-screen keyboard. Though these systems have the advantage of rapid calibration,11,12 their overall communication rates are limited by the speed at which the user can click on individual letters on a computer screen.

More recently, iBCIs and electrocorticography (ECoG)-based BCI systems have been shown to be effective at higher throughput communication, for example, through articulator-based speech decoding,13–26 character-based handwriting decoding,27,28 and finger-based typing.29,30 In these systems, neural activity pertaining to the intended movement, in the absence of actual movement of speech articulators or hand effectors, can be used to decode rapid sequences of phonemes or characters to construct language. High accuracy in recent work14–16,27,29 is maintained through the use of language models31 that take sequences of predicted phonemes or characters and select the most likely intended output sequence based on statistics of large English language datasets.32 Similarly to recent work,16 we displayed decoded words on a screen as they were decoded, after which a text-to-speech model33 was used to vocalize each decoded sentence in a voice similar to the participant’s original voice before ALS onset.

Prior research has shown that speech iBCIs can be used by people with dysarthria and other conditions that cause articulator paralysis to communicate accurately and with high throughput.14–16 Recent studies in particular demonstrate rapid calibration of such a system, achieving highly accurate communication within a single day16 and prolonged independent use of a speech neuroprosthesis for several months.18 In this work, we show that a speech iBCI neuroprosthesis can be used by a person with long-standing anarthria and locked-in syndrome. We report on the possibilities and challenges of developing a speech neuroprosthesis for an individual who has not spoken in over two years.

Previous work on speech decoding in anarthric participants has primarily involved ECoG recordings.34 This includes single word decoding constrained to a 50 word vocabulary35 and, more recently, large vocabulary sentence decoding.14 A previous attempt at communication with a locked-in person with advanced ALS, similar to our study participant, has demonstrated the feasibility of using an implanted BCI system for long-term, reliable communication, but throughput was limited to binary classification.36,37 In contrast, here we assess speech decoding performance with a comprehensive and thus much more pragmatically useful unconstrained vocabulary.

In this work, we explored the feasibility of speech decoding in iBCI research participant T17, who has locked-in syndrome. As the participant is anarthric, his enrollment in the clinical trial permitted an opportunity to assess whether articulator control is sufficiently represented in the underlying neural activity in ventral, middle and dorsal precentral gyrus of a person who has been unable to speak for more than 2 years. We posit that, although neural activity pertaining to orofacial movements persisted, control of behavior specifically pertaining to fine articulator movement for phoneme production may not have been completely preserved.

In this case, articulator-based decoding from ventral precentral gyrus caused several instances of overlapping neural signatures during generation of phonemes that entail combinations of similar orofacial movements. Given this barrier for participants with long-standing anarthria, we additionally present evidence demonstrating the feasibility of integrating the decoding of larger language units such as short sentences and full sentences from middle precentral gyrus into a real-time phonemic speech decoding pipeline to augment performance.

RESULTS

Neural tuning to phonemes varies along precentral gyrus

As a first assessment of neural population tuning, we asked participant T17 (Figure 1A) to perform a series of attempted isolated facial movements (Video S3) and phoneme articulations. In an instructed delay paradigm, individual trials of 13 distinct orofacial movements and 39 unique phonemes were cued (Figure 1B). Both arrays placed in area 6v had a similar number of strongly tuned (>0.5 fraction of variance accounted for [FVAF]) electrodes to the 13 orofacial movements, while the more dorsal of the two 6v arrays had more strongly tuned electrodes to the 39 phoneme articulations (Figure 1C). This is contrary to the tuning in 6v arrays found in another clinical trial participant (T12) with intracortical arrays,15 where the more dorsal 6v array electrodes were more tuned to orofacial movement conditions than phonemes and vice versa for the ventral 6v array. The area 55b arrays had 3 electrodes strongly tuned to orofacial movements (>0.5 FVAF), especially in the dorsal 55b array, while neither area 55b array had strongly tuned electrodes to the 39 phoneme articulations. Far fewer of T17’s 55b electrodes were tuned to either orofacial movements or phonemes compared to electrodes in T17’s 6v arrays, although the average threshold crossing rate was lower overall in electrodes in the 55b arrays (Figure S1). A small number of electrodes in area 6d were weakly (<0.2 FVAF) to moderately (>0.2 and <0.5 FVAF) tuned to the orofacial movements, with one electrode being weakly tuned to phoneme articulation.38,39 We observed the largest number of overall tuned single electrodes to conditions of both tasks in the ventral 6v array (Figure 1D).

Figure 1. Area 6v encodes orofacial movements and phonemes in a locked-in participant.

Figure 1.

(A) Participant T17 microelectrode array locations in the left hemisphere. Two arrays were placed in each of the following cortical areas: dorsal (6d), middle (55b), and ventral (6v) precentral gyrus.

(B) Instructed delay isolated cue task: the participant was instructed to attempt to speak a cued word (either with a single phoneme or a word with multiple phonemes) (top) or perform an orofacial movement (bottom).

(C) Tuning heatmaps for all six arrays; backgrounds are shaded by cortical area. Circles indicate that threshold crossing features on a given electrode varied significantly across all orofacial movements (top) or all phonemes (bottom) (p < 1 × 10−5 assessed with one-way analysis of variance). Shading indicates the fraction of variance accounted for (FVAF) by cross movement or cross phoneme differences in threshold crossing rate.

(D) Number of tuned electrodes where threshold crossing features on a given electrode varied significantly across all orofacial movements (top) and all spoken phonemes (bottom) (p < 1 × 10−5 assessed with one-way analysis of variance).

(E–G) Confusion matrices showing decoding errors when using a simple Gaussian Naive Bayes classifier to predict (E) isolated orofacial movements, (F) isolated phonemes, or (G) whole individual words from threshold crossing neural features from the 128 intracortical electrodes implanted in area 6v.

(H–J) Similarly, decoding accuracy using threshold crossing features from the 128 intracortical electrodes implanted in area 55b.

The high number of tuned electrodes in area 6v yielded high decoding accuracy when using a simple linear Gaussian Naive Bayes (GNB) classifier (with leave-one-out evaluation) to predict orofacial movements (91.5% [95% CI = (88.4, 94.6)]; Figure 1E), phonemes (50.4% [95% CI = (48.1, 52.8)]; Figure 1F) using just these 128 electrodes. Notably, decoding accuracy of orofacial movements is significantly higher than decoding accuracy of phonemes (p < 5.37 × 10–17 assessed with one-way analysis of variance) when respective chance levels are accounted for through phoneme class sub-sampling (see Methods). Moreover, similarly articulated consonant phonemes are less distinguishable by the decoder than with vowel phonemes, for example between labial phonemes such as P and B and between alveolar phonemes such as CH and JH. Similarly, individual vowels were most often confused with other vowels rather than consonants.

When decoding single words from a 50 word vocabulary,13 we found well above chance accuracy (51.1% [95% CI = (48.0, 54.2)]) using a GNB discriminator and threshold crossing features from area 6v electrodes (Figure 1G). This higher accuracy compared to isolated phoneme decoding (relative to chance) may be due to neural activity pertaining to single words, i.e., to a combination of several phonemes in a limited vocabulary, being somewhat easily separable, especially when these phoneme combinations are unique. Most word classification errors when decoding from area 6v electrodes arise from words with overlapping phonemes, such as “where” and “here” as well as “hello” and “tell,” which may suggest that the unit encoded in 6v is primarily phonemic or motoric as opposed to encoding entire words or larger language units such as short sentences or sentences.

In area 55b, there were relatively fewer and less strongly tuned electrodes to orofacial movements and phonemes, resulting in lower decoding accuracy in both instances (Figure 1I), 45.4% (95% CI = [40.0, 50.8]) and 17.4% (95% CI = [15.6, 19.2]), respectively. However, decoding of some orofacial movements, such as jaw and larynx-based movements (Figure 1H), was reliably accurate. Additionally, we noted that longer length words such as “comfortable,” “computer,” and “success” have the highest decoding accuracy using 55b area electrodes (Figure 1J), suggesting a potential role in encoding larger language units, as we explore below.

Despite the relatively high number of tuned 6d electrodes to orofacial movements, we did not see high accuracy when decoding isolated phonemes with a GNB decoder (5.0% [95% CI = (4.0, 6.1)]). Although previous work38,39 has shown 6d electrode tuning to spoken words, in this work we observed only chance-level accuracy when decoding words from a 50 word vocabulary with electrodes in area 6d (2.0% [95% CI = (1.1, 2.9)]).

6v and 55b neural activity cluster according to articulatory movements and to components of speech, respectively

To further explore the distribution of neural activity underlying the population tuning in these cortical areas, we calculated the mean Mahalanobis distance40,41 (a distance metric in multivariate space) between neural activity recorded during each attempted phoneme or orofacial gesture. We used these distances to run an unsupervised clustering algorithm yielding hierarchical dendrograms of neural tuning. We found that for area 6v, neural activity from the isolated phoneme task was neatly clustered by tongue position within the articulatory tract during phoneme production (Figure 2A). Groups of bilabial, labiodental, alveolar, dental, rhotic, and post-alveolar phonemes were each clustered together, with vowels and consonants distributed throughout the dendrogram.

Figure 2. Phonemic and orofacial representations cluster differently depending on array location in PCG.

Figure 2.

(A) Dendrogram based on the distance of neural activity recorded from the two 64 array electrode arrays in area 6v during the isolated phoneme production task. For area 6v activity, clustering depends most strongly on position of the tongue (i.e., phoneme place) within the articulatory tract during attempted phoneme production.

(B) For area 55b electrode neural activity, phonemes are most prominently clustered into consonants versus vowels, suggesting higher order classification.

(C) For the isolated orofacial movement task (which included twelve orofacial gestures and four phonemes), area 6v again demonstrates clustering based on tongue position.

(D) In area 55b, phonemes are clustered separately from the other orofacial movements, again suggesting stronger encoding of language than specific articulatory movements.

In contrast, the clustering of neural activity from area 55b most prominently separated consonants from vowels but otherwise was relatively insensitive to phoneme place (location of sound production in the mouth) or manner (how a sound is produced; Figure 2B). Similarly, when examining the neural activity from the orofacial movement task, activity from area 6v was again clustered most prominently by tongue position (Figure 2C). Interestingly, in area 55b, neural activity from the orofacial gesture task (which included four phoneme cues along with twelve other non-language cues) was clustered most prominently between phonemes and non-language orofacial gesture cues (Figure 2D).

Sentence speech decoding

We next asked participant T17 to attempt to speak whole sentences cued on screen (Figure 3A). Threshold crossing and spike power features from both area 6v and both area 55b microelectrode arrays (a total of 256 electrodes) were used to decode phonemes every 80 ms using a recurrent neural network (RNN). Similarly to recent work,15,16 we used a 5-gram language model performing Viterbi (beam) search to infer the most likely sentence spoken, given the complete sequence of RNN phoneme probabilities at all timesteps and the statistics of the English language (see STAR Methods for further decoding pipeline details). Upon completion, each completed decoded sentence was synthesized into audio using a text-to-speech model,33 which was personalized to the participant.

Figure 3. Online speech decoding.

Figure 3.

(A) Sentence decoding pipeline - phoneme probabilities are inferred by a recurrent neural network every 80 ms. Probabilities across all preceding time windows are considered by a 5-gram phoneme-based language model. The currently predicted sentence is displayed on screen. Each completed sentence decoded is synthesized into audio using a personalized text-to-speech model.

(B) Real-time closed-loop decoded phoneme and word error rates across trial days, separated by trial days where sentence vocabulary is limited to 50 words (green and blue markers) and sentences consisting of an unconstrained vocabulary (red and yellow markers). Each tick represents a block on a given trial day. Error bars indicate a 95% confidence interval.

(C) Real-time closed-loop speaking rate across trial days, measured in words per minute. Each tick represents a block on a given trial day. Shaded region indicates a 95% confidence interval.

In a sentence copy task, participant T17 was instructed to speak cued sentences from a 50-word vocabulary sentence corpus created using words from recent work.35 These are relatively simple sentences designed to be useful for communication in a care facility setting. In the first session in which sentence speech was decoded in closed loop using the above pipeline, raw phoneme error rate averaged 38%, then 27% in the second session. Accordingly, word error rate was an average of 30% in the first session then 22% in the second session (Figure 3B, Video S1).

Subsequently on trial day 33, we switched to sentence cues sourced from a conversational English corpus42 with an unconstrained vocabulary size. We initially observed a high raw phoneme error rate (56%) and a subsequent high word error rate (89%) when decoding conversational English sentences online. On trial day 68, online decoding on the first block yielded a lower average online raw phoneme error rate of 48%, with a 50% average word error rate. Phoneme and word error rate reached their lowest on the fourth real-time decoding block of trial day 68, with a raw average phoneme error rate of 41% and an average word error rate of 48% (Video S2). Phoneme error rate thereafter stayed consistently below 50% for the conversational English sentences, with word error rate usually slightly higher than this, sometimes much higher, as on trial day 75. Area 6d electrodes were excluded from sentence speech decoding as their inclusion was not found to improve phoneme error rate for either the 50-word or unconstrained vocabulary size.

Speaking rate averaged between 45 and 60 words per minute across all sessions (Figure 3C). Owing to the participant’s lack of volitional control of his articulators, all vocalized speech was attempted without any actual orofacial movement. Thus, the rate of speech was not limited by articulator movement as with prior studies.14–16 Nonetheless, the participant was asked to attempt to speak slower than a typical conversation speed of 150 words per minute to enable the decoder outlined in Figure 3A to more reliably delineate between phonemes.

Sources of error in phoneme decoding

Similar to previously reported studies,15,16 training decoders with an increasing number of open-loop spoken sentence trials collected over successive sessions decreased the average offline phoneme error rate (Figure 4A) from 49% on day 46 to 41% on day 67 (when including all cumulative trials in training the RNN model), although past this point, it was unclear whether more open-loop trials from an increasing number of proximal sessions would continue to improve the offline phoneme error rate or plateau past a certain number of training trials.

Figure 4. Training on sentences from several days improved decoding accuracy. Training on closed-loop data resulted in out-sized decoding improvement.

Figure 4.

(A) Sentences from cumulative sessions were used to train a recurrent neural network (RNN) offline, with raw phoneme error rate and word error rate post language model reported with all data up to and including the listed session. OL indicates all open-loop sentences (no real-time decoding) and CL indicates the further inclusion of closed-loop sentences, where online feedback of decoded sentences was presented to the participant. Shaded regions indicate a 95% confidence interval.

(B) Substitutions required to produce ground truth phoneme sequences from inferred phoneme sequences, using an RNN decoder trained offline with sentences from 5 sessions.

(C) Heatmap showing individual offline average phoneme error rates across days during inference with an RNN trained up to and including sentence trials of each row’s trial day.

We found that training additionally on closed-loop blocks, i.e., those where sentences are decoded and results presented in real time, had an outsized effect on improving the phoneme error rate compared to training just on open-loop blocks recorded on the same day, bringing the overall phoneme (word) error rate down to 34% (39%) offline on day 68 and down to 34% (43%) on day 75. This suggests that sentence trials from closed-loop blocks are more informative due to the feedback received by the participant, perhaps due to higher engagement with the interactive task. Furthermore, when training an RNN model on individual proximal sessions with a controlled number of open and closed-loop trials to isolate the effect of feedback, we see a significant decrease in word error rate when training a model with closed-loop trials compared to open loop trials alone (p < 6.72 × 10−5 assessed with one-way analysis of variance) (Figure S2).

While training an RNN decoder on an increasing number of sentence trials from proximal sessions decreased phoneme and word error rates (Figure 4A), the effects of non-stationarities in the neural data over trial days limited the impact of past data inclusion when evaluating on future trial days. We show that non-stationarities across several days are substantial, by training the RNN model with neural data from day 46 and evaluating on all subsequent session days without any further retraining (Figure S3). This non-stationarity across sessions may explain the reset in decoding performance seen with large gaps between session days, which is in contrast to local plateauing within nearby sessions. Most of the day-to-day fluctuations noted after performance plateaued on day 68 in Figure 3B may be due to participant-specific factors (e.g., level of fatigue on a given day, environmental distractions) beyond purely bioelectric considerations.

Interestingly, not all phonemes showed a similar improvement after RNN training with more data. Decoding of several consonants, such as M, DH, and NG, improved drastically when training on an increasing number of trials (Figure 4C). Post-alveolar consonant phonemes such as SH and CH also had marked improvements in offline decoding accuracy (Figure S4), especially when the RNN was trained on closed-loop sentences. However, vowel phonemes did not improve substantially in accuracy with an increase in training sentences. Decoding of isolated phonemes (Figure 1F) and decoding of phonemes from spoken sentences appear to have opposing confusion profiles, wherein, decoding of isolated phonemes caused greater confusion with consonants in isolated trials, whereas there was greater confusion with vowels in sentence decoding. This may be due to co-articulations during vowel production being more difficult to separate from adjacent phonemes when spoken in sequence, especially noting the participant’s long-standing anarthria.

To better understand the pattern of the decoding errors, we next examined offline substitution errors with an RNN decoder trained on trials from 5 sessions and tested on 50 held-out trials from the lattermost session. For consonant phonemes, sequence substitution errors were usually across similarly produced phonemes (e.g., V often substituted with F and D often substituted with T) (Figure 4B). Similarly, with vowel phonemes most substitution errors were with other vowels. Particular vowels such as AH and IH were consistently predicted as other vowels and conversely other vowels were consistently predicted as AH and IH. However, there were several errors where consonants were predicted as vowels and vice versa, notably with alveolar consonants such as T and D.

6v and 55b encode different language units

To better understand the functional localization in speech and language processing in the cortical areas from which we recorded, we conducted a three-phased reading, internal speech, and attempted vocalized speech task (Figure 5A), similarly to previous work.19,43 On each trial, a phoneme, word, short sentence, or sentence (a “language unit”) was first cued to be read during the “reading phase.” After a 2-second delay, the participant was then cued to say the language unit to himself from memory without vocalizing in the “internal speech” phase, and after another 2-second delay, he was instructed to then attempt to vocalize the language unit from memory in the “attempted speech” phase. Trials consisted of 5 repetitions of 10 unique conditions of each language unit (phoneme, word, short sentence, and sentence), for a total of 40 unique stimuli (see Table S3 for cue list). Stimuli from each language unit were of a similar length. Two of the ten word- and two of the ten short sentence stimuli were nonsensical (e.g., the nonsense word stimuli were “zelwog” and “cheldgup”).

Figure 5. 55b encodes short sentences and sentence level units of language.

Figure 5.

(A) In the three-phased task, each trial consisted of a reading, internal speech and attempted speech phase.

(B–D) Average decoding accuracy of phonemes, words, short sentences and sentences in each of the (B) reading, (C) internal speech, and (D) attempted speech phases, colored by implanted array. Shaded regions indicate a 95% confidence interval.

(E) Number of tuned electrodes across all task phases for each language unit in each of the four arrays in speech areas.

(F) Tuning heatmaps for the four arrays in speech areas, with backgrounds shaded by cortical location (6v in blue, 55b in yellow). Circles indicate that threshold crossing features on a given electrode varied significantly across all task phases for a given language unit (phonemes, words, short sentences, or sentences) (p < 1 × 10−5 assessed with one-way analysis of variance). Shading indicates the fraction of variance accounted for (FVAF) by cross phase differences in threshold crossing rate.

(G) Confusion matrices showing decoding errors when decoding short sentences in the reading (left) and imagined (right) phases of the task using a Gaussian Naive Bayes decoder on a subset ensemble of 20 contiguous electrodes from the dorsal 55b array. Nonsense short sentence stimuli are highlighted in orange.

When decoding with a Gaussian Naive Bayes decoder, we found that the ventral 6v array electrodes contained information preferentially about shorter units of language: neural activity measured from the ventral 6v array could decode phonemes, words, and short sentences well (short sentences with greater than 68% accuracy during the attempted speech phase (Figure 5D)), whereas sentence decoding was far less accurate. This relative pattern of lower sentence accuracy was observed across all three task phases. Neural decoding from the dorsal 6v array had lower but parallel accuracy to the ventral 6v array across all four language units in the reading and internal speech phases.

Meanwhile, the area 55b arrays, and the dorsal 55b array in particular, appeared to encode the longer units of language, short sentences and sentences (i.e., those with contextual information), much better than phonemes and words, especially during the reading phase (Figure 5B). Notably, phoneme and word decoding accuracy from dorsal 55b was similar across all three phases (at chance level), whereas average short sentence decoding accuracy was highest in the reading phase at 46%, decreasing to 30% in the internal speech phase, and lastly increasing to 34% in the attempted speech phase. Short sentence decoding was higher in all phases with a contiguous subset of dorsal 55b array electrodes which had sustained activity across all task phases (including the delay periods; Figure S5). The ventral 55b electrodes had similarly increased average accuracy when decoding words and short sentences as opposed to phonemes in both the reading and internal speech phases, but, sentence decoding accuracy was lower in these phases, similar to that of the 6v arrays. Notably, across all three phases, average word decoding accuracy was slightly higher with the ventral 55b electrodes than with the dorsal 55b electrodes.

Classification accuracy across language units was consistent with the number and strength of tuned electrodes within each array (Figure 5E and 5F), regardless of task phase. Across all task phases, the ventral 6v array had strongly tuned (>0.5 FVAF) electrodes encoding phonemes, words and short sentences, with these electrodes only weakly tuned (<0.2 FVAF) to sentences. This array had several moderately tuned electrodes encoding phonemes and words but these only weakly encode short sentences. A few dorsal 6v array electrodes only weakly encode phonemes, with minimal tuning to words, short sentences and sentences in the few electrodes encoding these language units. This is in contrast to the encoding of phonemes in the phoneme cued task (Figure 1C) where the participant was asked only to attempt to vocalize across all 39 English phonemes. In that task, there are a high number of moderate to strongly tuned electrodes in the dorsal 6v array. In this task, tuning in individual electrodes in the dorsal 6v array was less apparent in the reading and internal speech phases (Figures S6A and S6B), consistent with lower decoding accuracy in these phases when decoding from the dorsal 6v array electrodes (Figure 5B and 5C).

Conversely, the ventral and dorsal 55b arrays had no tuned electrodes to this subset of 10 English phonemes in any of the task phases (Figures S6A–S6C), with an increasing number and intensity of tuned electrodes as the language unit size increased (Figure 5E and 5F). There was a noticeable increase in the number and intensity of tuned electrodes encoding short sentences in both 55b arrays across all task phases, consistent with the subset of contiguous electrodes in dorsal 55b with sustained activity across all task phases. There was a further strengthening of tuning to sentences in two of the electrodes tuned to short sentences in the dorsal 55b array while almost all electrodes tuned to short sentences in the ventral 55b array were also similarly tuned to sentences.

Looking more closely at decoding errors when decoding using a GNB decoder with the contiguous subset of electrodes with sustained activity in the dorsal 55b array (Figure S5), we note a relatively high degree of confusion between the two nonsense short sentences, “Need how yes feel” and “Faith hungry computer,” and across all other short sentence pairs (Figure 5G), in both the reading and imagined speech phase of the task. However, there was less notable confusion between the two nonsense word stimuli: “zelwog” and “cheldgup” (Figure S7). This suggests that these electrodes in the dorsal 55b array may encode sensical semantic context44 at the short sentence level rather than articulatory features.

DISCUSSION

The speech and motor impairments caused by ALS adversely affect the ability to effectively communicate. Recently, high performance speech iBCIs14–16 have been demonstrated in dysarthric individuals with persistent phonation, however, the feasibility of such a system for individuals with complete and prolonged loss of motor control of the articulatory motor system has not previously been explored. Moreover, articulatory speech decoding had yet to be demonstrated in an individual with locked-in syndrome.45 While speech decoding did not yield high enough accuracy for effective communication, we were able to provide communication to T17 through a hand-motor decoding based QWERTY typing interface29 in lieu of high-performance speech decoding using T17’s 6d arrays. This iBCI system allowed T17 to communicate with high accuracy and relatively high speed with an unrestricted vocabulary.

With regards to speech decoding, we found that decoding accuracy of isolated orofacial movements was significantly higher than that of isolated articulated phonemes compared to respective chance levels. This may be due to the long-standing anarthria of participant T17, causing a decline in the learned behavior corresponding to the production of the complex series of orofacial movements required to generate a differentiable neural signature for similarly articulated phonemes. Under this hypothesis, isolated movements were decoded with higher accuracy because these simpler movements were easier for the participant to attempt.

In this work, we show speech sentence decoding with well above chance phoneme and word decoding accuracy in an iBCI clinical trial participant with long-standing anarthria and locked-in syndrome. Sentence-level speech decoding with a locked-in participant represents progress toward independent use of a BCI for a participant with ALS based on speech, albeit in only a copy task and not implemented as a long-term independent means of communication. Importantly, real-time sentence decoding word error rate is above the word error rate shown with other participants14,15 and above the word error rate that would be required for independent use for primary communication.16,18 In these recent studies,14–16 participants have been dysarthric but have maintained some volitional control of speech articulators, facial muscles, and muscles and apparatus of phonation. The improvement shown in this work in phoneme decoding accuracy when a decoder was trained on closed-loop sentences suggests a path toward improved articulator-based decoding with continued real-time feedback, though the extent to which such a learning-based strategy will help remains to be seen.

Electrodes in area 55b contributed greater word level accuracy to longer words, indicating higher than phoneme level language encoding. This is consistent with recent work outlining the role of area 55b in higher-order speech planning,46–48 suggesting that encoding of facial articulator and vocal tract muscle movements may be less prevalent than with area 6v. The short sentence and sentence level encoding observed, which was especially prevalent in 55b, suggests that while cortical area 6v should be used to decode phonemes,14–16 neural activity from cortical area 55b could also be incorporated into decoding at a higher language level.

Future work will focus on a path to independent use of such a system, even when phoneme decoding from intracortical electrodes implanted in speech motor areas such as 6v is suboptimal. Although real-time feedback may improve articulator-based decoding, we hypothesize that higher level encoding of intended language output at the short sentence and sentence level in middle precentral gyrus (area 55b) may be preserved even in the setting of long-standing anarthria, and propose ensemble model sentence decoding using phoneme encoding in ventral precentral gyrus (area 6v) and short sentence/sentence level encoding in middle precentral gyrus (area 55b) as a means to improve the accuracy of intended language decoding for participants with long-standing anarthria and locked-in syndrome. This could be achieved through informed language model discrimination across possible decoded sentence candidates inferred for a given utterance, using the sentence-level encoding in area 55b, such that the final decoded sentence at the end of a spoken sentence has an overall higher word accuracy.

Limitations of the study

Our study is only performed on one participant with the described patient profile, and so generalization to other participants with locked-in syndrome, anarthria, advanced ALS and ventilator dependence is unknown. This applies not only to phoneme based articulatory decoding performance but also to the prevalence of higher-level linguistic encoding that we see with T17. Array placement, T17’s residence, and other nuances of his ALS may affect observed results in future locked-in participants. Further, other participants may engage differently with real-time feedback presented during closed-loop speech decoding, thus improvements to decoding accuracy using this approach may vary.

RESOURCE AVAILABILITY

Lead contact

Requests for further information and resources should be directed to and will be fulfilled by the lead contact, Justin J. Jude (jjude@mgh.harvard.edu).

Materials availability

This study did not generate new reagents.

Data and code availability

STAR★METHODS

EXPERIMENTAL MODEL AND STUDY PARTICIPANT DETAILS

Permission for this study was granted by the U.S. Food and Drug Administration and the Institutional Review Boards of Massachusetts General Hospital, Brown University, and the VA Providence Healthcare System. Research sessions were conducted with participant T17 who is enrolled in the BrainGate2 clinical trial (ClinicalTrials.gov ID: NCT00912041). All research sessions were performed at the participant’s place of residence.

This manuscript does not report primary clinical trial outcomes; instead, it describes scientific and engineering discoveries that were made using data collected in the context of the ongoing clinical trial. All procedures were conducted in accordance with relevant guidelines and regulations. Trial Day numbers are relative to clinical trial enrollment on day 0, wherein the participant underwent surgical Utah array implantation.

Caution: Investigational Device. Limited by Federal Law to Investigational Use. The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health, or the Department of Veterans Affairs, or the United States Government.

Participant T17

T17 is a 34-year-old right-handed man with ALS. He was diagnosed with ALS 3 years prior to the start of this study. He has tetraplegia, anarthria, and ventilator dependence; his only remaining volitional motor control is over his extraocular muscles. Following enrollment, T17 had a preoperative structural MRI, resting state fMRI, and task-based fMRI to identify appropriate anatomical targets for microelectrode array (MEA) placement. Resting state fMRI was used to generate estimated parcellations of the relevant brain areas using a custom instantiation of the Human Connectome Project50,51 analysis pipeline that was modified for deployment on clinical MRI scanners able to accommodate a mechanical ventilator. Subsequently, T17 underwent placement of six 64-electrode microelectrode arrays (Blackrock Microsystems; 1.5 mm electrode length) in the left precentral gyrus; two arrays were placed in the dorsal precentral gyrus (area 6d), two arrays were placed in the ventral precentral gyrus (area 6v), and two arrays were placed in middle precentral gyrus (area 55b) (Figure 1A). The participant could not voluntarily close their eyelids, however, the participant was given regular eye drops to aid with eye lubrication. The visual display did not cause excess eye strain above that of the participant’s regular digital screen use (the participant used a similarly sized visual display to watch videos for much of his free time). Most of the visual display for our tasks had a black background with only the sentence text shown in white and the cue rectangle in red/green to minimize eye strain. T17 consented to publication of photographs and videos containing his likeness.

METHOD DETAILS

Informed consent

Participant T17 provided first-person, informed consent for this study. At the time of his enrollment, he still had sufficient volitional control of his extraocular muscles to use an eyegaze tracking communication tablet. Using this assistive technology, he was able to ask questions and express his clear understanding of the goals of the research, anticipated risks and benefits of participating, and alternatives. As with all participants in the BrainGate clinic trial, the informed consent was conducted in the presence of an impartial, third party witness (in this case a member of the Institutional Review Board (IRB)) who was able to verify that the informed consent meeting was conducted appropriately, that informed consent was voluntarily provided by T17, and that all questions and concerns were adequately addressed. As described above, in between the time that T17 provided informed consent and these research sessions were conducted, due to progression of his underlying ALS, his eye movements progressively weakened such that he was no longer able to use an eye-gaze tracking assistive device and instead relied on manual interpretation of subtle vertical and horizontal movements as his primary means of communication.

Communication with participant

Participant T17 was able to communicate with his carepartner and our research team through faint eye movements. Up-to-down eye movement signaled “Yes” and left-to-right eye movement signaled “No”. To signal the end of sentence speaking during data collection, the participant used the up-to-down eye movement (thus we were able to determine speaking rate). Through these movements the participant was able to answer binary questions about the tasks in this work and provide feedback on a completed recording block to report mistrials. The participant was further able to request medical treatment from his care team by iterating through letters of the latin alphabet (selecting letters through yes or no eye movements).

Neural signal processing

Voltage time series signals were recorded using the Neuroplex-E system (Blackrock Neurotech) attached to three percutaneous connectors on the participant’s head, and transmitted via three mini-HMDI cables (one to each percutaneous connector), attached to two Gemini hubs (Blackrock Neurotech), prior to final processing via a Neural Signal Processor (NSP) (Blackrock Neurotech). Neural data were streamed at 30 kHz from the NSP to our processing pipeline. Signals were analog filtered (4th order Butterworth with corners at 250 Hz to 5 kHz) using the Scipy python library (scipy.signal.filtfilt).

Linear regression referencing52 (LRR) is used to reduce cortically recorded artifacts in extracted neural features arising from environmental electrical noise, through the generation of electrode-specific reference signals composed of weighted sums of other electrodes. Intuitively, this ensures robustness to systematic noise at the electrode level, as individual signals are proportionally pegged to all other electrodes in a given array. LRR filter coefficients and subsequent electrode-specific thresholds were determined using filtered 30kHz data recorded from an initial reference block at the beginning of each session. LRR coefficients are computed by solving Y = W X where Y is the signal from a given electrode we require and X is the signal from all other electrodes. We solve for the LRR weight matrix through least squares calculation W: W = inv(XT X)XT Y where inv is matrix inversion.53 electrode specific thresholds are then calculated using filtered 30kHz data once these calculated references have been applied. Thresholds were set at −3.5 times the standard deviation of the voltage signal per electrode. The number of non-causal threshold crossing (ncTX) events were computed by counting the number of times the filtered neural time series crossed these pre-calculated thresholds. Spike band power was computed by taking the sum of squared voltages observed during each 10ms time bin. During closed-loop decoding blocks, feature normalization was employed to account for neural nonstationarities (drifts in mean firing rate) which could arise over the course of a block. Within each electrode, threshold crossing rates and spike band power were z-scored (mean subtracted and divided by standard deviation per electrode). Feature extraction (threshold crossings and spike band power), binning, decoding and task phase control were performed through the Python based, modular BRAND49 framework, where each process is instantiated as a self-contained node-based python program. Messaging between these nodes is performed using a Redis database.

Fraction of variance accounted for (FVAF)

To calculate electrode tuning heatmaps in Figure 1C, 5F, and S6, we used threshold crossing rates which were z-scored (mean subtracted and standard deviation divided) and time-averaged in a 1000ms window after the go-cue. Significance of tuning was then assessed via a 1-way ANOVA applied per electrode, where each ANOVA group corresponded to a different movement condition, and each observation was a scalar firing rate for a single trial. P-values from each ANOVA were used to define tuning significance (p < 1e−5) for the tuning heatmaps. The fraction of variance accounted for by movement tuning on a single electrode was defined as: =1-SSERRSSTOT, where SSTOT is the total sum of squared average firing rates over all trials. For computing SSTOT, squaring was performed after the mean across all trials was subtracted from each trial first, so that the overall mean firing rate did not contribute to the variance. SSERR is the sum of squared predictions over all trials. Prediction error was assessed with a cross-validated (5-fold) model which predicts the firing rate of each trial based only on the mean of the movement condition it belongs to. Condition specific means were estimated on the training set by taking the sample means across training trials, then then applied to the held-out test set. If there are large differences in mean firing rate between movement conditions (i.e., strong movement tuning), then SSERR will be small relative to SSTOT. For the electrode tuning counts and maps shown in Figure 5E and 5F, only the first 5000ms of each trial are considered in calculating fraction of variance accounted for.

Gaussian Naive Bayes

The Naive Bayes classifier is a simple probabilistic predictor which assumes conditional independence between features given the class. Gaussian Naive Bayes takes this a step further for continuous data and makes the additional assumption that values associated with each class are normally distributed. Offline single-trial classification results (reported in Figures 1E–1J; 5B, 5C, 5D, and 5F) were generated using a cross-validated (leave-one-out) Gaussian Naive Bayes classifier. All trials are z-scored (mean subtracted and standard deviation divided). For decoding results reported in Figures 1E–1J, trials are time-averaged in the entire go period(s) of each task. For decoding results reported in Figures 5B–5D, 5F, and 5G, trials are time-averaged in the entire go period of 2.5 s for phonemes and 4 s for words, then only the first 5 s of the go period for short sentences and sentences for each task phase. For every trial, the classifier is trained on all other trials and evaluated on the held-out trial. We use the Scikit-learn54 implementation of Gaussian Naive Bayes.

Orofacial vs. phoneme decoding accuracy ANOVA testing

In 100 iterations, 13 phoneme classes were randomly chosen from the 40 phoneme classes outlined in Figure 1F for each iteration. In each iteration, a Gaussian Naive Bayes (GNB) decoder was trained to discriminate between each of these 13 phoneme classes (leave one out evaluation). This was such that the GNB decoder was selecting between 13 phoneme classes overall. This mirrors the decoding task of discriminating between each of 13 isolated orofacial movements. We then performed one-way analysis of variance significance testing of all class sub-sampled phoneme trial decoding accuracies vs. orofacial trial decoding accuracies.

Mahalanobis distance and hierarchical clustering

The Mahalanobis distance is a measure of the distance between a point and a distribution, which, unlike the Euclidean distance, accounts for the correlations between variables and differences in scale, making it particularly useful in multivariate statistics. To quantify dissimilarity between population activity patterns across tasks, we computed pairwise Mahalanobis distances using both threshold-crossing (ncTX) and spike-band power features. Neural features were z-scored by subtracting the mean within each session and dividing by the global standard deviation across all sessions. For this analysis the cortical areas used were 6v and 55b, consisting of 128 electrodes each. From each electrode, two neural features (z-scored ncTX and spike-band power) were extracted, yielding 256 features per cortical area.

Neural signals were binned in 10-ms windows, and activity during the 1s period following the Go cue (100 time bins) was extracted for each trial. For each of 40 tasks and 15 trials per task, neural activity was averaged across the 100 time bins to yield a (task × trial × features) tensor. Pairwise Mahalanobis distance between task-averaged population vectors were then computed as:

Dij=μi-μjTSij-1μi-μj

where μi and μj denote the mean feature vectors for tasks i and j, and Sij represents the pooled covariance matrix defined as:

Sij=0.5Σi+Σj+10-3I

Averaging the two covariance matrices ensured symmetry of the resulting distance matrix, and the regularization term (10-3I, Tikhonov regularization) improved numerical stability and ensured matrix invertibility given the high feature dimensionality. The resulting (40×40) symmetric distance matrix was used for hierarchical clustering with Ward’s minimum-variance linkage (MATLAB R2026a, Statistics and Machine Learning Toolbox). Dendrogram leaf order was optimized to minimize cross-branch distances, and phoneme classes (e.g., vowels, labial, alveolar, etc.) were color-coded for visualization.

WFST language model

We use a 5-gram language model implemented using a weighted finite state transducer (WFST), built on the Kaldi system.31 This probabilistic model was trained using the WeNet framework55 and utilized the large OpenWebText232 corpus to compute conditional probabilities. When given lattices of RNN output probability vectors, the WeNet framework initiates several Viterbi56,57 searches through the most recent lattice, incorporating phoneme-by-phoneme transition probabilities when inferring the most likely typed sentence. Language model parameter details can be found in Table S2.

Sentence speech decoding

Sequence decoding of attempted articulatory movements was performed using a Recurrent Neural Network (RNN) decoder trained with a Connectionist Temporal Classification (CTC)58–60 loss function. This loss function allows the RNN to learn a mapping between two unaligned sequences (neural features and spoken phonemes) which are at varying temporal resolutions, similarly to recent work.14–16 A 5-gram language model then uses Viterbi (beam) search to infer the most likely sentence spoken given the complete sequence of RNN output probabilities at all timesteps and the statistics of the English language. Error rate, for both phonemes and words, is measured as the percent of incorrect insertions, deletions and erroneous substitutions to each phoneme when decoded in sequence.

Recurrent neural network

The recurrent neural network (RNN) used for phoneme sequence decoding is implemented in Tensorflow261 as a 5 layer gated recurrent unit (GRU) network, each with 512 units. A non-linear input layer is added per session to account for cross-session neural variability. Each non-linear layer contains the same number of units as the dimensionality of the neural features (512 features used). During training, batches of trials only from a given session are selected at random, such that the corresponding input layer is trained along with the 5 layer RNN. Training using backpropagation through time (BPTT) minimizes the Connectionist Temporal Classification loss function.58 Various regularization and data augmentation techniques are utilized: Dropout, Gaussian White noise, L2 weight norm. Hyperparameter details can be found in Table S1.

Much recent work focuses on reducing or eliminating the adverse decoding effects of neural nonstationarities.62–67 Here, we utilized training data across multiple days; a non-linear input layer is added per session to account for cross-session neural variability, as was used in recent work.15,16,62 Each non-linear layer contains the same number of units as the dimensionality of the neural features. During training, batches of trials only from a given session are selected at random, such that the corresponding input layer is trained along with the 5 layer RNN. This resulted in a cross-session ensemble decoder that was relatively robust to nonstationarities in the short to medium term (resulting in the reported online error rates in Figure 3B).

Phased task

Participant T17 was asked to memorize a phoneme, word, short sentences or sentence in the first reading phase of a phased task. In the second phase, the participant was asked to internally speak (without vocalizing) the memorized language unit. In the last phase, the participant was asked to attempt to speak with articulation the memorized language unit. A list of stimuli can be seen in Table S3. Each stimulus was repeated 5 times in a random order.

Personalized text-to-speech model

Similarly to previous work,16 we use a neural network based text-to-speech model33 which is first pre-trained on large spoken voice datasets such as VCTK68 containing numerous utterances in English across many speakers. The model was then fine-tuned with audio recordings of the participant’s voice prior to ALS onset to produce natural and relatively faithful voicing of decoded sentences.

QUANTIFICATION AND STATISTICAL ANALYSIS

In Figure 1C, 1D, 5E, and 5F Fraction of Variance Accounted For (FVAF) is used to ascertain the quantity and placement of tuned electrodes to each task. This is outlined in detail in the FVAF subsection of the method details section.

Significance testing of decoding accuracy of orofacial vs. decoding accuracy of phonemes assessed with one-way analysis of variance is outlined in detail in the orofacial vs. phoneme decoding accuracy ANOVA testing subsection of the method details section.

Significance testing of open vs. closed-loop trials to isolate the effect of feedback, assessed with one-way analysis of variance (p < 6.72 × 10−5) was performed with an equal number of decoded sentences. Described in the legend of Figure S2, we trained the RNN model with an equal number of open loop sentences on trial day 67 and an equal number (200) of either open loop or closed loop sentences on trial day 68. The word error rates of each held-out individual decoded sentence evaluated (40 each for open loop and closed loop) was considered for one-way analysis of variance significance testing to ascertain if held-out trial word decoding performance had significantly different means between the two test groups. In Figure S2, three asterisks denote p < 0.001, “NS” denotes no significance.

All one-way analysis of variance (ANOVA) tests were performed using the scipy Python library (scipy.stats.f_oneway https://docs.scipy.org/doc/scipy/reference/generated/scipy.stats.f_oneway.html).

Additional resources

Clinical trial registry number NCT00912041 (https://clinicaltrials.gov/study/NCT00912041?id=NCT00912041).

Supplementary Material

1
2
Download video file (96.9MB, mp4)
3
Download video file (147.9MB, mp4)
4
Download video file (39.6MB, mp4)

SUPPLEMENTAL INFORMATION

Supplemental information can be found online at https://doi.org/10.1016/j.celrep.2026.117162.

KEY RESOURCES TABLE.

REAGENT or RESOURCE SOURCE IDENTIFIER

Software and algorithms

MATLAB R2026a MathWorks Inc.https://www.mathworks.com/products/MATLAB.html RRID: SCR_001622
BRAND (BCI system software) Ali et al.49 https://github.com/brandbci/brand
Python 3.9 python.org RRID: SCR_008394
SciPy 1.11.3 scipy.org RRID: SCR_008058
NumPy 1.26.2 numpy.org RRID: SCR_008633
Pandas 2.2.2 pandas.pydata.org RRID: SCR_018214
scikit-learn 1.2.2 scikit-learn.org RRID: SCR_002577
matplotlib 3.8.4 matplotlib.org RRID: SCR_008624
seaborn 0.13.2 seaborn.pydata.org RRID: SCR_018132
Custom Code https://github.com/justin-jude/anarthria-speech-bci https://doi.org/10.5281/zenodo.18599809

Other

NeuroPort Neural Signal Processor Blackrock Neurotech https://blackrockneurotech.com/products/neuroport/
NeuroPlex E Blackrock Neurotech https://blackrockneurotech.com/products/neuroplex-e/
64 channel Utah Array Electrode Blackrock Neurotech https://blackrockneurotech.com/products/utah-array/

Deposited data

Neural data used to produce paper figures are publicly available Dryad https://doi.org/10.5061/dryad.vq83bk481

Highlights.

  • Articulatory speech can be decoded in a person with locked-in syndrome and anarthria

  • Speech decoding accuracy is improved through real-time feedback

  • Higher level encoding of sentences exists in middle precentral gyrus

  • Higher level features could be used to augment articulatory decoding

ACKNOWLEDGMENTS

We thank T17, their family, and carepartners for the time and effort they contributed to the BG2 trial. We thank Dr. Stephen Mernoff for his clinical monitoring of trial participants. We thank Dr. Gladys Hill for her illustrations in Figure 3. This work was supported by Office of Research and Development, Rehabilitation R&D Service, Department of Veterans Affairs (N2864C, A4820R, and A2295R); NIH NIDCD (U01DC017844, K23DC021297, and 1DP2DC021055); NIH NINDS R25NS065743, AHA (23SCEFIA1156586); and A.P. Giannini Postdoctoral Fellowship to N. Card. S.D.S. has a Career Award at the Scientific Interface from the Burroughs Wellcome Fund. Several analyses were performed using code from https://github.com/fwillett/speechBCI.

Footnotes

DECLARATION OF INTERESTS

The MGH Translational Research Center has a clinical research support agreement (CRSA) with Ability Neuro, Axoft, Neuralink, Neurobionics, Paradromics, Precision Neuro, Synchron, and Reach Neuro, for which L.R.H. provides consultative input. L.R.H. is a non-compensated member of the Board of Directors of a nonprofit assistive communication device technology foundation (Speak Your Mind Foundation). Mass General Brigham (MGB) is convening the Implantable Brain-Computer Interface Collaborative Community (iBCI-CC); charitable gift agreements to M.G.B., including those received to date from Paradromics, Synchron, Precision Neuro, Neuralink, and Blackrock Neurotech, support the iBCI-CC, for which L.R.H. provides effort. S.D.S. is an inventor on intellectual property owned by Stanford University that has been licensed to Blackrock Neurotech and Neuralink Corp. M.W., S.D.S., N.S.C., and D.M.B. have patent applications related to speech BCI owned by the Regents of the University of California including IP, which has been licensed to a neurotechnology startup. S.D.S. is an advisor to Sonera.

REFERENCES

  • 1.Hogden A, Greenfield D, Nugus P, and Kiernan MC (2015). Development of a model to guide decision making in amyotrophic lateral sclerosis multidisciplinary care. Health Expect. 18, 1769–1782. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Smidt A, and Pebdani RN (2023). Rethinking device abandonment: a capability approach focused model. Augment. Altern. Commun. 39, 198–206. [DOI] [PubMed] [Google Scholar]
  • 3.Waller A (2019). Telling tales: unlocking the potential of AAC technologies. Int. J. Lang. Commun. Disord 54, 159–169. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Hahn NV, Stein E, Consortium B, Donoghue JP, Simeral JD, Hochberg LR, and Willett FR (2025). Long-term performance of intracortical microelectrode arrays in 14 BrainGate clinical trial participants. Preprint at medRxiv. 10.1101/2025.07.02.25330310. [DOI] [Google Scholar]
  • 5.Deo DR, Okorokova EV, Pritchard AL, Hahn NV, Card NS, Nason-Tomaszewski SR, Jude J, Hosman T, Choi EY, Qiu D, et al. (2024). A mosaic of whole-body representations in human motor cortex. Preprint at bioRxiv. 10.1101/2024.09.14.613041. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Pandarinath C, Nuyujukian P, Blabe CH, Sorice BL, Saab J, Willett FR, Hochberg LR, Shenoy KV, and Henderson JM (2017). High performance communication by people with paralysis using an intracortical brain-computer interface. eLife 6, e18554. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Bacher D, Jarosiewicz B, Masse NY, Stavisky SD, Simeral JD, Newell K, Oakley EM, Cash SS, Friehs G, and Hochberg LR (2015). Neural point-and-click communication by a person with incomplete locked-in syndrome. Neurorehabil. Neural Repair 29, 462–471. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Jarosiewicz B, Sarma AA, Bacher D, Masse NY, Simeral JD, Sorice B, Oakley EM, Blabe C, Pandarinath C, Gilja V, et al. (2015). Virtual typing by people with tetraplegia using a self-calibrating intracortical brain-computer interface. Sci. Transl. Med. 7, 313ra179. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Ottenhoff MC, Verwoert M, Goulis S, Tousseyn S, van Dijk, Shanechi MM, Sani OG, Kubben P, and Herff C (2025). Decoding continuous goal-directed movement from human brain-wide intracranial recordings. Cell Reports 44, 116328. 10.1016/j.celrep.2025.116328. [DOI] [PubMed] [Google Scholar]
  • 10.Deo DR, Willett FR, Avansino DT, Hochberg LR, Henderson JM, and Shenoy KV (2024). Brain control of bimanual movement enabled by recurrent neural networks. Sci. Rep. 14, 1598. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Brandman DM, Hosman T, Saab J, Burkhart MC, Shanahan BE, Ciancibello JG, Sarma AA, Milstein DJ, Vargas-Irwin CE, Franco B, et al. (2018). Rapid calibration of an intracortical brain-computer interface for people with tetraplegia. J. Neural. Eng. 15, 026007. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Singer-Clark T, Hou X, Card NS, Wairagkar M, Iacobacci C, Peracha H, Hochberg LR, Stavisky SD, and Brandman DM (2025). Speech motor cortex enables BCI cursor control and click. Journal of Neural Engineering 22, 036015. 10.1088/1741-2552/add0e5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Chang EF, and Anumanchipalli GK (2020). Toward a Speech Neuroprosthesis. JAMA 323, 413–414. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Metzger SL, Littlejohn KT, Silva AB, Moses DA, Seaton MP, Wang R, Dougherty ME, Liu JR, Wu P, Berger MA, et al. (2023). A high-performance neuroprosthesis for speech decoding and avatar control. Nature 620, 1037–1046. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Willett FR, Kunz EM, Fan C, Avansino DT, Wilson GH, Choi EY, Kamdar F, Glasser MF, Hochberg LR, Druckmann S, et al. (2023). A high-performance speech neuroprosthesis. Nature 620, 1031–1036. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Card NS, Wairagkar M, Iacobacci C, Hou X, Singer-Clark T, Willett FR, Kunz EM, Fan C, Vahdati Nia M, Deo DR, et al. (2024). An Accurate and Rapidly Calibrating Speech Neuroprosthesis. N. Engl. J. Med. 391, 609–618. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Wairagkar M, Card NS, Singer-Clark T, Hou X, Iacobacci C, Miller LM, Hochberg LR, Brandman DM, and Stavisky SD (2025). An instantaneous voice-synthesis neuroprosthesis. Nature 644, 145–152. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Card NS, Singer-Clark T, Peracha H, Iacobacci C, Hou X, Wairagkar M, Fogg Z, Offenberg E, Hochberg LR, Brandman DM, and Stavisky SD (2025). Long-term independent use of an intracortical brain-computer interface for speech and cursor control. Preprint at bioRxiv. 10.1101/2025.06.26.661591. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Kunz EM, Abramovich Krasa B, Kamdar F, Avansino DT, Hahn N, Yoon S, Singh A, Nason-Tomaszewski SR, Card NS, Jude JJ, et al. (2025). Inner speech in motor cortex and implications for speech neuroprostheses. Cell 188, 4658–4673.e17. 10.1016/j.cell.2025.06.015. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Moses DA, Leonard MK, Makin JG, and Chang EF (2019). Real-time decoding of question-and-answer speech dialogue using human cortical activity. Nat. Commun. 10, 3096. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Anumanchipalli GK, Chartier J, and Chang EF (2019). Speech synthesis from neural decoding of spoken sentences. Nature 568, 493–498. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Angrick M, Herff C, Mugler E, Tate MC, Slutzky MW, Krusienski DJ, and Schultz T (2019). Speech synthesis from ECoG using densely connected 3D convolutional neural networks. J. Neural. Eng. 16, 036019. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Brumberg JS, Wright EJ, Andreasen DS, Guenther FH, and Kennedy PR (2011). Classification of intended phoneme production from chronic intracortical microelectrode recordings in speech-motor cortex. Front. Neurosci. 5, 65. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Guenther FH, Brumberg JS, Wright EJ, Nieto-Castanon A, Tourville JA, Panko M, Law R, Siebert SA, Bartels JL, Andreasen DS, et al. (2009). A wireless brain-machine interface for real-time speech synthesis. PLoS One 4, e8218. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Hou X, Iacobacci C, Card NS, Wairagkar M, Singer-Clark T, Kunz EM, Fan C, Kamdar F, Hahn N, Hochberg LR, et al. (2025). Error encoding in human speech motor cortex. Preprint at bioRxiv. 10.1101/2025.06.07.658426. [DOI] [Google Scholar]
  • 26.Srinivasan A, Wairagkar M, Iacobacci C, Hou X, Card NS, Jacques BG, Pritchard AL, Bechefsky PH, Hochberg LR, AuYong N, et al. (2025). Encoding of speech modes and loudness in ventral precentral gyrus. Preprint at bioRxiv. 10.1101/2025.06.07.658426. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Willett FR, Avansino DT, Hochberg LR, Henderson JM, and Shenoy KV (2021). High-performance brain-to-text communication via handwriting. Nature 593, 249–254. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Qi Y, Zhu X, Xiong X, Yang X, Ding N, Wu H, Xu K, Zhu J, Zhang J, and Wang Y (2025). Human motor cortex encodes complex handwriting through a sequence of stable neural states. Nat. Hum. Behav. 9, 1260–1271. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Jude JJ, Levi-Aharoni H, Acosta AJ, Allcroft SB, Nicolas C, Lacayo BE, Card NS, Wairagkar M, Levin AD, Brandman DM, et al. (2026). Restoring rapid natural bimanual typing with a neuroprosthesis after paralysis. Nature Neuroscience. 10.1038/s41593-026-02218-y. [DOI] [PubMed] [Google Scholar]
  • 30.Shah NP, Willsey MS, Hahn N, Kamdar F, Avansino DT, Hochberg LR, Shenoy KV, and Henderson JM (2023). A brain-computer typing interface using finger movements. In International IEEE/EMBS Conference on Neural Engineering, NER. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Povey D, Ghoshal A, Boulianne G, Burget L, Glembek O, Goel N, Hannemann M, Motlicek P, Qian Y, Schwarz P, et al. (2011). The Kaldi Speech Recognition Toolkit. IEEE Signal Process. Soc. [Google Scholar]
  • 32.Gao L, Biderman S, Black S, Golding L, Hoppe T, Foster C, Phang J, He H, Thite A, Nabeshima N, et al. (2020). The Pile: An 800GB Dataset of Diverse Text for Language Modeling. Preprint at arXiv. 10.48550/arXiv.2101.00027. [DOI] [Google Scholar]
  • 33.Li YA, Han C, Raghavan VS, Mischler G, and Mesgarani N (2023). StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models. Adv. Neural Inf. Process. Syst. 36, 19594–19621. [PMC free article] [PubMed] [Google Scholar]
  • 34.Rabbani Q, Milsap G, and Crone NE (2019). The Potential for a Speech Brain–Computer Interface Using Chronic Electrocorticography. Neurotherapeutics 16, 144–165. 10.1007/s13311-018-00692-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Moses DA, Metzger SL, Liu JR, Anumanchipalli GK, Makin JG, Sun PF, Chartier J, Dougherty ME, Liu PM, Abrams GM, et al. (2021). Neuroprosthesis for Decoding Speech in a Paralyzed Person with Anarthria. N. Engl. J. Med. 385, 217–227. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Vansteensel MJ, Pels EGM, Bleichner MG, Branco MP, Denison T, Freudenburg ZV, Gosselaar P, Leinders S, Ottens TH, Van Den Boom MA, et al. (2016). Fully Implanted Brain–Computer Interface in a Locked-In Patient with ALS. N. Engl. J. Med. 375, 2060–2066. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Vansteensel MJ, Leinders S, Branco MP, Crone NE, Denison T, Freudenburg ZV, Geukes SH, Gosselaar PH, Raemaekers M, Schippers A, et al. (2024). Longevity of a Brain–Computer Interface for Amyotrophic Lateral Sclerosis. N. Engl. J. Med. 391, 619–626. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Stavisky SD, Willett FR, Wilson GH, Murphy BA, Rezaii P, Avansino DT, Memberg WD, Miller JP, Kirsch RF, Hochberg LR, et al. (2019). Neural ensemble dynamics in dorsal motor cortex during speech in people with paralysis. eLife 8, e46015. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Wilson GH, Stavisky SD, Willett FR, Avansino DT, Kelemen JN, Hochberg LR, Henderson JM, Druckmann S, and Shenoy KV (2020). Decoding spoken English from intracortical electrode arrays in dorsal precentral gyrus. J. Neural. Eng. 17, 066007. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.De Maesschalck R, Jouan-Rimbaud D, and Massart DL (2000). The Mahalanobis distance. Chemomet. Intell. Lab. Syst. 50, 1–18. [Google Scholar]
  • 41.Natraj N, Seko S, Abiri R, Miao R, Yan H, Graham Y, Tu-Chan A, Chang EF, and Ganguly K (2025). Sampling representational plasticity of simple imagined movements across days enables long-term neuroprosthetic control. Cell 188, 1208–1225.e32. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Godfrey JJ, Holliman EC, and McDaniel J (1992). SWITCHBOARD: Telephone speech corpus for research and development. In ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing Proceedings. [Google Scholar]
  • 43.Wandelt SK, Bjånes DA, Pejsa K, Lee B, Liu C, and Andersen RA (2024). Representation of internal speech by single neurons in human supramarginal gyrus. Nat. Hum. Behav. 8, 1136–1149. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Tang J, LeBel A, Jain S, and Huth AG (2023). Semantic reconstruction of continuous language from non invasive brain recordings. Nat. Neurosci. 26, 858–866. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Silva AB, Littlejohn KT, Liu JR, Moses DA, and Chang EF (2024). The speech neuroprosthesis. Nat. Rev. Neurosci. 25, 473–492. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Silva AB, Liu JR, Zhao L, Levy DF, Scott TL, and Chang EF (2022). A Neurosurgical Functional Dissection of the Middle Precentral Gyrus during Speech Production. J. Neurosci. 42, 8416–8426. 10.1523/JNEUROSCI.1614-22.2022. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Liu JR, Zhao L, Hullett PW, and Chang EF (2025). Speech sequencing in the human precentral gyrus. Nat. Hum. Behav. 9, 2327–2344. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Khanna AR, Muñoz W, Kim YJ, Kfir Y, Paulk AC, Jamali M, Cai J, Mustroph ML, Caprara I, Hardstone R, et al. (2024). Single-neuronal elements of speech production in humans. Nature 626, 603–610. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Ali YH, Bodkin K, Rigotti-Thompson M, Patel K, Card NS, Bhaduri B, Nason-Tomaszewski SR, Mifsud DM, Hou X, Nicolas C, et al. (2024). BRAND: a platform for closed-loop experiments with deep network models. J. Neural. Eng. 21, 026046. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Glasser MF, Sotiropoulos SN, Wilson JA, Coalson TS, Fischl B, Andersson JL, Xu J, Jbabdi S, Webster M, Polimeni JR, et al. (2013). The minimal preprocessing pipelines for the Human Connectome Project. Neuroimage 80, 105–124. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Glasser MF, Coalson TS, Robinson EC, Hacker CD, Harwell J, Yacoub E, Ugurbil K, Andersson J, Beckmann CF, Jenkinson M, et al. (2016). A multi-modal parcellation of human cerebral cortex. Nature 536, 171–178. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Young D, Willett F, Memberg WD, Murphy B, Walter B, Sweet J, Miller J, Hochberg LR, Kirsch RF, and Ajiboye AB (2018). Signal processing methods for reducing artifacts in microelectrode brain recordings caused by functional electrical stimulation. J. Neural. Eng. 15, 026014. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Penrose R (1955). A generalized inverse for matrices. Math. Proc. Camb. Phil. Soc. 51, 406–413. [Google Scholar]
  • 54.Pedregosa F, Varoquaux G, Gramfort A, Michel V, Thirion B, Grisel O, Blondel M, Prettenhofer P, Weiss R, Dubourg V, et al. (2011). Scikit-learn: Machine learning in Python. J. Mach. Learn. Res. 12, 2825–2830. [Google Scholar]
  • 55.Yao Z, Wu D, Wang X, Zhang B, Yu F, Yang C, Peng Z, Chen X, Xie L, and Lei X (2021). WeNet: Production oriented streaming and non-streaming end-to-end speech recognition toolkit. In Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH. [Google Scholar]
  • 56.Hunt AJ, and Black AW (1996). Unit selection in a concatenative speech synthesis system using a large speech database. In ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings. [Google Scholar]
  • 57.Eddy SR (2011). Accelerated profile HMM searches. PLoS Comput. Biol. 7, e1002195. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58.Graves A, Fernández S, Gomez F, and Schmidhuber J (2006). Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks. ACM International Conference Proceeding Series. [Google Scholar]
  • 59.Graves A, Liwicki M, Fernández S, Bertolami R, Bunke H, and Schmidhuber J (2009). A novel connectionist system for unconstrained handwriting recognition. IEEE Trans. Pattern Anal. Mach. Intell. 31. [DOI] [PubMed] [Google Scholar]
  • 60.Graves A, Mohamed AR, and Hinton G (2013). Speech recognition with deep recurrent neural networks. In ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings. [Google Scholar]
  • 61.Abadi M, Barham P, Chen J, Chen Z, Davis A, Dean J, Devin M, Ghemawat S, Irving G, Isard M, et al. (2016). In Proceedings of the 12th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2016. [Google Scholar]
  • 62.Fan C, Hahn N, Kamdar F, Avansino D, Wilson GH, Hochberg L, Shenoy KV, Henderson JM, and Willett FR (2023). Plug-and-Play Stability for Intracortical Brain-Computer Interfaces: A One-Year Demonstration of Seamless Brain-to-Text Communication. Adv. Neural Info. Process. Syst. 36, 42258–42270. [PMC free article] [PubMed] [Google Scholar]
  • 63.Karpowicz BM, Ali YH, Wimalasena LN, Sedler AR, Keshtkaran MR, Bodkin K, Ma X, Rubin DB, Williams ZM, Cash SS, et al. (2025). Stabilizing brain-computer interfaces through alignment of latent dynamics. Nature Communications 16, 4662. 10.1038/s41467-025-59652-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64.Jude J, Perich MG, Miller LE, and Hennig MH (2022). Robust alignment of cross-session recordings of neural population activity by behaviour via unsupervised domain adaptation. Proc. Mach. Learn. Res. [Google Scholar]
  • 65.Jude J, Perich MG, Miller LE, and Hennig MH (2023). Capturing cross-session neural population variability through self-supervised identification of consistent neuron ensembles. Proc. Mach. Learn. Res. [Google Scholar]
  • 66.Pun TK, Khoshnevis M, Hosman T, Wilson GH, Kapitonava A, Kamdar F, Henderson JM, Simeral JD, Vargas-Irwin CE, Harrison MT, and Hochberg LR (2024). Measuring instability in chronic human intracortical neural recordings towards stable, long-term brain-computer interfaces. Commun. Biol. 7, 1363. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 67.Farshchian A, Gallego JA, Miller LE, Solla SA, Cohen JP, and Bengio Y (2019). Adversarial domain adaptation for stable brain-machine interfaces. In 7th International Conference on Learning Representations, ICLR 2019. [Google Scholar]
  • 68.Veaux C, Yamagishi J, and MacDonald K (2019). Others Superseded-Cstr Vctk Corpus: English Multi-Speaker Corpus for Cstr Voice Cloning Toolkit (University of Edinburgh. The Centre for Speech Technology Research (CSTR)). [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

1
2
Download video file (96.9MB, mp4)
3
Download video file (147.9MB, mp4)
4
Download video file (39.6MB, mp4)

Data Availability Statement

RESOURCES