bioRxiv Science⌕ Search

Biology subjects

Iacobacci, C.

Publications and source records attributed to Iacobacci, C..

12 recordsLinked to original sources

A generalizable speech neuroprosthesis

Intracortical brain-computer interfaces (BCIs) can restore communication to people with vocal tract paralysis by decoding cortical activity during attempted speech into text. State-of-the-art systems pairing neural-to-phoneme decoders with phoneme-to-word language models have achieved word error rates (WERs) as low as 1%, but only after collecting thousands of sentences of training data. Shortening the data collection process would facilitate scaling this new technology by reducing the time from device implant to high-accuracy communication. Here we introduce a transformer-based decoder model trained jointly across six intracortical speech BCI participants. For every participant -- regardless of sex, disease etiology, or attempted speaking strategy -- a multi-user model decoded speech more accurately (over 50% lower relative WER on average) than models trained on individual users data. Notably, the multi-user model could be finetuned on fewer than 200 sentences from a held-out user to achieve a WER below 7%. These results reveal how to pool intracortical data across people to yield more accurate, generalizable, and rapidly-deployable decoding models.

bioengineering↗

Brain2voice 2.0: High-performance voice synthesis brain-computer interface

Brain-computer interfaces (BCIs) offer a promising solution to speech loss due to neurological injury by decoding intended speech directly from brain activity. While recent BCIs have restored high-accuracy text-based communication, they fail to provide instantaneous voice output essential for the natural flow of conversation. Brain-to-voice BCIs address this gap by decoding voice directly from neural signals. However, even the state-of-the-art (SOTA) BCI-synthesized voice is not yet intelligible enough for real-world adoption. We introduce brain2voice 2.0, a new multimodal Transformer-based BCI decoder architecture capable of synthesizing highly intelligible voice from intracortical neural signals in real-time. Brain2voice 2.0 is trained on continuous and custom-tokenized acoustic targets and phoneme targets, leveraging their complementary speech information. We use self-supervised and adversarial training objectives that enhance acoustic feature quality and improve synthesis intelligibility. At each 10 ms timestep, the model causally outputs continuous and tokenized acoustic features for real-time voice synthesis as well as time-aligned phoneme predictions (raw phoneme error rate: 7%, comparable to the latest brain-to-text models). We evaluated this new approach on our prior intracortical brain-to-voice benchmark dataset (Wairagkar et al. 2025). Naive human listeners transcribed brain2voice 2.0 synthesized voice with a word error rate of 5.24%--an 8x improvement in intelligibility over previous SOTA results (43.75%). Brain2voice 2.0 demonstrates that highly intelligible real-time voice synthesis from neural signals is achievable, for the first time crossing the intelligibility threshold necessary for clinically viable brain-to-voice BCIs for people with paralysis.

neuroscience↗

Neural decoding of speech using deep neural ensembles

Speech brain-computer interfaces (BCIs) can restore rapid communication to people with paralysis, but decoding errors still limit performance. In recent brain-to-text decoding competitions, deep ensemble methods, which combine predictions from multiple independently trained decoders, have delivered striking accuracy improvements and account for the largest gains over baseline approaches. However, these methods have not previously been tested in real-time, require substantial computational resources, and their performance under various clinically relevant constraints remains poorly understood. Here, we present the first closed-loop test of deep ensembles in a participant with bilateral intracortical microelectrode arrays, demonstrating a reduction in word error rate from 33.7% to 26.0% on a large-vocabulary task. Using additional data from three participants, we then assess how these gains depend on baseline error rate, training dataset size, and ensemble size, including the resource-accuracy tradeoffs most relevant for real-world deployment. Finally, we introduce a computationally efficient pseudoensembling approach based on test-time augmentation that improves decoding accuracy while requiring only a single base decoder, greatly reducing the computational burden of ensembling. Together, these results show that the benefits of deep ensembling can be realized in real time and under practical resource constraints, bringing speech BCIs closer to broader clinical adoption.

neuroscience↗

Premotor cortex uses a compositional neural geometry to plan words

Speech requires precise serial ordering of words and phonemes into novel combinations. To accomplish this, the brain is believed to flexibly prepare utterances before producing them, even allowing pronunciation of never-before spoken words. To discover how neural populations achieve this, intracortical activity from premotor cortex was recorded while two speech neuroprosthesis pilot clinical trial participants attempted to speak factorially-balanced phoneme sequences. During preparation, activity encoded not only the next-phoneme, but multiple upcoming phoneme positions spanning whole words. We found that word-level plans were formed by compositionally combining phoneme representations, a mechanism that may enable efficient planning of novel sequences. When utterances contained more than one word, premotor cortex activity was largely limited to the first word, suggesting that articulatory planning is segmented by higher-order features. Together, these results reveal a compositional, hierarchically-segemented planning geometry, potentially a universal neural strategy for sequence organization across higher levels of language.

neuroscience↗

Cross-brain transfer of high-performance intracortical speech and handwriting BCIs

Intracortical brain-computer interfaces (BCIs) that decode complex movements, such as handwriting and speech, can require substantial training data to achieve high performance. We investigated whether leveraging the neural activity recordings of previous users could reduce this initial data collection burden for new BCI users (an approach we call "cross-brain transfer"). Using intracortical recordings from five BrainGate2 clinical trial participants, we tested cross-brain transfer for both speech and handwriting neural decoders trained and evaluated on general, unconstrained corpora of spoken and written English. We found that cross-brain transfer improved decoding performance when training data from the target user was limited (< 200 sentences), and that dataset-specific input layers to the decoder were critical for combining data across users. Without trainable input layers, transfer failed and performed worse than training from scratch on target user data only. Finally, we measured the effectiveness of cross-brain transfer relative to training with (1) more data from the same user and (2) more electrode-permuted data from the same user, which simulates sampling from another brain with identical neural latent structure. In some cases (T16 speech, T12 handwriting), cross-brain transfer appeared as effective as additional permuted data from the same user, while in others (T12 speech, T15 speech) electrode-permuted data was more beneficial. Our results successfully demonstrate and characterize cross-brain transfer learning between multiple intracortical BCI users, for both speech and handwriting, using a general open-ended dataset not restricted to small sets of words or phrases. This work highlights a promising path towards addressing a key barrier to the clinical translation of BCIs, while clarifying when cross-brain transfer may be most beneficial and the decoder design choices needed to realize those gains.

neuroscience↗

Long-term independent use of an intracortical brain-computer interface for speech and cursor control

Brain-computer interfaces (BCIs) can provide naturalistic communication and digital access to people with severe paralysis by decoding neural activity associated with attempted speech and movement. Recent work has demonstrated highly accurate intracortical BCIs for speech and cursor control, but two critical capabilities needed for practical viability were unmet: independent at-home operation without researcher assistance, and reliable long-term performance supporting accurate speech and cursor decoding. Here, we demonstrate the independent and near-daily use of a multimodal BCI with novel brain-to-text speech and computer cursor decoders by a man with paralysis and severe dysarthria due to amyotrophic lateral sclerosis (ALS). Over nearly two years, the participant used the BCI for more than 3,800 cumulative hours to maintain rich interpersonal communication with his family and friends, independently control his personal computer, and sustain full-time employment - despite being paralyzed. He communicated 183,060 sentences - totaling 1,960,163 words - at an average rate of 56.1 words per minute. He labeled 92.3% of sentences as being decoded at least mostly correctly. In formal quantifications of performance where he was asked to say words presented on a screen, attempted speech was consistently decoded with over 99% word accuracy (125,000 word vocabulary). The participant also used the speech BCI as keyboard input and the cursor BCI as mouse input to control his personal computer, enabling him to send text messages, emails, and to browse the internet. These results demonstrate that intracortical BCIs have the potential to support independent use in the home, marking a critical step toward practical assistive technology for people with severe motor impairment.

neuroscience↗

Error encoding in human speech motor cortex

Humans monitor their actions, including detecting errors during speech production. This self-monitoring capability also enables speech neuroprosthesis users to recognize mistakes in decoded output upon receiving visual or auditory feedback. However, it remains unknown whether neural activity related to error detection is present in the speech motor cortex. In this study, we demonstrate the existence of neural error signals in speech motor cortex firing rates during intracortical brain-to-text speech neuroprosthesis use. This activity could be decoded to enable the neuroprosthesis to identify its own errors with up to 86% accuracy. Additionally, we observed distinct neural patterns associated with specific types of mistakes, such as phonemic or semantic differences between the persons intended and displayed words. These findings reveal how feedback errors are represented within the speech motor cortex, and suggest strategies for leveraging these additional cognitive signals to improve neuroprostheses.

neuroscience↗

Encoding of speech modes and loudness in ventral precentral gyrus

The ability to vary the mode and loudness of speech is an important part of the expressive range of human vocal communication. However, the encoding of these behaviors in the ventral precentral gyrus (vPCG) has not been studied at the resolution of neuronal firing rates. We investigated this in two participants who had intracortical microelectrode arrays implanted in their vPCG as part of a speech neuroprosthesis clinical trial. Neuronal firing rates modulated strongly in vPCG as a function of attempted mimed, whispered, normal or loud speech. At the neural ensemble level, mode/loudness and phonemic content were encoded in distinct neural subspaces. Attempted mode/loudness could be decoded from vPCG with an accuracy of 94% and 89% for two participants respectively, and corresponding neural preparatory activity could be detected hundreds of milliseconds before speech onset. We then developed a closed-loop loudness decoder that achieved 94% online accuracy in modulating a brain-to-text speech neuroprosthesis output based on attempted loudness. These findings demonstrate the feasibility of decoding mode and loudness from vPCG, paving the way for speech neuroprostheses capable of synthesizing more expressive speech.

neuroscience↗

Speech motor cortex enables BCI cursor control and click

Decoding neural activity from ventral (speech) motor cortex is known to enable high-performance speech brain-computer interface (BCI) control. It was previously unknown whether this brain area could also enable computer control via neural cursor and click, as is typically associated with dorsal (arm and hand) motor cortex. We recruited a clinical trial participant with ALS and implanted intracortical microelectrode arrays in ventral precentral gyrus (vPCG), which the participant used to operate a speech BCI in a prior study. We developed a cursor BCI driven by the participants vPCG neural activity, and evaluated performance on a series of target selection tasks. The reported vPCG cursor BCI enabled rapidly-calibrating (40 seconds), accurate (2.90 bits per second) cursor control and click. The participant also used the BCI to control his own personal computer independently. These results suggest that placing electrodes in vPCG to optimize for speech decoding may also be a viable strategy for building a multi-modal BCI which enables both speech-based communication and computer control via cursor and click.

neuroscience↗

Representation of Verbal Thought in Motor Cortex and Implications for Speech Neuroprostheses

Speech brain-computer interfaces show great promise in restoring communication for people who can no longer speak1-3, but have also raised privacy concerns regarding their potential to decode private verbal thought4-6. Using multi-unit recordings in three participants with dysarthria, we studied the representation of inner speech in the motor cortex. We found a robust neural encoding of inner speech, such that individual words and continuously imagined sentences could be decoded in real-time This neural representation was highly correlated with overt and perceived speech. We investigated the possibility of "eavesdropping" on private verbal thought, and demonstrated that verbal memory can be decoded during a non-speech task. Nevertheless, we found a neural "overtness" dimension that can help to avoid any unintentional decoding. Together, these results demonstrate the strong representation of verbal thought in the motor cortex, and highlight important design considerations and risks that must be addressed as speech neuroprostheses become more widespread.

neuroscience↗

A mosaic of whole-body representations in human motor cortex

Understanding how the body is represented in motor cortex is key to understanding how the brain controls movement. The precentral gyrus (PCG) has long been thought to contain largely distinct regions for the arm, leg and face (represented by the "motor homunculus"). However, mounting evidence has begun to reveal a more intermixed, interrelated and broadly tuned motor map. Here, we revisit the motor homunculus using microelectrode array recordings from 20 arrays that broadly sample PCG across 8 individuals, creating a comprehensive map of human motor cortex at single neuron resolution. We found whole-body representations throughout all sampled points of PCG, contradicting traditional leg/arm/face boundaries. We also found two speech-preferential areas with a broadly tuned, orofacial-dominant area in between them, previously unaccounted for by the homunculus. Throughout PCG, movement representations of the four limbs were interlinked, with homologous movements of different limbs (e.g., toe curl and hand close) having correlated representations. Our findings indicate that, while the classic homunculus aligns with each areas preferred body region at a coarse level, at a finer scale, PCG may be better described as a mosaic of functional zones, each with its own whole-body representation.

neuroscience↗

An instantaneous voice synthesis neuroprosthesis

Brain computer interfaces (BCIs) have the potential to restore communication to people who have lost the ability to speak due to neurological disease or injury. BCIs have been used to translate the neural correlates of attempted speech into text1-3. However, text communication fails to capture the nuances of human speech such as prosody, intonation and immediately hearing ones own voice. Here, we demonstrate a "brain-to-voice" neuroprosthesis that instantaneously synthesizes voice with closed-loop audio feedback by decoding neural activity from 256 microelectrodes implanted into the ventral precentral gyrus of a man with amyotrophic lateral sclerosis and severe dysarthria. We overcame the challenge of lacking ground-truth speech for training the neural decoder and were able to accurately synthesize his voice. Along with phonemic content, we were also able to decode paralinguistic features from intracortical activity, enabling the participant to modulate his BCI-synthesized voice in real-time to change intonation, emphasize words, and sing short melodies. These results demonstrate the feasibility of enabling people with paralysis to speak intelligibly and expressively through a BCI.

neuroscience↗