bioRxiv Science⌕ Search

Biology subjects

Abramovich Krasa, B.

Publications and source records attributed to Abramovich Krasa, B..

3 recordsLinked to original sources

Neural decoding of speech using deep neural ensembles

Speech brain-computer interfaces (BCIs) can restore rapid communication to people with paralysis, but decoding errors still limit performance. In recent brain-to-text decoding competitions, deep ensemble methods, which combine predictions from multiple independently trained decoders, have delivered striking accuracy improvements and account for the largest gains over baseline approaches. However, these methods have not previously been tested in real-time, require substantial computational resources, and their performance under various clinically relevant constraints remains poorly understood. Here, we present the first closed-loop test of deep ensembles in a participant with bilateral intracortical microelectrode arrays, demonstrating a reduction in word error rate from 33.7% to 26.0% on a large-vocabulary task. Using additional data from three participants, we then assess how these gains depend on baseline error rate, training dataset size, and ensemble size, including the resource-accuracy tradeoffs most relevant for real-world deployment. Finally, we introduce a computationally efficient pseudoensembling approach based on test-time augmentation that improves decoding accuracy while requiring only a single base decoder, greatly reducing the computational burden of ensembling. Together, these results show that the benefits of deep ensembling can be realized in real time and under practical resource constraints, bringing speech BCIs closer to broader clinical adoption.

neuroscience↗

Premotor cortex uses a compositional neural geometry to plan words

Speech requires precise serial ordering of words and phonemes into novel combinations. To accomplish this, the brain is believed to flexibly prepare utterances before producing them, even allowing pronunciation of never-before spoken words. To discover how neural populations achieve this, intracortical activity from premotor cortex was recorded while two speech neuroprosthesis pilot clinical trial participants attempted to speak factorially-balanced phoneme sequences. During preparation, activity encoded not only the next-phoneme, but multiple upcoming phoneme positions spanning whole words. We found that word-level plans were formed by compositionally combining phoneme representations, a mechanism that may enable efficient planning of novel sequences. When utterances contained more than one word, premotor cortex activity was largely limited to the first word, suggesting that articulatory planning is segmented by higher-order features. Together, these results reveal a compositional, hierarchically-segemented planning geometry, potentially a universal neural strategy for sequence organization across higher levels of language.

neuroscience↗

Improved interpretability in LFADS models using a learned, context-dependent per-trial bias

The computation-through-dynamics perspective argues that biological neural circuits process information via the continuous evolution of their internal states. Inspired by this perspective, Latent Factor Activity using Dynamical systems (LFADS, [1]) identifies a generative model consistent with the neural activity recordings. LFADS models neural dynamics with a recurrent neural network (RNN) generator, which results in excellent fit to the data. However, it has been difficult to understand the dynamics of the LFADS generator. In this work, we show that this poor interpretability arises in part because the generator implements complex, multi-stable dynamics. We introduce a simple modification to LFADS that ameliorates issues with interpretability by providing an inferred per-trial bias (modeled as a constant input) to the RNN generator, enabling it to contextually adapt a simpler dynamical system to individual trials. In both simulated neural recordings from pendulum oscillations and real recordings during arm movements in nonhuman primates, we observed that the standard LFADS learned complex, multi-stable dynamics, whereas the modified LFADS learned easier-to-understand contextual dynamics. This enabled direct analysis of the generator, which reproduced at a single-trial level previous results shown only through more complex analyses at the trial average. Finally, we applied the per-trial inferred bias LFADS model to human intracortical brain computer interface recordings during attempted finger movements and speech. We show that modifying neural dynamics using linear operations of the per-trial bias addresses non-stationarity and identifies the extent of behavioral variability, problems known to plague BCI. We call our modification to LFADS as "contextual LFADS".

neuroscience↗