bioRxiv Science⌕ Search

Biology subjects

Fraczek, T.

Publications and source records attributed to Fraczek, T..

4 recordsLinked to original sources

Estimation of neuronal tuning for word meaning from passively recorded naturalistic speech

The ability to derive neural-level language coding models holds great scientific and clinical potential. Current approaches are limited by the scale and ethological validity of input data; applications requiring large, rare, or naturalistic samples in particular would benefit from the ability to infer neural coding from incidental everyday speech. Here we present a novel pipeline designed to leverage spontaneous and incidental naturalistic speech. This pipeline performs transcription, segmentation, and video-assisted diarization, as well as alignment and spike detection of neural data. We apply this pipeline to a dataset derived from 21 patients (6+ days each, over 800 hours and 5 million words total). We benchmark both encoding and decoding models against extensive and rare ground-truth control datasets consisting of human-curated word-level temporal alignment and manually sorted spikes. We further validate our approach by quantifying representational drift, effect of dataset size, and differences between six brain areas. Together, these findings demonstrate that incidental natural speech is sufficiently processed in the brain to enable the estimation neural-level embeddings.

neuroscience↗

A number simplex in the human medial temporal lobe

Humans handle numbers nimbly, suggesting a richer neural manifold structure than the prevalent mental number line model. In populations of medial temporal lobe (MTL) neurons in humans performing two simple tasks (dot counting and arithmetic), we find robust neural coding of numerosity that results in high dimensional, simplex-shaped manifolds. This shape affords more flexibility than a linear manifold due to its high shattering dimensionality and expressibility. Dot arrays and Arabic numerals evoked distinct simplicial population codes, yet they were linked by a linearly transferable latent structure within the same task. We find similar simplicial geometry of number representations in large language models (LLMs). Moreover, subjects internally computed arithmetic results were decodable during the calculation period, with decoding accuracy correlating with individual mathematical capacity. Finally, linear transformations of simplicial operand representations modeled the brains conversion of operands into decodable results, suggesting that the brains arithmetic procedures have some resemblance to the attention architecture of LLMs. Together, these findings establish a high dimensional representational foundation for numerical cognition in the brain.

neuroscience↗

Polysemanticity in human hippocampal neurons

To comprehend language, the brain must navigate a high-dimensional semantic landscape while seamlessly contextualizing meaning. Inspired by recent advances in the mechanistic interpretability of large language models (LLMs), we hypothesized that the brain utilizes polysemanticity, a coding strategy wherein individual neurons represent multiple semantically unrelated features through high-dimensional superposition (Elhage et al., 2022; Olah et al., 2020). We recorded single-unit activity from the human hippocampus during podcast listening. We found that hippocampal neurons exhibit dense semantic codes characterized by multiple tuning peaks with an overdispersed, isotropic geometry. This geometry satisfies the theoretical requirements for interference minimization in superimposed codes. Furthermore, semantic responses are strongly modulated by lexical and speaker-identity context; nonetheless, the underlying population geometry remains stable. This coding strategy permits rapid contextualization without requiring specialized, context-specific neurons. Indeed, we show clear pattern separation of similar terms, along with pattern completion for held-out words. Together, these results demonstrate that the human brain leverages superposition to solve a universal computational problem: maximizing semantic capacity within a constrained representational space.

neuroscience↗

Neural signatures of impaired semantic contextualization in Autism Spectrum Disorder

Social and communicative deficits are defining characteristics of autism spectrum disorder (ASD). Some theories suggest that these challenges, among other autistic traits, may arise from differences in predictive coding, or how the brain uses context to predict and interpret incoming information. This idea has the potential to link symptoms of autism to specific neurocomputational processes, and is especially promising for communication, whose impairment is a hallmark of ASD. Here we leveraged the ability of large language models (LLMs) to quantify semantic contextualization to analyze a unique dataset of responses from hippocampal neurons obtained during language listening in three mild-to-severe autistic individuals with comorbid epilepsy. Key elements of semantic coding were preserved in all three individuals with ASD: single-neuron response dynamics, representation of word-word semantic relationships, and patterns of context-dependent shifts in meaning. However, relative to controls, ASD resulted in reduced neural signatures of contextualization: (1) neuronal responses were aligned with earlier, less contextual layers of GPT-2, (2) ASD patients had lower effective dimensionality of the neural subspace predicting semantics, (3) neural representations of word meaning were less influenced by preceding context, and (4) neural signatures of lexical surprisal were reduced. Together, these results support theories of autism that emphasize impairments in contextualization, and highlight the power of LLMs as a tool for quantifying the computational basis of neurodevelopmental disorders.

neuroscience↗