bioRxiv Science⌕ Search

bioRxiv · 10.1101/2023.11.05.564832

Large language models implicitly learn to straighten neural sentence trajectories to construct a predictive representation of natural language

Abstract

Predicting upcoming events is critical to our ability to effectively interact with our environment and conspecifics. In natural language processing, transformer models, which are trained on next-word prediction, appear to construct a general-purpose representation of language that can support diverse downstream tasks. However, we still lack an understanding of how a predictive objective shapes such representations. Inspired by recent work in vision neuroscience Henaff et al. (2019), here we test a hypothesis about predictive representations of autoregressive transformer models. In particular, we test whether the neural trajectory of a sequence of words in a sentence becomes progressively more straight as it passes through the layers of the network. The key insight behind this hypothesis is that straighter trajectories should facilitate prediction via linear extrapolation. We quantify straightness using a 1-dimensional curvature metric, and present four findings in support of the trajectory straightening hypothesis: i) In trained models, the curvature progressively decreases from the first to the middle layers of the network. ii) Models that perform better on the next-word prediction objective, including larger models and models trained on larger datasets, exhibit greater decreases in curvature, suggesting that this improved ability to straighten sentence neural trajectories may be the underlying driver of better language modeling performance. iii) Given the same linguistic context, the sequences that are generated by the model have lower curvature than the ground truth (the actual continuations observed in a language corpus), suggesting that the model favors straighter trajectories for making predictions. iv) A consistent relationship holds between the average curvature and the average surprisal of sentences in the middle layers of models, such that sentences with straighter neural trajectories also have lower surprisal. Importantly, untrained models dont exhibit these behaviors. In tandem, these results support the trajectory straightening hypothesis and provide a possible mechanism for how the geometry of the internal representations of autoregressive models supports next word prediction.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Hosseini, E. A., Fedorenko, E.. 2023-11-06. Large language models implicitly learn to straighten neural sentence trajectories to construct a predictive representation of natural language. https://doi.org/10.1101/2023.11.05.564832

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Functional validation of allele-specific LMNB1 silencing in patient-derived astrocytes as a therapeutic option for Autosomal Dominant Leukodystrophy

Adult-onset Autosomal Dominant Leukodystrophy (ADLD) is a rare fatal leukodystrophy caused by increased LMNB1 gene dosage, most commonly resulting from duplication of the LMNB1 locus. Because ADLD is a gene dosage disorder, selective reduction of pathological LMNB1 expression represents a rational therapeutic strategy. Although allele-specific RNA interference has previously been shown to lower LMNB1 levels in patient-derived fibroblasts and directly reprogrammed neurons, its therapeutic effects have not been evaluated in disease-relevant human glial cells or using functional efficacy endpoints. Here, we established human induced pluripotent stem cell-derived astrocytes from ADLD patients as a human glial model in which to validate allele-specific LMNB1 silencing across molecular, cellular, and functional readouts. ADLD astrocytes recapitulated increased LMNB1 expression and characteristic nuclear abnormalities and displayed transcriptional alterations affecting extracellular matrix organization, calcium homeostasis, metabolism and RNA processing. Functionally, these cells also exhibited functional phenotypes suitable for therapeutic evaluation: astrocyte-conditioned medium impaired the viability of both murine and human oligodendroglial cultures, while conditioned-medium and direct astrocyte-seeding paradigms revealed impaired post-lesion myelin recovery in lysolecithin-treated cerebellar organotypic slices. Allele-specific LMNB1 silencing restored physiological LMNB1 levels, corrected nuclear abnormalities, attenuated astrocyte-mediated oligodendroglial toxicity, improved post-lesion myelin recovery, and was associated with selective transcriptional programs associated with extracellular support and cholesterol metabolism. Together, these findings provide molecular, cellular, and functional validation of allele-specific LMNB1 dosage correction in patient-derived human astrocytes and offer key support for LMNB1-lowering strategies in disease-relevant human glial cells.

neuroscience↗

Perceptual integration of multisensory haptic, visual, and auditory feedback for roughness discrimination in augmented reality

Understanding how our different senses interact to shape our perception is essential to design realistic and immersive virtual and augmented reality (VR/AR) experiences. The present study investigated how roughness perception can be modulated through haptic, visual, and auditory cues in AR using a vibrotactile wristband. Participants compared virtual textures varying in vibration frequency/amplitude, visual grain size, and friction sound. Results revealed strong linear relationships between stimulus parameters and perceived roughness, with haptic frequency and visual cues driving the highest discrimination performance. Adding non-informative sensory feedback reduced perceptual sensitivity, acting as noise. Individual differences emerged: participants who rated haptic as the easiest modality showed greater sensitivity to haptic variations, while visual-reliant participants performed better with visual cues. We conclude that roughness in AR can be systematically manipulated, but is vulnerable to perceptual interference from irrelevant inputs, where our work provides actionable insights for implementing optimized and adaptive AR/VR interfaces.

neuroscience↗

Structural and functional MRI signatures of Gambling Disorder: a case-control study

Gambling disorder (GD) is a behavioural addiction that may help identify addiction-related neural features without the direct neurobiological effects of a primary substance of dependence. We examined regional grey matter volume (GMV) and resting-state functional connectivity (rsFC) in the same well-characterised sample. Eighteen men with GD and 21 matched healthy controls underwent high-resolution structural and resting-state functional MRI. GMV was quantified across 214 cortical and subcortical regions, and seed-based rsFC analyses focused on striatal subdivisions and mesocorticolimbic regions. Group differences were evaluated using permutation testing and cluster-corrected mixed-effects modelling. GD was associated with lower GMV in the ventromedial prefrontal cortex, orbitofrontal regions and other cortical and subcortical areas, alongside higher GMV in a subset of limbic and default-mode regions. Participants with GD also showed lower connectivity between the limbic striatum and the hippocampus, thalamus and putamen. In exploratory analyses, somatomotor connectivity was positively associated with gambling severity (Problem Gambling Severity Index: Spearman's rho = 0.71, p = 0.003, false-discovery-rate-adjusted q = 0.016). Structural and functional findings overlapped spatially in regions associated with valuation, memory, reward and habit formation, but regional GMV did not mediate group differences in rsFC. These findings are broadly consistent with corticostriatal models of GD and identify candidate circuit-level differences for independent replication. Larger, more diverse and longitudinal samples are required to establish their reproducibility, temporal direction and clinical relevance.

neuroscience↗