bioRxiv Science⌕ Search

bioRxiv · 10.64898/2026.09.03.749141

Biologically grounded locality priors close the data gap for vision transformers in neural prediction

Abstract

For datasets with thousands of neurons and images, vision transformers have proven successful at predicting neural responses to stimuli. However, they are expected to underperform in low-data regimes, where CNNs and Gaussian processes are considered more effective. We ask whether transformers can be made competitive for small-scale neural prediction, and show that underperformance in this regime can be overturned with the right inductive bias. We equip a vision transformer with a differentiable per-neuron circular crop in feature space. The crop is centered on each neuron's receptive field, with a radius selected per neuron, so the model only sees the small image region that drives that neuron instead of the whole image. This makes the cost of attention scale with the size of the receptive field rather than with the size of the image. A single transformer stack is shared across all neurons: each neuron's specificity resides in the crop, not in the architecture. We call this model circular receptive-field vision transformer (CiRF-ViT). We evaluate it on small multi-electrode-array recordings of mouse and salamander retinas: a mouse preparation of 41 ganglion cells and two salamander preparations totaling 49 ganglion cells, each with only a few thousand stimulus-response pairs, two orders of magnitude below the scale at which transformers are typically trained. Against CNN and Gaussian-process baselines, CiRF-ViT reaches the highest mean explained variance on both datasets (0.94 on mouse, 0.95 on salamander). Probed with the local spike-triggered average (LSTA), a zero-shot test of context-dependent, nonlinear stimulus sensitivity, CiRF-ViT reproduces the qualitative polarity inversion that a linear model cannot capture by construction. A biologically grounded, per-neuron locality prior is therefore enough to make transformers competitive for neural prediction well below their usual data scale, while matching the reference models on an established functional signature of the retina's nonlinear computation.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Farina, M., zamberlan, p., Onken, A., Ferrari, U.. 2026-09-08. Biologically grounded locality priors close the data gap for vision transformers in neural prediction. https://doi.org/10.64898/2026.09.03.749141

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Attention Across Scales: From Individual Variation to Social Hierarchies and Brain Networks in Semi-Free-Ranging Macaques

Attention is a fundamental brain function supporting perception, decision-making, and social behavior, and its dysfunction profoundly impairs daily life. It is both dynamic and stable, varying across observations and individuals, changing across the lifespan, and being shaped by social and environmental experience. Yet capturing this complexity remains a central challenge in neuroscience. Here, we integrated longitudinal behavioral assessments of semi-free-ranging macaques living in naturalistic social groups with resting-state fMRI. We quantified performance across days, ages, and social hierarchies and related it to intrinsic brain organization. Distinct attentional phenotypes emerged, including individuals with reduced attentional control. Performance followed an inverted-U lifespan trajectory, improving from childhood to adulthood before declining. Social status modulated attentional performance. Critically, nonlinear lifespan trajectories and associations with individual attentional differences were most clearly expressed in frontoparietal connectivity. Together, these findings reveal how sustained attention is organized across scales, providing a biological framework for its individual diversity, social modulation, and neural basis.

neuroscience↗

Decoding natural scenes from patterned optogenetic responses in mouse visual cortex

A central challenge in developing visual cortical prostheses is to determine how visual stimuli should be transformed into effective patterns of cortical stimulation. Although advances in stimulation technologies, including optogenetics, provide increasingly precise control over cortical activity, it remains unclear whether artificially evoked activity can reproduce the information content of naturally evoked visual representations. Here we establish a quantitative framework for evaluating visual encoding strategies by decoding cortical responses evoked by natural vision and patterned optogenetic stimulation. We developed a novel dual-modal paradigm in awake mice to bridge the gap between endogenous photostimulation and artificial network driving. By co-expressing the high-performance calcium indicator GCaMP6s and the red-shifted, ultra-sensitive opsin rsChRmine-oScarlet in the primary visual cortex (V1), we successfully translated dynamic natural movie frames into patterned, spatiotemporal optogenetic stimulation. Quantitative comparisons of macro-scale dynamics demonstrated that this patterned optogenetic injection evokes cortical states highly comparable and representationally aligned with those driven by actual visual photostimulation. To systematically evaluate the fidelity of these responses, we developed STAR, a deep learning model featuring spatial and temporal attention mechanisms, and successfully reconstructed the frames of natural movies from V1 signals under both experimental modalities. Collectively, our results demonstrate that complex sensory information can be both naturally encoded and synthetically injected into V1 circuits with high decoding fidelity. This work provides an empirical and computational proof-of-concept for intelligent, closed-loop biomimetic encoders, establishing a robust framework for next-generation cortical visual neuroprostheses and bidirectional brain-machine interfaces.

neuroscience↗

Why Is Spontaneous Blink Timing Informative? An Adaptive Scheduling Perspective

Spontaneous eye blinks have long been linked to cognitive processing, yet how task demands shape blink timing and its relationship to behavioral performance remains unclear. We examined spontaneous blink behavior in 576 adults performing two variants of the Continuous Performance Task (CPT). Blink occurrence and timing were most strongly modulated by the experimental condition in the more demanding CPT-AX task, whereas their association with response time was stronger in the CPT-X task, where more consistent blink timing predicted faster responses. This dissociation suggests that task structure changes not only blink behavior but also the behavioral relevance of blink timing. These findings are consistent with an adaptive scheduling account of spontaneous blinking and provide a conceptual framework for understanding when and why blink timing contains chronometric information about ongoing cognition.

neuroscience↗