bioRxiv Science⌕ Search

Biology subjects

Dado, T.

Publications and source records attributed to Dado, T..

3 recordsLinked to original sources

PAM: Predictive attention mechanism for neural decoding of visual perception

In neural decoding, reconstruction seeks to create a literal image from information in brain activity, typically achieved by mapping neural responses to a latent representation of a generative model. A key challenge in this process is understanding how information is processed across visual areas to effectively integrate their neural signals. This requires an attention mechanism that selectively focuses on neural inputs based on their relevance to the task of reconstruction -- something conventional attention models, which capture only input-input relationships, cannot achieve. To address this, we introduce predictive attention mechanisms (PAMs), a novel approach that learns task-driven "output queries" during training to focus on the neural responses most relevant for predicting the latents underlying perceived images, effectively allocating attention across brain areas. We validate PAM with two datasets: (i) B2G, which contains GAN-synthesized images, their original latents and multiunit activity data; (ii) Shen-19, which includes real photographs, their inverted latents and functional magnetic resonance imaging data. Beyond achieving state-of-theart reconstructions, PAM offers a key interpretative advantage through the availability of (i) attention weights, revealing how the models focus was distributed across visual areas for the task of latent prediction, and (ii) values, capturing the stimulus information decoded from each area.

neuroscience↗

Brain2GAN: Feature-disentangled neural coding of visual perception in the primate brain

A challenging goal of neural coding is to characterize the neural representations underlying visual perception. To this end, multi-unit activity (MUA) of macaque visual cortex was recorded in a passive fixation task upon presentation of faces and natural images. We analyzed the relationship between MUA and latent representations of state-of-the-art deep generative models, including the conventional and feature-disentangled representations of generative adversarial networks (GANs) (i.e., z- and w-latents of StyleGAN, respectively) and language-contrastive representations of latent diffusion networks (i.e., CLIP-latents of Stable Diffusion). A mass univariate neural encoding analysis of the latent representations showed that feature-disentangled w representations outperform both z and CLIP representations in explaining neural responses. Further, w-latent features were found to be positioned at the higher end of the complexity gradient which indicates that they capture visual information relevant to high-level neural activity. Subsequently, a multivariate neural decoding analysis of the feature-disentangled representations resulted in state-of-the-art spatiotemporal reconstructions of visual perception. Taken together, our results not only highlight the important role of feature-disentanglement in shaping high-level neural representations underlying visual perception but also serve as an important benchmark for the future of neural coding. Author summaryNeural coding seeks to understand how the brain represents the world by modeling the relationship between stimuli and internal neural representations thereof. This field focuses on predicting brain responses to stimuli (neural encoding) and deciphering information about stimuli from brain activity (neural decoding). Recent advances in generative adversarial networks (GANs; a type of machine learning model) have enabled the creation of photorealistic images. Like the brain, GANs also have internal representations of the images they create, referred to as "latents". More recently, a new type of feature-disentangled "w-latent" of GANs has been developed that more effectively separates different image features (e.g., color; shape; texture). In our study, we presented such GAN-generated pictures to a macaque with cortical implants and found that the underlying w-latents were accurate predictors of high-level brain activity. We then used these w-latents to reconstruct the perceived images with high fidelity. The remarkable similarities between our predictions and the actual targets indicate alignment in how w-latents and neural representations represent the same stimulus, even though GANs have never been optimized on neural data. This implies a general principle of shared encoding of visual phenomena, emphasizing the importance of feature disentanglement in deeper visual areas.

neuroscience↗

Hyperrealistic neural decoding: Linear reconstruction of face stimuli from fMRI measurements via the GAN latent space

Neural decoding can be conceptualized as the problem of mapping brain responses back to sensory stimuli via a feature space. We introduce (i) a novel experimental paradigm which uses well-controlled yet highly naturalistic stimuli with a priori known feature representations and (ii) an implementation thereof for HYPerrealistic reconstruction of PERception (HYPER) of faces from brain recordings. To this end, we embrace the use of generative adversarial networks (GANs) at the earliest step of our neural decoding pipeline by acquiring fMRI data as subjects perceive face images synthesized by the generator network of a GAN. We show that the latent vectors used for generation effectively capture the same defining stimulus properties as the fMRI measurements. As such, GAN latent vectors can be used as features underlying the perceived images that can be predicted for (re-)generation, leading to the most accurate reconstructions of perception to date.

neuroscience↗