bioRxiv ScienceSearch

Biology subjects

Ghebreab, S.

Publications and source records attributed to Ghebreab, S..

2 recordsLinked to original sources

Scene complexity modulates degree of feedback activity during object recognition in natural scenes

Object recognition is thought to be mediated by rapid feed-forward activation of object-selective cortex, with limited contribution of feedback. However, disruption of visual evoked activity beyond feed-forward processing stages has been demonstrated to affect object recognition performance. Here, we unite these findings by reporting that the detection of target objects in natural scenes is selectively characterized by enhanced feedback when these objects are embedded in high complexity scenes. Human participants performed an animal target detection task on scenes with low, medium or high complexity as determined by a biologically plausible computational model of low-level contrast statistics. Three converging lines of evidence indicate that feedback was enhanced during categorization of scenes with high, but not low or medium complexity. First, functional magnetic resonance imaging (fMRI) activity in early visual cortex (V1) was selectively enhanced for target objects in scenes with high complexity. Second, event-related potentials (ERPs) evoked by high complexity scenes were selectively enhanced from 220 ms after stimulus-onset. Third, behavioral performance deteriorated for highly complex scenes when participants were pressed for time, but not when they could process the scenes fully and thereby benefit from the enhanced feedback. Formal modeling of the reaction time distributions revealed that object information accumulated more slowly for high complexity scenes (resulting in more errors especially for fast decisions), and directly related to the build-up of the feedback activity that was observed exclusively for high complexity scenes. Together, these results suggest that while feed-forward activity may suffice for simple scenes, the brain employs recurrent processing more adaptively in naturalistic settings, using minimal feedback for sparse, coherent scenes and increasing feedback for complex, fragmented scenes.\n\nAuthor summaryHow much neural processing is required to detect objects of interest in natural scenes? The astonishing speed of object recognition suggests that fast feed-forward buildup of perceptual activity is sufficient. However, this view is contradicted by findings that show that disruption of slower neural feedback leads to decreased detection performance. Our study unites these discrepancies by identifying scene complexity as a critical driver of neural feedback. We show how feedback is enhanced for complex, cluttered scenes compared to simple, well-organized scenes. Moreover, for complex scenes, more feedback is associated with better performances. These findings relate the flexibility of neural processes to perceptual decision-making by demonstrating that the brain dynamically directs neural resources based on the complexity of real-world visual inputs.

neuroscience

Characterizing the temporal dynamics of object recognition by deep neural networks: role of depth

Convolutional neural networks (CNNs) have recently emerged as promising models of human vision based on their ability to predict hemodynamic brain responses to visual stimuli measured with functional magnetic resonance imaging (fMRI). However, the degree to which CNNs can predict temporal dynamics of visual object recognition reflected in neural measures with millisecond precision is less understood. Additionally, while deeper CNNs with higher numbers of layers perform better on automated object recognition, it is unclear if this also results into better correlation to brain responses. Here, we examined 1) to what extent CNN layers predict visual evoked responses in the human brain over time and 2) whether deeper CNNs better model brain responses. Specifically, we tested how well CNN architectures with 7 (CNN-7) and 15 (CNN-15) layers predicted electro-encephalography (EEG) responses to several thousands of natural images. Our results show that both CNN architectures correspond to EEG responses in a hierarchical spatio-temporal manner, with lower layers explaining responses early in time at electrodes overlying early visual cortex, and higher layers explaining responses later in time at electrodes overlying lateral-occipital cortex. While the explained variance of neural responses by individual layers did not differ between CNN-7 and CNN-15, combining the representations across layers resulted in improved performance of CNN-15 compared to CNN-7, but only after 150 ms after stimulus-onset. This suggests that CNN representations reflect both early (feed-forward) and late (feedback) stages of visual processing. Overall, our results show that depth of CNNs indeed plays a role in explaining time-resolved EEG responses.

neuroscience