bioRxiv Science⌕ Search

Biology subjects

Snoek, C. G. M.

Publications and source records attributed to Snoek, C. G. M..

2 recordsLinked to original sources

One Hundred Neural Networks and Brains Watching Videos: Lessons from Alignment

What can we learn from comparing video models to human brains, arguably the most efficient and effective video processing systems in existence? Our work takes a step towards answering this question by performing the first large-scale benchmarking of deep video models on representational alignment to the human brain, using publicly available models and a recently released video brain imaging (fMRI) dataset. We disentangle four factors of variation in the models (temporal modeling, classification task, architecture, and training dataset) that affect alignment to the brain, which we measure by conducting Representational Similarity Analysis across multiple brain regions and model layers. We show that temporal modeling is key for alignment to brain regions involved in early visual processing, while a relevant classification task is key for alignment to higher-level regions. Moreover, we identify clear differences between the brain scoring patterns across layers of CNNs and Transformers, and reveal how training dataset biases transfer to alignment with functionally selective brain areas. Additionally, we uncover a negative correlation of computational complexity to brain alignment. Measuring a total of 99 neural networks and 10 human brains watching videos, we aim to forge a path that widens our understanding of temporal and semantic video representations in brains and machines, ideally leading towards more efficient video models and more mechanistic explanations of processing in the human brain.

neuroscience↗

Shape-Biased Learning by Thinking Inside the Box

Convolutional Neural Networks (CNNs) surpass human-level performance on visual object recognition and detection, but their behavior still differs from human behavior in important ways. One prominent example is that CNNs trained on ImageNet exhibit an image texture bias, while humans exhibit a strong bias toward object shape. Although CNN shape bias can be increased in various ways, e.g., using data augmentation or additional training techniques, it remains unclear what causes the strong discrepancy between human and CNN object recognition strategies. Developmental research suggests that one factor driving human shape bias is that during early childhood, toddlers tend to fill their field-of-view with close-up objects. Here, we operationalize this close-up as a zoom-in on objects during CNN training which we show increases shape bias without any additional training or data augmentation. We provide further evidence for the advantage of closeup object vision by systematically manipulating the background-object ratio during CNN training, and demonstrate a strong (inverse) correlation with shape bias. Moreover, zooming-in on objects, thereby more closely emulating child vision, not only increases shape bias but also concurrently aligns classification accuracy and shape bias between humans and CNNs. Finally, we achieve a near human-like shape bias when using a developmentally-inspired background-object ratio for training and shape bias assessment. In sum, from a simple adjustment to common image datasets - zooming-in on objects - human-like shape bias can emerge. These results suggest that taking inspiration from human learning strategies is a promising avenue for building human-aligned, efficient, and more robust vision CNNs.

neuroscience↗