bioRxiv Science⌕ Search

bioRxiv · 10.64898/2025.12.10.693611

Systematic comparison of color representations between humans and deep neural networks: towards predicting human color perception in a vast color space

Abstract

The representational structure of large-scale human color perception remains incompletely understood. While classical studies measured numerous color pairs, these measurements compared only similar colors, and exploring the global relationships among thousands of colors has been infeasible due to the time costs of psychophysical experiments. Given these constraints, deep neural networks (DNNs) have attracted attention as a promising tool for providing proxies or predictions of human perception beyond the scope of psychophysical experiments. However, it remains unclear which DNNs possess embeddings that geometrically align with human color perception. Furthermore, it is unclear which learning paradigm enables DNNs to acquire a color representation that aligns with that of humans. Here, we systematically investigate which learning paradigm enables DNNs to produce a color representation that is structurally congruent with that of humans, with a focus on three types: self-supervised learning (SSL) that trains on images alone, supervised learning (SL) that trains on images with category labels, and contrastive language-image pre-training (CLIP) that trains on image-text pairs. We compared the embeddings of DNNs with the human similarity judgments of 93 colors using a rigorous unsupervised method termed Gromov-Wasserstein Optimal Transport (GWOT). Our results show that, while each learning paradigm acquires color representations that strongly align with human data at the fine-item level in early layers, only CLIP sustains such a representation at the output. Furthermore, when we leveraged a key advantage of DNNs and investigated the representational structure of 4096 colors, the early layers of each learning paradigm and the output of CLIP consistently converged on their own characteristic structures. These structures present plausible predictions for the large-scale human color representation. Our work demonstrates an approach for exploring unknown territories of human perception through the use of computational models validated in a limited empirical space, and provides predictions for future large-scale psychophysical experiments. Author summaryHow do we perceive the vast world of color? Despite extensive research into human color perception, studies evaluating many colors have mostly captured differences between similar colors, while those mapping global relationships are restricted to a few dozen. Consequently, we still do not know the global structure of the massive "color map" that might underlie our perception of thousands of colors, as testing this directly is practically impossible. To explore this space, we turned to deep neural networks, a form of AI. Our first step was to identify models that "see" color in a way that matches humans. We compared models against human data capturing the global relationships among all possible pairs of 93 colors. Using a powerful geometric comparison method, we found the models that matched the human color map. This allowed us to use these models as reliable computational proxies. We then used them to do what human experiments currently cannot: chart a vast global map of 4,096 colors. The human-aligned models consistently converged on two distinct structures. Our work provides the first plausible, testable predictions for the large-scale structure of human color perception and demonstrates a new way to explore otherwise unreachable territories of our perceptual world.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Wickramanayaka, N. R., Oizumi, M.. 2025-12-13. Systematic comparison of color representations between humans and deep neural networks: towards predicting human color perception in a vast color space. https://doi.org/10.64898/2025.12.10.693611

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Attention Across Scales: From Individual Variation to Social Hierarchies and Brain Networks in Semi-Free-Ranging Macaques

Attention is a fundamental brain function supporting perception, decision-making, and social behavior, and its dysfunction profoundly impairs daily life. It is both dynamic and stable, varying across observations and individuals, changing across the lifespan, and being shaped by social and environmental experience. Yet capturing this complexity remains a central challenge in neuroscience. Here, we integrated longitudinal behavioral assessments of semi-free-ranging macaques living in naturalistic social groups with resting-state fMRI. We quantified performance across days, ages, and social hierarchies and related it to intrinsic brain organization. Distinct attentional phenotypes emerged, including individuals with reduced attentional control. Performance followed an inverted-U lifespan trajectory, improving from childhood to adulthood before declining. Social status modulated attentional performance. Critically, nonlinear lifespan trajectories and associations with individual attentional differences were most clearly expressed in frontoparietal connectivity. Together, these findings reveal how sustained attention is organized across scales, providing a biological framework for its individual diversity, social modulation, and neural basis.

neuroscience↗

Decoding natural scenes from patterned optogenetic responses in mouse visual cortex

A central challenge in developing visual cortical prostheses is to determine how visual stimuli should be transformed into effective patterns of cortical stimulation. Although advances in stimulation technologies, including optogenetics, provide increasingly precise control over cortical activity, it remains unclear whether artificially evoked activity can reproduce the information content of naturally evoked visual representations. Here we establish a quantitative framework for evaluating visual encoding strategies by decoding cortical responses evoked by natural vision and patterned optogenetic stimulation. We developed a novel dual-modal paradigm in awake mice to bridge the gap between endogenous photostimulation and artificial network driving. By co-expressing the high-performance calcium indicator GCaMP6s and the red-shifted, ultra-sensitive opsin rsChRmine-oScarlet in the primary visual cortex (V1), we successfully translated dynamic natural movie frames into patterned, spatiotemporal optogenetic stimulation. Quantitative comparisons of macro-scale dynamics demonstrated that this patterned optogenetic injection evokes cortical states highly comparable and representationally aligned with those driven by actual visual photostimulation. To systematically evaluate the fidelity of these responses, we developed STAR, a deep learning model featuring spatial and temporal attention mechanisms, and successfully reconstructed the frames of natural movies from V1 signals under both experimental modalities. Collectively, our results demonstrate that complex sensory information can be both naturally encoded and synthetically injected into V1 circuits with high decoding fidelity. This work provides an empirical and computational proof-of-concept for intelligent, closed-loop biomimetic encoders, establishing a robust framework for next-generation cortical visual neuroprostheses and bidirectional brain-machine interfaces.

neuroscience↗

Why Is Spontaneous Blink Timing Informative? An Adaptive Scheduling Perspective

Spontaneous eye blinks have long been linked to cognitive processing, yet how task demands shape blink timing and its relationship to behavioral performance remains unclear. We examined spontaneous blink behavior in 576 adults performing two variants of the Continuous Performance Task (CPT). Blink occurrence and timing were most strongly modulated by the experimental condition in the more demanding CPT-AX task, whereas their association with response time was stronger in the CPT-X task, where more consistent blink timing predicted faster responses. This dissociation suggests that task structure changes not only blink behavior but also the behavioral relevance of blink timing. These findings are consistent with an adaptive scheduling account of spontaneous blinking and provide a conceptual framework for understanding when and why blink timing contains chronometric information about ongoing cognition.

neuroscience↗