bioRxiv Science⌕ Search

bioRxiv · 10.1101/2025.04.03.646971

Human-like monocular depth biases in deep neural networks

Abstract

Human depth perception from 2D images is systematically distorted, yet the nature of these distortions is not fully understood. To gain insights into this fundamental problem, we compare human depth judgments with those of deep neural networks (DNNs), which have shown remarkable abilities in monocular depth estimation. Using a novel human-annotated dataset of natural indoor scenes and a systematic analysis of absolute depth judgments, we investigate error patterns in both humans and DNNs. Employing exponential-affine fitting, we decompose depth estimation errors into depth compression, per-image affine transformations (including scaling, shearing, and translation), and residual errors. Our analysis reveals that human depth judgments exhibit systematic and consistent biases, including depth compression, a vertical bias (perceiving objects in the lower visual field as closer), and consistent per-image affine distortions across participants. Intriguingly, we find that DNNs with higher accuracy partially recapitulate these human biases, demonstrating greater similarity in affine parameters and residual error patterns. This suggests that these seemingly suboptimal human biases may reflect efficient, ecologically adapted strategies for depth inference from inherently ambiguous monocular images. However, while DNNs capture metric-level residual error patterns similar to humans, they fail to reproduce human-level accuracy in ordinal depth perception within the affine-invariant space. These findings underscore the importance of evaluating error patterns beyond raw accuracy, providing new insights into how humans and computational models resolve depth ambiguity. Our dataset and methodology provide a framework for evaluating the alignment between computational models and human perceptual biases, thereby advancing our understanding of visual space representation and guiding the development of models that more faithfully capture human depth perception. Author summaryUnderstanding the characteristics of errors in depth judgments exhibited by humans and deep neural networks (DNNs) provides a foundation for developing functional models of human brain and artificial models with enhanced interpretability. To address this, we constructed a human depth judgment dataset using indoor photographs and compared human depth judgments with those of DNNs. Our results show that humans systematically compress far distances and exhibit distortions related to viewpoint shift, which remain remarkably consistent across observers. Strikingly, the better the DNNs were at depth estimation, the more they also exhibited human-like biases. This suggests that these seemingly suboptimal human biases could in fact reflect efficient strategies for inferring 3D structure from ambiguous 2D inputs. However, we also found a limit: while DNNs mimicked some human errors, they werent as good as humans at judging the relative order of objects in depth, especially when we accounted for viewpoint distortions. We believe that our dataset and discovery of multiple error factors will drive further comparative studies between humans and DNNs, facilitating model evaluations that go beyond simple accuracy to uncover how depth perception truly works--and how it might best be replicated in computational models.

Source connections

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Kubota, Y., Fukiage, T.. 2025-04-03. Human-like monocular depth biases in deep neural networks. https://doi.org/10.1101/2025.04.03.646971

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Cofilin Suppresses Tau-Induced Defects in Dense-Core Granule Formation and Aβ-Induced Neurodegeneration

Intracellular neurofibrillary tangles formed from hyperphosphorylated tau and extracellular amyloid plaques containing aggregated A{beta}-peptides, specific cleavage products of the Amyloid Precursor Protein (APP), are the primary histopathological hallmarks of Alzheimers Disease (AD), the leading cause of dementia in humans. However, the initiating steps that lead to these pathologies and early neurodegeneration, and the mechanisms by which tau- and A{beta}-induced effects might be linked remain unclear. Using the prostate-like secondary cell (SC) in Drosophila, we recently showed that A{beta} modulates normal APP- and membrane-associated protein aggregation in the dense-core granule (DCG) compartments of the regulated secretory pathway by interfering with subsequent membrane:DCG dissociation. This disrupts endolysosomal trafficking and propagates the resulting endolysosomal defects to other cells that endocytose the secreted abnormal DCG proteins. Here we show that overexpressing human tau also disrupts DCG aggregation and membrane:DCG dissociation inside SC secretory compartments, leading to increased endolysosomal targeting of these compartments. In a genetic screen, we find that knockdown of cofilin, which encodes an actin-severing protein required for dynamic remodelling of microfilaments, generates a similar phenotype. Consistent with this, overexpression of Cofilin, which is known to suppress tau-induced neurodegeneration in flies, reduces tau-induced DCG defects in SCs. Indeed, we find that Cofilin overexpression also suppresses A{beta}-induced degeneration in the fly eye. We conclude that membrane:DCG aggregate dissociation in DCG compartments is disrupted by both tau- and A{beta}-induced genetic changes that are relevant to AD, and this partially involves inhibition of actin cytoskeleton dynamics. Increasing actin remodelling activity can suppress neurodegeneration induced by both tau and A{beta}, suggesting that this process provides an important functional link between them that might be targeted therapeutically.

neuroscience↗

Lactate Promotes an Anti-Inflammatory Phenotype in Activated Microglia

Microglial activation is a central component of neuroinflammatory responses in many brain pathologies. Increasing evidence indicates that microglial phenotype is tightly linked to cellular metabolism, with pro-inflammatory activation associated with enhanced glycolytic flux. Lactate, traditionally considered a metabolic substrate, has recently emerged as a signaling molecule capable of modulating immune responses. However, its direct impact on microglial inflammatory activation remains incompletely understood. In the present study, we investigated the effects of lactate on microglial phenotype under inflammatory conditions using primary rat microglial cultures stimulated with lipopolysaccharide (LPS). Microglial activation was assessed through the expression of phenotypic markers, cytokine production, and secreted chemokine profiles. LPS stimulation induced a strong pro-inflammatory response characterized by increased CD86 expression, elevated TNF-alpha secretion, and enhanced release of several pro-inflammatory chemokines. Post-treatment with sodium L-lactate significantly attenuated these inflammatory responses, reducing pro-inflammatory marker expression and cytokine secretion, while restoring the anti-inflammatory marker CD206. To explore the relevance of these findings in a pathological context, the effects of lactate were further examined in a neonatal rat model of hypoxia-ischemia. Sodium L-lactate administration after injury reduced microglial activation and promoted a shift toward an anti-inflammatory phenotype in cortical regions, whereas hippocampal microglia showed a more limited response. Together, these results demonstrate that lactate directly modulates microglial inflammatory activation and cytokine production in vitro and suggest that lactate-mediated metabolic signaling may contribute in vivo to the regulation of neuroinflammatory responses.

neuroscience↗

Different hippocampal subfield volumes predict source memory performance and general cognitive ability in an adult lifespan sample

Modest positive associations between episodic memory performance and whole hippocampal and hippocampal subfield volumes have been reported in numerous prior studies. A smaller number of studies have reported associations between hippocampal volume and performance on tests of non-mnemonic cognition. The present study examined whether these associations were evident in a lifespan sample of cognitively healthy adults. Of particular interest was whether any identified associations were sensitive to age, and whether associations between subfield volumes and mnemonic and non-mnemonic performance were subfield dependent. We acquired high-resolution T1- and T2-weighted structural images from 163 adults (18-87 years of age). Participants also undertook a comprehensive neuropsychological test battery and an in-scanner test of source memory. Principal components analysis was employed to reduce the neuropsychological test scores to 5 cognitive components. Two components reflected memory performance while the other three reflected different aspects of non-mnemonic cognition. Hippocampal subfields (Cornu Ammonis (CA)1, CA2-3, dentate gyrus (DG) and subiculum) were segmented and measured with the Automated Segmentation of Hippocampus Subfields (ASHS) package. Source memory performance was selectively associated across participants with CA2-3 volume. By contrast, both mnemonic and non-mnemonic component scores derived from the test battery were associated exclusively with the volume of the DG. All associations were age-invariant. The findings indicate that different cognitive domains can be dissociated by virtue of their associations with different hippocampal subfields. Of importance, these associations appear to be life-long and hence are unlikely to reflect individual differences in age-related decline in structural integrity.

neuroscience↗