bioRxiv ScienceSearch

Biology subjects

Sabeti, P. C.

Publications and source records attributed to Sabeti, P. C..

4 recordsLinked to original sources

Comparative evidence for the independent evolution of hair and sweat gland traits in primates

Humans differ in many respects from other primates, but perhaps no derived human feature is more striking than our naked skin. Long purported to be adaptive, humans unique external appearance is characterized by changes in both the patterning of hair follicles and eccrine sweat glands, producing decreased hair cover and increased sweat gland density. Despite the conspicuousness of these features and their potential evolutionary importance, there is a lack of clarity regarding how they evolved within the primate lineage. We thus collected and quantified the density of hair follicles and eccrine sweat glands from five regions of the skin in three species of primates: macaque, chimpanzee and human. Although human hair cover is greatly attenuated relative to that of our close relatives, we find that humans have a chimpanzee-like hair density that is significantly lower than that of macaques. In contrast, eccrine gland density is on average 10-fold higher in humans compared to chimpanzees and macaques, whose density is strikingly similar. Our findings suggest that a decrease in hair density in the ancestors of humans and apes was followed by an increase in eccrine gland density and a reduction in fur cover in humans. This work answers longstanding questions about the traits that make human skin unique and substantiates a model in which the evolution of expanded eccrine gland density was exclusive to the human lineage.

evolutionary biology

Co-circulating mumps lineages at multiple geographic scales

Despite widespread vaccination, eleven thousand mumps cases were reported in the United States (US) in 2016-17, including hundreds in Massachusetts, primarily in college settings. We generated 203 whole genome mumps virus (MuV) sequences from Massachusetts and 15 other states to understand the dynamics of mumps spread locally and nationally, as well as to search for variants potentially related to vaccination. We observed multiple MuV lineages circulating within Massachusetts during 2016-17, evidence for multiple introductions of the virus to the state, and extensive geographic movement of MuV within the US on short time scales. We found no evidence that variants arising during this outbreak contributed to vaccine escape. Combining epidemiological and genomic data, we observed multiple co-circulating clades within individual universities as well as spillover into the local community. Detailed data from one well-sampled university allowed us to estimate an effective reproductive number within that university significantly greater than one. We also used publicly available small hydrophobic (SH) gene sequences to estimate migration between world regions and to place this outbreak in a global context, but demonstrate that these short sequences, historically used for MuV genotyping, are inadequate for tracing detailed transmission. Our findings suggest continuous, often undetected, circulation of mumps both locally and nationally, and highlight the value of combining genomic and epidemiological data to track viral disease transmission at high resolution.

genomics

Identifying Gene Expression Programs of Cell-type Identity and Cellular Activity with Single-Cell RNA-Seq

Identifying gene expression programs underlying both cell-type identity and cellular activities (e.g. life-cycle processes, responses to environmental cues) is crucial for understanding the organization of cells and tissues. Although single-cell RNA-Seq (scRNA-Seq) can quantify transcripts in individual cells, each cells expression profile may be a mixture of both types of programs, making them difficult to disentangle. Here we illustrate and enhance the use of matrix factorization as a solution to this problem. We show with simulations that a method that we call consensus non-negative matrix factorization (cNMF) accurately infers identity and activity programs, including the relative contribution of programs in each cell. Applied to published brain organoid and visual cortex scRNA-Seq datasets, cNMF refines the hierarchy of cell types and identifies both expected (e.g. cell cycle and hypoxia) and intriguing novel activity programs. We propose that one of the novel programs may reflect a neurosecretory phenotype and a second may underlie the formation of neuronal synapses. We make cNMF available to the community and illustrate how this approach can provide key insights into gene expression variation within and between cell types.

bioinformatics

Prognostic models for Ebola virus disease derived from data collected at five treatment units in Sierra Leone and Liberia: performance, external validation, and risk visualization

BackgroundWe created a family of prognostic models for Ebola virus disease from the largest dataset of EVD patients published to date. We incorporated these models into an app, \"Ebola Care Guidelines\", that provides access to recommended, evidence-based supportive care guidelines and highlights the signs/symptoms with the largest contribution to prognosis.\n\nMethodsWe applied multivariate logistic regression on 470 patients admitted to five Ebola treatment units in Liberia and Sierra Leone during the 2014-16 outbreak. We validated the models with two independent datasets from Sierra Leone.\n\nFindingsViral load and age were the most important predictors of death. We generated a parsimonious model including viral load, age, body temperature, bleeding, jaundice, dyspnea, dysphagia, and referral time recorded at triage. We also constructed fallback models for when variables in the parsimonious model are unavailable. The performance of the parsimonious model approached the predictive power of observational wellness assessments by experienced health workers, with Area Under the Curve (AUC) ranging from 0.7 to 0.8 and overall accuracy of 64% to 74%.\n\nInterpretationMachine-learning models and mHealth tools have the potential for improving the standard of care in low-resource settings and emergency scenarios, but data incompleteness and lack of generalizable models are major obstacles. We showed how harmonization of multiple datasets yields prognostic models that can be validated across different cohorts. Similar performance between the parsimonious model and those incorporating expert wellness assessments suggests that clinically-guided machine learning approaches can recapitulate clinical expertise, and thus be useful when such expertise is unavailable. We also demonstrated with our guidelines app how integration of those models with mobile technologies enables deployable clinical management support tools that facilitate access to comprehensive bodies of medical knowledge.\n\nFundingHoward Hughes Medical Institute, US National Institutes of Health

bioinformatics