bioRxiv ScienceSearch

Biology subjects

MacArthur, B. D.

Publications and source records attributed to MacArthur, B. D..

4 recordsLinked to original sources

GenePy - a score for estimating gene pathogenicity in individuals using next-generation sequencing data

NGS is a revolutionising diagnosis and treatment of rare diseases. However, its relatively modest application in common diseases is limited by analytical approaches.\n\nInstead of variant-level approaches, typical for rare disease or large cohort analyses, contemporary investigation of common polygenic disorders requires the development of tools combining mutational burden and biological impact of a personalised set of mutations into single gene scores. GenePy (https://github.com/UoS-HGIG/GenePy) is a gene score for transforming sequencing data capable of estimating whole-gene pathogenicity on a per-patient basis.\n\nGenePy implements known deleteriousness metrics, incorporates allele frequency and individual zygosity information. Individuals harbouring multiple rare highly deleterious mutations accumulate extreme gene scores while the majority of genes usually achieve very low scores. Following correction for gene length, GenePy intuitively prioritises genes within individuals and affords gene/pathway score comparison between groups of individuals. Herein, we generate GenePy scores from whole-exome sequencing data for [~]15,000 genes across a cohort of 508 individuals. We describe score attributes and model behaviour under various biological conditions.\n\nWe demonstrate proof of concept that GenePy sensitively identifies known causal genes by calculating GenePy scores for NOD2 (an established causal Crohns Disease gene), in a modest cohort of patients for comparison against controls. This test of GenePy using a positive control gene demonstrates markedly more significant results (p=1.37 x 10-4) compared to the most commonly applied tool for combining common and rare variation.\n\nIn addition to increasing the biological information content for each variant, the per gene-per individual nature of GenePy transforms the utility of sequencing data. GenePy scores are intuitive when assessing for individual patients or for comparing between groups. Because GenePy intrinsically reflects pathogenicity at the gene level, this specifically facilitates downstream data integration (e.g. into machine learning, network and topological analyses) with transcriptomic and proteomic data that also report at the gene level.\n\nAuthor SummaryRapid technological advances have made DNA sequencing an effective, economic tool for detecting genomic variation. Detecting rare variation at the individual level is proving very successful in identifying the genetic causes of disease when just single mutations are sufficient to manifest disease. However, interpreting genomic data is much less straightforward for common diseases such as asthma, arthritis or heart disease where many genetic changes across multiple genes combine with the environment to bring about disease symptoms.\n\nWe have developed a new scoring system called GenePy that generates whole gene pathogenicity scores for indiviual patients. The score corrects for the length of the gene and is intuitive to use. Unlike many mutation deleteriousness metrics, GenePy also takes into account the population frequency of the variant and the number of copies of any given mutation and combines data for as many variants as are present in a given gene for any one individual.\n\nIn this paper we apply the GenePy scoring system to a cohort of over 500 individuals for whom we have sequencing data across all genomics regions that code for protein. We descibe how GenePy performs and demonstrate superior sensitivity to detect known causal genes in a common autoimmune condition.

genomics

Pattern analysis of pluripotency reveals the molecular basis of naïve, primed and formative states

The molecular regulatory network underlying stem cell pluripotency has been intensively studied, and we now have a reliable ensemble model for the average pluripotent cell. However, evidence of significant cell-to-cell variability suggests that the activity of this network varies within individual stem cells, leading to differential processing of environmental signals and variability in cell fates. Here, we adapt a method originally designed for face recognition to infer regulatory network patterns within individual cells from single-cell expression data. Using this method we identify three distinct network configurations in cultured mouse embryonic stem cells - corresponding to naive and formative pluripotent states and an early primitive endoderm state - and associate these configurations with particular combinations of regulatory network activity archetypes that govern different aspects of the cells response to environmental stimuli, cell cycle status and core information processing circuitry. These results show how variability in cell identities arise naturally from alterations in underlying regulatory network dynamics and demonstrate how methods from machine learning may be used to better understand single cell biology, and the collective dynamics of cell communities.

systems biology

Information Theory and Stem Cell Biology

Purpose of ReviewTo outline how ideas from Information Theory may be used to analyze single cell data and better understand stem cell behaviour.\n\nRecent findingsRecent technological breakthroughs in single cell profiling have made it possible to interrogate cell-to-cell variability in a multitude of contexts, including the role it plays in stem cell dynamics. Here we review how measures from information theory are being used to extract biological meaning from the complex, high-dimensional and noisy datasets that arise from single cell profiling experiments. We also discuss how concepts linking information theory and statistical mechanics are being used to provide insight into cellular identity, variability and dynamics.\n\nSummaryWe provide a brief introduction to some basic notions from information theory and how they may be used to understand stem cell identities at the single cell level. We also discuss how work in this area might develop in the near future.

cell biology

Stem cell differentiation is a stochastic process with memory

Pluripotent stem cells are able to self-renew indefinitely in culture and differentiate into all somatic cell types in vivo. While much is known about the molecular basis of pluripotency, the molecular mechanisms of lineage commitment are complex and only partially understood. Here, using a combination of single cell profiling and mathematical modeling, we examine the differentiation dynamics of individual mouse embryonic stem cells (ESCs) as they progress from the ground state of pluripotency along the neuronal lineage. In accordance with previous reports we find that cells do not transit directly from the pluripotent state to the neuronal state, but rather first stochastically permeate an intermediate primed pluripotent state, similar to that found in the maturing epiblast in development. However, analysis of rate at which individual cells enter and exit this intermediate metastable state using a hidden Markov model reveals that the observed ESC and epiblast-like macrostates conceal a chain of unobserved cellular microstates, which individual cells transit through stochastically in sequence. These hidden microstates ensure that individual cells spend well-defined periods of time in each functional macrostate and encode a simple form of epigenetic memory that allows individual cells to record their position on the differentiation trajectory. To examine the generality of this model we also consider the differentiation of mouse hematopoietic stem cells along the myeloid lineage and observe remarkably similar dynamics, suggesting a general underlying process. Based upon these results we suggest a statistical mechanics view of cellular identities that distinguishes between functionally-distinct macrostates and the many functionally-similar molecular microstates associated with each macrostate. Taken together these results indicate that differentiation is a discrete stochastic process amenable to analysis using the tools of statistical mechanics.

systems biology