bioRxiv ScienceSearch

Biology subjects

Reid, J.

Publications and source records attributed to Reid, J..

5 recordsLinked to original sources

Projection layers improve deep learning models of regulatory DNA function

With the increasing application of deep learning methods to the modelling of regulatory DNA sequences has come an interest in exploring what types of architecture are best suited to the domain. Networks designed to predict many functional characteristics of noncoding DNA in a multitask framework have to recognise a large number of motifs and as a result benefit from large numbers of convolutional filters in the first layer. The use of large first layers in turn motivates an exploration of strategies for addressing the sparsity of output and possibility for overfitting that result. To this end we propose the use of a dimensionality-reducing linear projection layer after the initial motif-recognising convolutions. In experiments with a reduced version of the DeepSEA dataset we find that inserting this layer in combination with dropout into convolutional and convolutional-recurrent architectures can improve predictive performance across a range of first layer sizes. We further validate our approach by incorporating the projection layer into a new convolutional-recurrent architecture which achieves state of the art performance on the full DeepSEA dataset. Analysis of the learned projection weights shows that the inclusion of this layer simplifies the networks internal representation of the occurrence of motifs, notably by projecting features representing forward and reverse-complement motifs to similar positions in the lower dimensional feature space output by the layer.

bioinformatics

Branch-recombinant Gaussian processes for analysis of perturbations in biological time series

MotivationA common class of behaviour encountered in the biological sciences involves branching and recombination. During branching, a statistical process bifurcates resulting in two or more potentially correlated processes that may under-go further branching; the contrary is true during recombination, where two or more statistical processes converge into one. A key objective is to identify the time of this bifurcation (branch time) from time series measurements e.g., comparing a control time series with a perturbed time series. Whilst statistical treatments for the two branch (control versus treatment) case exists, the ability to infer more complex branching structure from time series data remains open. Gaussian processes (GPs) represents an ideal framework for such analysis, allowing for nonlinear regression that includes a rigorous treatment of uncertainty. Currently, however, GP models only exist for two-branch systems. Here we highlight how arbitrarily complex branching processes can be built using the correct composition of covariance functions within a GP framework, thus outlining a general framework for the treatment of branching and recombination in the form of branch-recombinant Gaussian processes (B-RGPs). We first demonstrate the performance of B-RGPs compared to a variety of existing regression approaches, and demonstrate robustness to model misspecification. B-RGPs are then used to investigate the branching patterns of Arabidopsis thaliana gene expression following inoculation with the hemibotrophic bacteria, Pseudomonas syringae DC3000, and a disarmed mutant strain, hrpA. By grouping genes according to the number of branches, we could naturally separate out genes involved in basal immune response from those subverted by the virulent strain, and show enrichment for targets of pathogen protein effectors. Finally, we identify two early branching genes WRKY11 and WRKY17, and showed that groups of genes that branched at similar times to WRKY11/17 were enriched for W-box binding motifs, and overrepresented for genes differentially expressed in WRKY11/17 knockouts, suggesting that branch time could be used for identifying direct and indirect binding targets of key transcription factors. Software is available from: https://github.com/cap76/BranchingGPs.

bioinformatics

Nonparametric Bayesian inference of transcriptional branching and recombination identifies regulators of early human germ cell development

During embryonic development, cells undertake a series of fate decisions to form a complete organism comprised of various cell types, epitomising a branching process. A striking example of branching occurs in humans around the time of implantation, when primordial germ cells (PGCs), precursors of sperm and eggs, and somatic lineages are specified. Due to inaccessibility of human embryos at this stage of development, understanding the mechanisms of PGC specification remains difficult. The integrative modelling of single cell transcriptomics data from embryos and appropriate in vitro models should prove to be a useful resource for investigating this system, provided that the cells can be suitably ordered over a developmental axis. Unfortunately, most methods for inferring cell ordering were not designed with structured (time series) data in mind. Although some probabilistic approaches address these limitations by incorporating prior information about the developmental stage (capture time) of the cell, they do not allow the ordering of cells over processes with more than one terminal cell fate. To investigate the mechanisms of PGC specification, we develop a probabilistic pseudotime approach, branch-recombinant Gaussian process latent variable models (B-RGPLVMs), that use an explicit model of transcriptional branching in individual marker genes, allowing the ordering of cells over developmental trajectories with arbitrary numbers of branches. We use first demonstrate the advantage of our approach over existing pseudotime algorithms and subsequently use it to investigate early human development, as primordial germ cells (PGCs) and somatic cells diverge. We identify known master regulators of human PGCs, and predict roles for a variety of signalling pathways, transcription factors, and epigenetic modifiers. By concentrating on the earliest branched signalling events, we identified an antagonistic role for FGF receptor (FGFR) signalling pathway in the acquisition of competence for human PGC fate, and identify putative roles for PRC1 and PRC2 in PGC specification. We experimentally validate our predictions using pharmacological blocking of FGFR or its downstream effectors (MEK, PI3K and JAK), and demonstrate enhanced competency for PGC fate in vitro, whilst small molecule inhibition of the enzymatic component of PRC1/PRC2 reveals reduced capacity of cells to form PGCs in vitro. Thus, B-RGPLVMs represent a powerful and flexible data-driven approach for dissecting the temporal dynamics of cell fate decisions, providing unique insights into the mechanisms of early embryogenesis. Scripts relating to this analysis are available from: https://github.com/cap76/PGCPseudotime

developmental biology

Mutual Information Estimation For Transcriptional Regulatory Network Inference

Mutual information-based network inference algorithms are an important tool in the reverse-engineering of transcriptional regulatory networks, but all rely on estimates of the mutual information between the expression of pairs of genes. Various methods exist to compute estimates of the mutual information, but none have been firmly established as optimal for network inference. The performance of 9 mutual information estimation methods are compared using three popular network inference algorithms: CLR, MRNET and ARACNE. The performance of the estimators is compared on one synthetic and two real datasets. For estimators that discretise data, the effect of discretisation parameters are also studied in detail. Implementations of 5 estimators are provided in parallelised C++ with an R interface. These are faster than alternative implementations, with reductions in computation time up to a factor of 3,500.\n\nResultsThe B-spline estimator consistently performs well on real and synthetic datasets. CLR was found to be the best performing inference algorithm, corroborating previous results indicating that it is the state of the art mutual inference algorithm. It is also found to be robust to the mutual information estimation method and their parameters. Furthermore, when using an estimator that discretises expression data, using N1/3 bins for N samples gives the most accurate inferred network. This contradicts previous findings that suggested using N 1/2 bins.

genetics

Mutation spectrum of NOD2 reveals recessive inheritance as a main driver of Early Onset Crohn’s Disease

Inflammatory bowel disease (IBD), clinically defined as Crohns disease (CD), ulcerative colitis (UC), or IBD-unclassified, results in chronic inflammation of the gastrointestinal tract in genetically susceptible hosts. Pediatric onset IBD represents [≥]25% of all IBD diagnoses and often presents with intestinal stricturing, perianal disease, and failed response to conventional treatments. NOD2 was the first and is the most replicated locus associated with adult IBD, to date. To determine the role of NOD2 and other genes in pediatric IBD, we performed whole-exome sequencing on a cohort of 1,183 patients with pediatric onset IBD (ages 0-18.5 years). We identified 92 probands who were homozygous or compound heterozygous for rare and low frequency NOD2 variants accounting for approximately 8% of our cohort, suggesting a Mendelian recessive inheritance pattern of disease. Additionally, we investigated the contribution of recessive inheritance of NOD2 alleles in adult IBD patients from the Regeneron Genetics Center (RGC)-Geisinger Health System DiscovEHR study, which links whole exome sequences to longitudinal electronic health records (EHRs) from 51,289 participants. We found that ~7% of cases in this adult IBD cohort, including ~10% of CD cases, can be attributed to recessive inheritance of NOD2 variants, confirming the observations from our pediatric IBD cohort. Exploration of EHR data showed that 14% of these adult IBD patients obtained their initial IBD diagnosis before 18 years of age, consistent with early onset disease. Collectively, our findings show that recessive inheritance of rare and low frequency deleterious NOD2 variants account for 7-10% of CD cases and implicate NOD2 as a Mendelian disease gene for early onset Crohns Disease.\n\nAuthor SummaryPediatric onset inflammatory bowel disease (IBD) represents [≥]25% of IBD diagnoses; yet the genetic architecture of early onset IBD remains largely uncharacterized. To investigate this, we performed whole-exome sequencing and rare variant analysis on a cohort of 1,183 pediatric onset IBD patients. We found that 8% of patients in our cohort were homozygous or compound heterozygous for rare or low frequency deleterious variants in the nucleotide binding and oligomerization domain containing 2 (NOD2) gene. Further investigation of whole-exome sequencing of a large clinical cohort of adult IBD patients uncovered recessive inheritance of rare and low frequency NOD2 variants in 7% of cases and that the relative risk for NOD2 variant homozygosity has likely been underestimated. While it has been reported that having >1 NOD2 risk alleles is associated with increased susceptibility to Crohns Disease (CD), our data formally demonstrate what has long been suspected: recessive inheritance of NOD2 alleles is a mechanistic driver of early onset IBD, specifically CD, likely due to loss of NOD2 protein function. Our data suggest that a subset of IBD-CD patients with early disease onset is characterized by recessive inheritance of NOD2 alleles, which has important implications for the screening, diagnosis, and treatment of IBD.

genetics