bioRxiv ScienceSearch

Biology subjects

Moses, A. M.

Publications and source records attributed to Moses, A. M..

4 recordsLinked to original sources

Variational Infinite Heterogeneous Mixture Model for Semi-supervised Clustering of Heart Enhancers

MotivationPMammalian genomes can contain thousands of enhancers but only a subset are actively driving gene expression in a given cellular context. Integrated genomic datasets can be harnessed to predict active enhancers. One challenge in integration of large genomic datasets is the increasing heterogeneity: continuous, binary and discrete features may all be relevant. Coupled with the typically small numbers of training examples, semi-supervised approaches for heterogeneous data are needed; however, current enhancer prediction methods are not designed to handle heterogeneous data in the semi-supervised paradigm.\n\nResultsWe implemented a Dirichlet Process Heterogeneous Mixture model that infers Gaussian, Bernoulli and Poisson distributions over features. We derived a novel variational inference algorithm to handle semi-supervised learning tasks where certain observations are forced to cluster together. We applied this model to enhancer candidates in mouse heart tissues based on heterogeneous features. We constrained a small number of known active enhancers to appear in the same cluster, and 47 additional regions clustered with them. Many of these are located near heart-specific genes. The model also predicted 1176 active promoters, suggesting that it can discover new enhancers and promoters.\n\nAvailabilityWe created the dphmix Python package: https://pypi.org/project/dphmix/\n\nContactalan.moses@utoronto.ca

bioinformatics

Learning unsupervised feature representations for single cell microscopy images with paired cell inpainting

Cellular microscopy images contain rich insights about biology. To extract this information, researchers use features, or measurements of the patterns of interest in the images. Here, we introduce a convolutional neural network (CNN) to automatically design features for fluorescence microscopy. We use a self-supervised method to learn feature representations of single cells in microscopy images without labelled training data. We train CNNs on a simple task that leverages the inherent structure of microscopy images and controls for variation in cell morphology and imaging: given one cell from an image, the CNN is asked to predict the fluorescence pattern in a second different cell from the same image. We show that our method learns high-quality features that describe protein expression patterns in single cells both yeast and human microscopy datasets. Moreover, we demonstrate that our features are useful for exploratory biological analysis, by capturing high-resolution cellular components in a proteome-wide cluster analysis of human proteins, and by quantifying multi-localized proteins and single-cell variability. We believe paired cell inpainting is a generalizable method to obtain feature representations of single cells in multichannel microscopy images.\n\nAuthor SummaryTo understand the cell biology captured by microscopy images, researchers use features, or measurements of relevant properties of cells, such as the shape or size of cells, or the intensity of fluorescent markers. Features are the starting point of most image analysis pipelines, so their quality in representing cells is fundamental to the success of an analysis. Classically, researchers have relied on features manually defined by imaging experts. In contrast, deep learning techniques based on convolutional neural networks (CNNs) automatically learn features, which can outperform manually-defined features at image analysis tasks. However, most CNN methods require large manually-annotated training datasets to learn useful features, limiting their practical application. Here, we developed a new CNN method that learns high-quality features for single cells in microscopy images, without the need for any labeled training data. We show that our features surpass other comparable features in identifying protein localization from images, and that our method can generalize to diverse datasets. By exploiting our method, researchers will be able to automatically obtain high-quality features customized to their own image datasets, facilitating many downstream analyses, as we highlight by demonstrating many possible use cases of our features in this study.

bioinformatics

An analog to digital converter creates nuclear localization pulses in yeast calcium signaling

Several examples of transcription factors that show stochastic, unsynchronized pulses of nuclear localization have been described. Here we show that under constant calcium stress, nuclear localization pulses of the transcription factor Crz1 follow stochastic variations in cytoplasmic calcium concentration. We find that the size of the stochastic calcium pulses is positively correlated with the number of subsequent Crz1 pulses. Based on our observations, we propose a simple stochastic model of how the signaling pathway converts a constant external calcium concentration into a digital number of Crz1 pulses in the nucleus, due to the time delay from nuclear transport and the stochastic decoherence of individual Crz1 molecule dynamics. We find support for several additional predictions of the model and conclude that stochastic input to nuclear transport may produce digital responses to analog signals in other signaling systems.

systems biology

Short linear motifs in intrinsically disordered regions modulate HOG signaling capacity

The effort to characterize intrinsically disordered regions of signaling proteins is rapidly expanding. An important class of disordered interaction modules are ubiquitous and functionally diverse elements known as short linear motifs (SLiMs). To further examine the role of SLiMs in signal transduction, we used a previously devised bioinformatics method to predict evolutionarily conserved SLiMs within a well-characterized pathway in S. cerevisiae. Using a single cell, reporter-based flow cytometry assay in conjunction with a fluorescent reporter driven by a pathway-specific promoter, we quantitatively assessed pathway output via systematic deletions of individual motifs. We found that, when deleted, 34% (10/29) of predicted SLiMs displayed a significant decrease in pathway output, providing evidence that these motifs play a role in signal transduction. In addition, we show that perturbations of parameters in a previously published stochastic model of HOG signaling could reproduce the quantitative effects of 4 out of 7 mutations in previously unknown SLiMs. Our study suggests that, even in well-characterized pathways, large numbers of functional elements remain undiscovered, and that challenges remain for application of systems biology models to interpret the effects of mutations in signalling pathways.\n\nOne-sentence SummaryMutations of short conserved elements in disordered regions have quantitative effects on a model signaling pathway.

systems biology