bioRxiv ScienceSearch

Biology subjects

Mao, Q.

Publications and source records attributed to Mao, Q..

8 recordsLinked to original sources

Towards inferring causal gene regulatory networks from single cell expression measurements

Single-cell transcriptome sequencing now routinely samples thousands of cells, potentially providing enough data to reconstruct causal gene regulatory networks from observational data. Here, we present Scribe, a toolkit for detecting and visualizing causal regulatory interactions between genes and explore the potential for single-cell experiments to power network reconstruction. Scribe employs Restricted Directed Information to determine causality by estimating the strength of information transferred from a potential regulator to its downstream target. We apply Scribe and other leading approaches for causal network reconstruction to several types of single-cell measurements and show that there is a dramatic drop in performance for \"pseudotime\" ordered single-cell data compared to true time series data. We demonstrate that performing causal inference requires temporal coupling between measurements. We show that methods such as \"RNA velocity\" restore some degree of coupling through an analysis of chromaffin cell fate commitment. These analyses therefore highlight an important shortcoming in experimental and computational methods for analyzing gene regulation at single-cell resolution and point the way towards overcoming it.

genomics

Predict plant-derived xenomiRs from plant miRNA sequences using random forest and one-dimensional convolutional neural network models

BackgroundAn increasing number of studies reported that exogenous miRNAs (xenomiRs) can be detected in animal bodies, however, some others reported negative results. Some attributed this divergence to the selective absorption of plant-derived xenomiRs by animals.\n\nResultsHere, we analyzed 166 plant-derived xenomiRs reported in our previous study and 942 non-xenomiRs extracted from miRNA expression profiles of four species of commonly consumed plants. Employing statistics analysis and cluster analysis, our study revealed the potential sequence specificity of plant-derived xenomiRs. Furthermore, a random forest model and a one-dimensional convolutional neural network model were trained using miRNA sequence features and raw miRNA sequences respectively and then employed to predict unlabeled plant miRNAs in miRBase. A total of 241 possible plant-derived xenomiRs were predicted by both models. Finally, the potential functions of these possible plant-derived xenomiRs along with our previously reported ones in human body were analyzed.\n\nConclusionsOur study, for the first time, presents the systematic plant-derived xenomiR sequences analysis and provides evidence for selective absorption of plant miRNA by human body, which could facilitate the future investigation about the mechanisms underlying the transference of plant-derived xenomiR.

bioinformatics

Single tube bead-based DNA co-barcoding for cost effective and accurate sequencing, haplotyping, and assembly

Obtaining accurate sequences from long DNA molecules is very important for genome assembly and other applications. Here we describe single tube long fragment read (stLFR), a technology that enables this a low cost. It is based on adding the same barcode sequence to sub-fragments of the original long DNA molecule (DNA co-barcoding). To achieve this efficiently, stLFR uses the surface of microbeads to create millions of miniaturized barcoding reactions in a single tube. Using a combinatorial process up to 3.6 billion unique barcode sequences were generated on beads, enabling practically non-redundant co-barcoding with 50 million barcodes per sample. Using stLFR, we demonstrate efficient unique co-barcoding of over 8 million 20-300 kb genomic DNA fragments. Analysis of the genome of the human genome NA12878 with stLFR demonstrated high quality variant calling and phasing into contigs up to N50 34 Mb. We also demonstrate detection of complex structural variants and complete diploid de novo assembly of NA12878. These analyses were all performed using single stLFR libraries and their construction did not significantly add to the time or cost of whole genome sequencing (WGS) library preparation. stLFR represents an easily automatable solution that enables high quality sequencing, phasing, SV detection, scaffolding, cost-effective diploid de novo genome assembly, and other long DNA sequencing applications.

genomics

Multicenter validation of a sepsis prediction algorithm using only vital sign data in the emergency department, general ward and ICU

ObjectivesWe validate a machine learning-based sepsis prediction algorithm (InSight) for detection and prediction of three sepsis-related gold standards, using only six vital signs. We evaluate robustness to missing data, customization to site-specific data using transfer learning, and generalizability to new settings.\n\nDesignA machine learning algorithm with gradient tree boosting. Features for prediction were created from combinations of only six vital sign measurements and their changes over time.\n\nSettingA mixed-ward retrospective data set from the University of California, San Francisco (UCSF) Medical Center (San Francisco, CA) as the primary source, an intensive care unit data set from the Beth Israel Deaconess Medical Center (Boston, MA) as a transfer learning source, and four additional institutions datasets to evaluate generalizability.\n\nParticipants684,443 total encounters, with 90,353 encounters from June 2011 to March 2016 at UCSF.\n\nInterventionsnone\n\nPrimary and secondary outcome measuresArea under the receiver operating characteristic curve (AUROC) for detection and prediction of sepsis, severe sepsis, and septic shock.\n\nResultsFor detection of sepsis and severe sepsis, InSight achieves an area under the receiver operating characteristic (AUROC) curve of 0.92 (95% CI 0.90 - 0.93) and 0.87 (95% CI 0.86 - 0.88), respectively. Four hours before onset, InSight predicts septic shock with an AUROC of 0.96 (95% CI 0.94 -0.98), and severe sepsis with an AUROC of 0.85 (95% CI 0.79 - 0.91).\n\nConclusionsInSight outperforms existing sepsis scoring systems in identifying and predicting sepsis, severe sepsis, and septic shock. This is the first sepsis screening system to exceed an AUROC of 0.90 using only vital sign inputs. InSight is robust to missing data, can be customized to novel hospital data using a small fraction of site data, and retained strong discrimination across all institutions.\n\nStrengths and limitations of this studyO_LIMachine learning is applied to the detection and prediction of three separate sepsis standards in the emergency department, general ward and intensive care settings.\nC_LIO_LIOnly six commonly measured vital signs are used as input for the algorithm.\nC_LIO_LIThe algorithm is robust to randomly missing data.\nC_LIO_LITransfer learning successfully leverages large dataset information to a target dataset.\nC_LIO_LIRetrospective nature of the study does not predict clinician reaction to information.\nC_LI

bioinformatics

Pediatric Severe Sepsis Prediction Using Machine Learning

Early detection of pediatric severe sepsis is necessary in order to administer effective treatment. In this study, we assessed the efficacy of a machine-learning-based prediction algorithm applied to electronic healthcare record (EHR) data for the prediction of severe sepsis onset. The resulting prediction performance was compared with the Pediatric Logistic Organ Dysfunction score (PELOD-2) and pediatric Systemic Inflammatory Response Syndrome score (SIRS) using cross-validation and pairwise t-tests. EHR data were collected from a retrospective set of de-identified pediatric inpatient and emergency encounters drawn from the University of California San Francisco (UCSF) Medical Center, with encounter dates between June 2011 and March 2016. Patients (n = 11,127) were 2-17 years of age and 103 [0.93%] were labeled severely septic. In four-fold cross-validation evaluations, the machine learning algorithm achieved an AUROC of 0.912 for discrimination between severely septic and control pediatric patients at onset and AUROC of 0.727 four hours before onset. Under the same measure, the prediction algorithm also significantly outperformed PELOD-2 (p < 0.05) and SIRS (p < 0.05) in the prediction of severe sepsis four hours before onset. This machine learning algorithm has the potential to deliver high-performance severe sepsis detection and prediction for pediatric inpatients.

bioinformatics

Significant abundance of cis configurations of mutations in diploid human genomes

To fully understand human genetic variation, one must assess the specific distribution of variants between the two chromosomal homologues of genes, and any functional units of interest, as the phase of variants can significantly impact gene function and phenotype. To this end, we have systematically analyzed 18,121 autosomal protein-coding genes in 1,092 statistically phased genomes from the 1000 Genomes Project, and an unprecedented number of 184 experimentally phased genomes from the Personal Genome Project. Here we show that mutations predicted to functionally alter the protein, and coding variants as a whole, are not randomly distributed between the two homologues of a gene, but do occur significantly more frequently in cis-than trans-configurations, with cis/trans ratios of [~]60:40. Significant cis-abundance was observed in virtually all individual genomes in all populations. Nearly all variable genes exhibited either cis, or trans configurations of protein-altering mutations in significant excess, allowing distinction of cis- and trans-abundant genes. These common patterns of phase were largely constituted by a shared, global set of phase-sensitive genes. We show significant enrichment of this global set with gene sets indicating its involvement in adaptation and evolution. Moreover, cis- and trans-abundant genes were found functionally distinguishable, and exhibited strikingly different distributional patterns of protein-altering mutations. This work establishes common patterns of phase as key characteristics of diploid human exomes and provides evidence for their potential functional significance. Thus, it highlights the importance of phase for the interpretation of protein-coding genetic variation, challenging the current conceptual and functional interpretation of autosomal genes.

genomics

Advanced whole genome sequencing and analysis of fetal genomes from amniotic fluid

Amniocentesis is typically performed to identify large chromosomal abnormalities within the fetus. Here we demonstrate that it is feasible to generate an accurate whole genome sequence (WGS) of a fetus from an amniotic sample. DNA from cells and the amniotic fluid were isolated and sequenced from 31 amniocenteses. Concordance of variant calls between the two DNA sources and with parental libraries was high. Two fetal genomes were found to harbor potentially detrimental variants in CHD8 and LRP1, variations in these genes have been associated with Autism Spectrum Disorder (ASD) and Keratosis pilaris atrophicans, respectively. We also discovered drug sensitivities and carrier information of fetuses for a variety of diseases. In this study, we demonstrate for the first time the sequencing of the whole genome of fetuses from amniotic fluid and show that much more information than large chromosomal abnormalities can be gained from an amniocentesis.

genomics

Reversed graph embedding resolves complex single-cell developmental trajectories

Organizing single cells along a developmental trajectory has emerged as a powerful tool for understanding how gene regulation governs cell fate decisions. However, learning the structure of complex single-cell trajectories with two or more branches remains a challenging computational problem. We present Monocle 2, which uses reversed graph embedding to reconstruct single-cell trajectories in a fully unsupervised manner. Monocle 2 learns an explicit principal graph to describe the data, greatly improving the robustness and accuracy of its trajectories compared to other algorithms. Monocle 2 uncovered a new, alternative cell fate in what we previously reported to be a linear trajectory for differentiating myoblasts. We also reconstruct branched trajectories for two studies of blood development, and show that loss of function mutations in key lineage transcription factors diverts cells to alternative branches on the a trajectory. Monocle 2 is thus a powerful tool for analyzing cell fate decisions with single-cell genomics.

genomics