bioRxiv ScienceSearch

Biology subjects

Zhai, J.

Publications and source records attributed to Zhai, J..

3 recordsLinked to original sources

Exact Tests of Zero Variance Component in Presence of Multiple VarianceComponents with Application to Longitudinal Microbiome Study

In the metagenomics studies, testing the association of microbiome composition and clinical conditions translates to testing the nullity of variance components. Computationally efficient score tests have been the major tools. But they can only apply to the null hypothesis with a single variance component and when sample sizes are large. Therefore, they are not applicable to longitudinal microbiome studies. In this paper, we propose exact tests (score test, likelihood ratio test, and restricted likelihood ratio test) to solve the problems of (1) testing the association of the overall microbiome composition in a longitudinal design and (2) detecting the association of one specific microbiome cluster while adjusting for the effects from related clusters. Our approach combines the exact tests for null hypothesis with a single variance component with a strategy of reducing multiple variance components to a single one. Simulation studies demonstrate that our method has correct type I error rate and superior power compared to existing methods at small sample sizes and weak signals. Finally, we apply our method to a longitudinal pulmonary microbiome study of human immunodeficiency virus (HIV) infected patients and reveal two interesting genera Prevotella and Veillonella associated with forced vital capacity. Our findings shed lights on the impact of lung microbiome to HIV complexities. The method is implemented in the open source, high-performance computing language Julia and is freely available at https://github.com/JingZhai63/VCmicrobiome.

genomics

The basis of accumulation differences in plant 21-nt reproductive phasiRNAs, and their cis-directed activity

O_LIPost-transcriptional gene silencing in plants results from independent activities of diverse small RNA types. In anthers of grasses, hundreds of loci yield non-coding RNAs that are processed into 21- and 24-nt phased small interfering RNAs (phasiRNAs); these are triggered by miR2118 and miR2275.\nC_LIO_LIWe characterized these \"reproductive phasiRNAs\" from rice panicles and anthers across seven developmental stages. Our computational analysis identified characteristics of the 21-nt reproductive phasiRNAs that impact their biogenesis, stability, and potential functions.\nC_LIO_LIWe demonstrate that 21-nt reproductive phasiRNAs can function in cis to target their own precursors. We observed evidence of this cis regulatory activity in both rice (Oryza sativa) and maize (Zea mays). We validated this activity with evidence of cleavage and a resulting shift in the pattern of phasiRNA production.\nC_LIO_LIWe characterize biases in phasiRNA biogenesis, demonstrating that the Pol II-derived \"top\" strand phasiRNAs are consistently higher abundance than the bottom strand. The first phasiRNA from each precursor overlaps the miR2118 target site, and this impacts phasiRNA accumulation or stability, evident in the weak accumulation of this phasiRNA position. Additional influences on this first phasiRNA duplex include the sequence composition and length, and we show that these factors impact Argonaute loading.\nC_LI

plant biology

PEA: an integrated R toolkit for plant epitranscriptome analysis

MotivationThe epitranscriptome, also known as chemical modifications of RNA (CMRs), is a newly discovered layer of gene regulation, the biological importance of which emerged through analysis of only a small fraction of CMRs detected by high-throughput sequencing technologies. Understanding of the epitranscriptome is hampered by the absence of computational tools for the systematic analysis of epitranscriptome sequencing data. In addition, no tools have yet been designed for accurate prediction of CMRs in plants, or to extend epitranscriptome analysis from a fraction of the transcriptome to its entirety.\n\nResultsHere, we introduce PEA, an integrated R toolkit to facilitate the analysis of plant epitranscriptome data. The PEA toolkit contains a comprehensive collection of functions required for read mapping, CMR calling, motif scanning and discovery, and gene functional enrichment analysis. PEA also takes advantage of machine learning technologies for transcriptome-scale CMR prediction, with high prediction accuracy, using the Positive Samples Only Learning algorithm, which addresses the two-class classification problem by using only positive samples (CMRs), in the absence of negative samples (non-CMRs). Hence PEA is a versatile epitranscriptome analysis pipeline covering CMR calling, prediction, and annotation, and we describe its application to predict N6-methyladenosine (m6A) modifications in Arabidopsis thaliana. Experimental results demonstrate that the toolkit achieved 71.6% sensitivity and 73.7% specificity, which is superior to existing m6A predictors. PEA is potentially broadly applicable to the in-depth study of epitranscriptomics.\n\nAvailabilityPEA is implemented using R and available at https://github.com/cma2015/PEA.

bioinformatics