bioRxiv ScienceSearch

Biology subjects

Jia, Z.

Publications and source records attributed to Jia, Z..

5 recordsLinked to original sources

Inference of Chromosome-length Haplotypes Using Genomic Data of Three to Five Single Gametes

Knowledge of chromosome-length haplotypes will not only advance our understanding of the relationship between DNA and phenotypes, but also promote a variety of genetic applications. Here we present Hapi, an innovative method for chromosomal haplotype inference using only 3 to 5 gametes. Hapi outperformed all existing haploid-based phasing methods in terms of accuracy, reliability, and cost efficiency in both simulated and real gamete datasets. This highly cost-effective phasing method will make large-scale haplotype studies feasible to facilitate human disease studies and plant/animal breeding. In addition, Hapi can detect meiotic crossovers in gametes, which has promise in the diagnosis of abnormal recombination activity in human reproductive cells.

bioinformatics

Structure-guided disruption of pseudopilus tip inhibits Type II secretion in Pseudomonas aeruginosa

Pseudomonas aeruginosa utilizes the Type II secretion system (T2SS) to translocate a wide range of large, structured protein virulence factors through the periplasm to the extracellular environment for infection. In the T2SS, five pseudopilins assemble into the pseudopilus that acts as a piston to extrude exoproteins out of cells. Through structure determination of the pseudopilin complexes of XcpVWX and XcpVW and function analysis, we have confirmed that two minor pseudopilins, XcpV and XcpW, constitute a core complex indispensable to the pseudopilus tip. The absence of either XcpV or -W resulted in the non-functional T2SS. Our small-angle X-ray scattering experiment for the first time revealed the architecture of the entire pseudopilus tip and established the working model. Based on the interaction interface of complexes, we have developed inhibitory peptides. The structure-based peptides not only disrupted of the XcpVW core complex and the entire pseudopilus tip in vitro but also inhibited the T2SS in vivo. More importantly, these peptides effectively reduced the virulence of P. aeruginosa towards Caenorhabditis elegans.

microbiology

Optimizing Trait Predictability in Hybrid Rice Using Superior Prediction Models and Selective Omic Datasets

Hybrid breeding has dramatically boosted yield and its stability in rice. Genomic prediction further benefits rice breeding by increasing selection intensity and accelerating breeding cycles. With the rapid advancement of technology, other omic data, such as metabolomic data and transcriptomic data, are readily available for predicting genetic values (or breeding values) for agronomically important traits. In the current study, we searched for the best prediction strategy for four traits (yield, 1000 grain weight, number of grains per panicle and number of tillers per plant) of hybrid rice by evaluating all possible combinations of omic datasets with different prediction methods. We conclude that, in rice, the predictions using the combination of genomic and metabolomic data generally produce better results than single-omics predictions or predictions based on other combined omic data. Inclusion of transcriptomic data does not improve predictability possibly because transcriptome does not provide more information for the trait than the sum of genome and metabolome; rather, the computational complexity is substantially increased if transcriptomic data is included in the models. Best linear unbiased prediction (BLUP) appears to be the most efficient prediction method compared to the other commonly used approaches, including LASSO, SSVS, SVM-RBF, SVP-POLY and PLS. Our study has provided a guideline for selection of hybrid rice in terms of which types of omic datasets and which method should be used to achieve higher trait predictability.

genetics

GDCRNATools: an R/Bioconductor package for integrative analysis of lncRNA, miRNA, and mRNA data in GDC

The large-scale multidimensional omics data in the Genomic Data Commons (GDC) provides opportunities to investigate the crosstalk among different RNA species and their regulatory mechanisms in cancers. Easy-to-use bioinformatics pipelines are needed to facilitate such studies. We have developed a user-friendly R/Bioconductor package, named GDCRNATools, to facilitate downloading, organizing, and analyzing RNA data in GDC with an emphasis on deciphering the lncRNA-mRNA related competing endogenous RNAs (ceRNAs) regulatory network in cancers. Many widely used bioinformatics tools and databases are utilized in our package. Users can easily pack preferred downstream analysis pipelines or integrate their own pipelines into the workflow. Interactive shiny web apps built in GDCRNATools greatly improve visualization of results from the analysis.\n\nAvailabilityGDCRNATools is an R/Bioconductor package that is freely available at https://github.com/Jialab-UCR/GDCRNATools

bioinformatics

Karyotype stability and unbiased fractionation in the paleo-allotetraploid Cucurbita genomes

The Cucurbita genus contains several economically important species in the Cucurbitaceae family. Interspecific hybrids between C. maxima and C. moschata are widely used as rootstocks for other cucurbit crops. We report high-quality genome sequences of C. maxima and C. moschata and provide evidence supporting an allotetraploidization event in Cucurbita. We are able to partition the genome into two homoeologous subgenomes based on different genetic distances to melon, cucumber and watermelon in the Benincaseae tribe. We estimate that the two diploid progenitors successively diverged from Benincaseae around 31 and 26 million years ago (Mya), and the allotetraploidization happened earlier than 3 Mya, when C. maxima and C. moschata diverged. The subgenomes have largely maintained the chromosome structures of their diploid progenitors. Such long-term karyotype stability after polyploidization is uncommon in plant polyploids. The two subgenomes have retained similar numbers of genes, and neither subgenome is globally dominant in gene expression. Allele-specific expression analysis in the C. maxima x C. moschata interspecific F1 hybrid and the two parents indicates the predominance of trans-regulatory effects underlying expression divergence of the parents, and detects transgressive gene expression changes in the hybrid correlated with heterosis in important agronomic traits. Our study provides insights into plant genome evolution and valuable resources for genetic improvement of cucurbit crops.

genomics