bioRxiv ScienceSearch

Biology subjects

Xie, J.

Publications and source records attributed to Xie, J..

10 recordsLinked to original sources

QUBIC2: A novel biclustering algorithm for large-scale bulk RNA-sequencing and single-cell RNA-sequencing data analysis

The combination of biclustering and large-scale gene expression data holds a promising potential for inference of the condition specific functional pathways/networks. However, existing biclustering tools do not have satisfied performance on high-resolution RNA-sequencing (RNA-Seq) data, majorly due to the lack of (i) a consideration of high sparsity of RNA-Seq data, e.g., the massive zeros or lowly expressed genes in the data, especially for single-cell RNA-Seq (scRNA-Seq) data, and (ii) an understanding of the underlying transcriptional regulation signals of the observed gene expression values. Here we presented a novel biclustering algorithm namely QUBIC2, for the analysis of large-scale bulk RNA-Seq and scRNA-Seq data. Key novelties of the algorithm include (i) used a truncated model to handle the unreliable quantification of genes with low or moderate expression, (ii) adopted the mixture Gaussian distribution and an information-divergency objective function to capture shared transcriptional regulation signals among a set of genes, (iii) utilized a Core-Dual strategy to identify biclusters and optimize relevant parameters, and (iv) developed a size-based P-value framework to evaluate the statistical significances of all the identified biclusters. Our method validation on comprehensive data sets of bulk and single cell RNA-seq data suggests that QUBIC2 had superior performance in functional modules detection and cell type classification compared with the other five widely-used biclustering tools. In addition, the applications of temporal and spatial data demonstrated that QUBIC2 can derive meaningful biological information from scRNA-Seq data. The source code for QUBIC2 can be freely accessed at https://github.com/maqin2001/qubic2.

bioinformatics

Widespread separation of the polypyrimidine tract from 3’ AG by G tracts in association with alternative exons in metazoa and plants

At the end of introns, the polypyrimidine tract (Py) is often close to the 3 AG in a consensus (Y)20NCAGgt in humans. Interestingly, we have found that they could also be separated by purine-rich elements including G tracts in thousands of human genes. These regulatory elements between the Py and 3AG (REPA) mainly regulate alternative 3 splice sites (3SS) and intron retention. Here we show their widespread distribution and special properties across kingdoms. The purine-rich 3SS are found in up to about 60% of the introns among more than 1000 species/lineages by whole genome analysis, and up to 18% of these introns contain the REPA G tracts in about 2.4 millions of 3SS in total. In particular, they are significantly enriched over their 3SS and genome backgrounds in metazoa and plants, and highly associated with alternative splicing of genes in diverse functional clusters. They are also highly enriched (3-6 folds) in the canonical as well as aberrantly used 3 splice sites in cancer patients carrying mutations of the branch point factor SF3B1 or the 3AG binding factor U2AF35. Moreover, the REPA G tract-harbouring 3SS have significantly reduced occurrences of branch point (BP) motifs between the -24 and -4 positions, in particular absent from the -7 - -5 positions in several model organisms examined. The more distant branch points are associated with increased occurrences of alternative splicing in human and zebrafish. The branch points, REPA G tracts and associated 3SS motifs appear to have emerged differentially in a phylum- or species-specific way during evolution. Thus, there is widespread separation of the Py and 3AG by REPA G tracts, likely evolved among different species or branches of life. This special 3SS arrangement contributes to the generation of diverse transcript or protein isoforms in biological functions or diseases through alternative or aberrant splicing.

molecular biology

The Big Five, Self-efficacy, and Self-control in Boxers

Inviting 210 boxers of national athletes in China as participants, this study applied the NEO Five-Factor Inventory and self-control and self-efficacy scales for athletes to examine the relationship between personality traits and self-control, as well as any effect of self-efficacy as a mediator between the two variables. The data analysis indicated that, firstly, the boxers overall level of self-control is high, and the higher the competitive level, the higher the level of self-control. Secondly, there were significant correlations among the Big Five, self-control, and self-efficacy. Thirdly, the mediation model showed that self-efficacy has a significant mediating effect between the Big Five and self-control. These results suggest that formulating training and intervention programs based on the personality traits of boxers and focusing on training their self-efficacy (1) to help them enhance their self-control ability, thereby improving athletic performance and promoting physical and mental health, and (2) to support the inclusion of personality traits, self-efficacy, and self-control among psychological indicators to be assessed in boxers.

scientific communication and education

Bacterial Glycosyltransferase-mediated Cell-surface Chemoenzymatic Glycan Editing: Methods and Applications

AbstractChemoenzymatic glycan editing that modifies glycan structures directly on the cell surface has emerged as a complementary tool to metabolic oligosaccharide engineering. In this article, we report the discovery that three bacterial enzymes--Pasteurella multocida 2-3-sialyltransferase M144D mutant (Pm2,3ST-M144D), Photobacterium damsel 2-6-sialyltransferase (Pd2,6ST) and Helicobacter mustelae 1-2-fucosyltransferase (Hm1,2FT)--can serve as highly efficient tools for cell-surface glycan editing. Among these three enzymes, the two sialyltransferases were also found to be tolerant to large substituents introduced to the C-5 position of the cytidine monophosphate N-acetylneuraminic acid donor, including biotin and fluorescent dyes. Combining these enzymes with our previously discovered Helicobacter pylori 1-3-FT, we developed a live cell-based assay to probe host-cell glycan-mediated influenza A virus (IAV) infection including both wild-type and mutant strains of human H1N1 and H3N2 influenza subtypes. At high SiaNAc2-6-Gal levels, the ability of a viral strain to induce the host cell death is positively correlated with the SiaNAc2-6-Gal binding affinity of its haemagglutinin. Surprisingly, the creation of sLeX on the host cell surface via in situ 1-3-Fuc editing also exacerbated the killing induced by several wild-type IAV strains as well as a mutant known as HK68-MTA. Structural alignment of HAs from the wild-type HK68 and HK68-MTA revealed the formation of a putative hydrogen bond between Trp222 of HA-HK68-MTA and the C-4 hydroxyl group of the 1-3-linked fucose of sLeX. This interaction is likely to be responsible for the better binding affinity of HA-HK68-MTA to sLeX and accordingly the enhanced host-cell killing compared with the wild-type HK68.

biochemistry

GeneQC: A quality control tool for gene expression estimation based on RNA-sequencing reads mapping

MotivationOne of the main benefits of using modern RNA-sequencing (RNA-Seq) technology is the more accurate gene expression estimations compared with previous generations of expression data, such as the microarray. However, numerous issues can result in the possibility that an RNA-Seq read can be mapped to multiple locations on the reference genome with the same alignment scores, which occurs in plant, animal, and metagenome samples. Such a read is so-called a multiple-mapping read (MMR). The impact of these MMRs is reflected in gene expression estimation and all downstream analyses, including differential gene expression, functional enrichment, etc. Current analysis pipelines lack the tools to effectively test the reliability of gene expression estimations, thus are incapable of ensuring the validity of all downstream analyses.\n\nResultsOur investigation into 95 RNA-Seq datasets from seven species (totaling 1,951GB) indicates an average of roughly 22% of all reads are MMRs for plant and animal species. Here we present a tool called GeneQC (Gene expression Quality Control), which can accurately estimate the reliability of each genes expression level. The underlying algorithm is designed based on extracted genomic and transcriptomic features, which are then combined using elastic-net regularization and mixture model fitting to provide a clearer picture of mapping uncertainty for each gene. GeneQC allows researchers to determine reliable expression estimations and conduct further analysis on the gene expression that is of sufficient quality. This tool also enables researchers to investigate continued re-alignment methods to determine more accurate gene expression estimates for those with low reliability.\n\nAvailabilityGeneQC is freely available at http://bmbl.sdstate.edu/GeneQC/home.html.\n\nContactqin.ma@sdstate.edu\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

RMalign: an RNA structural alignment tool based on a size independent scoring function

RNA-protein 3D complex structure prediction is still challenging. Recently, a template-based approach PRIME is proposed in our team to build RNA-protein complex 3D structure models with a higher success rate than computational docking software. However, scoring function of RNA alignment algorithm SARA in PRIME is size-dependent, which limits its ability to detect templates in some cases. Herein, we developed a novel RNA 3D structural alignment approach RMalign, which is based on a size-independent scoring function RMscore. The parameter in RMscore is then optimized in randomly selected RNA pairs and phase transition points (from dissimilar to similar) are determined in another randomly selected RNA pairs. In tRNA benchmarking, the precision of RMscore is higher than that of SARAscore (0.8771 and 0.7766, respectively) with phase transition points. In balance-FSCOR benchmarking, RMalign performed as good as ESA-RNA with a non-normalized score measuring RNA structure similarity. In balance-x-FSCOR benchmarking, RMalign achieves much better than a state-of-the-art RNA 3D structural alignment approach SARA due to a size-independent scoring function. Taking the advantage of RMalign, we update our RNA-protein modeling approach PRIME to version 2.0. The PRIME2.0 significantly improves about 10% success rate than PRIME.\n\nAuthor summaryRNA structures are important for RNA functions. With the increasing of RNA structures in PDB, RNA 3D structure alignment approaches have been developed. However, the scoring function which is used for measuring RNA structural similarity is still length dependent. This shortcoming limits its ability to detect RNA structure templates in modeling RNA structure or RNA-protein 3D complex structure. Thus, we developed a length independent scoring function RMscore to enhance the ability to detect RNA structure homologs. The benchmarking data shows that RMscore can distinct the similar and dissimilar RNA structure effectively. RMscore should be a useful scoring function in modeling RNA structures for the biological community. Based on RMscore, we develop an RNA 3D structure alignment RMalign. In both RNA structure and function classification benchmarking, RMalign obtains as good as or even better performance than the state-of-the-art approaches. With a length independent scoring function RMscore, RMalign should be useful for the modeling RNA structures. Based on above results, we update PRIME to PRIME2.0. We provide a more accurate RNA-protein 3D complex structure modeling tool PRIME2.0 which should be useful for the biological community.

bioinformatics

Deep-RBPPred: Predicting RNA binding proteins in the proteome scale based on deep learning

RNA binding protein (RBP) plays an important role in cell processes. Identifying RBPs by computation and experiment are both essential. Recently, RBPPred is proposed in our group to predict RBP with a high performance. However, RBPPred is too slow for that it will generate PSSM matrix as its feature. Herein, we develop a deep learning model called Deep-RBPPred. The model has three advantages comparing to previous models. 1. Deep-RBPPred only needs few physicochemical properties. 2. Deep-RBPPred runs much faster. 3. Deep-RBPPred has a good generalization ability. In the meantime, the performance is still as good as the stats-of-the-art method. In the testing in A. thaliana, S. cerevisiae and H. sapiens proteomics, MCC (AUC) are 0.6077 (0.9421), 0.573 (0.9034) and 0.8141(0.9515) respectively when the score cutoff is set to 0.5. In the verifying in Gerstberger-1538, the SN of our model is 90.38%. The running times are 9s, 7s, 8s and 10s, respectively, for H.sapiens, A.thaliana, S.cerevisiae and Gerstberger-1538 when it is tested in GPU. Deep-RBPPred forecasts 94.65% of 299 new RBP and about 8% higher sensitivity than RBPPred. We also apply deep-RBPPred in 19 eukaryotes proteomics and 11 bacteria proteomics downloaded from Uniprot. The result shows that rate of RBPs in eukaryotes proteome are much higher than bacteria proteome. Testing in 6 proteomics shows the many RBPs may be still undiscovered so far.

bioinformatics

Replacing reprogramming factors with antibodies selected from autocrine antibody libraries

Signaling pathways initiated at the membrane establish and maintain cell fate during development and can be harnessed in the nucleus to generate induced pluripotent stem cells (iPSCs) from differentiated cells. Yet, the impact of extracellular signaling on reprogramming to pluripotency has not been systematically addressed. Here, we screen a lentiviral library encoding {small tilde}100 million secreted and membrane-bound antibodies and identify multiple antibodies that can replace Sox2/c-Myc or Oct4 during reprogramming. We show that one Sox2-replacing antibody initiates reprogramming by antagonizing the membrane-associated protein Basp1, thereby inducing nuclear factors WT1 and Esrrb/Lin28 independent of Sox2. By successively manipulating this pathway we identify three new methods to generate iPSCs. This study expands current knowledge of reprogramming methods and mechanisms and establishes unbiased selection from autocrine antibody libraries as a powerful orthogonal platform to discover new biologics and pathways regulating pluripotency and cell fate.

cell biology

Molecular Mapping Of YrTZ2, A Stripe Rust Resistance Gene In Wild Emmer Accession TZ-2 And Its Comparative Analyses With Aegilops tauschii

Wheat stripe rust, caused by Puccinia striiformis f. sp. tritici (Pst), is a devastating disease that can cause severe yield losses. Identification and utilization of stripe rust resistance genes are essential for effective breeding against the disease. Wild emmer accession TZ-2, originally collected from Mount Hermon, Israel, confers near-immunity resistance against several prevailing Pst races in China. A set of 200 F6:7 recombinant inbred lines (RILs) derived from a cross between susceptible durum wheat cultivar Langdon and TZ-2 was used for stripe rust evaluation. Genetic analysis indicated that the stripe rust resistance of TZ-2 to Pst race CYR34 was controlled by a single dominant gene, temporarily designated YrTZ2. Through bulked segregant analysis (BSA) and SSR mapping, YrTZ2 was located on chromosome arm 1BS and flanked by SSR markers Xwmc230 and Xgwm413 with genetic distance of 0.8 cM (distal) and 0.3 cM (proximal), respectively. By applying wheat 90K iSelect SNP genotyping assay, 11 polymorphic loci (consist of 250 SNP markers) closely linked with YrTZ2 were identified. YrTZ2 was further delimited into a 0.8 cM genetic interval between SNP marker IWB19368 and SSR marker Xgwm413, and co-segregated with SNP marker IWB28744 (attached with 28 SNP markers). Comparative genomics analyses revealed high level of collinearity between the YrTZ2 genomic region and the orthologous region of Aegilops tauschii 1DS. The genomic region between loci IWB19368 and IWB31649 harboring YrTZ2 is orthologous to a 24.5 Mb genomic region between AT1D0112 and AT1D0150, spanning 15 contigs on chromosome 1DS. The genetic and comparative maps of YrTZ2 provide framework for map-based cloning and marker-assisted selection (MAS) of YrTZ2.

plant biology

Mapping Human Hematopoietic Hierarchy At Single Cell Resolution By Microwell-seq

The classical hematopoietic hierarchy, which is mainly built with fluorescence-activated cell sorting (FACS) technology, proves to be inaccurate in recent studies. Single cell RNA-seq (scRNA-seq) analysis provides a solution to overcome the limit of FACS-based cell type definition system for the dissection of complex cellular hierarchy. However, large-scale scRNA-seq is constrained by the throughput and cost of traditional methods. Here, we developed Microwell-seq, a high-throughput and low-cost scRNA-seq platform using extremely simple devices. Using Microwell-seq, we constructed a single-cell resolution transcriptome atlas of human hematopoietic differentiation hierarchy by profiling more than 50,000 single cells throughout adult human hematopoietic system. We found that adult human hematopoietic stem and progenitor cell (HSPC) compartment is dominated by progenitors primed with lineage specific regulators. Our analysis revealed differentiation pathways for each cell types, through which HSPCs directly progress to lineage biased progenitors before differentiation. We propose a revised adult human hematopoietic hierarchy independent of oligopotent progenitors. Our study also demonstrates the broad applicability of Microwell-seq technology.

cell biology