bioRxiv ScienceSearch

Biology subjects

Cao, S.

Publications and source records attributed to Cao, S..

10 recordsLinked to original sources

LTMG (Left truncated mixture Gaussian) based modeling of transcriptional regulatory heterogeneities in single cell RNA-seq data - a perspective from the kinetics of mRNA metabolism

A key challenge in modeling single-cell RNA-seq (scRNA-seq) data is to capture the diverse gene expression states regulated by different transcriptional regulatory inputs across single cells, which is further complicated by a large number of observed zero and low expressions. We developed a left truncated mixture Gaussian (LTMG) model that stems from the kinetic relationships between the transcriptional regulatory inputs and metabolism of mRNA and gene expression abundance in a cell. LTMG infers the expression multi-modalities across single cell entities, representing a genes diverse expression states; meanwhile the dropouts and low expressions are treated as left truncated, specifically representing an expression state that is under suppression. We demonstrated that LTMG has significantly better goodness of fitting on an extensive number of single-cell data sets, comparing to three other state of the art models. In addition, our systems kinetic approach of handling the low and zero expressions and correctness of the identified multimodality are validated on several independent experimental data sets. Application on data of complex tissues demonstrated the capability of LTMG in extracting varied expression states specific to cell types or cell functions. Based on LTMG, a differential gene expression test and a co-regulation module identification method, namely LTMG-DGE and LTMG-GCR, are further developed. We experimentally validated that LTMG-DGE is equipped with higher sensitivity and specificity in detecting differentially expressed genes, compared with other five popular methods, and that LTMG-GCR is capable to retrieve the gene co-regulation modules corresponding to perturbed transcriptional regulations. A user-friendly R package with all the analysis power is available at https://github.com/zy26/LTMGSCA.

bioinformatics

ICTD: Inference of cell types and deconvolution -- a next-generation deconvolution method for accurate assess cell population and activities in tumor microenvironment.

We developed a novel deconvolution method, namely Inference of Cell Types and Deconvolution (ICTD) that addresses the fundamental issue of identifiability and robustness in current tissue data deconvolution problem. ICTD provides substantially new capabilities for omics data based characterization of a tissue microenvironment, including (1) maximizing the resolution in identifying resident cell and sub types that truly exists in a tissue, (2) identifying the most reliable marker genes for each cell type, which are tissue and data set specific, (3) handling the stability problem with co-linear cell types, (4) co-deconvoluting with available matched multi-omics data, and (5) inferring functional variations specific to one or several cell types. ICTD is empowered by (i) rigorously derived mathematical conditions of identifiable cell type and cell type specific functions in tissue transcriptomics data and (ii) a semi supervised approach to maximize the knowledge transfer of cell type and functional marker genes identified in single cell or bulk cell data in the analysis of tissue data, and (iii) a novel unsupervised approach to minimize the bias brought by training data. Application of ICTD on real and single cell simulated tissue data validated that the method has consistently good performance for tissue data coming from different species, tissue microenvironments, and experimental platforms. Other than the new capabilities, ICTD outperformed other state-of-the-art devolution methods on prediction accuracy, the resolution of identifiable cell, detection of unknown sub cell types, and assessment of cell type specific functions. The premise of ICTD also lies in characterizing cell-cell interactions and discovering cell types and prognostic markers that are predictive of clinical outcomes.

bioinformatics

Asparagine availability is an essential limiting factor for poxvirus protein synthesis

Virus actively interfaces with host metabolism because viral replication relies on host cells to provide nutrients and energy. For efficient viral replication in culture, vaccinia virus (VACV; the prototype poxvirus) prefers glutamine to glucose, to the extent that in glutamine-free medium, VACV replication is inefficient. Remarkably, VACV replication can be fully rescued from glutamine depletion by asparagine supplementation. By global metabolic profiling, genetic and chemical intervening of asparagine supply, we provide evidence demonstrating that the requirement of asparagine for efficient viral replication accounts for VACVs preference of glutamine to glucose, rather than because glutamine is superior to glucose in feeding the tricarboxylic acid (TCA) cycle. Further, we show that asparagine availability is a critical factor for efficient viral protein synthesis. Our study highlights that the asparagine metabolism, whose regulation has been evolutionarily tailored in mammalian cells, presents a critical barrier to poxvirus replication, suggesting new directions of anti-viral strategy development.

microbiology

QUBIC2: A novel biclustering algorithm for large-scale bulk RNA-sequencing and single-cell RNA-sequencing data analysis

The combination of biclustering and large-scale gene expression data holds a promising potential for inference of the condition specific functional pathways/networks. However, existing biclustering tools do not have satisfied performance on high-resolution RNA-sequencing (RNA-Seq) data, majorly due to the lack of (i) a consideration of high sparsity of RNA-Seq data, e.g., the massive zeros or lowly expressed genes in the data, especially for single-cell RNA-Seq (scRNA-Seq) data, and (ii) an understanding of the underlying transcriptional regulation signals of the observed gene expression values. Here we presented a novel biclustering algorithm namely QUBIC2, for the analysis of large-scale bulk RNA-Seq and scRNA-Seq data. Key novelties of the algorithm include (i) used a truncated model to handle the unreliable quantification of genes with low or moderate expression, (ii) adopted the mixture Gaussian distribution and an information-divergency objective function to capture shared transcriptional regulation signals among a set of genes, (iii) utilized a Core-Dual strategy to identify biclusters and optimize relevant parameters, and (iv) developed a size-based P-value framework to evaluate the statistical significances of all the identified biclusters. Our method validation on comprehensive data sets of bulk and single cell RNA-seq data suggests that QUBIC2 had superior performance in functional modules detection and cell type classification compared with the other five widely-used biclustering tools. In addition, the applications of temporal and spatial data demonstrated that QUBIC2 can derive meaningful biological information from scRNA-Seq data. The source code for QUBIC2 can be freely accessed at https://github.com/maqin2001/qubic2.

bioinformatics

Metastable contacts and structural disorder in the estrogen receptor transactivation domain

The N-terminal transactivation domain (NTD) of estrogen receptor alpha, a well-known member of the family of intrinsically disordered proteins (IDPs), mediates the receptors transactivation function to regulate gene expression. However, an accurate molecular dissection of NTDs structure-function relationships remains elusive. Here, using small-angle X-ray scattering (SAXS), nuclear magnetic resonance (NMR), circular dichroism, and hydrogen exchange mass spectrometry, we show that NTD adopts a mostly disordered, unexpectedly compact conformation that undergoes structural expansion upon chemical denaturation. By combining SAXS, hydroxyl radical protein footprinting and computational modeling, we derive the ensemble-structures of the NTD and determine its ensemble-contact map that reveals metastable regional and long-range contacts, including interactions between residues I33 and S118. We show that mutation at S118, a known phosphorylation site, promotes conformational changes and increases coactivator binding. We further demonstrate via fluorine-19 (19F) NMR that mutations near residue I33 alter 19F chemical shifts at residue S118, confirming the proposed I33-S118 contact in the ensemble of structural disorder. These findings extend our understanding of IDPs structure-function relationship, and how specific metastable contacts mediate critical functions of disordered proteins.\n\nHighlightsO_LIA compact disorder is observed for the N-terminal domain (NTD) of estrogen receptor\nC_LIO_LIMulti-technique modeling elucidates the NTD ensemble structures\nC_LIO_LIEnsemble-based contact map reveals metastable contacts between I33 and S118\nC_LIO_LI19F-NMR data validate the proposed I33-S118 contact in the IDP\nC_LI

biophysics

K-Ras G-domain binding with signaling lipid phosphoinositides: PIP2 association, orientation, function

Ras genes are potent drivers of human cancers, with mutated K-Ras4B being the most abundant isoform. Targeted inhibition of oncogenic gene products is considered the holy grail of present-day cancer therapy, and recent discoveries of small molecule inhibitors for K-Ras4B greatly benefited from a deeper understanding of the protein structure and dynamics of the GTPase. Since interactions with biological membranes are key for Ras function, the details of Ras - lipid interactions have become a major focus of study, especially since it is becoming clear that such interactions not only involve the Ras C-terminus for lipid anchoring, but also the G-protein domain. Here we investigated the interaction between K-Ras4B with the signaling lipid phosphatidyl inositol (4,5) phosphate (PIP2) using NMR spectroscopy and molecular dynamics simulations, complemented by biophysical and cell biology assays. We discovered that the {beta}2 and {beta}3 strands as well as helices 4 and 5 of the GTPase G-domain bind to PIP2, and that these secondary structural elements employ specific residues for these interactions. These likely occur in two orientation states of the protein relative to the membrane. Importantly, we found that some of these residues, which are known to be oncogenic when mutated (D47K, D92N, K104M and D126N), are critical for K-Ras-mediated transformation of fibroblast cells, while not substantially affecting basal and assisted nucleotide hydrolysis and exchange. We further showed that mutation K104M can indeed abolish localization of mutant K-Ras to the plasma membrane. These findings suggest that specific G-domain residues play an important, previously-unknown role in regulating Ras function by mediating interactions with membrane PIP2 lipids. Thus, a detailed description of the novel K-Ras-PIP2 binding surfaces is likely to inform the future design of therapeutic reagents.

biophysics

B7-H1(PD-L1) confers chemoresistance through ERK and p38 MAPK pathway in tumor cells

Development of resistance to chemotherapy and immunotherapy is a major obstacle in extending the survival of patients with cancer. Although several molecular mechanisms have been identified that can contribute to chemoresistance, the role of immune checkpoint molecules in tumor chemoresistance remains underestimated. It has been recently observed that overexpression of B7-H1(PD-L1) confers chemoresistance in human cancers, however the underlying mechanisms are unclear. Here we show that the development of chemoresistance depends on the increased activation of ERK pathway in tumor cells overexpressing B7-H1. Conversely, B7-H1 deficiency renders tumor cells susceptible to chemotherapy in a cell-context dependent manner through activation of the p38 MAPK pathway. B7-H1 in tumor cells associates with the catalytic subunit of a DNA-dependent serine / threonine protein kinase (DNA-PKcs). DNA-PKcs is required for the activation of ERK or p38 MAPK in tumors expressing B7-H1, but not in B7-H1 negative or B7-H1 deficient tumors. Ligation of B7-H1 by anti-B7-H1 monoclonal antibody (H1A) increased the sensitivity of human triple negative breast tumor cells to cisplatin therapy in vivo. Our results suggest that B7-H1(PD-L1) expression in cancer cells modifies their chemosensitivity towards certain drugs and targeting B7-H1 intracellular signaling pathway is a new way to overcome cancer chemoresistance.

cancer biology

The interplay of immune components and ECM in oral cancer

We report that in Oral Squamous Cell Carcinoma (OSCC), extracellular matrix (ECM) plays a vitally important role in defining the characteristics of cancer vs. normal, as it is a compartment with significant enrichment of known OSCC biomarkers, and the number of genes constituting ECM are more prominently upregulated in OSCC than almost all the rest of the cancer types. This is probably due to the constant exposure of oral cavity to external stimuli, resulting in the ECM remodeling, which further is a key player in tumor invasion. While we showed ECM molecules alone could well distinguish oral cancer from normal tissue samples, a significant portion of these predictive ECM molecules, share the same transcriptional regulator, NFKB1, a master regulator of immune response. We further studied the level of involvement of the immune system in OSCC, and found that the immune composition in OSCC is distinctly different from the other cancer types. OSCC has a higher level of infiltration of adaptive immune cells, including B cell, T cell and neutrophil, compared with other cancer types, while a lower level of infiltration of innate immune, including macrophage and monocyte. Previous studies have revealed the roles of ECM and immune system in OSCC development, and our study showed that ECM plays a very prominent role in OSCC, subject to the complex microenvironment in oral cavity, particularly the immune system profile, and our association analysis revealed it is likely the interactions between ECM and immune cells that define the highly invasive property of OSCC.

bioinformatics

A probabilistic model-based bi-clustering method for single-cell transcriptomic data analysis

We present here novel computational techniques for tackling four problems related to analyses of single-cell RNA-Seq data: (1) a mixture model for coping with multiple cell types in a cell population; (2) a truncated model for handling the unquantifiable errors caused by large numbers of zeros or low-expression values; (3) a bi-clustering technique for detection of sub-populations of cells sharing common expression patterns among subsets of genes; and (4) detection of small cell sub-populations with distinct expression patterns. Through case studies, we demonstrated that these techniques can derive high-resolution information from single-cell data that are not feasible using existing techniques.

bioinformatics

Transcriptome Deconvolution of Heterogeneous Tumor Samples with Immune Infiltration

Transcriptomic deconvolution in cancer and other heterogeneous tissues remains challenging. Available methods lack the ability to estimate both component-specific proportions and expression profiles for individual samples. We present DeMixT, a new tool to deconvolve high dimensional data from mixtures of more than two components. DeMixT implements an iterated conditional mode algorithm and a novel gene-set-based component merging approach to improve accuracy. In a series of experimental validation studies and application to TCGA data, DeMixT showed high accuracy. Improved deconvolution is an important step towards linking tumor transcriptomic data with clinical outcomes. An R package, scripts and data are available: https://github.com/wwylab/DeMixT/.

bioinformatics