bioRxiv ScienceSearch

Biology subjects

Sinha, S.

Publications and source records attributed to Sinha, S..

16 recordsLinked to original sources

A systematic genome-wide mapping of the oncogenic risks associated with CRISPR-Cas9 editing

Recent studies have reported that CRISPR-Cas9 gene editing induces a p53-dependent DNA damage response in primary cells, which may select for cells with oncogenic p53 mutations11,12. It is unclear whether these CRISPR-induced changes are applicable to different cell types, and whether CRISPR gene editing may select for other oncogenic mutations. Addressing these questions, we analyzed genome-wide CRISPR and RNAi screens to systematically chart the mutation selection potential of CRISPR knockouts across the whole exome. Our analysis suggests that CRISPR gene editing can select for mutants of KRAS and VHL, at a level comparable to that reported for p53. These predictions were further validated in a genome-wide manner by analyzing independent CRISPR screens and patients tumor data. Finally, we performed a new set of pooled and arrayed CRISPR screens to evaluate the competition between CRISPR-edited isogenic p53 WT and mutant cell lines, which further validated our predictions. In summary, our study systematically charts and points to the potential selection of specific cancer driver mutations during CRISPR-Cas9 gene editing.

cancer biology

Regularization Improves the Robustness of Learned Sequence-to-Expression Models

Understanding of the gene regulatory activity of enhancers is a major problem in regulatory biology. The nascent field of sequence-to-expression modelling seeks to create quantitative models of gene expression based on regulatory DNA (cis) and cellular environmental (trans) contexts. All quantitative models are defined partially by numerical parameters, and it is common to fit these parameters to data provided by existing experimental results. However, the relative paucity of experimental data appropriate for this task, and lacunae in our knowledge of all components of the systems, results in problems often being under-specified, which in turn may lead to a situation where wildly different model parameterizations perform similarly well on training data. It may also lead to models being fit to the idiosyncrasies of the training data, without representing the more general process (overfitting).\n\nIn other contexts where parameter-fitting is performed, it is common to apply regularization to reduce overfitting. We systematically evaluated the efficacy of three types of regularization in improving the generalizability of trained sequence-to-expression models. The evaluation was performed in two types of cross-validation experiments: one training on D. melanogaster data and predicting on orthologous enhancers from related species, and the other cross-validating between four D. melanogaster neurogenic ectoderm enhancers, which are thought to be under control of the same transcription factors. We show that training with a combination of noise-injection, L1, and L2 regularization can drastically reduce overfitting and improve the generalizability of learned sequence-to-expression models. These results suggest that it may be possible to mitigate the tendency of sequence-to-expression models to overfit available data, thus improving predictive power and potentially resulting in models that provide better insight into underlying biological processes.

systems biology

Inference of phenotype-relevant transcriptional regulatory networks elucidates cancer type-specific regulatory mechanisms in a pan-cancer study

Reconstruction of transcriptional regulatory networks (TRNs) is a powerful approach to unravel the gene expression programs involved in healthy and disease states of a cell. However, these networks are usually reconstructed independent of the phenotypic properties of the samples and therefore cannot identify regulatory mechanisms that are related to a phenotypic outcome of interest. In this study, we developed a new method called InPheRNo to identify phenotype-relevant transcriptional regulatory networks. This method is based on a probabilistic graphical model whose conditional probability distributions model the simultaneous effects of multiple transcription factors (TFs) on their target genes as well as the statistical relationship between target gene expression and phenotype. Extensive comparison of InPheRNo with related approaches using primary tumor samples of 18 cancer types from The Cancer Genome Atlas revealed that InPheRNo can accurately reconstruct cancer type-relevant TRNs and identify cancer driver TFs. In addition, survival analysis revealed that the activity level of TFs with many target genes could distinguish patients with good prognosis from those with poor prognosis.

bioinformatics

A competence-regulated toxin-antitoxin system in Haemophilus influenzae

Natural competence allows bacteria to respond to environmental and nutritional cues by taking up free DNA from their surroundings, thus gaining nutrients and genetic information. In the Gram-negative bacterium Haemophilus influenae, the DNA uptake machinery is induced by the CRP and Sxy transcription factors in response to lack of preferred carbon sources and nucleotide precursors. Here we show that HI0659--which is absolutely required for DNA uptake-- encodes the antitoxin of a competence-regulated toxin-antitoxin operon ( toxTA), likely acquired by horizontal gene transfer from a Streptococcus species. Deletion of the toxin restores uptake to the antitoxin mutant. In addition to the expected Sxy-and CRP-dependent-competence promoter, transcript analysis using RNA-seq identified an internal antitoxin-repressed promoter whose transcription starts within toxT and will yield nonfunctional protein. We present evidence that the most likely effect of unopposed toxin expression is non-specific cleavage of mRNAs and arrest or death of competent cells in the culture, and we show that the toxin gene has been inactivated by deletion in many H. influenzae strains. We suggest that this competence-regulated toxin-antitoxin system may facilitate downregulation of protein synthesis and recycling of nucleotides under starvation conditions, or alternatively be a simple genetic parasite.

microbiology

Evolutionary changes in DNA accessibility and sequence predict divergence of transcription factor binding and enhancer activity

Transcription factor (TF) binding is determined by sequence as well as chromatin accessibility. While the role of accessibility in shaping TF-binding landscapes is well recorded, its role in evolutionary divergence of TF binding, which in turn can alter cis-regulatory activities, is not well understood. In this work, we studied the evolution of genome-wide binding landscapes of five major transcription factors (TFs) in the core network of mesoderm specification, between D. melanogaster and D. virilis, and examined its relationship to accessibility and sequence-level changes. We generated chromatin accessibility data from three important stages of embryogenesis in both D. melanogaster and D. virilis, and recorded conservation and divergence patterns. We then used multi-variable models to correlate accessibility and sequence changes to TF binding divergence. We found that accessibility changes can in some cases, e.g., for the master regulator Twist and for earlier developmental stages, more accurately predict binding change than is possible using TF binding motif changes between orthologous enhancers. Accessibility changes also explain a significant portion of the co-divergence of TF pairs. We noted that accessibility and motif changes offer complementary views of the evolution of TF binding, and developed a combined model that captures the evolutionary data much more accurately than either view alone. Finally, we trained machine learning models to predict enhancer activity from TF binding, and used these functional models to argue that motif and accessibility-based predictors of TF binding change can substitute for experimentally measured binding change, for the purpose of predicting evolutionary changes in enhancer activity.

evolutionary biology

An information theoretic treatment of sequence-to-expression modeling

Studying a genes regulatory mechanisms is a tedious process that involves identification of candidate regulators by transcription factor (TF) knockout or over-expression experiments, delineation of enhancers by reporter assays, and demonstration of direct TF influence by site mutagenesis, among other approaches. Such experiments are often chosen based on the biologists intuition, from several testable hypotheses. We pursue the goal of making this process systematic by using ideas from information theory to reason about experiments in gene regulation, in the hope of ultimately enabling rigorous experiment design strategies. For this, we make use of a state-of-the-art mathematical model of gene expression, which provides a way to formalize our current knowledge of cis- as well as trans-regulatory mechanisms of a gene. Ambiguities in such knowledge can be expressed as uncertainties in the model, which we capture formally by building an ensemble of plausible models that fit the existing data and defining a probability distribution over the ensemble. We then characterize the impact of a new experiment on our understanding of the genes regulation based on how the ensemble of plausible models and its probability distribution changes when challenged with results from that experiment. This allows us to assess the value of the experiment retroactively as the reduction in entropy of the distribution (information gain) resulting from the experiments results. We fully formalize this novel approach to reasoning about gene regulation experiments and use it to evaluate a variety of perturbation experiments on two developmental genes of D. melanogaster. We also provide objective and biologist-friendly descriptions of the information gained from each such experiment. The rigorously defined information theoretic approaches presented here can be used in the future to formulate systematic strategies for experiment design pertaining to studies of gene regulatory mechanisms.\n\nAuthor summaryIn-depth studies of gene regulatory mechanisms employ a variety of experimental approaches such as identifying a genes enhancer(s) and testing its variants through reporter assays, followed by transcription factor mis-expression or knockouts, site mutagenesis, etc. The biologist is often faced with the challenging problem of selecting the ideal next experiment to perform so that its results provide novel mechanistic insights, and has to rely on their intuition about what is currently known on the topic and which experiments may add to that knowledge. We seek to make this intuition-based process more systematic, by borrowing ideas from the mature statistical field of experiment design. Towards this goal, we use the language of mathematical models to formally describe what is known about a genes regulatory mechanisms, and how an experiments results enhance that knowledge. We use information theoretic ideas to assign a value to an experiment as well as explain objectively what is learned from that experiment. We demonstrate use of this novel approach on two extensively studied developmental genes in fruitfly. We expect our work to lead to systematic strategies for selecting the most informative experiments in a study of gene regulation.

systems biology

Finding reliable phenotypes and detecting artefacts among in vivo and in vitro assays to characterize the refractory transcriptional activator Sxy (TfoX) in Escherichia coli

The Sxy (TfoX) protein is required for expression of a distinct subset of the genes regulated by the cAMP receptor protein (CRP) in the model organisms Escherichia coli, Haemophilus influenzae, and Vibrio cholerae. Genetic studies have established that CRP and Sxy co-activate transcription at gene promoters containing DNA binding sites called CRP-S sites. In contrast, CRP acts without Sxy at gene promoters containing canonical CRP-N sites, suggesting that Sxy makes physical contacts with CRP and/or DNA to assist in transcriptional activation at CRP-S promoters. Despite growing interest in Sxys activity as a transcription factor, Sxy remains poorly characterized due to a lack of reliable phenotypes in E. coli. Experiments are further hampered by growth inhibition and formation of inclusion bodies when Sxy is overexpressed. In this study we applied diverse phenotypic and molecular assays to test for postulated Sxy functions and interactions. Mutations in conserved regions of Sxy and truncations in the Sxy C-terminus abolish transcriptional activation of a CRP-S promoter, and a 37 amino acid truncation of the C-terminus relieves the growth inhibition normally caused by Sxy overexpression. Sxy was unable to augment weakened CRP interactions to restore carbon metabolism phenotypes. Bandshift analysis and chromatin pull-down assays of Sxy-CRP-DNA interactions yielded intriguing evidence of CRP-Sxy and Sxy-DNA physical interactions. However, despite the careful application of standard protein purification protocols and quality control steps for nickel affinity column purification, protein mass spectrometry revealed the enrichment of additional DNA-binding proteins in nickel column eluates, presenting a probable source of artefactual protein-protein and protein-DNA interaction results. These findings highlight the importance of extensive controls and phenotypic assays for the study of poorly characterized and recalcitrant proteins like Sxy.

microbiology

Emergent memory in cell signaling: Persistent adaptive dynamics in cascades can arise from the diversity of relaxation time-scales

The mitogen-activated protein kinase (MAPK) signaling cascade, an evolutionarily conserved motif present in all eukaryotic cells, is involved in coordinating critical cell-fate decisions, regulating protein synthesis, and mediating learning and memory. While the steady-state behavior of the pathway stimulated by a time-invariant signal is relatively well-understood, we show using a computational model that it exhibits a rich repertoire of transient adaptive responses to changes in stimuli. When the signal is switched on, the response is characterized by long-lived modulations in frequency as well as amplitude. On withdrawing the stimulus, the activity decays over timescales much longer than that of phosphorylation-dephosphorylation processes, exhibiting reverberations characterized by repeated spiking in the activated MAPK concentration. The long-term persistence of such post-stimulus activity suggests that the cascade retains memory of the signal for a significant duration following its removal, even in the absence of any explicit feedback or cross-talk with other pathways. We find that the molecular mechanism underlying this behavior is related to the existence of distinct relaxation rates for the different cascade components. This results in the imbalance of fluxes between different layers of the cascade, with the repeated reuse of activated kinases as enzymes when they are released from sequestration in complexes leading to one or more spike events following the removal of the stimulus. The persistent adaptive response reported here, indicative of a cellular \"short-term\" memory, suggests that this ubiquitous signaling pathway plays an even more central role in information processing by eukaryotic cells.

cell biology

An epithelial-mesenchymal-amoeboid transition gene signature reveals molecular subtypes of breast cancer progression and metastasis

Cancer cells within a tumor are known to display varying degrees of metastatic propensity but the molecular basis underlying such heterogeneity remains unclear. We analyzed genome-wide gene expression data obtained from primary tumors of lymph node-negative breast cancer patients using a novel metastasis biology-based Epithelial-Mesenchymal-Amoeboid Transition (EMAT) gene signature, and identified subtypes associated with distinct prognostic profiles. EMAT subtype status improved prognosis accuracy of clinical parameters and statistically outperformed traditional breast cancer intrinsic subtypes even after adjusting for treatment variables. Additionally, analysis of 3D spheroids from an in vitro isogenic model of breast cancer progression reveals that EMAT subtypes display progression from premalignant to malignant and pre-invasive to invasive cancer. EMAT classification is a biologically informed method to assess metastasis risk in early stage, lymph node-negative breast cancer patients.

cancer biology

Cross-species systems analyses reveal a conserved brain transcriptional response to social challenge

Social challenges like territorial intrusions evoke behavioral responses in widely diverging species. Recent work has revealed that evolutionary \"toolkits\" - genes and modules with lineage-specific variations but deep conservation of function - participate in the behavioral response to social challenge. Here, we develop a multi-species computational-experimental approach to characterize such a toolkit at a systems level. Brain transcriptomic responses to social challenge was probed via RNA-seq profiling in three diverged species - honey bees, mice, and three-spined stickleback fish - following a common methodology, allowing fair comparisons across species. Data were collected from multiple brain regions and multiple time points after social challenge exposure, achieving anatomical and temporal resolution substantially greater than previous work. We developed statistically rigorous analyses equipped to find homologous functional groups among these species at the levels of individual genes, functional and coexpressed gene modules, and transcription factor sub-networks. We identified six orthogroups involved in response to social challenge, including groups represented by mouse genes Npas4 and Nr4a1, as well as common modulation of systems such as transcriptional regulators, ion channels, G-protein coupled receptors, and synaptic proteins. We also identified conserved coexpression modules enriched for mitochondrial fatty acid metabolism and heat shock that constitute the shared neurogenomic response. Our analysis suggests a toolkit wherein nuclear receptors, interacting with chaperones, induce transcriptional changes in mitochondrial activity, neural cytoarchitecture, and synaptic transmission after social challenge. It reveals systems-level mechanisms that have been repeatedly co-opted during evolution of analogous behaviors, thus advancing the genetic toolkit concept beyond individual genes.

genomics

A closer look at cross-validation for assessing the accuracy of gene regulatory networks and models

Cross-validation (CV) is a technique to assess the generalizability of a model to unseen data. This technique relies on assumptions that may not be satisfied when studying genomics datasets. For example, random CV (RCV) assumes that a randomly selected set of samples, the test set, well represents unseen data. This assumption does not hold true where samples are obtained from different experimental conditions, and the goal is to learn regulatory relationships among the genes that generalize beyond the observed conditions. In this study, we investigated how the CV procedure affects the assessment of methods used to learn gene regulatory networks. We compared the performance of a regression-based method for gene expression prediction, estimated using RCV with that estimated using a clustering-based CV (CCV) procedure. Our analysis illustrates that RCV can produce over-optimistic estimates of generalizability of the model compared to CCV. Next, we defined the distinctness of a test set from a training set and showed that this measure is predictive of the performance of the regression method. Finally, we introduced a simulated annealing method to construct partitions with gradually increasing distinctness and showed that performance of different gene expression prediction methods can be better evaluated using this method.

bioinformatics

Sensitivity analysis based ranking reveals unknown biological hypotheses for down regulated genes in time buffer during administration of PORCN-WNT inhibitor ETC-1922159 in CRC

In a recent development of the PORCN-WNT inhibitor ETC-1922159 for colorectal cancer, a list of down-regulated genes were recorded in a time buffer after the administration of the drug. The regulation of the genes were recorded individually but it is still not known which higher ([≥] 2) order interactions might be playing a greater role after the administration of the drug. In order to reveal the priority of these higher order interactions among the down-regulated genes or the likely unknown biological hypotheses, a search engine was developed based on the sensitivity indices of the higher order interactions that were ranked using a support vector ranking algorithm and sorted. For example, LGR family (Wnt signal enhancer) is known to neutralize RNF43 (Wnt inhibitor). After the administration of ETC-1922159 it was found that using HSIC (and rbf, linear and laplace variants of kernel) the rankings of the interaction between LGR5-RNF43 were 61, 114 and 85 respectively. Rankings for LGR6-RNF43 were 1652, 939 and 805 respectively. The down-regulation of LGR family after the drug treatment is evident in these rankings as it takes bottom priorities for LGR5-RNF43 interaction. The LGR6-RNF43 takes higher ranking than LGR5-RNF43, indicating that it might not be playing a greater role as LGR5 during the Wnt enhancing signals. These rankings confirm the efficacy of the proposed search engine design. Conclusion: Prioritized unknown biological hypothesis form the basis of further wet lab tests with the aim to reduce the cost of (1) wet lab experiments (2) combinatorial search and (3) lower the testing time for biologist who search for influential interactions in a vast combinatorial search forest. From in silico perspective, a framework for a search engine now exists which can generate rankings for nth order interactions in Wnt signaling pathway, thus revealing unknown/untested/unexplored biological hypotheses and aiding in understanding the mechanism of the pathway. The generic nature of the design can be applied to any signaling pathway or phenomena under investigation where a prioritized order of interactions among the involved factors need to be investigated for deeper understanding. Future improvements of the design are bound to facilitate medical specialists/oncologists in their respective investigations.\n\nSignificanceRecent development of PORCN-WNT inhibitor enantiomer ETC-1922159 cancer drug show promise in suppressing some types of colorectal cancer. However, the search and wet lab testing of unknown/unexplored/untested biological hypotheses in the form of combinations of various intra/ extracellular factors/genes/proteins affected by ETC-1922159 is not known. Currently, a major problem in biology is to cherry pick the combinations based on expert advice, literature survey or guesses to investigate a particular combinatorial hypothesis. A search engine has be developed to reveal and prioritise these unknown/untested/unexplored combinations affected by the inhibitor. These ranked unknown biological hypotheses facilitate in narrowing down the investigation in a vast combinatorial search forest of ETC-1922159 affected synergistic-factors.

cancer biology

Cell growth rate dictates the onset of glass to fluid-like transition and long time super-diffusion in an evolving cell colony

Collective migration dominates many phenomena, from cell movement in living systems to abiotic self-propelling particles. Focusing on the early stages of tumor evolution, we enunciate the principles involved in cell dynamics and highlight their implications in understanding similar behavior in seemingly unrelated soft glassy materials and possibly chemokine-induced migration of CD8+ T cells. We performed simulations of tumor invasion using a minimal three dimensional model, accounting for cell elasticity and adhesive cell-cell interactions as well as cell birth and death to establish that cell growth rate-dependent tumor expansion results in the emergence of distinct topological niches. Cells at the periphery move with higher velocity perpendicular to the tumor boundary, while motion of interior cells is slower and isotropic. The mean square displacement, {Delta}(t), of cells exhibits glassy behavior at times comparable to the cell cycle time, while exhibiting super-diffusive behavior, {Delta}(t) {approx} t ( > 1), at longer times. We derive the value of {approx} 1.33 using a field theoretic approach based on stochastic quantization. In the process we establish the universality of super-diffusion in a class of seemingly unrelated non-equilibrium systems. Super diffusion at long times arises only if there is an imbalance between cell birth and death rates. Our findings for the collective migration, which also suggests that tumor evolution occurs in a polarized manner, are in quantitative agreement with in vitro experiments. Although set in the context of tumor invasion the findings should also hold in describing collective motion in growing cells and in active systems where creation and annihilation of particles play a role.

biophysics

Identification of Pathways Associated with Chemosensitivity through Network Embedding

Basal gene expression levels have been shown to be predictive of cellular response to cytotoxic treatments. However, such analyses do not fully reveal complex genotype-phenotype relationships, which are partly encoded in highly interconnected molecular networks. Biological pathways provide a complementary way of understanding drug response variation among individuals. In this study, we integrate chemosensitivity data from a recent pharmacogenomics study with basal gene expression data from the CCLE project and prior knowledge of molecular networks to identify specific pathways mediating chemical response. We first develop a computational method called PACER, which ranks pathways for enrichment in a given set of genes using a novel network embedding method. It examines known relationships among genes as encoded in a molecular network along with gene memberships of all pathways to determine a vector representation of each gene and pathway in the same low-dimensional vector space. The relevance of a pathway to the given gene set is then captured by the similarity between the pathway vector and gene vectors. To apply this approach to chemosensitivity data, we identify genes with basal expression levels in a panel of cell lines that are correlated with cytotoxic response to a compound, and then rank pathways for relevance to these response-correlated genes using PACER. Extensive evaluation of this approach on benchmarks constructed from databases of compound target genes, compound chemical structure, as well as large collections of drug response signatures demonstrates its advantages in identifying compound-pathway associations, compared to existing statistical methods of pathway enrichment analysis. The associations identified by PACER can serve as testable hypotheses about chemosensitivity pathways and help further study the mechanism of action of specific cytotoxic drugs. More broadly, PACER represents a novel technique of identifying enriched properties of any gene set of interest while also taking into account networks of known gene-gene relationships and interactions.

pharmacology and toxicology

Principled Multi-Omic Analysis Reveals Gene Regulatory Mechanisms Of Phenotype Variation

Recent studies have analyzed large scale data sets of gene expression to identify genes associated with inter-individual variation in phenotypes ranging from cancer sub-types to drug sensitivity, promising new avenues of research in personalized medicine. However, gene expression data alone is limited in its ability to reveal cis-regulatory mechanisms underlying phenotypic differences. In this study, we develop a new probabilistic model, called pGENMi, that integrates multi-omics data to investigate the transcriptional regulatory mechanisms underlying inter-individual variation of a specific phenotype - that of cell line response to cytotoxic treatment. In particular, pGENMi simultaneously analyzes genotype, DNA methylation, gene expression and transcription factor (TF)-DNA binding data, along with phenotypic measurements, to identify TFs regulating the phenotype. It does so by combining statistical information about expression quantitative trait loci (eQTLs) and expression-correlated methylation marks (eQTMs) located within TF binding sites, as well as observed correlations between gene expression and phenotype variation. Application of pGENMi to data from a panel of lymphoblastoid cell lines treated with 24 drugs, in conjunction with ENCODE TF ChIP data, yielded a number of known as well as novel TF-drug associations. Experimental validations by TF knock-down confirmed 41% of the predicted and tested associations, compared to a 12% confirmation rate of tested non-associations (controls). Extensive literature survey also corroborated 62% of the predicted associations above a stringent threshold. Moreover, associations predicted only when combining eQTL and eQTM data showed higher precision compared to an eQTL-only or eQTM-only analysis with the same method, further demonstrating the value of multi-omic integrative analysis.

genomics

Knowledge-Guided Prioritization of Genes Determinant of Drug Response using ProGENI

BackgroundIdentification of genes whose basal mRNA expression predicts the sensitivity of tumor cells to cytotoxic treatments can play an important role in individualized cancer medicine. It enables detailed characterization of the mechanism of action of drugs. Furthermore, screening the expression of these genes in the tumor tissue may suggest the best course of chemotherapy or a combination of drugs to overcome drug resistance.\n\nResultsWe developed a computational method called ProGENI to identify genes most associated with the variation of drug response across different individuals, based on gene expression data. In contrast to existing methods, ProGENI also utilizes prior knowledge of protein-protein and genetic interactions, using random walk techniques. Analysis of two relatively new and large datasets including gene expression data on hundreds of cell lines and their cytotoxic responses to a large compendium of drugs reveals a significant improvement in prediction of drug sensitivity using genes identified by ProGENI compared to other methods. Our siRNA knockdown experiments on ProGENI-identified genes confirmed the role of many new genes in sensitivity to three chemotherapy drugs: cisplatin, docetaxel and doxorubicin. Based on such experiments and extensive literature survey, we demonstrate that about 73% our top predicted genes modulate drug response in selected cancer cell lines. In addition, global analysis of genes associated with groups of drugs uncovered pathways of cytotoxic response shared by each group.\n\nConclusionsOur results suggest that knowledge-guided prioritization of genes using ProGENI gives new insight into mechanisms of drug resistance and identifies genes that may be targeted to overcome this phenomenon.

bioinformatics