bioRxiv ScienceSearch

Biology subjects

Krishnaswamy, S.

Publications and source records attributed to Krishnaswamy, S..

6 recordsLinked to original sources

Protein Docking using Constrained Self-adaptive Differential Evolution Algorithm

The objective of protein docking is to achieve a relative orientation and an optimized conformation between two proteins that results in a stable structure with the minimized potential energy. Constrained Self-adaptive Differential Evolution (Cons_SaDE) algorithm is used to find the minimum energy conformation using proposed constraints such as boundary surface complementary interactions, non-bonded inter-atomic allowed distances, and finding of interaction and non-interaction sites. With these constraints, Cons_SaDE is efficient enough to explore the promising solutions by gradually self-adapting the strategies and parameters learnt from their previous experiences. Modified sampling scheme called Rotate Only Representation is used to represent a docking conformation. GROMOS53A6 force field is used to find the potential energy. To test the performance of this algorithm, few bound and unbound complexes from Protein Data Bank (PDB) and few easy, medium and difficult complexes from Zlab benchmark4.0 are used. Buried Surface Area, Root Mean Square Deviation (RMSD) and Correlation Coefficient are some of the metrics applied to evaluate the best docked conformations. RMSD values of the best docked conformations obtained from five popular docking web servers are compared with Cons_SaDE results and nonparametric statistical tests for multiple comparisons with control method are implemented to show the performance of this algorithm. Cons_SaDE has produced good quality solutions for the most of the data sets considered.

bioinformatics

Exploring Single-Cell Data with Multitasking Deep Neural Networks

Biomedical researchers are generating high-throughput, high-dimensional single-cell data at a staggering rate. As costs of data generation decrease, experimental design is moving towards measurement of many different single-cell samples in the same dataset. These samples can correspond to different patients, conditions, or treatments. While scalability of methods to datasets of these sizes is a challenge on its own, dealing with large-scale experimental design presents a whole new set of problems, including batch effects and sample comparison issues. Currently, there are no computational tools that can both handle large amounts of data in a scalable manner (many cells) and at the same time deal with many samples (many patients or conditions). Moreover, data analysis currently involves the use of different tools that each operate on their own data representation, not guaranteeing a synchronized analysis pipeline. For instance, data visualization methods can be disjoint and mismatched with the clustering method. For this purpose, we present SAUCIE, a deep neural network that leverages the high degree of parallelization and scalability offered by neural networks, as well as the deep representation of data that can be learned by them to perform many single-cell data analysis tasks, all on a unified representation.\n\nA well-known limitation of neural networks is their interpretability. Our key contribution here are newly formulated regularizations (penalties) that render features learned in hidden layers of the neural network interpretable. When large multi-patient datasets are fed into SAUCIE, the various hidden layers contain denoised and batch-corrected data, a low dimensional visualization, unsupervised clustering, as well as other information that can be used to explore the data. We show this capability by analyzing a newly generated 180-sample dataset consisting of T cells from dengue patients in India, measured with mass cytometry. We show that SAUCIE, for the first time, can batch correct and process this 11-million cell data to identify cluster-based signatures of acute dengue infection and create a patient manifold, stratifying immune response to dengue on the basis of single-cell measurements.

bioinformatics

Sequence fingerprints distinguish erroneous from correct predictions of Intrinsically Disordered Protein Regions

More than sixty prediction methods for intrinsically disordered proteins (IDPs) have been developed over the years, many of which are accessible on the world-wide web. Nearly, all of these predictors give balanced accuracies in the ~65% to ~80% range. Since predictors are not perfect, further studies are required to uncover the role of amino acid residues in native IDP as compared to predicted IDP regions. In the present work, we make use of sequences of 100% predicted IDP regions, false positive disorder predictions, and experimentally determined IDP regions to distinguish the characteristics of native versus predicted IDP regions. A higher occurrence of asparagine is observed in sequences of native IDP regions but not in sequences of false positive predictions of IDP regions. The occurrences of certain combinations of amino acids at the pentapeptide level provide a distinguishing feature in the IDPs with respect to globular proteins. The distinguishing features presented in this paper provide insights into the sequence fingerprints of amino acid residues in experimentally-determined as compared to predicted IDP regions. These observations and additional work along these lines should enable the development of improvements in the accuracy of disorder prediction algorithm.

bioinformatics

Learning Edge Rewiring in EMT from Single Cell Data

Cellular regulatory networks are not static, but continuously reconfigure in response to stimuli via alterations in gene expression and protein confirmations. However, typical computational approaches treat them as static interaction networks derived from a single experimental time point. Here, we provide a method for learning the dynamic modulation, or rewiring of pairwise relationships (edges) from a static single-cell data. We use the epithelial-to-mesenchymal transition (EMT) in murine breast cancer cells as a model system, and measure mass cytometry data three days after induction of the transition by TGF{beta}. We take advantage of transitional rate variability between cells in the data by deriving a pseudo-time EMT trajectory. Then we propose methods for visualizing and quantifying time-varying edge behavior over the trajectory and use these methods: TIDES (Trajectory Imputed DREMI scores), and measure of edge dynamism (3DDREMI) to predict and validate the effect of drug perturbations on EMT.

systems biology

PHATE: A Dimensionality Reduction Method for Visualizing Trajectory Structures in High-Dimensional Biological Data

With the advent of high-throughput technologies measuring high-dimensional biological data, there is a pressing need for visualization tools that reveal the structure and emergent patterns of data in an intuitive form. We present PHATE, a visualization method that captures both local and global nonlinear structure in data by an information-geometric distance between datapoints. We perform extensive comparison between PHATE and other tools on a variety of artificial and biological datasets, and find that it consistently preserves a range of patterns in data including continual progressions, branches, and clusters. We define a manifold preservation metric DEMaP to show that PHATE produces quantitatively better denoised embeddings than existing visualization methods. We show that PHATE is able to gain unique insight from a newly generated scRNA-seq dataset of human germ layer differentiation. Here, PHATE reveals a dynamic picture of the main developmental branches in unparalleled detail, including the identification of three novel subpopulations. Finally, we show that PHATE is applicable to a wide variety of datatypes including mass cytometry, single-cell RNA-sequencing, Hi-C, and gut microbiome data, where it can generate interpretable insights into the underlying systems.

bioinformatics

MAGIC: A diffusion-based imputation method reveals gene-gene interactions in single-cell RNA-sequencing data

Single-cell RNA-sequencing is fast becoming a major technology that is revolutionizing biological discovery in fields such as development, immunology and cancer. The ability to simultaneously measure thousands of genes at single cell resolution allows, among other prospects, for the possibility of learning gene regulatory networks at large scales. However, scRNA-seq technologies suffer from many sources of significant technical noise, the most prominent of which is dropout due to inefficient mRNA capture. This results in data that has a high degree of sparsity, with typically only ~10% non-zero values. To address this, we developed MAGIC (Markov Affinity-based Graph Imputation of Cells), a method for imputing missing values, and restoring the structure of the data. After MAGIC, we find that two- and three-dimensional gene interactions are restored and that MAGIC is able to impute complex and non-linear shapes of interactions. MAGIC also retains cluster structure, enhances cluster-specific gene interactions and restores trajectories, as demonstrated in mouse retinal bipolar cells, hematopoiesis, and our newly generated epithelial-to-mesenchymal transition dataset.

bioinformatics