bioRxiv ScienceSearch

Biology subjects

Nagarajan, N.

Publications and source records attributed to Nagarajan, N..

8 recordsLinked to original sources

Structure mapping of dengue and Zika viruses reveals new functional long-range interactions

Dengue and Zika are clinically important members of the Flaviviridae family that utilizes an 11kb positive strand RNA for genome regulation. While structures have been mapped primarily in the UTRs, much remains to be learnt about how the rest of the genome folds to enable function. Here, we performed secondary structure and pair-wise interaction mapping on four dengue serotypes and four Zika strains in their native virus particles and infected cells. Comparative analysis of SHAPE reactivities across serotypes nominated potentially functional regions that are highly structured, show structure conservation, and low synonymous mutation rates, including a structure associated with ribosome pausing. Pair-wise interaction mapping by SPLASH further reveals new pair-wise interactions, in addition to the known circularization sequence. 40% of pair-wise interactions form alternative structures, suggesting extensive structural heterogeneity. Analysis of shared pair-wise interactions between serotypes revealed macro-organization whereby interactions are preserved at their physical locations, beyond their sequence identities. In addition, structure mapping of virus genomes released in solution-as well as inside host cells-showed that other helicases, in addition to the ribosome, play a role in unwinding viral structures inside cells. Mutational experiments that disrupt in cell and in virion pair-wise interactions result in virus attenuation, demonstrating their importance during the virus life-cycle.

genomics

Gut microbiome recovery after antibiotic usage is mediated by specific bacterial species

Dysbiosis in the gut microbiome due to antibiotic usage can persist for extended periods of time, impacting host health and increasing the risk for pathogen colonization. The specific factors associated with variability in gut microbiome recovery remain unknown. Using data from 4 different cohorts in 3 continents comprising >500 microbiome profiles from 117 subjects, we identified 20 bacterial species exhibiting robust association with gut microbiome recovery post antibiotic therapy. Functional and growth analysis showed that microbiome recovery is supported by enrichment in carbohydrate degradation and energy production capabilities. Association rule mining on 782 microbiome profiles from the MEDUSA database enabled reconstruction of the gut microbial food-web, identifying many recovery-associated bacteria (RABs) as primary colonizing species, with the ability to use both host and diet-derived energy sources, and to break down complex carbohydrates to support the growth of other bacteria. Experiments in a mouse model recapitulated the ability of RABs (Bacteroides thetaiotamicron and Bifidobacterium adolescentis) to promote microbiome recovery with synergistic effects, providing a two orders of magnitude boost to microbial abundance in early time-points and faster maturation of microbial diversity. The identification of specific microbial factors promoting microbiome recovery opens up opportunities for rationally fine-tuning pre- and probiotic formulations that prevent pathogen colonization and promote gut health.

genomics

System Biology Modeling with Compositional Microbiome Data Reveals Personalized Gut Microbial Dynamics and Keystone Species

A growing body of literature points to the important roles that different microbial communities play in diverse natural environments and the human body. The dynamics of these communities is driven by a range of microbial interactions from symbiosis to predator-prey relationships, the majority of which are poorly understood, making it hard to predict the response of the community to different perturbations. With the increasing availability of high-throughput sequencing based community composition data, it is now conceivable to directly learn models that explicitly define microbial interactions and explain community dynamics. The applicability of these approaches is however affected by several experimental limitations, particularly the compositional nature of sequencing data. We present a new computational approach (BEEM) that addresses this key limitation in the inference of generalised Lotka-Volterra models (gLVMs) by coupling biomass estimation and model inference in an expectation maximization like algorithm (BEEM). Surprisingly, BEEM outperforms state-of-the-art methods for inferring gLVMs, while simultaneously eliminating the need for additional experimental biomass data as input. BEEMs application to previously inaccessible public datasets (due to the lack of biomass data) allowed us for the first time to analyse microbial communities in the human gut on a per individual basis, revealing personalised dynamics and keystone species.

systems biology

A MinION-based pipeline for fast and cost-effective DNA barcoding

DNA barcodes are useful for species discovery and species identification, but obtaining barcodes currently requires a well-equipped molecular laboratory, is time-consuming, and/or expensive. We here address these issues by developing a barcoding pipeline for Oxford Nanopore MinION and demonstrate that one flowcell can generate barcodes for [~]500 specimens despite high base-call error rates of MinION. The pipeline overcomes the errors by first summarizing all reads for the same tagged amplicon as a consensus barcode. These barcodes are overall mismatch-free but retain indel errors that are concentrated in homopolymeric regions. We thus complement the barcode caller with an optional error correction pipeline that uses conserved amino-acid motifs from publicly available barcodes to correct the indel errors. The effectiveness of this pipeline is documented by analysing reads from three MinION runs that represent three different stages of MinION development. They generated data for (1) 511 specimens of a mixed Diptera sample, (2) 575 specimens of ants, and (3) 50 specimens of Chironomidae. The run based on the latest chemistry yielded MinION barcodes for 490 specimens which were assessed against reference Sanger barcodes (N=471). Overall, the MinION barcodes have an accuracy of 99.3%-100% and the number of ambiguities ranges from <0.01-1.5% depending on which correction pipeline is used. We demonstrate that it requires only 2 hours of sequencing to gather all information that is needed for obtaining reliable barcodes for most specimens (>90%). We estimate that up to 1000 barcodes can be generated in one flowcell and that the cost of a MinION barcode can be <USD 2.

bioinformatics

Predicting Cancer Drug Response Using a Recommender System

MotivationAs we move towards an era of precision medicine, the ability to predict patient-specific drug responses in cancer based on molecular information such as gene expression data represents both an opportunity and a challenge. In particular, methods are needed that can accommodate the high-dimensionality of data to learn interpretable models capturing drug response mechanisms, as well as providing robust predictions across datasets.\n\nResultsWe propose a method based on ideas from \"recommender systems\" (CaDRReS) that predicts cancer drug responses for unseen cell-lines/patients based on learning projections for drugs and cell-lines into a latent \"pharmacogenomic\" space. Comparisons with other proposed approaches for this problem based on large public datasets (CCLE, GDSC) shows that CaDRReS provides consistently good models and robust predictions even across unseen patient-derived cell-line datasets. Analysis of the pharmacogenomic spaces inferred by CaDRReS also suggests that they can be used to understand drug mechanisms, identify cellular subtypes, and further characterize drug-pathway associations.\n\nAvailabilitySource code and datasets are available at https://github.com/CSB5/CaDRReS\n\nContactnagarajann@gis.a-star.edu.sg\n\nSupplementary informationSupplementary data are available online.

bioinformatics

Single-Virion Sequencing Of Lamivudine Treated HBV Populations Reveal Population Evolution Dynamics And Demographic History

Viral populations are complex, dynamic, and fast evolving. The evolution of groups of closely related viruses in a competitive environment is termed quasispecies. To fully understand the role that quasispecies play in viral evolution, characterizing the trajectories of viral genotypes in an evolving population is the key. In particular, long-range haplotype information for thousands of individual viruses is critical; yet generating this information is non-trivial. Popular deep sequencing methods generate relatively short reads that do not preserve linkage information, while third generation sequencing methods have higher error rates that make detection of low frequency mutations a bioinformatics challenge. Here we applied BAsE-Seq, an Illumina-based single-virion sequencing technology, to eight samples from four chronic hepatitis B (CHB) patients - once before antiviral treatment and once after viral rebound due to resistance. We obtained 248-8,796 single-virion sequences per sample, which allowed us to find evidence for both hard and soft selective sweeps. We were also able to reconstruct population demographic history that was independently verified by clinically collected data. We further verified four of the samples independently on PacBio and Illumina sequencers. Overall, we showed that single-virion sequencing yields insight into viral evolution and population dynamics in an efficient and high throughput manner. We believe that single-virion sequencing is widely applicable to the study of viral evolution in the context of drug resistance, differentiating between soft or hard selective sweeps, and the reconstruction of intra-host viral population demographic history.

evolutionary biology

ConsensusDriver Improves Upon Individual Algorithms For Predicting Driver Alterations In Different Cancer Types And Individual Patients -- A Toolbox For Precision Oncology

BackgroundIn recent years, several large-scale cancer genomics studies have helped generate detailed molecular profiling datasets for many cancer types and thousands of patients. These datasets provide a unique resource for studying cancer driver prediction methods and their utility for precision oncology, both to predict driver genetic alterations in patient subgroups (e.g. defined by histology or clinical phenotype) or even individual patients.\n\nMethodsWe performed the most comprehensive assessment to date of 18 driver gene prediction methods, on more than 3,400 tumour samples, from 15 cancer types, to determine their suitability in guiding precision medicine efforts. These methods have diverse approaches, which can be classified into five categories: functional impact on proteins in general (FI) or specific to cancer (FIC), cohort-based analysis for recurrent mutations (CBA), mutations with expression correlation (MEC) and methods that use gene interaction network-based analysis (INA).\n\nResultsThe performance of driver prediction methods varies considerably, with concordance with a gold-standard varying from 9% to 68%. FI methods show relatively poor performance (concordance <22%) while CBA methods provide conservative results, but require large sample sizes for high sensitivity. INA methods, through the integration of genomic and transcriptomic data, and FIC methods, by training cancer-specific models, provide the best trade-off between sensitivity and specificity. As the methods were found to predict different subsets of drivers, we propose a novel consensus-based approach, ConsensusDriver, which significantly improves the quality of predictions (20% increase in sensitivity). This tool can be applied to predict driver alterations in patient subgroups (e.g. defined by histology or clinical phenotype) or even individual patients.\n\nConclusionExisting cancer driver prediction methods are based on very different assumptions and each of them can only detect a particular subset of driver events. Consensus-based methods, like ConsensusDriver, are thus a promising approach to harness the strengths of different driver prediction paradigms.

bioinformatics

Critical Assessment of Metagenome Interpretation - a benchmark of computational metagenomics software

In metagenome analysis, computational methods for assembly, taxonomic profiling and binning are key components facilitating downstream biological data interpretation. However, a lack of consensus about benchmarking datasets and evaluation metrics complicates proper performance assessment. The Critical Assessment of Metagenome Interpretation (CAMI) challenge has engaged the global developer community to benchmark their programs on datasets of unprecedented complexity and realism. Benchmark metagenomes were generated from ~700 newly sequenced microorganisms and ~600 novel viruses and plasmids, including genomes with varying degrees of relatedness to each other and to publicly available ones and representing common experimental setups. Across all datasets, assembly and genome binning programs performed well for species represented by individual genomes, while performance was substantially affected by the presence of related strains. Taxonomic profiling and binning programs were proficient at high taxonomic ranks, with a notable performance decrease below the family level. Parameter settings substantially impacted performances, underscoring the importance of program reproducibility. While highlighting current challenges in computational metagenomics, the CAMI results provide a roadmap for software selection to answer specific research questions.

bioinformatics