bioRxiv ScienceSearch

Biology subjects

Kumar, V.

Publications and source records attributed to Kumar, V..

29 records · Page 2Linked to original sources

Full-Length Envelope Analyzer (FLEA): A tool for longitudinal analysis of viral amplicons

Next generation sequencing of viral populations has advanced our understanding of viral population dynamics, the development of drug resistance, and escape from host immune responses. Many applications require complete gene sequences, which can be impossible to reconstruct from short reads. HIV-1 env, the protein of interest for HIV vaccine studies, is exceptionally challenging for long-read sequencing and analysis due to its length, high substitution rate, and extensive indel variation. While long-read sequencing is attractive in this setting, the analysis of such data is not well handled by existing methods. To address this, we introduce FLEA (Full-Length Envelope Analyzer), which performs end-to-end analysis and visualization of long-read sequencing data.\n\nFLEA consists of both a pipeline (optionally run on a high-performance cluster), and a client-side web application that provides interactive results. The pipeline transforms FASTQ reads into high-quality consensus sequences (HQCSs) and uses them to build a codon-aware multiple sequence alignment. The resulting alignment is then used to infer phylogenies, selection pressure, and evolutionary dynamics. The web application provides publication-quality plots and interactive visualizations, including an annotated viral alignment browser, time series plots of evolutionary dynamics, visualizations of gene-wide selective pressures (such as dN /dS) across time and across protein structure, and a phylogenetic tree browser.\n\nWe demonstrate how FLEA may be used to process Pacific Biosciences HIV-1 env data and describe recent examples of its use. Simulations show how FLEA dramatically reduces the error rate of this sequencing platform, providing an accurate portrait of complex and variable HIV-1 env populations.\n\nA public instance of FLEA is hosted at http://flea.datamonkey.org. The Python source code for the FLEA pipeline can be found at https://github.com/veg/flea-pipeline. The client-side application is available at https://github.com/veg/flea-web-app. A live demo of the P018 results can be found at http://flea.murrell.group/view/P018.

bioinformatics

Atypical domain communication and domain functions of a Hsp110 chaperone

Hsp110s are well recognized nucleotide exchange factors (NEFs) of Hsp70s, in addition they are implicated in various aspects of cellular proteostasis as discrete chaperones with yet enigmatic molecular mechanism. Stark similarity in domain organization and structure between Hsp110s and Hsp70s, is easily discernible although the nature of domain communication and domain functions of Hsp110s are still puzzling. Here, we report atypical domain communication of yeast Hsp110, Sse1 using single molecule FRET, small angle X-ray scattering measurements (SAXS) and Molecular Dynamic simulations. Our data show that Sse1 lacks typical domain movements as exhibited by Hsp70s, albeit it undergoes unique structural alteration upon nucleotide and substrate binding. Hsp70-like domain-movements can be artificially salvaged in chimeric constructs of Hsp110-Hsp70 although such salvaging proves detrimental for the NEF activity of the protein. Furthermore, we show that substrate binding domain (SBD) of Hsp110, chaperones self, as well as foreign nucleotide binding domains (NBD). Interestingly, the substrate binding specificity of Hsp110 is largely determined by its NBD rather than SBD, the latter being the foremost substrate binding region for Hsp70s.

biochemistry

Discovering genetic interactions bridging pathways in genome-wide association studies

Genetic interactions have been reported to underlie phenotypes in a variety of systems, but the extent to which they contribute to complex disease in humans remains unclear. In principle, genome-wide association studies (GWAS) provide a platform for detecting genetic interactions, but existing methods for identifying them from GWAS data tend to focus on testing individual locus pairs, which undermines statistical power. Importantly, the global genetic networks mapped for a model eukaryotic organism revealed that genetic interactions often connect genes between compensatory functional modules in a highly coherent manner. Taking advantage of this expected structure, we developed a computational approach called BridGE that identifies pathways connected by genetic interactions from GWAS data. Applying BridGE broadly, we discovered significant interactions in Parkinsons disease, schizophrenia, hypertension, prostate cancer, breast cancer, and type 2 diabetes. Our novel approach provides a general framework for mapping complex genetic networks underlying human disease from genome-wide genotype data.

genetics

Insights into regeneration from the genome, transcriptome and metagenome analysis of Eisenia fetida

Earthworms show a wide spectrum of regenerative potential with certain species like Eisenia fetida capable of regenerating more than two-thirds of their body while other closely related species, such as Paranais litoralis seem to have lost this ability. Earthworms belong to the phylum annelida, in which the genomes of the marine oligochaete Capitella telata, and the freshwater leech Helobdella robusta have been sequenced and studied. The terrestrial annelids, in spite of their ecological relevance and unique biochemical repertoire, are represented by a single rough genome draft of Eisenia fetida (North American isolate), which suggested that extensive duplications have led to a large number of HOX genes in this annelid. Herein, we report the draft genome sequence of Eisenia fetida (Indian isolate), a terrestrial redworm widely used for vermicomposting assembled using short reads and mate-pair reads. An in-depth analysis of the miRNome of the worm, showed that many miRNA gene families have also undergone extensive duplications. Genes for several important proteins such as sialidases and neurotrophins were identified by RNA sequencing of tissue samples. We also used de novo assembled RNA-Seq data to identify genes that are differentially expressed during regeneration, both in the newly regenerating cells and in the adjacent tissue. Sox4, a master regulator of TGF-beta induced epithelial-mesenchymal transition was induced in the newly regenerated tissue. The regeneration of the ventral nerve cord was also accompanied by the induction of nerve growth factor and neurofilament genes. The metagenome of the worm, characterized using 16S rRNA sequencing, revealed the identity of several bacterial species that reside in the nephridia of the worm. Comparison of the bodywall and cocoon metagenomes showed exclusion of hereditary symbionts in the regenerated tissue. In summary, we present extensive genome, transcriptome and metagenome data to establish the transcriptome and metagenome dynamics during regeneration.

genomics

Forward Genetic ENU Mutagenesis Screen for Mouse Models of Chronic Fatigue Identifies a Novel Mutation in Slc2a4 (GLUT4)

In a screen of voluntary wheel-running behavior designed to identify genetic mouse models of chronic fatigue in ENU mutagenized C57BL/6J mice, we discovered two lines that showed aberrant wheel-running patterns. These lines both stem from a single original founder identified as a low body-weight candidate in a recessive screen. Progeny from both of these lines showed the abnormal wheel-running behavior, with affected mice showing significantly lower daily activity levels than unaffected mice. They also exhibited low amplitude circadian rhythms, consisting of lower activity levels during the normal active phase, and increased levels of activity during the rest or light phase, but only a modest alteration in free-running period. Their activity is not consolidated into longer bouts, but is frequently interrupted with periods of inactivity throughout the dark phase of the light-dark (LD) cycle. As seen with the low body weight, expression of the behavioral phenotypes in offspring of strategic crosses was consistent with a recessive heritance pattern. Mapping of these phenotypic abnormalities showed linkage to a single locus on chromosome 11, and whole exome sequencing (WES) identified a single point mutation in the Slc2a4 gene encoding the GLUT4 insulin-responsive glucose transporter. The single nucleotide change (A to T) was found in the distal end of exon 10, and results in a premature stop (Y440*). To our knowledge, this is the first time a mutation in this gene has been shown to result in extensive changes in general behavioral patterns.\n\nSIGNIFICANCE STATEMENTChronic fatigue is a debilitating and devastating disorder with widespread consequences for both the patient and the persons around them, but effective treatment strategies are lacking. The identification of novel genetic mouse models of chronic fatigue may prove invaluable for the study of its underlying physiological mechanisms and for the testing of treatments and interventions. A novel mutation in Slc2a4 (GLUT4) was identified in a forward mutagenesis screen because affected mice showed abnormal daily patterns and levels of wheel running consistent with chronic fatigue. This new mouse model may shed light on the pathophysiology of chronic fatigue.

genetics

Reference Quality Assembly of the 3.5 Gb genome of Capsicum annuum from a Single Linked-Read Library

BackgroundLinked-Read sequencing technology has recently been employed successfully for de novo assembly of multiple human genomes, however the utility of this technology for complex plant genomes is unproven. We evaluated the technology for this purpose by sequencing the 3.5 gigabase (Gb) diploid pepper (Capsicum annuum) genome with a single Linked-Read library. Plant genomes, including pepper, are characterized by long, highly similar repetitive sequences. Accordingly, significant effort is used to ensure the sequenced plant is highly homozygous and the resulting assembly is a haploid consensus. With a phased assembly approach, we targeted a heterozygous F1 derived from a wide cross to assess the ability to derive both haplotypes for a pungency gene characterized by a large insertion/deletion.\n\nResultsThe Supernova software generated a highly ordered, more contiguous sequence assembly than all currently available C. annuum reference genomes. Eighty-four percent of the final assembly was anchored and oriented using four de novo linkage maps. A comparison of the annotation of conserved eukaryotic genes indicated the completeness of assembly. The validity of the phased assembly is further demonstrated with the complete recovery of both 2.5 kb insertion/deletion haplotypes of the PUN1 locus in the F1 sample that represents pungent and non-pungent peppers.\n\nConclusionsThe most contiguous pepper genome assembly to date has been generated through this work which demonstrates that Linked-Read library technology provides a rapid tool to assemble de novo complex highly repetitive heterozygous plant genomes. This technology can provide an opportunity to cost-effectively develop high-quality reference genome assemblies for other complex plants and compare structural and gene differences through accurate haplotype reconstruction.

genomics

Mutational sequencing for accurate count and long-range assembly

We introduce a new protocol, mutational sequencing or muSeq, which randomly deaminates unmethylated cytosines at a fixed and tunable rate. The muSeq protocol marks each initial template molecule with a unique mutation signature that is present in every copy of the template, and in every fragmented copy of a copy. In the sequenced read data, this signature is observed as a unique pattern of C-to-T or G-to-A nucleotide conversions. Clustering reads with the same conversion pattern enables accurate count and long-range assembly of initial template molecules from short-read sequence data. We explore count and low-error sequencing by profiling a 135,000 fragment PstI representation, demonstrating that muSeq improves copy number inference and significantly reduces sporadic sequencer error. We explore long-range assembly in the context of cDNA, generating contiguous transcript clusters greater than 3,000 bp in length. The muSeq assemblies reveal transcriptional diversity not observable from short-read data alone.

genomics

Improved de novo Genome Assembly: Synthetic long read sequencing combined with optical mapping produce a high quality mammalian genome at relatively low cost

Current short-read methods have come to dominate genome sequencing because they are cost-effective, rapid, and accurate. However, short reads are most applicable when data can be aligned to a known reference. Two new methods for de novo assembly are linked-reads and restriction-site labeled optical maps. We combined commercial applications of these technologies for genome assembly of an endangered mammal, the Hawaiian Monk seal.\n\nWe show that the linked-reads produced with 10X Genomics Chromium chemistry and assembled with Supernova v1.1 software produced scaffolds with an N50 of 22.23 Mbp with the longest individual scaffold of 84.06 Mbp. When combined with Bionano Genomics optical maps using Bionano RefAligner, the scaffold N50 increased to 29.65 Mbp for a total of 170 hybrid scaffolds, the longest of which was 84.78 Mbp. These results were 161X and 215X, respectively, improved over DISCOVAR de novo assemblies. The quality of the scaffolds was assessed using conserved synteny analysis of both the DNA sequence and predicted seal proteins relative to the genomes of humans and other species. We found large blocks of conserved synteny suggesting that the hybrid scaffolds were high quality. An inversion in one scaffold complementary to human chromosome 6 was found and confirmed by optical maps.\n\nThe complementarity of linked-reads and optical maps is likely to make the production of high quality genomes more routine and economical and, by doing so, significantly improve our understanding of comparative genome biology.

genomics

High-spatial-resolution transcriptome profiling reveals uncharacterized regulatory complexity underlying cambial growth and wood formation in Populus tremula

Trees represent the largest terrestrial carbon sink and a renewable source of ligno-cellulose. There is significant scope for yield and quality improvement in these largely undomesticated species, and efforts to engineer new, elite varieties will benefit from an improved understanding of the transcriptional network underlying cambial growth and wood formation. We generated high-spatial-resolution RNA Sequencing data spanning the secondary phloem, vascular cambium and wood forming tissues. The transcriptome comprised 28,294 expressed, previously annotated genes, 78 novel protein-coding genes and 567 long intergenic non-coding RNAs. Most paralogs originating from the Salicaceae whole genome duplication were found to have diverged expression, with the notable exception of those with high expression during secondary cell wall deposition. Co-expression network analysis revealed that the regulation of the transcriptome underlying cambial growth and wood formation comprises numerous modules forming a continuum of active processes across the tissues. The high spatial resolution enabled identification of novel roles for characterised genes involved in xylan and cellulose biosynthesis, regulators of xylem vessel and fiber differentiation and lignification. The associated web resource (AspWood, http://aspwood.popgenie.org) integrates the data within a set of interactive tools for exploring the expression profiles and co-expression network.

plant biology

Genome Wide Computational Prediction of miRNAs in Kyasanur Forest Disease Virus and their Targeted Genes in Human

RNAs are versatile biomolecules and can be coding or non-coding. Among the non-coding RNAs, miRNAs are small endogenous molecules that play important role in posttranscriptional gene regulation. miRNAs are identified in viruses too and involved in down regulation of host genes. Flavivirus family members are classified in to two groups: mosquito-borne flaviviruses (MBFV) and tick-borne flaviviruses (TBFV). Kyasanur forest disease virus (KFDV) found in India in 1957 (Karnataka) relates to TBFV. Virus has been diffuse to new areas in India and needs attention as it can cause severe hemorrhagic fever. Here in this study, we scanned the virus genome for prediction of miRNAs that can inhibit host target genes. VMir, tool was used for extraction of pre-miRNAs. A total of four miRNAs were found and submitted to ViralMir for classification in to real or pseudo. Interestingly, all four pre-miRNAs were classified as real. Eight mature miRNAs were located in pre-miRNAs by Mature Bayes. A total of 539 human target genes has been identified by using miRDB but ANGPT1 (angiopoietin 1) and TFRC (transferrin receptor) genes were screened to play role in hemorrhagic fever and neurological problems. GO analysis of target genes also supported the evidences.

bioinformatics

The evolutionary history of bears is shaped by gene flow across species

Bears are iconic mammals with a complex evolutionary history. Natural bear hybrids and studies of few nuclear genes indicate that gene flow among bears may be more common than expected and not limited to the closely related polar and brown bears. Here we present a genome analysis of the bear family with representatives of all living species. Phylogenomic analyses of 869 mega base pairs divided into 18,621 genome fragments yielded a well-resolved coalescent species tree despite signals for extensive gene flow across species. However, genome analyses using three different statistical methods show that gene flow is not limited to closely related species pairs. Strong ancestral gene flow between the Asiatic black bear and the ancestor to polar, brown and American black bear explains numerous uncertainties in reconstructing the bear phylogeny. Gene flow across the bear clade may be mediated by intermediate species such as the geographically wide-spread brown bears leading to massive amounts of phylogenetic conflict. Genome-scale analyses lead to a more complete understanding of complex evolutionary processes. The increasing evidence for extensive inter-specific gene flow, found also in other animal species, necessitates shifting the attention from speciation processes achieving genome-wide reproductive isolation to the selective processes that maintain species divergence in the face of gene flow.

evolutionary biology