bioRxiv ScienceSearch

SEARCH · bioRxiv Science

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,153 records · Page 64Linked to original sources

Millennia of genomic stability within the invasive Para C Lineage of Salmonella enterica

Salmonella enterica serovar Paratyphi C is the causative agent of enteric (paratyphoid) fever. While today a potentially lethal infection of humans that occurs in Africa and Asia, early 20th century observations in Eastern Europe suggest it may once have had a wider-ranging impact on human societies. We recovered a draft Paratyphi C genome from the 800-year-old skeleton of a young woman in Trondheim, Norway, who likely died of enteric fever. Analysis of this genome against a new, significantly expanded database of related modern genomes demonstrated that Paratyphi C is descended from the ancestors of swine pathogens, serovars Choleraesuis and Typhisuis, together forming the Para C Lineage. Our results indicate that Paratyphi C has been a pathogen of humans for at least 1,000 years, and may have evolved after zoonotic transfer from swine during the Neolithic period.\n\nOne Sentence SummaryThe combination of an 800-year-old Salmonella enterica Paratyphi C genome with genomes from extant bacteria reshapes our understanding of this pathogens origins and evolution.

microbiology

Validation and Implementation of CLIA-Compliant Whole Genome Sequencing (WGS) in Public Health Laboratory

BackgroundPublic health microbiology laboratories (PHL) are at the cusp of unprecedented improvements in pathogen identification, antibiotic resistance detection, and outbreak investigation by using whole genome sequencing (WGS). However, considerable challenges remain due to the lack of common standards.\n\nObjectives1) Establish the performance specifications of WGS applications used in PHL to conform with CLIA (Clinical Laboratory Improvements Act) guidelines for laboratory developed tests (LDT), 2) Develop quality assurance (QA) and quality control (QC) measures, 3) Establish reporting language for end users with or without WGS expertise, 4) Create a validation set of microorganisms to be used for future validations of WGS platforms and multi-laboratory comparisons and, 5) Create modular templates for the validation of different sequencing platforms.\n\nMethodsMiSeq Sequencer and Illumina chemistry (Illumina, Inc.) were used to generate genomes for 34 bacterial isolates with genome sizes from 1.8 to 4.7 Mb and wide range of GC content (32.1%-66.1%). A customized CLCbio Genomics Workbench - shell script bioinformatics pipeline was used for the data analysis.\n\nResultsWe developed a validation panel comprising ten Enterobacteriaceae isolates, five gram-positive cocci, five gram-negative non-fermenting species, nine Mycobacterium tuberculosis, and five miscellaneous bacteria; the set represented typical workflow in the PHL. The accuracy of MiSeq platform for individual base calling was >99.9% with similar results shown for reproducibility/repeatability of genome-wide base calling. The accuracy of phylogenetic analysis was 100%. The specificity and sensitivity inferred from MLST and genotyping tests were 100%. A test report format was developed for the end users with and without WGS knowledge.\n\nConclusionWGS was validated for routine use in PHL according to CLIA guidelines for LDTs. The validation panel, sequencing analytics, and raw sequences will be available for future multi-laboratory comparisons of WGS in PHL. Additionally, the WGS performance specifications and modular validation template are likely to be adaptable for the validation of other platforms and reagents kits.

microbiology

Genomic contingencies and beak shape variation in a hybrid species

Hybridization is increasingly recognized as a potent evolutionary force. Though additive genetic variation and novel combinations of parental genes theoretically increase the potential for hybrid species to adapt, few empirical studies have investigated the adaptive potential within a hybrid species. Here, we investigate factors promoting phenotypic divergence using genomically diverged island populations of the homoploid hybrid Italian sparrow Passer italiae from Crete, Corsica, and Sicily. We address whether genomic contingencies, adaptation to climate or diet best explain divergence in beak morphology. Populations vary significantly in beak morphology, both between and within islands of origin. Temperature seasonality best explains population divergence in beak size. Interestingly, beak shape along all significant dimensions of variation was best explained by annual precipitation, genomic composition and their interaction, suggesting a role for contingencies. Moreover, beak shape similarity to a parent species correlates with proportion of the genome inherited from that species, consistent with the presence of contingencies. In conclusion, adaptation to local conditions and genomic contingencies arising from putatively independent hybridization events jointly explain beak morphology in the Italian sparrow. Hence, hybridization may induce contingencies and restrict evolution in certain directions dependent on the genetic background.

evolutionary biology

Intervene: a tool for intersection and visualization of multiple gene or genomic region sets

BackgroundA common task for scientists relies on comparing lists of genes or genomic regions derived from high-throughput sequencing experiments. While several tools exist to intersect and visualize sets of genes, similar tools dedicated to the visualization of genomic region sets are currently limited.\n\nResultsTo address this gap, we have developed the Intervene tool, which provides an easy and automated interface for the effective intersection and visualization of genomic region or list sets, thus facilitating their analysis and interpretation. Intervene contains three modules: venn to generate Venn diagrams of up to six sets, upset to generate UpSet plots of multiple sets, and pairwise to compute and visualize intersections of multiple sets as clustered heat maps. Intervene, and its interactive web ShinyApp companion, generate publication-quality figures for the interpretation of genomic region and list sets.\n\nConclusionsIntervene and its web application companion provide an easy command line, and an interactive web interface to compute intersections of multiple genomic and list sets. They also have the capacity to plot intersections using easy-to-interpret visual approaches. Intervene is developed and designed to meet the needs of both computer scientists and biologists. The source code is freely available at https://bitbucket.org/CBGR/intervene, with the web application available at https://asntech.shinyapps.io/intervene.

bioinformatics

Genome-wide regulatory model from MPRA data predicts functional regions, eQTLs, and GWAS hits

Massively-parallel reporter assays (MPRA) enable unprecedented opportunities to test for regulatory activity of thousands of regulatory sequences. However, MPRA only assay a subset of the genome thus limiting their applicability for genome-wide functional annotations. To overcome this limitation, we have used existing MPRA datasets to train a machine learning model that uses DNA sequence information, regulatory motif annotations, evolutionary conservation, and epigenomic information to predict genomic regions that show enhancer activity when tested in MPRA assays. We used the resulting model to generate global predictions of regulatory activity at single-nucleotide resolution across 14 million common variants. We find that genetic variants with stronger predicted regulatory activity show significantly lower minor allele frequency, indicative of evolutionary selection within the human population. They also show higher over-lap with eQTL annotations across multiple tissues relative to the background SNPs, indicating that their perturbations in vivo more frequently result in changes in gene expression. In addition, they are more frequently associated with trait-associated SNPs from genome-wide association studies (GWAS), enabling us to prioritize genetic variants that are more likely to be causal based on their predicted regulatory activity. Lastly, we use our model to compare MPRA inferences across cell types and platforms and to prioritize the assays most predictive of MPRA assay results, including cell-dependent DNase hypersensitivity sites and transcription factors known to be active in the tested cell types. Our results indicate that high-throughput testing of thousands of putative regions, coupled with regulatory predictions across millions of sites, presents a powerful strategy for systematic annotation of genomic regions and genetic variants.

bioinformatics

REPARATION: Ribosome Profiling Assisted (Re-)Annotation of Bacterial genomes.

Prokaryotic genome annotation is highly dependent on automated methods, as manual curation cannot keep up with the exponential growth of sequenced genomes. Current automated methods depend heavily on sequence context and often underestimate the complexity of the proteome. We developed REPARATION (RibosomeE Profiling Assisted (Re-)AnnotaTION), a de novo algorithm that takes advantage of experimental protein translation evidence from ribosome profiling (Ribo-seq) to delineate translated open reading frames (ORFs) in bacteria, independent of genome annotation. REPARATION evaluates all possible ORFs in the genome and estimates minimum thresholds based on a growth curve model to screen for spurious ORFs. We applied REPARATION to three annotated bacterial species to obtain a more comprehensive mapping of their translation landscape in support of experimental data. In all cases, we identified hundreds of novel (small) ORFs including variants of previously annotated ORFs. Our predictions were supported by matching mass spectrometry (MS) proteomics data, sequence composition and conservation analysis. REPARATION is unique in that it makes use of experimental translation evidence to perform de novo ORF delineation in bacterial genomes irrespective of the sequence context of the reading frame.

microbiology

GAMtools: an automated pipeline for analysis of Genome Architecture Mapping data

Genome Architecture Mapping (GAM) is a recently developed method for mapping chromatin interactions genome-wide. GAM is based on sequencing genomic DNA extracted from thin cryosections of cell nuclei. As a new approach, GAM datasets require specialized analytical tools and approaches. Here we present GAMtools, a pipeline for analysing GAM datasets. GAMtools covers the automated mapping of raw next-generation sequencing data generated by GAM, detection of genomic regions present in each nuclear slice, calculation of quality control metrics, generation of inferred proximity matrices, plotting of heatmaps and detection of genomic features for which chromatin interactions are enriched/depleted.

bioinformatics

The genomic footprint of climate adaptation in Chironomus riparius

The gradual heterogeneity of climatic factors pose varying selection pressures across geographic distances that leave signatures of clinal variation in the genome. Separating signatures of clinal adaptation from signatures of other evolutionary forces, such as demographic processes, genetic drift, and adaptation to non-clinal conditions of the immediate local environment is a major challenge. Here, we examine climate adaptation in five natural populations of the harlequin fly Chironomus riparius sampled along a climatic gradient across Europe. Our study integrates experimental data, individual genome resequencing, Pool-Seq data, and population genetic modelling. Common-garden experiments revealed a positive correlation of population growth rates corresponding to the population origin along the climate gradient, suggesting thermal adaptation on the phenotypic level. Based on a population genomic analysis, we derived empirical estimates of historical demography and migration. We used an FST outlier approach to infer positive selection across the climate gradient, in combination with an environmental association analysis. In total we identified 162 candidate genes as genomic basis of climate adaptation. Enriched functions among these candidate genes involved the apoptotic process and molecular response to heat, as well as functions identified in other studies of climate adaptation in other insects. Our results show that local climate conditions impose strong selection pressures and lead to genomic adaptation despite strong gene flow. Moreover, these results imply that selection to different climatic conditions seems to converge on a functional level, at least between different insect species.

evolutionary biology

Discovering Complete Quasispecies In Bacterial Genomes

Mobile genetic elements can be found in almost all genomes. Possibly the most common non-autonomous mobile genetic elements in bacteria are REPINs that can occur hundreds of times within a genome. The sum of all REPINs within a genome are an evolving populations because they replicate and mutate. We know the exact composition of this population and the sequence of each member of a REPIN population, in contrast to most other biological populations. Here, we model the evolution of REPINs as quasispecies. We fit our quasispecies model to ten different REPIN populations from ten different bacterial strains and estimate duplication rates. We find that our estimated duplication rates range from about 5 x 10-9 to 37 x 10-9 duplications per generation per genome. The small range and the low level of the REPIN duplication rates suggest a universal trade-off between the survival of the REPIN population and the reduction of the mutational load for the host genome. The REPIN populations we investigated also possess features typical of other natural populations. One population shows hallmarks of a population that is going extinct, another population seems to be growing in size and we also see an example of competition between two REPIN populations.

evolutionary biology

HiPiler: Visual Exploration Of Large Genome Interaction Matrices With Interactive Small Multiples

This paper presents an interactive visualization interface--HiPiler--for the exploration and visualization of regions-of-interest in large genome interaction matrices. Genome interaction matrices approximate the physical distance of pairs of regions on the genome to each other and can contain up to 3 million rows and columns with many sparse regions. Regions of interest (ROIs) can be defined, e.g., by sets of adjacent rows and columns, or by specific visual patterns in the matrix. However, traditional matrix aggregation or pan-and-zoom interfaces fail in supporting search, inspection, and comparison of ROIs in such large matrices. In HiPiler, ROIs are first-class objects, represented as thumbnail-like \"snippets\". Snippets can be interactively explored and grouped or laid out automatically in scatterplots, or through dimension reduction methods. Snippets are linked to the entire navigable genome interaction matrix through brushing and linking. The design of HiPiler is based on a series of semi-structured interviews with 10 domain experts involved in the analysis and interpretation of genome interaction matrices. We describe six exploration tasks that are crucial for analysis of interaction matrices and demonstrate how HiPiler supports these tasks. We report on a user study with a series of data exploration sessions with domain experts to assess the usability of HiPiler as well as to demonstrate respective findings in the data.

bioinformatics

Uncovering The Repertoire Of Endogenous Flaviviral Elements In Aedes Mosquito Genomes

Endogenous viral elements derived from non-retroviral RNA viruses were described in various animal genomes. Whether they have a biological function such as host immune protection against related viruses is a field of intense study. Here, we investigated the repertoire of endogenous flaviviral elements (EFVEs) in Aedes mosquitoes, the vectors of arboviruses such as dengue and chikungunya viruses. Previous studies identified three EFVEs from Ae. albopictus and one from Ae. aegypti cell lines. However, in-depth characterization of EFVEs in wild-type mosquito populations and individuals in vivo has not been performed. We detected the full-length DNA sequence of the previously described EFVEs and their respective transcripts in several Ae. albopictus and Ae. aegypti populations from geographically distinct areas. However, EFVE-derived proteins were not detected by mass spectrometry. Using deep sequencing, we detected the production of piRNA-like small RNAs in antisense orientation, targeting the EFVEs and their flanking regions in vivo. The EFVEs were integrated in repetitive regions of the mosquito genomes, and their flanking sequences varied among mosquito populations from different geographical regions. We bioinformatically predicted several new EFVEs from a Vietnamese Ae. albopictus population and observed variation in the occurrence of those elements among mosquito populations. Phylogenetic analysis of an Ae. aegypti EFVE suggested that it integrated prior to the global expansion of the species and subsequently diverged among and within populations. Together, this study revealed substantial structural and nucleotide diversity of flaviviral integrations in Aedes genomes. Unraveling this diversity will help to elucidate the potential biological function of these EFVEs.\n\nImportanceEndogenous viral elements (EVEs) are whole or partial viral sequences integrated in host genomes. Interestingly, some EVEs have important functions for host fitness and antiviral defense. Because mosquitoes also have EVEs in their genomes, we decided to thoroughly characterized them to lay the foundation of the potential use of these EVEs to manipulate the mosquito antiviral response. Here, we focused on EVEs related to the Flavivirus genus, to which dengue and Zika viruses belong, in Aedes mosquito individuals from geographically distinct areas. We showed the existence in vivo of flaviviral EVEs previously identified in mosquito cell lines and we detected new ones. We showed that EVEs have evolved differently in each mosquito population. They produced transcripts and small RNAs, but not proteins, suggesting a function at the RNA level. Our study uncovers the diverse repertoire of flaviviral EVEs in Aedes mosquito populations and suggests a role in the host antiviral system.

microbiology

A Standardized Framework For Representation Of Ancestry Data In Genomics Studies

BackgroundThe accurate description of ancestry is essential to interpret and integrate human genomics data, and to ensure that advances in the field of genomics benefit individuals from all ancestral backgrounds. However, there are no established guidelines for the consistent, unambiguous and standardized description of ancestry. To fill this gap, we provide a framework, designed for the representation of ancestry in GWAS data, but with wider application to studies and resources involving human subjects.\n\nResultHere we describe our framework and its application to the representation of ancestry data in a widely-used publically available genomics resource, the NHGRI-EBI GWAS Catalog. We present the first analyses of GWAS data using our ancestry categories, demonstrating the validity of the framework to facilitate the tracking of ancestry in big data sets. We exhibit the broader relevance and integration potential of our method by its usage to describe the well-established HapMap and 1000 Genomes reference populations. Finally, to encourage adoption, we outline recommendations for authors to implement when describing samples.\n\nConclusionsWhile the known bias towards inclusion of European ancestry individuals in GWA studies persists, African and Hispanic or Latin American ancestry populations contribute a disproportionately high number of associations, suggesting that analyses including these groups may be more effective at identifying new associations. We believe the widespread adoption of our framework will increase standardization of ancestry data, thus enabling improved analysis, interpretation and integration of human genomics data and furthering our understanding of disease.

genetics

Identification And Prioritisation Of Variants In The Short Open-Reading Frame Regions Of The Human Genome

As whole-genome sequencing technologies improve and accurate maps of the entire genome are assembled, short open-reading frames (sORFs) are garnering interest as functionally important regions that were previously overlooked. However, there is a paucity of tools available to investigate variants in sORF regions of the genome. Here we investigate the performance of commonly used tools for variant calling and variant prioritisation in these regions, and present a framework for optimising these processes. First, the performance of four widely used germline variant calling algorithms is systematically compared. Haplotype Caller is found to perform best across the whole genome, but FreeBayes is shown to produce the most accurate variant set in sORF regions. An accurate set of variants is found by taking the intersection of called variants. The potential deleteriousness of each variant is then predicted using a pathogenicity scoring algorithm developed here, called sORF-c. This algorithm uses supervised machine-learning to predict the pathogenicity of each variant, based on a holistic range of functional, conservation-based and region-based scores defined for each variant. By training on a dataset of over 130,000 variants, sORF-c outperforms other comparable pathogenicity scoring algorithms on a test set of variants in sORF regions of the human genome.\n\nList of Abbreviations

genetics

SUMO E3 ligase Mms21 prevents spontaneous DNA damage induced genome rearrangements

Mms21, a subunit of the Smc5/6 complex, possesses an E3 ligase activity for the Small Ubiquitin-like MOdifier (SUMO), which has a major, but poorly understood role in genome maintenance. Here we show mutations that inactivate the E3 ligase activity of Mms21 cause Rad52- and Pol32-dependent break-induced replication (BIR), which specifically requires the Rrm3 DNA helicase. Interestingly, mutations affecting both Mms21 and the Sgs1 helicase, but not sumoylation of Sgs1, cause further accumulation of genome rearrangements, indicating the distinct roles of Mms21 and Sgs1 in suppressing genome rearrangements. Whole genome sequencing further revealed that the Mre11 endonuclease prevents microhomology-mediated translocations and hairpin-mediated inverted duplications in the mms21 mutant. Consistent with the accumulation of endogenous DNA lesions, mms21 cells accumulate spontaneous Ddc2 foci and display a hyper-activated DNA damage checkpoint. Together, these findings support a new paradigm that Mms21 prevents the accumulation of spontaneous DNA lesions that cause diverse genome rearrangements.

genetics

ANNOgesic: A Pipeline To Translate Bacterial/Archaeal RNA-Seq Data Into High-Resolution Genome Annotations

To understand the gene regulation of an organism of interest, a comprehensive genome annotation is essential. While some features, such as coding sequences, can be computationally predicted with high accuracy based purely on the genomic sequence, others, such as promoter elements or non-coding RNAs are harder to detect. RNA-Seq has proven to be an efficient method to identify these genomic features and to improve genome annotations. However, processing and integrating RNA-Seq data in order to generate high-resolution annotations is challenging, time consuming and requires numerous different steps. We have constructed a powerful and modular tool called ANNOgesic that provides the required analyses and simplifies RNA-Seq-based bacterial and archaeal genome annotation. It can integrate data from conventional RNA-Seq and dRNA-Seq, predicts and annotates numerous features, including small non-coding RNAs, with high precision. The software is available under an open source license (ISCL) at https://pypi.org/project/ANNOgesic/.

bioinformatics

Imputation-Based Genomic Coverage Assessments of Current Genotyping Arrays: Illumina HumanCore, OmniExpress, Multi-Ethnic global array and sub-arrays, Global Screening Array, Omni2.5M, Omni5M, and Affymetrix UK Biobank

Genotyping arrays have been widely adopted as an efficient means to interrogate variation across the human genome. Genetic variants may be observed either directly, via genotyping, or indirectly, through linkage disequilibrium with a genotyped variant. The total proportion of genomic variation captured by an array, either directly or indirectly, is referred to as \"genomic coverage.\" Here we use genotype imputation and Phase 3 of the 1000 Genomes Project to assess genomic coverage of several modern genotyping arrays. We find that in general, coverage increases with increasing array density. However, arrays designed to cover specific populations may yield better coverage in those populations compared to denser arrays not tailored to the given population. Ultimately, array choice involves trade-offs between cost, density, and coverage, and our work helps inform investigators weighing these choices and trade-offs.

genetics

Genome Architecture Leads a Bifurcation in Cell Identity

Genome architecture is important in transcriptional regulation and study of its features is a critical part of fully understanding cell identity. Altering cell identity is possible through overexpression of transcription factors (TFs); for example, fibroblasts can be reprogrammed into muscle cells by introducing MYOD1. How TFs dynamically orchestrate genome architecture and transcription as a cell adopts a new identity during reprogramming is not well understood. Here we show that MYOD1-mediated reprogramming of human fibroblasts into the myogenic lineage undergoes a critical transition, which we refer to as a bifurcation point, where cell identity definitively changes. By integrating knowledge of genome-wide dynamical architecture and transcription, we found significant chromatin reorganization prior to transcriptional changes that marked activation of the myogenic program. We also found that the local architectural and transcriptional dynamics of endogenous MYOD1 and MYOG reflected the global genomic bifurcation event. These TFs additionally participate in entrainment of biological rhythms. Understanding the system-level genome dynamics underlying a cell fate decision is a step toward devising more sophisticated reprogramming strategies that could be used in cell therapies.

cell biology

Project MinE: study design and pilot analyses of a large-scale whole-genome sequencing study in amyotrophic lateral sclerosis

The most recent genome-wide association study in amyotrophic lateral sclerosis (ALS) demonstrates a disproportionate contribution from low-frequency variants to genetic susceptibility of disease. We have therefore begun Project MinE, an international collaboration that seeks to analyse whole-genome sequence data of at least 15,000 ALS patients and 7,500 controls. Here, we report on the design of Project MinE and pilot analyses of newly whole-genome sequenced 1,264 ALS patients and 611 controls drawn from the Netherlands. As has become characteristic of sequencing studies, we find an abundance of rare genetic variation (minor allele frequency < 0.1 %), the vast majority of which is absent in public data sets. Principal component analysis reveals local geographical clustering of these variants within The Netherlands. We use the whole-genome sequence data to explore the implications of poor geographical matching of cases and controls in a sequence-based disease study and to investigate how ancestry-matched, externally sequenced controls can induce false positive associations. Also, we have publicly released genome-wide minor allele counts in cases and controls, as well as results from genic burden tests.

genetics