bioRxiv ScienceSearch

SEARCH · bioRxiv Science

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23Linked to original sources

Chicago and Dovetail Hi-C proximity ligation yield chromosome length scaffolds of Ixodes scapularis genome

A high-quality genome sequence is essential for understanding an organism on molecular level. However, the larger genomes with substantial repetitive sequences are challenging to assemble with the sequencing technologies. Hi-C technique is changing the genome architecture landscape by providing links across a variety of length scales, spanning even whole chromosomes. Ixodes scapularis haploid genome is 2.1 gbp and the current assembly consists of 369,495 scaffolds representing 57% of the genome. The fragmented genome poses challenges with functional gene analysis and an improved assembly is needed. We therefore used the Hi C technique to achieve chromosomal level assembly of tick genome. With Chicago and Dovetail Hi C assemblies, we were able to achieve 28 >10Mb sequences that correspond to 28 chromosomes in I. scapularis.

genomics

Functional and Comparative Genomics of Niche-Specific Adapted Actinomycetes Kocuria rhizophila Strain D2 Isolated from Healthy Human Gut

Incidences of infection and occurrence of Kocuria rhizophila in human gut are prominent but certainly no reports on the species ability to withstand human gastrointestinal dynamics. Kocuria rhizophila strain D2 isolated from healthy human gut was comprehensively characterized. The functional analysis revealed the ability to produce various gastric enzymes and sensitive to major clinical antibiotics. It also exhibited tolerance to acidic pH and bile salts. Strain D2 displayed bile-salt hydrolytic (BSH) activity, strong cell surface traits such as hydrophobicity, auto-aggregation capacity and adherence to human HT-29 cell line. Prominently, it showed no hemolytic activity and was susceptible to the human serum. Exploration of the genome led to the discovery of the genes for the above said properties and has ability to produce various essential amino acids and vitamins. Further, comparative genomics have identified core, accessory and unique genetic features. The core genome has given insights into the phylogeny while the accessory and unique genes has led to the identification of niche specific genes. Bacteriophage, virulence factors and biofilm formation genes were absent with this species. Housing CRISPR and antibiotic resistance gene was strain specific. The integrated approach of functional, genomic and comparative analysis denotes the niche specific adaption to gut dynamics of strain D2. Moreover the study has comprehensively characterized genome sequence of each strain to know the genetic difference and intern recognize the effects of on phenotype and functionality complexity. The evolutionary relationship among strains along and adaptation strategies has been included in this study.\n\nSignificanceReports of Kocuria rhizophila isolation from various sources have been reported but the few disease outbreaks in humans and fishes have been prominent, but no supportive evidence about the survival ability of Kocuria spp. within human GIT. Here, we report the gut adaption potential of K. rhizophila strain D2 by functional and genomic analysis. Further; comparative genomics reveals this adaption to be strain specific (Gluten degradation). Genetic difference, evolutionary relationship and adaptation strategies have been including in this study.

genomics

Self-Organized Critical Control of Genome Expression: Novel Scenario on Cell-Fate Decision

In our current studies on whole genome expression in several biological processes, we have demonstrated the actual existence of self-organized critical control (SOC) of gene expression at both population and single cell level. SOC allows for cell-fate change by critical transition encompassing the entire genome expression that, in turn, is partitioned into distinct response domains (critical states).\n\nIn this paper, we go more in depth into the elucidation of SOC control of genome expression focusing on the determination of critical point (CP) and associated distinct critical states in single-cell genome expression. This leads us to the proposal of a potential universal model with genome-engine mechanism for cell-fate change. Our findings suggest that the CP is fixed point in terms of temporal expression variance, where the CP (set of critical genes) becomes active (ON) for cell-fate change ( super-critical in genome-state) or else inactive (OFF) state ( sub-critical in genome-state); this may lead to a novel scenario of the cell-fate control through activating or inactivating CP.

genomics

Automated ensemble assembly and validation of microbial genomes

BackgroundThe continued democratization of DNA sequencing has sparked a new wave of development of genome assembly and assembly validation methods. As individual research labs, rather than centralized centers, begin to sequence the majority of new genomes, it is important to establish best practices for genome assembly. However, recent evaluations such as GAGE and the Assemblathon have concluded that there is no single best approach to genome assembly. Instead, it is preferable to generate multiple assemblies and validate them to determine which is most useful for the desired analysis; this is a labor-intensive process that is often impossible or unfeasible.\n\nResultsTo encourage best practices supported by the community, we present iMetAMOS, an automated ensemble assembly pipeline; iMetAMOS encapsulates the process of running, validating, and selecting a single assembly from multiple assemblies. iMetAMOS packages several leading open-source tools into a single binary that automates parameter selection and execution of multiple assemblers, scores the resulting assemblies based on multiple validation metrics, and annotates the assemblies for genes and contaminants. We demonstrate the utility of the ensemble process on 225 previously unassembled Mycobacterium tuberculosis genomes as well as a Rhodobacter sphaeroides benchmark dataset. On these real data, iMetAMOS reliably produces validated assemblies and identifies potential contamination without user intervention. In addition, intelligent parameter selection produces assemblies of R. sphaeroides that exceed the quality of those from the GAGE-B evaluation, affecting the relative ranking of some assemblers.\n\nConclusionsEnsemble assembly with iMetAMOS provides users with multiple, validated assemblies for each genome. Although computationally limited to small or mid-sized genomes, this approach is the most effective and reproducible means for generating high-quality assemblies and enables users to select an assembly best tailored to their specific needs.

Bioinformatics

Genomic Repeat Element Analyzer for Mammals (GREAM)

Background: Understanding the mechanism behind the transcriptional regulation of genes is still a challenge. Recent findings indicate that the genomic repeat elements (such as LINES, SINES and LTRs) could play an important role in the transcription control. Hence, it is important to further explore the role of genomic repeat elements in the gene expression regulation, and perhaps in other molecular processes. Although many computational tools exists for repeat element analysis, almost all of them simply identify and/or classifying the genomic repeat elements within query sequence(s); none of them facilitate identification of repeat elements that are likely to have a functional significance, particularly in the context of transcriptional regulation.\n\nResult: We developed the Genomic Repeat Element Analyzer for Mammals (GREAM) to allow gene-centric analysis of genomic repeat elements in 17 mammalian species, and validated it by comparing with some of the existing experimental data. The output provides a categorized list of the specific type of transposons, retro-transposons and other genome-wide repeat elements that are statistically over-represented across specific neighborhood regions of query genes. The position and frequency of these elements, within the specified regions, are displayed as well. The tool also offers queries for position-specific distribution of repeat elements within chromosomes. In addition, GREAM facilitates the analysis of repeat element distribution across the neighborhood of orthologous genes.\n\nConclusion: GREAM allows researchers to short-list the potentially important repeat elements, from the genomic neighborhood of genes, for further experimental analysis. GREAM is free and available for all at http://resource.ibab.ac.in/GREAM/

Bioinformatics

Divide and Conquer approach for Genome Classification based on subclass characterization

Classification of large grass genome sequences has major challenges in functional genomes. The presence of motifs in grass genome chains can make the prediction of the functional behavior of grass genome possible. The correlation between grass genome properties and their motifs is not always obvious, since more than one motif may exist within a genome chain. Due to the complexity of this association most pattern classification algorithms are either vain or time consuming. Attempted to a reduction of high dimensional data that utilizes DAC technique is presented. Data are disjoining into equal multiple sets while preserving the original data distribution in each set. Then, multiple modules are created by using the data sets as independent training sets and classified into respective modules. Finally, the modules are combined to produce the final classification rules, containing all the previously extracted information. The methodology is tested using various grass genome data sets. Results indicate that the time efficiency of our algorithm is improved compared to other known data mining algorithms.

Bioinformatics

Measuring the Contribution of Genomic Predictors to Improving Estimator Precision in Randomized trials

The use of genomic data in the clinic has not been as widespread as was envisioned when sequencing and genomic analysis became common techniques. An underlying difficulty is the direct assessment of how much additional information genomic data are providing beyond standard clinical measurements. This is hard to quantify in the clinical setting where laboratory tests based on genomic signatures are fairly new and there are not sufficient data collected to determine how valuable these tests have been in practice. Here we focus on the potential precision gain from using the popular MammaPrint genomic signature in a covariate-adjusted, randomized clinical trial. We describe how adjustment of an estimator for the average treatment effect using baseline measurements can improve precision. This precision gain can be translated directly into sample size reduction and corresponding cost savings. We conduct a simulation study using genomic and clinical data gathered for breast cancer patients and find that adjusting for clinical factors alone provides a gain in precision of 5-6%, adjusting for genomic factors alone provides a similar gain (5%), and combining the two yields a 2-3% additional gain over only adjusting for clinical covariates.

Bioinformatics

Metagenome-assembled genomes uncover a global brackish microbiome

Microbes are main drivers of biogeochemical cycles in oceans and lakes, yet surprisingly few bacterioplankton genomes have been sequenced, partly due to difficulties in cultivating them. Here we used automatic binning to reconstruct a large number of bacterioplankton genomes from a metagenomic time-series from the Baltic Sea. The genomes represent novel species within freshwater and marine clades, including clades not previously genome-sequenced. Their seasonal dynamics followed phylogenetic patterns, but with fine-grained lineage specific adaptations. Signs of streamlining were evident in most genomes, and estimated genome sizes correlated with abundance variation across filter size fractions. Comparing the genomes with globally distributed aquatic metagenomes suggested the existence of a global brackish metacommunity whose populations diverged from freshwater and marine relatives >100,000 years ago, hence long before the Baltic Sea was formed (8000 years). This markedly contrasts to most Baltic Sea multicellular organisms that are locally adapted populations of fresh- or marine counterparts.

Microbiology

OPERA-LG: Efficient and exact scaffolding of large, repeat-rich eukaryotic genomes with performance guarantees

The assembly of large, repeat-rich eukaryotic genomes continues to represent a significant challenge in genomics. While long-read technologies have made the high-quality assembly of small, microbial genomes increasingly feasible, data generation can be prohibitively expensive for larger genomes. Advances in assembly algorithms are thus essential to exploit the characteristics of short and long-read sequencing technologies to consistently and reliably provide high-quality assemblies in a cost-efficient manner. OPERA-LG is a scalable, exact algorithm for the scaffold assembly of large, repeat-rich genomes, with consistent improvement over state-of-the-art programs for scaffold correctness and contiguity. It provides a rigorous framework for scaffolding of repetitive sequences and a systematic approach for combining data from different second-generation (Illumina, Ion Torrent) and third-generation (PacBio, ONT) sequencing technologies. OPERA-LG efficiently scaffolds large genomes with provable scaffold properties, providing an avenue for systematic augmentation and improvement of 1000s of existing draft eukaryotic genome assemblies.

Bioinformatics

Maize pan-transcriptome provides novel insights into genome complexity and quantitative trait variation

Variation in gene expression contributes to the diversity of phenotype. The construction of the pan-transcriptome is especially necessary for species with complex genomes, such as maize. However, knowledge of the regulation mechanisms and functional consequences of the pan-transcriptome is limited. In this study, we identified 13,382 nuclear expression presence and absence variation candidates (ePAVs, expressed in 5%~95% lines; based on the reference genome) by re-analyzing the RNA sequencing data from the kernels (15 days after pollination) of 368 maize diverse inbreds. It was estimated that only ~1% of the ePAVs are explained by DNA sequence presence and absence variations (PAV). The ePAV genes tend to be regulated by distant eQTLs when compared with non-ePAV genes (called here core expression genes, expressed in more than 95% lines). When the expression presence/absence status was used as the \" genotype\" to perform genome-wide association study, 56 (0.42%) ePAVs were significantly associated with 15 agronomic traits and 1,967 (14.74%) with 526 metabolic traits, measured from the mature kernels. While the above was majorly based on the reference genome, by using a modified assemble-then-align strategy, 2,355 high confidence novel sequences with a total length of 1.9Mb were found absent in the current B73 reference genome (v2). Ten randomly selected novel sequences were validated with genomic PCR. A simulation analysis suggested that the pan-transcriptome of the maize whole kernel is approaching a maximum value of 63,000 genes. Two novel validated sequences annotated as NBS_LRR like genes were found to associate with flavonoid content and their homologs in rice were also found to affect flavonoids and disease-resistance. Novel sequences absent in the present reference genome might be functionally important and deserve more attentions. This study provides novel perspectives and resources to discover maize quantitative trait variations and help us to better understand the kernel regulation networks, thus enhancing maize breeding.

Genetics

Scalable multi whole-genome alignment using recursive exact matching

The emergence of third generation sequencing technologies has brought near perfect de-novo genome assembly within reach. This clears the way towards reference-free detection of genomic variations.\n\nIn this paper, we introduce a novel concept for aligning whole-genomes which allows the alignment of multiple genomes. Alignments are constructed in a recursive manner, in which alignment decisions are statistically supported. Computational performance is achieved by splitting an initial indexing data structure into a multitude of smaller indices.\n\nWe show that our method can be used to detect high resolution structural variations between two human genomes, and that it can be used to obtain a high quality multiple genome alignment of at least nineteen Mycobacterium tuberculosis genomes.\n\nAn implementation of the outlined algorithm called REVEAL is available on: https://github.com/jasperlinthorst/REVEAL

Bioinformatics

Genome-Wide Prediction of cis-Regulatory Regions Using Supervised Deep Learning Methods

Identifying active cis-regulatory regions in the human genome is critical for understanding gene regulation and assessing the impact of genetic variation on phenotype. Based on rich data resources such as the Encyclopedia of DNA Elements (ENCODE) and the Functional Annotation of the Mammalian Genome (FANTOM) projects, we introduce DECRES, the first supervised deep learning approach for the identification of enhancer and promoter regions in the human genome. Due to their ability to discover patterns in large and complex data, the introduction of deep learning methods enables a significant advance in our knowledge of the genomic locations of cis-regulatory regions. Using models for well-characterized cell lines, we identify key experimental features that contribute to the predictive performance. Applying DECRES, we delineate locations of 300,000 candidate enhancers genome wide (6.8% of the genome, of which 40,000 are supported by bidirectional transcription data) and 26,000 candidate promoters (0.6% of the genome).

Bioinformatics

From genomes to phenotypes: Traitar, the microbial trait analyzer

The number of sequenced genomes is growing exponentially, profoundly shifting the bottleneck from data generation to genome interpretation. Traits are often used to characterize and distinguish bacteria, and are likely a driving factor in microbial community composition, yet little is known about the traits of most microbes. We describe Traitar, the microbial trait analyzer, which is a fully automated software package for deriving phenotypes from the genome sequence. Traitar provides phenotype classifiers to predict 67 traits related to the use of various substrates as carbon and energy sources, oxygen requirement, morphology, antibiotic susceptibility, proteolysis and enzymatic activities. Furthermore, it suggests protein families associated with the presence of particular phenotypes. Our method uses L1-regularized L2-loss support vector machines for phenotype assignments based on phyletic patterns of protein families and their evolutionary histories across a diverse set of microbial species. We demonstrate reliable phenotype assignment for Traitar to bacterial genomes from 572 species of 8 phyla, also based on incomplete single-cell genomes and simulated draft genomes. We also showcase its application in metagenomics by verifying and complementing a manual metabolic reconstruction of two novel Clostridiales species based on draft genomes recovered from commercial biogas reactors. Traitar is available at https://github.com/hzi-bifo/traitar.

Bioinformatics

Absence of Genome Reduction In Diverse, Facultative Endohyphal Bacteria

Fungi interact closely with bacteria both on the surfaces of hyphae, and within their living tissues (i.e., endohyphal bacteria, EHB). These EHB can be obligate or facultative symbionts, and can mediate a diverse phenotypic traits in their hosts. Although EHB have been observed in many major lineages of fungi, it remains unclear how widespread and general these associations are, and whether there are unifying ecological and genomic features found across all EHB strains. We cultured 11 bacterial strains after they emerged from the hyphae of diverse Ascomycota that were isolated as foliar endophytes of cupressaceous trees, and generated nearly complete genome sequences for all. Unlike the genomes of largely obligate EHB, genomes of these facultative EHB resemble those of closely related strains isolated from environmental sources. Although all analyzed genomes encode structures that can be used to interact with eukaryotic hosts, we find no known pathways that facilitate intimate EHB-fungal interactions in all strains. We isolated two strains with nearly identical genomes from different classes of fungi, consistent with previous suggestions of horizontal transfer of EHB across endophytic hosts. Because bacteria are differentially present during the fungal life cycle, these genomes could shed light on the mechanisms of plant growth promotion by fungal endophytes during the symbiotic phase as well as degradation of plant material during saprotrophic and reproductive phases. Given the capacity of EHB to influence fungal phenotypes, these findings illuminate a new dimension of fungal biodiversity.

Microbiology

Uplift and erosion of genomic islands with standing genetic variation

Details of the processes that generate biological diversity have long been of interest to evolutionary biologists. A common theme in nature is diversification via divergent selection with gene flow. Empirical studies on this topic find variable genetic differentiation throughout the genome, that genetic differentiation is non-randomly distributed, and that loci of adaptive significance are often found clustered together within \"genomic islands of divergence\". Theoretical models based on new mutations show how these genomic islands can arise and grow as a result of a complex interaction of various evolutionary and genic processes. In the current study, I ask if such genomic islands can alternatively arise from divergent selection from standing genetic variation and I tested this using a simple two locus model of selection. There are numerous ways in which standing genetic variation can be partitioned (e.g., between alleles, between loci, and between populations) and I tested which of these scenarios can give rise to an island pattern compared to no genomic differentiation or complete genomic differentiation. I found that divergent selection, even without reciprocal gene exchange between populations, following a bout of admixture can relatively quickly produce an island pattern. Moreover, I found two pathways in which islands can form from divergence from standing variation: 1) through the build up of islands and 2) through the breakdown of larger, genome-wide differentiation. Lastly, similar to new mutation theory, I found that the frequency of recombination is an important determinant of island formation from standing genetic variation such that mating behavior of a species (e.g., facultative or obligate sexual) can impact the likelihood of island formation.

Evolutionary Biology

PBrowse: A web-based platform for real-time collaborative exploration of genomic data

SummaryThe central task of a genome browser is to enable easy visual exploration of large genomic data to gain biological insight. Most existing genome browsers were designed for data exploration by individual users, while a few allow some limited forms of collaboration among multiple users, such as file sharing and wiki-style collaborative editing of gene annotations. Our works premise is that allowing sharing of genome browser views instantaneously in real-time enables the exchange of ideas and insight in a collaborative project, thus harnessing the wisdom of the crowd. PBrowse is a parallel-access real-time collaborative web-based genome browser that provides both an integrated, real-time collaborative platform and a comprehensive file sharing system. PBrowse also allows real-time track comment and has integrated group chat to facilitate interactive discussion among multiple users. Through the Distributed Annotation Server protocol, PBrowse can easily access a wide range of publicly available genomic data, such as the ENCODE data sets. We argue that PBrowse, with the re-designed user management, data management and novel collaborative layer based on Biodalliance, represents a paradigm shift from seeing genome browser merely as a tool of data visualisation to a tool that enables real-time human-human interaction and knowledge exchange in a collaborative setting.\n\nAvailabilityPBrowse is available at http://pbrowse.victorchang.edu.au, and its source code is available via the open source BSD 3 license at http://github.com/VCCRI/PBrowse.\n\nContactj.ho@victorchang.edu.au\n\nSupplementary InformationSupplementary video demonstrating collaborative feature of pbrowse is available in https://www.youtube.com/watch?v=ROvKXZoXiIc.

Bioinformatics

panX: pan-genome analysis and exploration

Horizontal transfer, gene loss, and duplication result in dynamic bacterial genomes shaped by a complex mixture of different modes of evolution. Closely related strains can differ in the presence or absence of many genes, and the total number of distinct genes found in a set of related isolates - the pan-genome - is often many times larger than the genome of individual isolates. We have developed a pipeline that efficiently identifies orthologous gene clusters in the pan-genome. This pipeline is coupled to a powerful yet easy-to-use web-based visualization software for interactive exploration of the pan-genome. The visualization consists of connected components that allow rapid filtering and searching of genes and inspection of their evolutionary history. For each gene cluster, panX displays an alignment, a phylogenetic tree, maps mutations within that cluster to the branches of the tree and infers gain and loss of genes on the core-genome phylogeny. PanX is available at pangenome.de. Custom pan-genomes can be visualized either using a webserver or by serving panX locally as a browser-based application.

Evolutionary Biology

The hidden elasticity of avian and mammalian genomes

Genome size in mammals and birds shows remarkably little interspecific variation compared to other taxa. Yet, genome sequencing has revealed that many mammal and bird lineages have experienced differential rates of transposable element (TE) accumulation, which would be predicted to cause substantial variation in genome size between species. Thus, we hypothesize that there has been co-variation between the amount of DNA gained by transposition and lost by deletion during mammal and avian evolution, resulting in genome size homeostasis. To test this model, we develop a computational pipeline to quantify the amount of DNA gained by TE expansion and lost by deletion over the last 100 million years (My) in the lineages of 10 species of eutherian mammals and 24 species of birds. The results reveal extensive variation in the amount of DNA gained via lineage-specific transposition, but that DNA loss counteracted this expansion to various extent across lineages. Our analysis of the rate and size spectrum of deletion events implies that DNA removal in both mammals and birds has proceeded mostly through large segmental deletions (>10 kb). These findings support a unified accordion model of genome size evolution in eukaryotes whereby DNA loss counteracting TE expansion is a major determinant of genome size. Furthermore, we propose that extensive DNA loss, and not necessarily a dearth of TE activity, has been the primary force maintaining the greater genomic compaction of flying birds and bats relative to their flightless relatives.

evolutionary biology