bioRxiv ScienceSearch

SEARCH · bioRxiv Science

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

Allele-specific control of replication timing and genome organization during development

DNA replication occurs in a defined temporal order known as the replication-timing (RT) program. RT is regulated during development in discrete chromosomal units, coordinated with transcriptional activity and 3D genome organization. Here, we derived distinct cell types from F1 hybrid musculus X castaneus mouse crosses and exploited the high single nucleotide polymorphism (SNP) density to characterize allelic differences in RT (Repli-seq), genome organization (Hi-C and promoter-capture Hi-C), gene expression (nuclear RNA-seq) and chromatin accessibility (ATAC-seq). We also present HARP: a new computational tool for sorting SNPs in phased genomes to efficiently measure allele-specific genome-wide data. Analysis of 6 different hybrid mESC clones with different genomes (C57BL/6, 129/sv and CAST/Ei), parental configurations and gender revealed significant RT asynchrony between alleles across ~12 % of the autosomal genome linked to sub-species genomes but not to parental origin, growth conditions or gender. RT asynchrony in mESCs strongly correlated with changes in Hi-C compartments between alleles but not SNP density, gene expression, imprinting or chromatin accessibility. We then tracked mESC RT asynchronous regions during development by analyzing differentiated cell types including extraembryonic endoderm stem (XEN) cells, 4 male and female primary mouse embryonic fibroblasts (MEFs) and neural precursors (NPCs) differentiated in vitro from mESCs with opposite parental configurations. Surprisingly, we found that RT asynchrony and allelic discordance in Hi-C compartments seen in mESCs was largely lost in all differentiated cell types, coordinated with a more uniform Hi-C compartment arrangement, suggesting that genome organization of homologues converges to similar folding patterns during cell fate commitment.

cell biology

Reconstructing phylogeny from reduced-representation genome sequencing data without assembly or alignment

Although genome sequencing is becoming cheaper and faster, reducing the quantity of data by only sequencing part of the genome lowers both sequencing costs and computational burdens. One popular genome-reduction approach is restriction site associated DNA sequencing, or RADseq. RADseq was initially designed for studying genetic variation across genomes usually at the population level, and it has also proved to be suitable for interspecific phylogeny reconstruction. RADseq data pose challenges for standard phylogenomic methods, however, due to incomplete coverage of the genome and large amounts of missing data. Alignment-free methods are both efficient and accurate for phylogenetic reconstructions with whole genomes and are especially practical for non-model organisms; nonetheless, alignment-free methods have only been applied with whole genome sequences. Here, we test a full-genome assembly and alignment-free method, AAF, in application to RADseq data and propose two procedures for reads selection to remove missing data. We validate these methods using both simulations and a real dataset. Reads selection improved the accuracy of phylogenetic construction in every simulated scenario and the real dataset, making AAF comparable to or better than alignment-based method with much lower computation burdens. We also investigated the sources of missing data in RADseq and their effects on phylogeny reconstruction using AAF. The AAF pipeline modified for RADseq data, phyloRAD, is available on github (https://github.com/fanhuan/phyloRAD).

bioinformatics

Improving the annotation of the Heterorhabditis bacteriophora genome

Genome assembly and annotation remains an exacting task. As the tools available for these tasks improve, it is useful to return to data produced with earlier instances to assess their credibility and correctness. The entomopathogenic nematode Heterorhabditis bacteriophora is widely used to control insect pests in horticulture. The genome sequence for this species was reported to encode an unusually high proportion of unique proteins and a paucity of secreted proteins compared to other related nematodes. We revisited the H. bacteriophora genome assembly and gene predictions to ask whether these unusual characteristics were biological or methodological in origin. We mapped an independent resequencing dataset to the genome and used the blobtools pipeline to identify potential contaminants. While present (0.2% of the genome span, 0.4% of predicted proteins), assembly contamination was not significant. Re-prediction of the gene set using BRAKER1 and published transcriptome data generated a predicted proteome that was very different from the published one. The new gene set had a much reduced complement of unique proteins, better completeness values that were in line with other related species genomes, and an increased number of proteins predicted to be secreted. It is thus likely that methodological issues drove the apparent uniqueness of the initial H. bacteriophora genome annotation and that similar contamination and misannotation issues affect other published genome assemblies.

bioinformatics

How well can we create phased, diploid, human genomes?: An assessment of FALCON-Unzip phasing using a human trio

Long read sequencing technology has allowed researchers to create de novo assemblies with impressive continuity[1,2]. This advancement has dramatically increased the number of reference genomes available and hints at the possibility of a future where personal genomes are assembled rather than resequenced. In 2016 Pacific Biosciences released the FALCON-Unzip framework, which can provide long, phased haplotype contigs from de novo assemblies. This phased genome algorithm enhances the accuracy of highly heterozygous organisms and allows researchers to explore questions that require haplotype information such as allele-specific expression and regulation. However, validation of this technique has been limited to small genomes or inbred individuals[3].\n\nAs a roadmap to personal genome assembly and phasing, we assess the phasing accuracy of FALCON-Unzip in humans using publicly available data for the Ashkenazi trio from the Genome in a Bottle Consortium[4]. To assess the accuracy of the Unzip algorithm, we assembled the genome of the son using FALCON and FALCON Unzip, genotyped publicly available short read data for the mother and the father, and observed the inheritance pattern of the parental SNPs along the phased genome of the son. We found that 72.8% of haplotype contigs share SNPs with only one parent suggesting that these contigs are correctly phased. Most mis-phased SNPs are random but present in high frequency toward the end of haplotype contigs. Approximately 20.7% of mis-phased haplotype contigs contain clusters of mis-phased SNPs, suggesting that haplotypes were mis-joined by FALCON-Unzip. Mis-joined boundaries in those contigs are located in areas of low SNP density. This research demonstrates that the FALCON-Unzip algorithm can be used to create long and accurate haplotypes for humans and identifies problematic regions that could benefit in future improvement.

bioinformatics

HISAT-genotype: Next Generation Genomic Analysis Platform on a Personal Computer

Rapid advances in next-generation sequencing technologies have dramatically changed our ability to perform genome-scale analyses of human genomes. The human reference genome used for most genomic analyses represents only a small number of individuals, limiting its usefulness for genotyping. We designed a novel method, HISAT-genotype, for representing and searching an expanded model of the human reference genome, in which a comprehensive catalogue of known genomic variants and haplotypes is incorporated into the data structure used for searching and alignment. This strategy for representing a population of genomes, along with a very fast and memory-efficient search algorithm, enables more detailed and accurate variant analyses than previous methods. We demonstrate HISAT-genotypes accuracy for HLA typing, a critical task in human organ transplantation, and for the DNA fingerprinting tests widely used in forensics. In both applications, HISAT-genotype not only improves upon earlier computational methods, but matches or exceeds the accuracy of laboratory-based assays.\n\nOne Sentence SummaryHISAT-genotype is a software platform that has the ability to genotype all the genes in an individuals genome within a few hours on a desktop computer.

bioinformatics

Functional equivalence of genome sequencing analysis pipelines enables harmonized variant calling across human genetics projects

Hundreds of thousands of human whole genome sequencing (WGS) datasets will be generated over the next few years to interrogate a broad range of traits, across diverse populations. These data are more valuable in aggregate: joint analysis of genomes from many sources increases sample size and statistical power for trait mapping, and will enable studies of genome biology, population genetics and genome function at unprecedented scale. A central challenge for joint analysis is that different WGS data processing and analysis pipelines cause substantial batch effects in combined datasets, necessitating computationally expensive reprocessing and harmonization prior to variant calling. This approach is no longer tenable given the scale of current studies and data volumes. Here, in a collaboration across multiple genome centers and NIH programs, we define WGS data processing standards that allow different groups to produce \"functionally equivalent\" (FE) results suitable for joint variant calling with minimal batch effects. Our approach promotes broad harmonization of upstream data processing steps, while allowing for diverse variant callers. Importantly, it allows each group to continue innovating on data processing pipelines, as long as results remain compatible. We present initial FE pipelines developed at five genome centers and show that they yield similar variant calling results - including single nucleotide (SNV), insertion/deletion (indel) and structural variation (SV) - and produce significantly less variability than sequencing replicates. Residual inter-pipeline variability is concentrated at low quality sites and repetitive genomic regions prone to stochastic effects. This work alleviates a key technical bottleneck for genome aggregation and helps lay the foundation for broad data sharing and community-wide \"big-data\" human genetics studies.

bioinformatics

The yeast core spliceosome maintains genome integrity through R-loop prevention and alpha-tubulin expression

To achieve genome stability cells must coordinate the action of various DNA transactions including DNA replication, repair, transcription and chromosome segregation. How transcription and RNA processing enable genome stability is only partly understood. Two predominant models have emerged: one involving changes in gene expression that perturb other genome maintenance factors, and another in which genotoxic DNA:RNA hybrids, called R-loops, impair DNA replication. Here we characterize genome instability phenotypes in a panel yeast splicing factor mutants and find that mitotic defects, and in some cases R-loop accumulation, are causes of genome instability. Genome instability in splicing mutants is exacerbated by loss of the spindle-assembly checkpoint protein Mad1. Moreover, removal of the intron from the -tubulin gene TUB1 restores genome integrity. Thus, while R-loops contribute in some settings, defects in yeast splicing predominantly lead to genome instability through effects on gene expression.

cell biology

segment_liftover: a Python tool to convert segments between genome assemblies

The process of assembling a species reference genome may be performed in a number of iterations, with subsequent genome assemblies differing in the coordinates of mapped elements. The conversion of genome coordinates between different assemblies is required for many integrative and comparative studies. While currently a number of bioinformatics tools are available to accomplish this task, most of them are tailored towards the conversion of single genome coordinates. When converting the boundary positions of segments spanning larger genome regions, segments may be mapped into smaller subsegments if the original segments continuity is disrupted in the target assembly. Such a conversion may lead to a relevant degree of data loss in some circumstances such as copy number variation (CNV) analysis, where the quantitative representation of a genomic region takes precedence over base-specific accuracy. segment_liftover aims at continuity-preserving remapping of genome segments between assemblies and provides features such as approximate locus conversion, automated batch processing and comprehensive logging to facilitate processing of datasets containing large numbers of structural genome variation data.

bioinformatics

sppIDer: a species identification tool to investigate hybrid genomes with high-throughput sequencing

The genomics era has expanded our knowledge about the diversity of the living world, yet harnessing high-throughput sequencing data to investigate alternative evolutionary trajectories, such as hybridization, is still challenging. Here we present sppIDer, a pipeline for the characterization of interspecies hybrids and pure species,that illuminates the complete composition of genomes. sppIDer maps short-read sequencing data to a combination genome built from reference genomes of several species of interest and assesses the genomic contribution and relative ploidy of each parental species, producing a series of colorful graphical outputs ready for publication. As a proof-of-concept, we use the genus Saccharomyces to detect and visualize both interspecies hybrids and pure strains, even with missing parental reference genomes. Through simulation, we show that sppIDer is robust to variable reference genome qualities and performs well with low-coverage data. We further demonstrate the power of this approach in plants, animals, and other fungi. sppIDer is robust to many different inputs and provides visually intuitive insight into genome composition that enables the rapid identification of species and their interspecies hybrids. sppIDer exists as a Docker image, which is a reusable, reproducible, transparent, and simple-to-run package that automates the pipeline and installation of the required dependencies (https://github.com/GLBRC/sppIDer).

evolutionary biology

Understanding trivial challenges of microbial genomics: An assembly example

The perceived \"simplicity\" of bacterial genomics (these genomes are small and easy to assemble) feeds the decentralized state of the field where computational analysis standards have been slow to evolve. This situation has a historical explanation. In cases of human, mouse, fly, worm and other model organisms there have been large sustained multinational genome sequencing efforts and analysis consortia such as the 1,000 genomes, ENCODE, modENCODE, GTEx and others. These resulted in development and proliferation of common tools, workflows, and data standards. This is not the case in microbiology. After the development of highly parallel sequencing methodologies in mid-2000s bacterial genomes no longer required initiatives of such scale. The flipside of this is the extreme heterogeneity of approaches to many well established microbial genomic analysis problems such as genome assembly. While competition amongst different methods is good, we argue that the quality of data analyses will improve if cutting edge tools are more accessible and microbiologists become more computationally savvy. Here we use genome assembly as an example to highlight current challenges and to provide a possible solution.

microbiology

Brd4 and P300 regulate zygotic genome activation through histone acetylation

The awakening of the zygote genome, signaling the transition from maternal transcriptional control to zygotic control, is a watershed in embryonic development, but the factors and mechanisms controlling this transition are still poorly understood. By combining CRISPR-Cas9-mediated live imaging of the first transcribed genes (miR-430), chromatin and transcription analysis during zebrafish embryogenesis, we observed that genome activation is gradual and stochastic, and the active state is inherited in daughter cells. We discovered that genome activation is regulated through both translation of maternal mRNAs and the effects of these factors on the chromatin. We show that chemical inhibition of H3K27Ac writer (P300) and reader (Brd4) block genome activation, while induction of a histone acetylation prematurely activates transcription, and restore genome activation in embryos where translation of maternal mRNAs is impaired, demonstrating that they are limiting factors for the activation of the genome. In contrast to current models, we do not observe triggering of genome activation by a reduction of the nuclear-cytoplasmic (N/C) ratio or slower cell division. We conclude that genome activation is controlled by a time-dependent mechanism involving the translation of maternal mRNAs and the regulation of histone acetylation through P300 and Brd4. This mechanism is critical to initiating zygotic development and developmental reprogramming.

developmental biology

A genus definition for Bacteria and Archaea based on genome relatedness and taxonomic affiliation.

Genus assignment is fundamental in the characterization of microbes, yet there is currently no unambiguous way to demarcate genera solely using standard genomic relatedness indices. Here, we propose an approach to demarcate genera that relies on the combined use of the average nucleotide identity, genome alignment fraction, and the distinction between type species and non-type species. More than 750 genomes representing type strains of species from 10 different phyla, and 19 different taxonomic orders/families in Gram-positive/negative, bacterial and archaeal lineages were tested. Overall, all 19 analyzed taxa conserved significant genomic differences between members of a genus and type species of other genera in the same taxonomic family. Bacillus, Flavobacterium, Hydrogenovibrio, Lactococcus, Methanosarcina, Thiomicrorhabdus, Thiomicrospira, Shewanella, and Vibrio are discussed in detail. Less than 1% of the type strains analyzed need reclassification, highlighting that the adoption of the 16S rRNA gene as a taxonomic marker has provided consistency to the classification of microorganisms in recent decades. One exception to this is the genus Bacillus with 61% of type strains needing reclassification, including the human pathogens B. cereus and B. anthracis. The results provide a first line of evidence that the combination of genomic indices provides appropriate resolution to effectively demarcate genera within the current taxonomic framework that is based on the 16S rRNA gene. We also identify the emergence of natural breakpoints at the genome level that can further help in the circumscription of genera. Altogether, these results show that a distinct difference between distant relatives and close relatives at the genome level (i.e., genomic coherence) is an emergent property of genera in Bacteria and Archaea.

microbiology

VAPiD: a lightweight cross platform viral annotation pipeline and identification tool to facilitate virus genome submissions to NCBI GenBank

BackgroundWith sequencing technologies becoming cheaper and easier to use, more groups are able to obtain whole genome sequences of viruses of public health and scientific importance. Submission of genomic data to NCBI GenBank is a requirement prior to publication and plays a critical role in making scientific data publicly available.\n\nGenBank currently has automatic prokaryotic and eukaryotic genome annotation pipelines but has no viral annotation pipeline beyond influenza virus. Annotation and submission of viral genome sequence is a non-trivial task, especially for groups that do not routinely interact with GenBank for data submissions.\n\nResultsWe present Viral Annotation Pipeline and iDentification (VAPiD), a portable and lightweight command-line tool for annotation and GenBank deposition of viral genomes. VAPiD supports annotation of nearly all unsegmented viral genomes. The pipeline has been validated on human immunodeficiency virus, human parainfluenza virus 1-4, human metapneumovirus, human coronaviruses (229E/OC43/NL63/HKU1/SARS/MERS), human enteroviruses/rhinoviruses, measles virus, mumps virus, Hepatitis A-E Virus, Chikungunya virus, dengue virus, and West Nile virus, as well the human polyomaviruses BK/JC/MCV, human adenoviruses, and human papillomaviruses. The program can handle individual or batch submissions of different viruses to GenBank and correctly annotates multiple viruses, including those that contain ribosomal slippage or RNA editing without prior knowledge of the virus to be annotated. VAPiD is programmed in Python and is compatible with Windows, Linux, and Mac OS systems.\n\nConclusionsWe have created a portable, lightweight, user-friendly, internet-enabled, open-source, command-line genome annotation and submission package to facilitate virus genome submissions to NCBI GenBank. Instructions for downloading and installing VAPiD can be found at https://github.com/rcs333/VAPiD.

bioinformatics

Using QC-Blind for quality control and contamination screening of bacteria DNA sequencing data without reference genome

Quality control in next generation sequencing has become increasingly important as the technique becomes widely used. Tools have been developed for filtering possible contaminants in the sequencing data of species with known reference genome. Unfortunately, reference genomes for all the species involved, including the contaminants, are required for these tools to work. This precludes many real-life samples that have no information about the complete genome of the target species, and are contaminated with unknown microbial species.\n\nIn this work we propose QC-Blind, a novel quality control pipeline for removing contaminants without any use of reference genomes. The pipeline requires only very little information from the marker genes of the target species. The entire pipeline consists of unsupervised read assembly, contig binning, read clustering and marker gene assignment.\n\nWhen evaluated on in silico, ab initio and in vivo datasets, QC-Blind proved effective in removing unknown contaminants with high specificity and accuracy, while preserving most of the genomic information of the target bacterial species. Therefore, QC-Blind could serve well in situations where limited information is available for both target and contamination species.\n\nIMPORTANCEAt present, many sequencing projects are still performed on potentially contaminated samples, which bring into question their accuracies. However, current reference-based quality control method are limited as they need either the genome of target species or contaminations. In this work we propose QC-Blind, a novel quality control pipeline for removing contaminants without any use of reference genomes. When evaluated on in silico, ab initio and in vivo datasets, QC-Blind proved effective in removing unknown contaminants with high specificity and accuracy, while preserving most of the genomic information of the target bacterial species. Therefore, QC-Blind is suitable for real-life samples where limited information is available for both target and contamination species.

bioinformatics

High-resolution structural genomics reveals new therapeutic vulnerabilities in glioblastoma

We investigated the role of 3D genome architecture in instructing functional properties of glioblastoma stem cells (GSCs) by generating the highest-resolution 3D genome maps to-date for this cancer. Integration of DNA contact maps with chromatin and transcriptional profiles identified specific mechanisms of gene regulation, including individual physical interactions between regulatory regions and their target genes. Residing in structurally conserved regions in GSCs was CD276, a gene known to play a role in immuno-modulation. We show that, unexpectedly, CD276 is part of a stemness network in GSCs and can be targeted with an antibody-drug conjugate to curb self-renewal, a key stemness property. Our results demonstrate that integrated structural genomics datasets can be employed to rationally identify therapeutic vulnerabilities in self-renewing cells.\n\nSIGNIFICANCEIn adult GBM, GSCs act as therapy-resistant reservoirs to nucleate tumor recurrence. New therapeutic approaches that target these cell populations hold the potential of significantly improving patient care and overall prognosis for this always-lethal cancer. Our work describes new links between 3D genome architecture and stemness properties in GSCs. In particular, through integration of multiple genomics and structural genomics datasets, we found an unexpected connection between immune-related genes and self-renewal programs in GBM. Among these, we show that targeting CD276 with knockdown strategies or specific antibody-drug conjugates achieve suppression of self-renewal. Strategies to target CD276+ cells are currently in clinical trials for solid tumors. Our results indicate that CD276-targeting agents could be deployed in GBM to specifically target GSC populations.\n\nHIGHLIGHTSO_LIWe generated high (sub-5 kb) resolution Hi-C maps for stem-like cells from GBM patients.\nC_LIO_LIIntegration of Hi-C and genomics datasets dissects mechanisms of gene regulation.\nC_LIO_LI3D genomes poise immune-related genes, including CD276, for expression.\nC_LIO_LITargeting CD276 curbs self-renewal properties of GBM cells.\nC_LI

cancer biology

A direct comparison of genome alignment and transcriptome pseudoalignment

MotivationGenome alignment of reads is the first step of most genome analysis workflows. In the case of RNA-Seq, transcriptome pseudoalignment of reads is a fast alternative to genome alignment, but the different \"coordinate systems\" of the genome and transcriptome have made it difficult to perform direct comparisons between the approaches.\n\nResultsWe have developed tools for converting genome alignments to transcriptome pseudoalignments, and conversely, for projecting transcriptome pseudoalignments to genome alignments. Using these tools, we performed a direct comparison of genome alignment with transcriptome pseudoalignment. We find that both approaches produce similar quantifications. This means that for many applications genome alignment and transcriptome pseudoalignment are interchangeable.\n\nAvailability and Implementationbam2tcc is a C++14 software for converting alignments in SAM/BAM format to transcript compatibility counts (TCCs) and is available at https://github.com/pachterlab/bam2tcc. kallisto genomebam is a user option of kallisto that outputs a sorted BAM file in genome coordinates as part of transcriptome pseudoalignment. The feature has been released with kallisto v0.44.0, and is available at https://pachterlab.github.io/kallisto/.\n\nSupplementary MaterialN/A\n\nContactLior Pachter (lpachter@caltech.edu)

bioinformatics

Inference of Gorilla demographic and selective history from whole genome sequence data

While population-level genomic sequence data have been gathered extensively for humans, similar data from our closest living relatives are just beginning to emerge. Examination of genomic variation within great apes offers many opportunities to increase our understanding of the forces that have differentially shaped the evolutionary history of hominid taxa. Here, we expand upon the work of the Great Ape Genome Project by analyzing medium to high coverage whole genome sequences from 14 western lowland gorillas (Gorilla gorilla gorilla), 2 eastern lowland gorillas (G. beringei graueri), and a single Cross River individual (G. gorilla diehli). We infer that the ancestors of western and eastern lowland gorillas diverged from a common ancestor [~]261 thousand years ago (kya), and that the ancestors of the Cross River population diverged from the western lowland gorilla lineage [~]68 kya. Using a diffusion approximation approach to model the genome-wide site frequency spectrum, we infer a history of western lowland gorillas that includes an ancestral population expansion of [~]1.4-fold around [~]970 kya and a recent [~]5.6-fold contraction in population size [~]23 kya. The latter may correspond to a major reduction in African equatorial forests around the Last Glacial Maximum. We also analyze patterns of variation among western lowland gorillas to identify several genomic regions with strong signatures of recent selective sweeps. We find that processes related to taste, pancreatic and saliva secretion, sodium ion transmembrane transport, and cardiac muscle function are overrepresented in genomic regions predicted to have experienced recent positive selection.

Genomics

The projection of a test genome onto a reference population and applications to humans and archaic hominins

We introduce a method for comparing a test genome with numerous genomes from a reference population. Sites in the test genome are given a weight w that depends on the allele frequency x in the reference population. The projection of the test genome onto the reference population is the average weight for each x, [Formula]. The weight is assigned in such a way that if the test genome is a random sample from the reference population, [Formula]. Using analytic theory, numerical analysis, and simulations, we show how the projection depends on the time of population splitting, the history of admixture and changes in past population size. The projection is sensitive to small amounts of past admixture, the direction of admixture and admixture from a population not sampled (a ghost population). We compute the projection of several human and two archaic genomes onto three reference populations from the 1000 Genomes project, Europeans (CEU), Han Chinese (CHB) and Yoruba (YRI) and discuss the consistency of our analysis with previously published results for European and Yoruba demographic history. Including higher amounts of admixture between Europeans and Yoruba soon after their separation and low amounts of admixture more recently can resolve discrepancies between the projections and demographic inferences from some previous studies.

Genomics