bioRxiv ScienceSearch

SEARCH · bioRxiv Science

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 775 records · Page 43Linked to original sources

CONSTRUCTION OF WHOLE GENOMES FROM SCAFFOLDS USING SINGLE CELL STRAND-SEQ DATA

Accurate reference genome sequences provide the foundation for modern molecular biology and genomics as the interpretation of sequence data to study evolution, gene expression and epigenetics depends heavily on the quality of the genome assembly used for its alignment. Correctly organising sequenced fragments such as contigs and scaffolds in relation to each other is a critical and often challenging step in the construction of robust genome references. We previously identified misoriented regions in the mouse and human reference assemblies using Strand-seq, a single cell sequencing technique that preserves DNA directionality1, 2. Here we demonstrate the ability of Strand-seq to build and correct full-length chromosomes, by identifying which scaffolds belong to the same chromosome and determining their correct order and orientation, without the need for overlapping sequences. We demonstrate that Strand-seq exquisitely maps assembly fragments into large related groups and chromosome-sized clusters without using new assembly data. Using template strand inheritance as a bi-allelic marker, we employ genetic mapping principles to cluster scaffolds that are derived from the same chromosome and order them within the chromosome based solely on directionality of DNA strand inheritance. We prove the utility of our approach by generating improved genome assemblies for several model organisms including the ferret, pig, Xenopus, zebrafish, Tasmanian devil and the Guinea pig.

genomics

Evaluation of Whole Exome Sequencing as an Alternative of BeadChip and Whole Genome Sequencing in Human Population Genetic Analysis

Understanding the underlying genetic structure of human populations is of fundamental interest to both biological and social sciences. Advances in high-throughput genotyping technology have markedly improved our understanding of global patterns of human genetic variation. The most widely used methods for collecting variant information at the DNA-level include whole genome sequencing, which continues to remain costly, and the more economical solution of array-based techniques, as these are capable of simultaneously genotyping a pre-selected set of variable DNA sites in the human genome. The largest publicly accessible set of human genomic sequence data available today originates from exome sequencing that comprises around 1.2% of the whole genome (approximately 30 million base pairs). In this study, we compared the application of the exome dataset to the array-based dataset and to the gold standard whole genome dataset using the same population genetic analysis methods. Our results draw attention to some of the inherent problems that arise from using pre-selected SNP sets for population genetic analysis. Additionally, we demonstrate that exome sequencing provides a better alternative to the array-based methods for population genetic analysis. In this study, we propose a strategy for unbiased variant collection from exome data and offer a bioinformatics protocol for proper data processing.

genomics

Association of whole-genome and NETRIN1 signaling pathway-derived polygenic risk scores for Major Depressive Disorder and thalamic radiation white matter microstructure in UK Biobank

BackgroundMajor Depressive Disorder (MDD) is a clinically heterogeneous psychiatric disorder with a polygenic architecture. Genome-wide association studies have identified a number of risk-associated variants across the genome, and growing evidence of NETRIN1 pathway involvement. Stratifying disease risk by genetic variation within the NETRIN1 pathway may provide an important route for identification of disease mechanisms by focusing on a specific process excluding heterogeneous risk-associated variation in other pathways. Here, we sought to investigate whether MDD polygenic risk scores derived from the NETRIN1 signaling pathway (NETRIN1-PRS) and the whole genome excluding NETRIN1 pathway genes (genomic-PRS) were associated with white matter integrity.\n\nMethodsWe used two diffusion tensor imaging measures, fractional anisotropy (FA) and mean diffusivity (MD), in the most up-to-date UK Biobank neuroimaging data release (FA: N = 6,401; MD: N = 6,390).\n\nResultsWe found significantly lower FA in the superior longitudinal fasciculus ({beta} = -0.035, pcorrected = 0.029) and significantly higher MD in a global measure of thalamic radiations ({beta} = 0.029, pcorrected = 0.021), as well as higher MD in the superior ({beta} = 0.034, pcorrected = 0.039) and inferior ({beta} = 0.029, pcorrected = 0.043) longitudinal fasciculus and in the anterior ({beta} = 0.025, pcorrected = 0.046) and superior ({beta} = 0.027, pcorrected = 0.043) thalamic radiation associated with NETRIN1-PRS. Genomic-PRS was also associated with lower FA and higher MD in several tracts.\n\nConclusionsOur findings indicate that variation in the NETRIN1 signaling pathway may confer risk for MDD through effects on thalamic radiation white matter microstructure.

genetics

Immuno-genomic PanCancer Landscape Reveals Diverse Immune Escape Mechanisms and Immuno-Editing Histories

Immune reactions in the tumor micro-environment are one of the cancer hallmarks and emerging immune therapies have been proven effective in many types of cancer. To investigate cancer genome-immune interactions and the role of immuno-editing or immune escape mechanisms in cancer development, we analyzed 2,834 whole genomes and RNA-seq datasets across 31 distinct tumor types from the PanCancer Analysis of Whole Genomes (PCAWG) project with respect to key immuno-genomic aspects. We show that selective copy number changes in immune-related genes could contribute to immune escape. Furthermore, we developed an index of the immuno-editing history of each tumor sample based on the information of mutations in exonic regions and pseudogenes. Our immuno-genomic analyses of pan-cancer analyses have the potential to identify a subset of tumors with immunogenicity and diverse background or intrinsic pathways associated with their immune status and immuno-editing history.

genomics

Three invariant Hi-C interaction patterns: applications to genome assembly

Assembly of reference-quality genomes from next-generation sequencing data is a key challenge in genomics. Recently, we and others have shown that Hi-C data can be used to address several outstanding challenges in the field of genome assembly. This principle has since been developed in academia and industry, and has been used in the assembly of several major genomes. In this paper, we explore the central principles underlying Hi-C-based assembly approaches, by quantitatively defining and characterizing three invariant Hi-C interaction patterns on which these approaches can build: Intrachromosomal interaction enrichment, distance-dependent interaction decay and local interaction smoothness. Specifically, we evaluate to what degree each invariant pattern holds on a single locus level in different species, cell types and Hi-C map resolutions. We find that these patterns are generally consistent across species and cell types but are affected by sequencing depth, and that matrix balancing improves consistency of loci with all three invariant patterns. Finally, we overview current Hi-C-based assembly approaches in light of these invariant patterns and demonstrate how local interaction smoothness can be used to easily detect scaffolding errors in extremely sparse Hi-C maps. We suggest that simultaneously considering all three invariant patterns may lead to better Hi-C-based genome assembly methods.

genomics

EquCab3, an Updated Reference Genome for the Domestic Horse

EquCab2, a high-quality reference genome for the domestic horse, was released in 2007. Since then, it has served as the foundation for nearly all genomic work done in equids. Recent advances in genomic sequencing technology and computational assembly methods have allowed scientists to improve reference assemblies of large animal and plant genomes in terms of contiguity and composition. In 2014, the equine genomics research community began a project to improve the reference sequence for the horse, building upon the solid foundation of EquCab2 and incorporating new short-read data, long-read data, and proximity ligation data. The result, EquCab3, is presented here. The count of non-N bases in the incorporated chromosomes is improved from 2.33Gb in EquCab2 to 2.41Gb from EquCab3. Contiguity has also been improved nearly 40-fold with a contig N50 of 4.5Mb and scaffold contiguity enhanced to where all but one of the 32 chromosomes is comprised of a single scaffold.

genomics

TSA-Seq Mapping of Nuclear Genome Organization

While nuclear compartmentalization is an essential feature of three-dimensional genome organization, no genomic method exists for measuring chromosome distances to defined nuclear structures. Here we describe TSA-Seq, a new mapping method able to estimate mean chromosomal distances from nuclear speckles genome-wide and predict several Mbp chromosome trajectories between nuclear compartments without sophisticated computational modeling. Ensemble-averaged results reveal a clear nuclear lamina to speckle axis correlated with a striking spatial gradient in genome activity. This gradient represents a convolution of multiple, spatially separated nuclear domains, including two types of transcription \"hot-zones\". Transcription hot-zones protruding furthest into the nuclear interior and positioning deterministically very close to nuclear speckles have higher numbers of total genes, the most highly expressed genes, house-keeping genes, genes with low transcriptional pausing, and super-enhancers. Our results demonstrate the capability of TSA-Seq for genome-wide mapping of nuclear structure and suggest a new model for nuclear spatial organization of transcription.

genomics

Genomic prediction using individual-level data and summary statistics from multiple populations

This study presents a method for genomic prediction that uses individual-level data and summary statistics from multiple populations. Genome-wide markers are nowadays widely used to predict complex traits, and genomic prediction using multi-population data is an appealing approach to achieve higher prediction accuracies. However, sharing of individual-level data across populations is not always possible. We present a method that enables integration of summary statistics from separate analyses with the available individual-level data. The data can either consist of individuals with single or multiple (weighted) phenotype records per individual. We developed a method based on a hypothetical joint analysis model and absorption of population specific information. We show that population specific information is fully captured by estimated allele substitution effects and the accuracy of those estimates, i.e. the summary statistics. The method gives identical result as the joint analysis of all individual-level data when complete summary statistics are available. We provide a series of easy-to-use approximations that can be used when complete summary statistics are not available or impractical to share. Simulations show that approximations enables integration of different sources of information across a wide range of settings yielding accurate predictions. The method can be readily extended to multiple-traits. In summary, the developed method enables integration of genome-wide data in the individual-level or summary statistics form from multiple populations to obtain more accurate estimates of allele substitution effects and genomic predictions.

genomics

Livestock genome annotation: transcriptome and chromatin structure profiling in cattle, goat, chicken and pig.

BackgroundFunctional annotation of livestock genomes is a critical step to decipher the genotype-to-phenotype relationship underlying complex traits. As part of the Functional Annotation of Animal Genomes (FAANG) action, the FR-AgENCODE project (http://www.fragencode.org) aimed to profile the landscape of transcription (RNA-seq), chromatin accessibility (ATAC-seq) and conformation (Hi-C) in four livestock species representing ruminants (cattle, goat), monogastrics (pig) and birds (chicken), using three target samples related to metabolism (liver) and immunity (CD4+ and CD8+ T cells).\n\nResultsRNA-seq assays considerably extended the available catalog of annotated transcripts and identified differentially expressed genes with unknown function, including new syntenic lncRNAs. ATAC-seq highlighted an enrichment for transcription factor binding sites in differentially accessible regions of the chromatin. Comparative analyses revealed a core set of conserved regulatory regions across species. Topologically Associating Domains (TADs) and epigenetic A/B compartments annotated from Hi-C data were consistent with RNA-seq and ATAC-seq data. Multi-species comparisons showed that conserved TAD boundaries had stronger insulation properties than species-specific ones and that the genomic distribution of orthologous genes in A/B compartments was significantly conserved across species.\n\nConclusionsWe report the first multi-species and multi-assay genome annotation results obtained by a FAANG project. Beyond the generation of reference annotations and the confirmation of previous findings on model animals, the integrative analysis of data from multiple assays and species sheds a new light on the multi-scale selective pressure shaping genome organization from birds to mammals. Overall, these results emphasize the value of FAANG for research on domesticated animals and reinforces the importance of future meta-analyses of the reference datasets being generated by this community on different species.

genomics

Viral Taxonomy Derived From Evolutionary Genome Relationships

We describe a new genome alignment-based model for classification of viruses based on evolutionary genetic relationships. This approach uses information theory and a physical model to determine the information shared by the genes in two genomes. Pairwise comparisons of genes from the viruses are created from alignments using NCBI BLAST, and their match scores are combined to produce a metric between genomes, which is in turn used to determine a global classification using the 5,817 viruses on RefSeq. In cases where there is no measurable alignment between any genes, the method falls back to a coarser measure of genome relationship: the mutual information of k-mer frequency. This results in a principled model which depends only on the genome sequence, which captures many interesting relationships between viral families, and which creates clusters which correlate well with both the Baltimore and ICTV classifications. The incremental computational cost of classifying a novel virus is low and therefore newly discovered viruses can be quickly identified and classified.

genomics

The African Bullfrog (Pyxicephalus adspersus) genome unites the two ancestral ingredients formaking vertebrate sex chromosomes

Heteromorphic sex chromosomes have evolved repeatedly among vertebrate lineages despite largely deleterious reductions in gene dose. Understanding how this gene dose problem is overcome is hampered by the lack of genomic information at the base of tetrapods and comparisons across the evolutionary history of vertebrates. To address this problem, we produced a chromosome-level genome assembly for the African Bullfrog (Pyxicephalus adspersus)--an amphibian with heteromorphic ZW sex chromosomes--and discovered that the Bullfrog Z is surprisingly homologous to substantial portions of the human X. Using this new reference genome, we identified ancestral synteny among the sex chromosomes of major vertebrate lineages, showing that non-mammalian sex chromosomes are strongly associated with a single vertebrate ancestral chromosome, while mammals are associated with another that displays increased haploinsufficiency. The sex chromosomes of the African Bullfrog however, share genomic blocks with both humans and non-mammalian vertebrates, connecting the two ancestral chromosome sequences that repeatedly characterize vertebrate sex chromosomes. Our results highlight the consistency of sex-linked sequences despite sex determination system lability and reveal the repeated use of two major genomic sequence blocks during vertebrate sex chromosome evolution.

genomics

Extended regions of suspected mis-assembly in the rat reference genome

We performed whole-genome sequencing for eight inbred rat strains commonly used in genetic mapping studies, and they are the founders of the NIH heterogeneous stock (HS) outbred colony. We provide their sequences and variant calls to the rat genomics community. When analyzing the variant calls we identified regions with unusually high heterozygosity. We show that these regions are consistent across the eight inbred strains, including the BN strain, which was the basis of the rat reference genome. These regions show significantly higher read depth than other regions in the genome. The evidence suggests that these regions may contain segmental duplications that are incorrectly overlaid in the reference genome. We provide masks for these suspected regions of mis-assembly as a resource for the community to flag potentially false interpretations of mapping results or functional data.

genomics

Genomic variants among threatened Acropora corals

Genomic sequence data for non-model organisms are increasingly available requiring the development of efficient and reproducible workflows. Here, we develop the first genomic resources and reproducible workflows for two threatened members of the reef-building coral genus Acropora. We generated genomic sequence data from multiple samples of the Caribbean A. cervicornis (staghorn coral) and A. palmata (elkhorn coral), and predicted millions of nucleotide variants among these two species and the Pacific A. digitifera. A subset of predicted nucleotide variants were verified using restriction length polymorphism assays and proved useful in distinguishing the two Caribbean Acroporids and the hybrid they form (\"A. prolifera\"). Nucleotide variants are freely available from the Galaxy server (usegalaxy.org), and can be analyzed there with computational tools and stored workflows that require only an internet browser. We describe these data and some of the analysis tools, concentrating on fixed differences between A. cervicornis and A. palmata. In particular, we found that fixed amino acid differences between these two species were enriched in proteins associated with development, cellular stress response and the hosts interactions with associated microbes, for instance in the Wnt pathway, ABC transporters and superoxide dismutase. Identified candidate genes may underlie functional differences in the way these threatened species respond to changing environments. Users can expand the presented analyses easily by adding genomic data from additional species as they become available.\n\nArticle SummaryWe provide the first comprehensive genomic resources for two threatened Caribbean reef-building corals in the genus Acropora. We identified genetic differences in key pathways and genes known to be important in the animals response to the environmental disturbances and larval development. We further provide a list of candidate loci for large scale genotyping of these species to gather intra- and interspecies differences between A. cervicornis and A. palmata across their geographic range. All analyses and workflows are made available and can be used as a resource to not only analyze these corals but other non-model organisms.

genomics

Nuclear and mitochondrial genomic resources for the meltwater stonefly, Lednia tumana (Plecoptera: Nemouridae)

AbstractWith more than 3,700 described species, stoneflies (Order Plecoptera) are an important component of global aquatic biodiversity. The meltwater stonefly Lednia tumana (Family Nemouridae) is endemic to alpine streams of Glacier National Park and has been petitioned for listing under the U.S. Endangered Species Act (ESA) due to climate change-induced loss of alpine glaciers and snowfields. Here, we present de novo assemblies of the nuclear (~520 million base pairs [bp]) and mitochondrial (13,752 bp) genomes for L. tumana. The L. tumana nuclear genome is the most complete stonefly genome reported to date, with ~71% of genes present in complete form and > 4,600 contigs longer than 10 kilobases (kb). The L. tumana mitochondrial genome is the second for the family Nemouridae and the first from North America. Together, both genomes represent important foundational resources, setting the stage for future efforts to understand the evolution of L. tumana, stoneflies, and aquatic insects worldwide.

genomics

A Chromosome-Scale Assembly of the Enormous (32 Gb) Axolotl Genome

The axolotl (Ambystoma mexicanum) provides critical models for studying regeneration, evolution and development. However, its large genome (~32 gigabases) presents a formidable barrier to genetic analyses. Recent efforts have yielded genome assemblies consisting of thousands of unordered scaffolds that resolve gene structures, but do not yet permit large scale analyses of genome structure and function. We adapted an established mapping approach to leverage dense SNP typing information and for the first time assemble the axolotl genome into 14 chromosomes. Moreover, we used fluorescence in situ hybridization to verify the structure of these 14 scaffolds and assign each to its corresponding physical chromosome. This new assembly covers 27.3 gigabases and encompasses 94% of annotated gene models on chromosomal scaffolds. We show the assemblys utility by resolving genome-wide orthologies between the axolotl and other vertebrates, identifying the footprints of historical introgression events that occurred during the development of axolotl genetic stocks, and precisely mapping several phenotypes including a large deletion underlying the cardiac mutant. This chromosome-scale assembly will greatly facilitate studies of the axolotl in biological research.

genomics

Repeats Of Unusual Size in Plant Mitochondrial Genomes: Identification, Incidence and Evolution

Plant mitochondrial genomes have excessive size relative to coding capacity, a low mutation rate in genes and a high rearrangement rate. They also have non-tandem repeats in two size groups: a few large repeats which cause isomerization of the genome by recombination, and numerous repeats longer than 50bp, often found in exactly two copies per genome. It appears that repeats in the size range from several hundred to a few thousand base pair are underrepresented. The repeats are not well-conserved between species, and are infrequently annotated in mitochondrial sequence assemblies. Because they are much larger than expected by chance we call them Repeats Of Unusual Size (ROUS). The repeats consist of two functional classes, those that are involved in genome isomerization through frequent crossing over, and those for which crossovers are rare unless there are mutations in DNA repair genes, or the rate of double-strand breakage is increased. We systematically described and compared these repeats, which are important clues to mechanisms of DNA maintenance in mitochondria. We developed a tool to find non-tandem repeats and analyzed the complete mitochondrial sequences from 135 plant species. We observed an interesting difference between taxa: the repeats are larger and more frequent in the vascular plants. Analysis of closely related species also shows that plant mitochondrial genomes evolve in dramatic bursts of breakage and rejoining, complete with DNA sequence gain and loss, and the repeats are included in these events. We suggest an adaptive explanation for the existence of the repeats and their evolution.

genomics

New de novo assembly of the Atlantic bottlenose dolphin (Tursiops truncatus) improves genome completeness and provides haplotype phasing.

High quality genomes are essential to resolve challenges in breeding, comparative biology, medicine and conservation planning. New library preparation techniques along with better assembly algorithms result in continued improvements in assemblies for non-model organisms, moving them toward reference quality genomes. We report on the latest genome assembly of the Atlantic bottlenose dolphin leveraging Illumina sequencing data coupled with a combination of several library preparation techniques. These include Linked-Reads (Chromium, 10x Genomics), mate pairs, long insert paired ends and standard paired ends. Data were assembled with the commercial DeNovoMAGICTM assembly software resulting in two assemblies, a traditional \"haploid\" assembly (Tur_tru_Illumina_hap_v1) that is a mosaic of the two parental haplotypes and a phased assembly (Tur_tru_Illumina_phased_v1) where each scaffold has sequence from a single homologous chromosome. We show that Tur_tru_Illumina_hap_v1 is more complete and accurate compared to the current best reference based on the amount and composition of sequence, the consistency of the mate pair alignments to the assembled scaffolds, and on the analysis of conserved single-copy mammalian orthologs. The phased de novo assembly Tur_tru_Illumina_phased_v1 is the first publicly available for this species and provides the community with novel and accurate ways to explore the heterozygous nature of the dolphin genome.

genomics

The Genomic Basis of Arthropod Diversity

BackgroundArthropods comprise the largest and most diverse phylum on Earth and play vital roles in nearly every ecosystem. Their diversity stems in part from variations on a conserved body plan, resulting from and recorded in adaptive changes in the genome. Dissection of the genomic record of sequence change enables broad questions regarding genome evolution to be addressed, even across hyper-diverse taxa within arthropods.\n\nResultsUsing 76 whole genome sequences representing 21 orders spanning more than 500 million years of arthropod evolution, we document changes in gene and protein domain content and provide temporal and phylogenetic context for interpreting these innovations. We identify many novel gene families that arose early in the evolution of arthropods and during the diversification of insects into modern orders. We reveal unexpected variation in patterns of DNA methylation across arthropods and examples of gene family and protein domain evolution coincident with the appearance of notable phenotypic and physiological adaptations such as flight, metamorphosis, sociality and chemoperception.\n\nConclusionsThese analyses demonstrate how large-scale comparative genomics can provide broad new insights into the genotype to phenotype map and generate testable hypotheses about the evolution of animal diversity.

genomics