bioRxiv ScienceSearch

SEARCH · bioRxiv Science

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,495 records · Page 83Linked to original sources

Improved chromosome level genome assembly of the Glanville fritillary butterfly (Melitaea cinxia) based on SMRT Sequencing and linkage map.

The Glanville fritillary (Melitaea cinxia) butterfly is a long-term model system for metapopulation dynamics research in fragmented landscapes. Here, we provide a chromosome level assembly of the butterflys genome produced from Pacific Biosciences sequencing of a pool of males, combined with a linkage map from population crosses. The final assembly size of 484 Mb is an increase of 94 Mb on the previously published genome. Estimation of the completeness of the genome with BUSCO, indicates that the genome contains 93 - 95% of the BUSCO genes in complete and single copies. We predicted 14,830 gene models using the MAKER pipeline and manually curated 1,232 of these gene models. The genome and its annotated gene models are a valuable resource for future comparative genomics, molecular biology, transcriptome and genetics studies on this species.

genomics

Analysis of Genomes of Bacterial Isolates from Lameness Outbreaks in Broilers

We investigated lameness outbreaks at commercial broiler farms in Arkansas. From Bacterial Chondronecrosis with Osteomyelitis (BCO) lesions, we obtained different isolates of distinct bacterial species. Genome assemblies for Escherichia coli and Staphylococcus aureus isolates show that BCO-lameness pathogens on farms can differ significantly. Genomes assembled from Escherichia coli isolates from three different farms were quite different from each other, and more similar to isolates from different hosts and geographical locations. The S aureus genomes were closely related to chicken isolates from Europe, and appear to have been restricted to chicken hosts for more than 40 years. Detailed analyses of genomes from this clade of chicken isolates with a sister clade of human isolates, suggests the acquisition of a particular pathogenicity island in the transition from human to chicken pathogen and that pathogenesis in chickens may depend on this mobile element. Phylogenomics is consistent with more frequent host shifts for E. coli, while S. aureus appears to be highly host restricted. Isolate-specific genome characterizations will help further our understanding of the disease mechanisms and spread of BCO-lameness, a significant animal welfare issue. ImportanceDetailed inspection of the genome sequences of different bacterial species associated with causing lameness in broiler chickens reveals that one species, E. coli, appears to easily switch hosts from humans to chickens and other host species. Conversely, isolates of S. aureus appear to be restricted to specific hosts. One potential mobile DNA element has been identified that may be critical for causing disease in chickens for S. aureus.

genomics

Lost in translation: the pitfalls of Ensembl Gene annotations between human genome assemblies and their impact on diagnostics

BackgroundThe GRCh37 human genome assembly is still widely used in genomics despite the fact an updated human genome assembly (GRCh38) has been available for many years. A particular issue with relevant ramifications for clinical genetics currently is the case of the GRCh37 Ensembl gene annotations which has been archived, and thus not updated, since 2013. These Ensembl GRCh37 gene annotations are just as ubiquitous as the former assembly and are the default gene models used and preferred by the majority of genomic projects internationally. In this study, we highlight the issue of genes with discrepant annotations, that have been recognized as protein coding in the new but not the old assembly. These genes are ignored by all genomic resources that still rely on the archived and outdated gene annotations. Moreover, the majority if not all of these discrepant genes (DGs) are automatically discarded and ignored by all variant prioritization tools that rely on the GRCh37 Ensembl gene annotations. MethodsWe performed bioinformatics analysis identifying Ensembl genes with discrepant annotations between the two most recent human genome assemblies, hg37, hg38, respectively. Clinical and phenotype gene curations have been obtained and compared for this gene set. Furthermore, matching RefSeq transcripts have also been collated and analyzed. ResultsWe found hundreds of genes (N=267) that were reclassified as "protein-coding" in the new hg38 assembly. Notably, 169 of these genes also had a discrepant HGNC gene symbol between the two assemblies. Most genes had RefSeq matches (N=199/267) including all the genes with defined phenotypes in Ensembl genes GRCh38 assembly (N=10). However, many protein-coding genes remain missing from the current known RefSeq gene models (N=68) ConclusionWe found many clinically relevant genes in this group of neglected genes and we anticipate that many more will be found relevant in the future. For these genes, the inaccurate label of "non-protein-coding" hinders the possibility of identifying any causal sequence variants that overlap them. In addition, Important additional annotations such as evolutionary constraint metrics are also not calculated for these genes for the same reason, further relegating them into oblivion.

genomics

Rapid evolution and horizontal gene transfer in the genome of a male-killing Wolbachia

Wolbachia are widespread bacterial endosymbionts that infect a large proportion of insect species. While some strains of this bacteria do not cause observable host phenotypes, many strains of Wolbachia have some striking effects on their hosts. In some cases, these symbionts manipulate host reproduction to increase the fitness of infected, transmitting females. Here we examine the genome and population genomics of a male-killing Wolbachia strain, wInn, that infects Drosophila innubila mushroom-feeding flies. We compared wInn to other closely-related Wolbachia genomes to understand the evolutionary dynamics of specific genes. The wInn genome is similar in overall gene content to wMel, but also contains many unique genes and repetitive elements that indicate distinct gene transfers between wInn and non-Drosophila hosts. We also find that genes in the Wolbachia prophage and Octomom regions are particularly rapidly evolving, including those putatively or empirically confirmed to be involved in host pathogenicity. Of the genes that rapidly evolve, many also show evidence of recent horizontal transfer among Wolbachia symbiont genomes, suggesting frequent movement of rapidly evolving regions among individuals. These dynamics of rapid evolution and horizontal gene transfer across the genomes of several Wolbachia strains and divergent host species may be important underlying factors in Wolbachias global success as a symbiont.

genomics

Accurate Genomic Variant Detection in Single Cells with Primary Template-Directed Amplification

Improvements in whole genome amplification (WGA) would enable new types of basic and applied biomedical research, including studies of intratissue genetic diversity that require more accurate single-cell genotyping. Here we present primary template-directed amplification (PTA), a new isothermal WGA method that reproducibly captures >95% of the genomes of single cells in a more uniform and accurate manner than existing approaches, resulting in significantly improved variant calling sensitivity and precision. To illustrate the new types of studies that are enabled by PTA, we developed direct measurement of environmental mutagenicity (DMEM), a new tool for mapping genome-wide interactions of mutagens with single living human cells at base pair resolution. In addition, we utilized PTA for genome-wide off-target indel and structural variant detection in cells that had undergone CRISPR-mediated genome editing, establishing the feasibility for performing single-cell evaluations of biopsies from edited tissues. The improved precision and accuracy of variant detection with PTA overcomes the current limitations of accurate whole genome amplification, which is the major obstacle to studying genetic diversity and evolution at cellular resolution.

genomics

A Draft Reference Genome for Hirudo verbana, the Medicinal Leech

The medicinal leech, Hirudo verbana, is a powerful model organism for investigating fundamental neurobehavioral processes. The well-documented arrangement and properties of H. verbanas nervous system allows changes at the level of specific neurons or synapses to be linked to physiological and behavioral phenomena. Juxtaposed to the extensive knowledge of H. verbanas nervous system is a limited, but recently expanding, portfolio of molecular and multi-omics tools. Together, the advancement of genetic databases for H. verbana will complement existing pharmacological and electrophysiological data by affording targeted manipulation and analysis of gene expression in neural pathways of interest. Here, we present the first draft genome assembly for H. verbana, which is approximately 250 Mbp in size and consists of 61,282 contigs. Whole genome sequencing was conducted using an Illumina sequencing platform followed by genome assembly with CLC-Bio Genomics Workbench and subsequent functional annotation. Ultimately, the diversity of organisms for which we have genomic information should parallel the availability of next generation sequencing technologies to widen the comparative approach to understand the involvement and discovery of genes in evolutionarily conserved processes. Results of this work hope to facilitate comparative studies with H. verbana and provide the foundation for future, more complete, genome assemblies of the leech.

genomics

Manual Annotation of Genes within Drosophila Species: the Genomics Education Partnership protocol

Annotating the genomes of multiple organisms allows us to study their genes as well as the evolution of those genes. While many eukaryotic genome assemblies already include computational gene predictions, these predictions can benefit from review and refinement through manual gene annotation. The Genomics Education Partnership (GEP; thegep.org) has developed an annotation protocol for protein-coding genes that enables undergraduate students and other researchers to create high-quality gene annotations that can be utilized in subsequent scientific investigations. For example, this protocol has been utilized by the GEP faculty to engage undergraduate students in the comparative annotation of genes involved in the insulin signaling pathway in 28 Drosophila species, using D. melanogaster as the informant genome. Students construct gene models using multiple lines of computational and experimental evidence including expression data (e.g., RNA-Seq), sequence similarity (e.g., BLAST, multiple sequence alignments), and computational gene predictions. For quality control, each gene is annotated by at least two students working independently, followed by reconciliation of the submitted gene models by a more experienced student. This article provides an overview of the annotation protocol and describes how discrepancies in student submitted gene models are resolved to produce a final, high-quality gene set suitable for subsequent analyses. This annotation protocol can be adapted to other scientific questions (e.g., expansion of the Drosophila Muller F element) and other species (e.g., parasitoid wasps) to provide additional opportunities for undergraduate students to participate in genomics research. These student annotation efforts can substantially improve the quality of gene annotations in publicly available genomic databases.

genomics

Genomic surveillance framework and global population structure for Klebsiella pneumoniae

K. pneumoniae is a leading cause of antimicrobial-resistant (AMR) healthcare-associated infections, neonatal sepsis and community-acquired liver abscess, and is associated with chronic intestinal diseases. Its diversity and complex population structure pose challenges for analysis and interpretation of K. pneumoniae genome data. Here we introduce Kleborate, a tool for analysing genomes of K. pneumoniae and its associated species complex, which consolidates interrogation of key features of proven clinical importance. Kleborate provides a framework to support genomic surveillance and epidemiology in research, clinical and public health settings. To demonstrate its utility we apply Kleborate to analyse publicly available Klebsiella genomes, including clinical isolates from a pan-European study of carbapenemase-producing Klebsiella, highlighting global trends in AMR and virulence as examples of what could be achieved by applying this genomic framework within more systematic genomic surveillance efforts. We also demonstrate the application of Kleborate to detect and type K. pneumoniae from gut metagenomes.

genomics

Chromosome-scale genome assembly of Japanese pear (Pyrus pyrifolia) variety 'Nijisseiki

AimThe Japanese pear (P. pyrifolia) variety Nijisseiki is valued for its superior flesh texture, which has led to its use as a breeding parent for most Japanese pear cultivars. However, in the absence of genomic resources for Japanese pear, the parents of the Nijisseiki cultivar remain unknown, as does the genetic basis of its favorable texture. The genomes of pear and related species are complex due to ancestral whole genome duplication and high heterozygosity, and long-sequencing technology was used to address this. Methods and ResultsDe novo assembly of long sequence reads covered 136x of the Japanese pear genome and generated 503.9 Mb contigs consisting of 114 sequences with an N50 value of 7.6{square}Mb. Contigs were assigned to Japanese pear genetic maps to establish 17 chromosome-scale sequences. In total, 44,876 protein-encoding genes were predicted, 84.3% of which were supported by predicted genes and transcriptome data from Japanese pear relatives. As expected, evidence of whole genome duplication was observed, consistent with related species. Conclusion and PerspectiveThis is the first genome sequence analysis reported for Japanese pear, and this resource will support breeding programs and provide new insights into the physiology and evolutionary history of Japanese pear.

genomics

Genome-wide Copy Number Variations in a Large Cohort of Bantu African Children

BackgroundCopy number variations (CNVs) account for a substantial proportion of inter-individual genomic variation. However, a majority of genomic variation studies have focused on single-nucleotide variations (SNVs), with limited genome-wide analysis of CNVs in large cohorts, especially in populations that are under-represented in genetic studies including people of African descent. ResultsIn this study, we carried out a genome-wide analysis in > 3400 healthy Bantu Africans from Tanzania using high density (> 2.5 million probes) genotyping arrays. We identified over 400000 CNVs larger than 1 kilobase (kb), for an average of 120 CNVs (SE = 2.57) per individual. We detected 866 large CNVs ([≥] 300 kb), some of which overlapped genomic regions previously associated with multiple congenital anomaly syndromes, including Prader-Willi/Angelman syndrome (Type1) and 22q11.2 deletion syndrome. Furthermore, several of the common CNVs seen in our cohort ([≥] 5%) overlap genes previously associated with developmental disorders. ConclusionThese findings may help refine the phenotypic outcomes and penetrance of variations affecting genes and genomic regions previously implicated in diseases. Our study provides one of the largest datasets of CNVs from individuals of African ancestry, enabling improved clinical evaluation and disease association of CNVs observed in research and clinical studies in African populations.

genomics

Mutation rate variations in the human genome are encoded in DNA shape

Single nucleotide mutation rates have critical implications for human evolution and genetic diseases. Accurate modeling of these mutation rates has long remained an open problem since the rates vary substantially across the human genome. A recent model, however, explained much of the variation by considering higher order nucleotide interactions in the local (7-mer) sequence context around mutated nucleotides. Despite this models predictive value, we still lack a biophysically-grounded understanding of genome-wide mutation rate variations. DNA shape features are geometric measurements of DNA structural properties, such as helical twist and tilt, and are known to capture information on interactions between neighboring nucleotides within a local context. Motivated by this characteristic of DNA shape features, we used them to model mutation rates in the human genome. The DNA shape feature based models show up to 15% higher accuracy than the current nucleotide sequence-based models and pinpoint DNA structural properties predictive of mutation rates in the human genome. Further analyzing the mutation rates of individual positions of transcription factor (TF) binding sites in the human genome, we found a strong association between DNA shape and the position-specific mutation rates. The trend holds for hundreds of TFs and is even stronger in evolutionarily conserved regions. To our knowledge, this is the first attempt that demonstrates the structural underpinnings of nucleotide mutations in the human genome and lays the groundwork for future studies to incorporate DNA shape information in modeling genetic variations.

genomics

Improvements in the Sequencing and Assembly of Plant Genomes

BackgroundAdvances in DNA sequencing have reduced the difficulty of sequencing and assembling plant genomes. A range of methods for long read sequencing and assembly have been recently compared and we now extend the earlier study and report a comparison with more recent methods. ResultsUpdated Oxford Nanopore Technology software supported improved assemblies. The use of more accurate sequences produced by repeated sequencing of the same molecule (PacBio HiFi) resulted in much less fragmented assembly of sequencing reads. The use of more data to give increased genome coverage resulted in longer contigs (higher N50) but reduced the total length of the assemblies and improved genome completeness (BUSCO). The original model species, Macadamia jansenii, a basal eudicot, was also compared with the 3 other Macadamia species and with avocado (Persea americana), a magnoliid, and jojoba (Simmondsia chinensis) a core eudicot. In these phylogenetically diverse angiosperms, increasing sequence data volumes also caused a highly linear increase in contig size, decreased assembly length and further improved already high completeness. Differences in genome size and sequence complexity apparently influenced the success of assembly from these different species. ConclusionsAdvances in long read sequencing technology have continued to significantly improve the results of sequencing and assembly of plant genomes. However, results were consistently improved by greater genome coverage (using an increased number of reads) with the amount needed to achieve a particular level of assembly being species dependant.

genomics

Drosophila Evolution over Space and Time (DEST) - A New Population Genomics Resource

Drosophila melanogaster is a leading model in population genetics and genomics, and a growing number of whole-genome datasets from natural populations of this species have been published over the last 20 years. A major challenge is the integration of these disparate datasets, often generated using different sequencing technologies and bioinformatic pipelines, which hampers our ability to address questions about the evolution and population structure of this species. Here we address these issues by developing a bioinformatics pipeline that maps pooled sequencing (Pool-Seq) reads from D. melanogaster to a hologenome consisting of fly and symbiont genomes and estimates allele frequencies using either a heuristic (PoolSNP) or a probabilistic variant caller (SNAPE-pooled). We use this pipeline to generate the largest data repository of genomic data available for D. melanogaster to date, encompassing 271 population samples from over 100 locations in >20 countries on four continents based on a combination of 121 unpublished and 150 previously published genomic datasets. Several of these locations have been sampled at different seasons across multiple years. This dataset, which we call Drosophila Evolution over Space and Time (DEST), is coupled with sampling and environmental meta-data. A web-based genome browser and web portal provide easy access to the SNP dataset. Our aim is to provide this scalable platform as a community resource which can be easily extended via future efforts for an even more extensive cosmopolitan dataset. Our resource will enable population geneticists to analyze spatio-temporal genetic patterns and evolutionary dynamics of D. melanogaster populations in unprecedented detail.

genomics

Copy number variation profile-based genomic subtyping of premenstrual dysphoric disorder in Chinese

Premenstrual dysphoric disorder (PMDD) affects nearly 5% women of reproductive age. The symptomatic heterogeneity, along with largely unknown genetics, of PMDD have greatly hindered its effective treatment. In the present study, 127 Chinese PMDD patients of the invasion and depression subtypes clinically differentiated by us earlier were analyzed together with 108 non-PMDD controls for genome-wide copy number variations (CNVs). Germline genomic DNA samples from white blood cells were subjected to AluScan sequencing-based CNV profiling, which enabled clustering of patient samples readily into the V and D groups, dominated by the "invasion" and "depression" clinical subtypes, respectively; the CNVs obtained with 100-kb windows yielded two clusters that were correlated with these subtypes with a consistency of up to 89.8%. Diagnostic correlation- and frequency-based CNV features of either CNV-gain (CNVG) or CNV-loss (CNVL) that could differentiate between V and D subtypes were selected and analyzed. CNVG features located preferentially in S2-phase replicating regions and enriched with steroid hormone biosynthesis pathway of genes were found protective against PMDD. Moreover, machine learning employing the correlation-based CNV features could predict with >80% accuracy whether a genomic sample was D-type, V-type or control. In terms of their CNV profiles, the D- and V-types differed more from one another than from the controls, thereby providing a genomic basis for the clinical D-V subtyping of PMDD. Genome-wide profiling of CNVs, as a new approach to complex disease genetics, has revealed recurrent CNVs and genomic features beyond individual genes and mutations underlying PMDD clinical diversity.

genomics

A high-quality Genome and Comparison of Short versus Long Read Transcriptome of the Palaearctic duck Aythya fuligula (Tufted Duck)

BackgroundThe tufted duck is a non-model organism that suffers high mortality in highly pathogenic avian influenza out-breaks. It belongs to the same bird family (Anatidae) as the mallard, one of the best-studied natural hosts of low-pathogenic avian influenza viruses. Studies in non-model bird species are crucial to disentangle the role of the host response in avian influenza virus infection in the natural reservoir. Such endeavour requires a high-quality genome assembly and transcriptome. ResultsThis study presents the first high-quality, chromosome-level reference genome assembly of the tufted duck using the Vertebrate Genomes Project pipeline. We sequenced RNA (cDNA) from brain, ileum, lung, ovary, spleen and testis using Illumina short-read and PacBio long-read sequencing platforms, which was used for annotation. We found 34 autosomes plus Z and W sex chromosomes in the curated genome assembly, with 99.6% of the sequence assigned to chromosomes. Functional annotation revealed 14,099 protein-coding genes that generate 111,934 transcripts, which implies an average of 7.9 isoforms per gene. We also identified 246 small RNA families. ConclusionsThis annotated genome contributes to continuing research into the host response in avian influenza virus infections in a natural reservoir. Our findings from a comparison between short-read and long-read reference transcriptomics contribute to a deeper understanding of these competing options. In this study, both technologies complemented each other. We expect this annotation to be a foundation for further comparative and evolutionary genomic studies, including many waterfowl relatives with differing susceptibilities to the avian influenza virus.

genomics

The evolutionary history of a gammaretrovirus currently colonizing the mule deer genome is marked by extensive recombination

All vertebrate genomes have been colonized by retroviruses along their evolutionary trajectory. While endogenous retroviruses (ERVs) can contribute important physiological functions to contemporary hosts, such benefits are attributed to long-term co-evolution of ERV and host because germline infections are rare and expansion is slow, because the host effectively silences them. The genomes of several outbred species including mule deer (Odocoileus hemionus) are currently being colonized by ERVs, which provides an opportunity to study ERV dynamics at a time when few are fixed. Because we have locus-specific data on the distribution of cervid endogenous retrovirus (CrERV) in populations of mule deer, in this study we determine the molecular evolutionary processes acting on CrERV at each locus in the context of phylogenetic origin, genome location, and population prevalence. A mule deer genome was de novo assembled from short and long insert mate pair reads and CrERV sequence generated at each locus. CrERV composition and diversity have recently measurably increased by horizontal acquisition of a new retrovirus lineage. This new lineage has further expanded CrERV burden and CrERV genomic diversity by activating and recombining with existing CrERV. Resulting inter-lineage recombinants endogenized and subsequently retrotransposed. CrERV loci are significantly closer to genes than expected if integration were random and gene proximity might explain the recent expansion by retrotransposition of one recombinant CrERV lineage. Thus, in mule deer, retroviral colonization is a dynamic period in the molecular evolution of CrERV that also provides a burst of genomic diversity to the host population.

genomics

Insights from the first genome assembly of Onion (Allium cepa)

Onion is an important vegetable crop with an estimated genome size of 16Gb. We describe the de novo assembly and ab initio annotation of the genome of a doubled haploid onion line DHCU066619, which resulted in a final assembly of 14.9 Gb with a N50 of 461 Kb. Of this, 2.2 Gb was ordered into 8 pseudomolecules using five genetic linkage maps. The remainder of the genome is available in 89.8 K scaffolds. Only 72.4% of the genome could be identified as repetitive sequences and consist, to a large extent, of (retro) transposons. In addition, an estimated 20% of the putative (retro) transposons had accumulated a large number of mutations, hampering their identification, but facilitating their assembly. These elements are probably already quite old. The ab initio gene prediction indicated 540,925 putative gene models, which is far more than expected, possibly due to the presence of pseudogenes. Of these models, 86,073 showed similarity to published proteins (UNIPROT). No gene rich regions were found, genes are uniformly distributed over the genome. Analysis of synteny with A. sativum (garlic) showed collinearity but also major rearrangements between both species. This assembly is the first high-quality genome sequence available for the study of onion and will be a valuable resource for further research.

genomics

Complete genome sequence of Xylella taiwanensis and comparative analysis of virulence gene content with Xylella fastidiosa

The bacterial genus Xylella contains plant pathogens that are major threats to agriculture in America and Europe. Although extensive research was conducted to characterize different subspecies of Xylella fastidiosa (Xf), comparative analysis at above-species levels were lacking due to the unavailability of appropriate data sets. Recently, a bacterium that causes pear leaf scorch (PLS) in Taiwan was described as the second Xylella species (i.e., Xylella taiwanensis; Xt). In this work, we report the complete genome sequence of Xt type strain PLS229T. The genome-scale phylogeny provided strong support that Xf subspecies pauca (Xfp) is the basal lineage of this species and Xylella was derived from the paraphyletic genus Xanthomonas. Quantification of genomic divergence indicated that different Xf subspecies share [~]87-95% of their chromosomal segments, while the two Xylella species share only [~]66-70%. Analysis of overall gene content suggested that Xt is most similar to Xf subspecies sandyi (Xfs). Based on the existing knowledge of Xf virulence genes, the homolog distribution among 28 Xylella representatives was examined. Among the 11 functional categories, those involved in secretion and metabolism are the most conserved ones with no copy number variation. In contrast, several genes related to adhesins, hydrolytic enzymes, and toxin-antitoxin systems are highly variable in their copy numbers. Those virulence genes with high levels of conservation or variation may be promising candidates for future studies. In summary, the new genome sequence and analysis reported in this work contributed to the study of several important pathogens in the family Xanthomonadaceae. Contribution to the FieldXylella fastidiosa is a plant-pathogenic bacterium with multiple subspecies that are major threats to agriculture in America and Europe. Although extensive research has been conducted, comparative analysis of this species with other bacteria is lacking due to the unavailability of known close relatives. In this work, we report the complete genome sequence of Xylella taiwanensis, a newly described species within the same genus. This new data set and our focused analysis helped to better understand the evolutionary relationships among different Xylella lineages and their genomic diversity. Moreover, detailed examination of their virulence genes identified those that are either highly conserved or variable, providing promising candidates for future studies to further investigate the molecular mechanisms of Xylella virulence.

genomics