bioRxiv ScienceSearch

SEARCH · bioRxiv Science

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

Direct estimate of the spontaneous mutation rate uncovers the effects of drift and recombination in the Chlamydomonas reinhardtii plastid genome

Plastids perform crucial cellular functions, including photosynthesis, across a wide variety of eukaryotes. Since endosymbiosis, plastids have maintained independent genomes that now display a wide diversity of gene content, genome structure, gene regulation mechanisms, and transmission modes. The evolution of plastid genomes depends on an input of de novo mutation, but our knowledge of mutation in the plastid is limited to indirect inference from patterns of DNA divergence between species. Here, we use a mutation accumulation experiment, where selection acting on mutations is rendered ineffective, combined with whole-plastid genome sequencing to directly characterize de novo mutation in Chlamydomonas reinhardtii. We show that the mutation rates of the plastid and nuclear genomes are similar, but that the base spectra of mutations differ significantly. We integrate our measure of the mutation rate with a population genomic dataset of 20 individuals, and show that the plastid genome is subject to substantially stronger genetic drift than the nuclear genome. We also show that high levels of linkage disequilibrium in the plastid genome are not due to restricted recombination, but are instead a consequence of increased genetic drift. One likely explanation for increased drift in the plastid genome is that there are stronger effects of genetic hitchhiking. The presence of recombination in the plastid is consistent with laboratory studies in C. reinhardtii and demonstrates that although the plastid genome is thought to be uniparentally inherited, it recombines in nature at a rate similar to the nuclear genome.

Genomics

FGMP: assessing fungal genome completeness and gene content

BackgroundInexpensive high-throughput DNA sequencing has democratized access to genetic information for most organisms so that research utilizing a genome or transcriptome of an organism is not limited to model systems. However, the quality of the assemblies of sampled genomes can vary greatly which hampers utility for comparisons and meaningful interpretation. The uncertainty of the completeness of a given genome sequence can limit feasibility of asserting patterns of high rates of gene loss reported in many lineages.\n\nResultsWe propose a computational framework and sequence resource for assessing completeness of fungal genomes called FGMP (Fungal Genome Mapping Project). Our approach is based on evolutionary conserved sets of proteins and DNA elements and is applicable to various types of genomic data. We present a comparison of FGMP and state-of-the-art methods for genome completeness assessment utilizing 246 genome assemblies of fungi. We discuss genome assembly improvements/degradations in 57 cases where assemblies have been updated, as recorded by NCBI assembly archive.\n\nConclusionFGMP is an accurate tool for quantifying level of completion from fungal genomic data. It is particularly useful for non-model organisms without reference genomes and can be used directly on unassembled reads, which can help reducing genome sequencing costs.

Genomics

Ultrasensitive capture of human herpes simplex virus genomes directly from clinical samples reveals extraordinarily limited evolution in cell culture

Herpes simplex viruses (HSV) are difficult to sequence due to their large DNA genome, high GC content, and the presence of repeats. To date, most HSV genomes have been recovered from culture isolates, raising concern that these genomes may not accurately represent circulating clinical strains. We report the development and validation of a DNA oligonucleotide hybridization panel to recover near complete HSV genomes at abundances up to 50,000-fold lower than previously reported. Using copy number information on herpesvirus and host DNA background via quantitative PCR, we developed a protocol for pooling for cost-effective recovery of more than 50 HSV-1 or HSV-2 genomes per MiSeq run. We demonstrate the ability to recover >99% of the HSV genome at >100X coverage in 72 hours at viral loads that allow whole genome recovery from latently-infected ganglia. We also report a new computational pipeline for rapid HSV genome assembly and annotation. Using the above tools and a series of 17 HSV-1-positive clinical swabs sent to our laboratory for viral isolation, we show limited evolution of HSV-1 during viral isolation in human fibroblast cells compared to the original clinical samples. Our data indicate that previous studies using low passage clinical isolates of herpes simplex viruses are reflective of the viral sequences present in the lesion and thus can be used in phylogenetic analyses. We also detect superinfection within a single sample with unrelated HSV-1 strains recovered from separate oral lesions in an immunosuppressed patient during a 2.5-week period, illustrating the power of direct-from-specimen sequencing of HSV.\n\nImportanceHerpes simplex viruses affect more than 4 billion people across the globe, constituting a large burden of disease. Understanding global diversity of herpes simplex viruses is important for diagnostics and therapeutics as well as cure research and tracking transmission among humans. To date, most HSV genomics has been performed on culture isolates and DNA swabs with high quantities of virus. We describe the development of wet-lab and computational tools that enable the accurate sequencing of near-complete genomes of HSV-1 and HSV-2 directly from clinical specimens at abundances >50,000-fold lower than previously sequenced and at significantly reduced cost. We use these tools to profile circulating HSV-1 strains in the community and illustrate limited changes to the viral genome during the viral isolation process. These techniques enable cost-effective, rapid sequencing of HSV-1 and HSV-2 genomes that will help enable improved detection, surveillance, and control of this human pathogen.

genomics

A Hybrid de novo Assembly of the Sea Pansy (Renilla muelleri) Genome

BackgroundOver 3,000 species of octocorals (Cnidaria, Anthozoa) inhabit an expansive range of environments, from shallow tropical seas to the deep-ocean floor. They are important foundation species that create coral \"forests\" which provide unique niches and three-dimensional living space for other organisms. The octocoral genus Renilla inhabits sandy, continental shelves in the subtropical and tropical Atlantic and eastern Pacific Oceans. Renilla is especially interesting because it produces secondary metabolites for defense, exhibits bioluminescence, and produces a luciferase that is widely used in dual-reporter assays in molecular biology. Although several cnidarian genomes are currently available, the majority are from hexacorals. Here, we present a de novo assembly of the R. muelleri genome, making this the first complete draft genome from an octocoral.\n\nFindingsWe generated a hybrid de novo assembly using the Maryland Super-Read Celera Assembler v.3.2.6 (MaSuRCA). The final assembly included 4,825 scaffolds and a haploid genome size of 172 Mb. A BUSCO assessment found 88% of metazoan orthologs present in the genome. An Augustus ab initio gene prediction found 23,660 genes, of which 66% (15,635) had detectable similarity to annotated genes from the starlet sea anemone, Nematostella vectensis, or to the Uniprot database. Although the R. muelleri genome is smaller (172 Mb) than other publicly available, hexacoral genomes (256-448 Mb), the R. muelleri genome is similar to the hexacoral genomes in terms of the number of complete metazoan BUSCOs and predicted gene models.\n\nConclusionsThe R. muelleri hybrid genome provides a novel resource for researchers to investigate the evolution of genes and gene families within Octocorallia and more widely across Anthozoa. It will be a key resource for future comparative genomics with other corals and for understanding the genomic basis of coral diversity.

genomics

ISMapper: Identifying insertion sequences in bacterial genomes from short read sequence data

BackgroundInsertion sequences (IS) are small transposable elements, commonly found in bacterial genomes. Identifying the location of IS in bacterial genomes can be useful for a variety of purposes including epidemiological tracking and predicting antibiotic resistance. However IS are commonly present in multiple copies in a single genome, which complicates genome assembly and the identification of IS insertion sites. Here we present ISMapper, a mapping-based tool for identification of the site and orientation of IS insertions in bacterial genomes, direct from paired-end short read data.\n\nResultsISMapper was validated using three types of short read data: (i) simulated reads from a variety of species, (ii) Illumina reads from 5 isolates for which finished genome sequences were available for comparison, and (iii) Illumina reads from 7 Acinetobacter baumannii isolates for which predicted IS locations were tested using PCR. A total of 20 genomes, including 13 species and 32 distinct IS, were used for validation. ISMapper correctly identified 96% of known IS insertions in the analysis of simulated reads, and 98% in real Illumina reads. Subsampling of real Illumina reads to lower depths indicated ISMapper was reliable for average genome-wide read depths >20x. All ISAba1 insertions identified by ISMapper in the A. baumannii genomes were confirmed by PCR. In each A. baumannii genome, ISMapper successfully identified an IS insertion upstream of the ampC beta-lactamase that could explain phenotypic resistance to third-generation cephalosporins. The utility of ISMapper was further demonstrated by profiling genome-wide IS6110 insertions in 138 publicly available Mycobacterium tuberculosis genomes, revealing lineage-speific inserction and multi inserction hotspot.\n\nConclusionsISMapper provides a rapid and robust method for identifying IS insertion sites direct from short read data, with a high degree of accuracy demonstrated across a wide range of bacteria.

Bioinformatics

Genome ARTIST: a robust, high-accuracy aligner tool for mapping transposon insertions and self-insertions

A critical topic of insertional mutagenesis experiments performed on model organisms is mapping the hits of artificial transposons (ATs) at nucleotide level accuracy. Obviously, mapping errors may occur when sequencing artifacts or mutations as SNPs and small indels are present very close to the junction between a genomic sequence and a transposon inverted repeat (TIR). Another particular item of insertional mutagenesis is mapping of the transposon self-insertions and, to our best knowledge, there is no publicly available mapping tool designed to analyze such molecular events. We developed Genome ARTIST, a pairwise gapped aligner tool which works out both issues by means of an original, robust mapping strategy. Genome ARTIST is not designed to use NGS data but to analyze ATs insertions obtained in small to medium-scale mutagenesis experiments. Genome ARTIST employs a heuristic approach to find DNA sequence similarities and harnesses a multi-step implementation of a Smith-Waterman adapted algorithm to compute the mapping alignments. The experience is enhanced by easily customizable parameters and a user-friendly interface that describes the genomic landscape surrounding the insertion. Genome ARTIST deals with many genomes of bacteria and eukaryotes available in Ensembl and GenBank repositories. Our tool specifically harnesses/exploits the sequence annotation data provided by FlyBase for Drosophila melanogaster (the fruit fly), which enables mapping of insertions relative to various genomic features such as natural transposons. Genome ARTIST was tested against other alignment tools using relevant query sequences derived from the D. melanogaster and Mus musculus (mouse) genomes. Real and simulated query sequences were also comparatively inquired, revealing that Genome ARTIST is a very robust solution for mapping transposon insertions.\n\nGenome ARTIST is a stand-alone user-friendly application, designed for high-accuracy mapping of transposon insertions and self-insertions. The tool is also useful for routine aligning assessments like detection of SNPs or checking the specificity of primers and probes. Genome ARTIST is an open source software and is available for download at www.genomeartist.ro and at www.bioinformatics.org.

Bioinformatics

DNA from dust: comparative genomics of large DNA viruses in field surveillance samples

Mareks disease (MD) is a lymphoproliferative disease of chickens caused by airborne gallid herpesvirus type 2 (GaHV-2, aka MDV-1). Mature virions are formed in the feather follicle epithelium cells of infected chickens from which the virus is shed as fine particles of skin and feather debris, or poultry dust. Poultry dust is the major source of virus transmission between birds in agricultural settings. Despite both clinical and laboratory data that show increased virulence in field isolates of MDV-1 over the last 40 years, we do not yet understand the genetic basis of MDV-1 pathogenicity. Our present knowledge on genome-wide variation in the MDV-1 genome comes exclusively from laboratory-grown isolates. MDV-1 isolates tend to lose virulence with increasing passage number in vitro, raising concerns about their ability to accurately reflect virus in the field. The ability to rapidly and directly sequence field isolates of MDV-1 is critical to understanding the genetic basis of rising virulence in circulating wild strains. Here we present the first complete genomes of uncultured, field-isolated MDV-1. These five consensus genomes were derived directly from poultry dust or single chicken feather follicles without passage in cell culture. These sources represent the shed material that is transmitted to new hosts, vs. the virus produced by a point source in one animal. We developed a new procedure to extract and enrich viral DNA, while reducing host and environmental contamination. DNA was sequenced using Illumina MiSeq high-throughput approaches and processed through a recently described bioinformatics workflow for de novo assembly and curation of herpesvirus genomes. We comprehensively compared these genomes to one another and also to previously described MDV-1 genomes. The field-isolated genomes had remarkably high DNA identity when compared to one another, with few variant proteins between them. In an analysis of genetic distance, the five new field genomes grouped separately from all previously described genomes. Each consensus genome was also assessed to determine the level of polymorphisms within each sample, which revealed that MDV-1 exists in the wild as a polymorphic population. By tracking a new polymorphic locus in ICP4 over time, we found that MDV-1 genomes can evolve in short period of time. Together these approaches advance our ability to assess MDV-1 variation within and between hosts, over time, and during adaptation to changing conditions.

Microbiology

Ecophysiology of freshwater Verrucomicrobia inferred from genomes recovered through time-series metagenomics

Microbes are critical in carbon and nutrient cycling in freshwater ecosystems. Members of the Verrucomicrobia are ubiquitous in such systems, yet their roles and ecophysiology are not well understood. In this study, we recovered 19 Verrucomicrobia draft genomes by sequencing 184 time-series metagenomes from a eutrophic lake and a humic bog that differ in carbon source and nutrient availabilities. These genomes span four of the seven previously defined Verrucomicrobia subdivisions, and greatly expand the known genomic diversity of freshwater Verrucomicrobia. Genome analysis revealed their potential role as (poly)saccharide-degraders in freshwater, uncovered interesting genomic features for this life style, and suggested their adaptation to nutrient availabilities in their environments. Between the two lakes, Verrucomicrobia populations differ significantly in glycoside hydrolase gene abundance and functional profiles, reflecting the autochthonous and terrestrially-derived allochthonous carbon sources of the two ecosystems respectively. Interestingly, a number of genomes recovered from the bog contained gene clusters that potentially encode a novel porin-multiheme cytochrome c complex and might be involved in extracellular electron transfer in the anoxic humic-rich environment. Notably, most epilimnion genomes have large numbers of so-called \"Planctomycete-specific\" cytochrome c-containing genes, which exhibited nearly opposite distribution patterns with glycoside hydrolase genes, probably associated with the different environmental oxygen availability and carbohydrate complexity between lakes/layers. Overall, the recovered genomes are a major step towards understanding the role, ecophysiology and distribution of Verrucomicrobia in freshwater.\n\nIMPORTANCEFreshwater Verrucomicrobia are cosmopolitan in lakes and rivers, yet their roles and ecophysiology are not well understood, as cultured freshwater Verrucomicrobia are restricted to one subdivision of this phylum. Here, we greatly expand the known genomic diversity of this freshwater lineage by recovering 19 Verrucomicrobia draft genomes from 184 metagenomes collected from a eutrophic lake and a humic bog across multiple years. Most of these genomes represent first freshwater representatives of several Verrucomicrobia subdivisions. Genomic analysis revealed Verrucomicrobia as potential (poly)saccharide-degraders, and suggested their adaptation to carbon source of different origins in the two contrasting ecosystems. We identified putative extracellular electron transfer genes and so-called \"Planctomycete-specific\" cytochrome c-containing genes, and found their distinct distribution patterns between the lakes/layers. Overall, our analysis greatly advances the understanding of the function, ecophysiology and distribution of freshwater Verrucomicrobia, while highlighting their potential role in freshwater carbon cycling.

microbiology

xenoGI: reconstructing the history of genomic island insertions in clades of closely related bacteria

BackgroundGenomic islands play an important role in microbial genome evolution, providing a mechanism for strains to adapt to new ecological conditions. A variety of computational methods, both genome-composition based and comparative have been developed to identify them. Some of these methods are explicitly designed to work in single strains, while others make use of multiple strains. In general, existing methods do not identify islands in the context of the phylogeny in which they evolved. Even multiple strain approaches are best suited to identifying genomic islands that are present in one strain but absent in others. They do not automatically recognize islands which are shared between some strains in the clade or determine the branch on which these islands inserted within the phylogenetic tree.\n\nResultsWe have developed a software package, xenoGI, that identifies genomic islands and maps their origin within a clade of closely related bacteria, determining which branch they inserted on. It takes as input a set of sequenced genomes and a tree specifying their phylogenetic relationships. Making heavy use of synteny information, the package builds gene families in a species-tree-aware way, and then attempts to combine into islands those families whose members are adjacent and whose most recent common ancestor is shared. The package provides a variety of text-based analysis functions, as well as the ability to export genomic islands into formats suitable for viewing in a genome browser. We demonstrate the capabilities of the package with several examples from enteric bacteria, including an examination of the evolution of the acid fitness island in the genus Escherichia. In addition we use output from simulations and a set of known genomic islands from the literature to show that xenoGI can accurately identify genomic islands and place them on a phylogenetic tree.\n\nConclusionsxenoGI is an effective tool for studying the history of genomic island insertions in a clade of microbes. It identifies genomic islands, and determines which branch they inserted on within the phylogenetic tree for the clade. Such information is valuable because it helps us understand the adaptive path that has produced living species. Given the large and growing number of sequenced microbial genomes, this sort of analysis will become increasingly useful in the future.

bioinformatics

RNAs as proximity labeling media for identifying nuclear speckle positions relative to the genome

Nuclear speckles are interchromatin structures enriched in RNA splicing factors. Determining their relative positions with respect to the folded nuclear genome could provide critical information on co-and post-transcriptional regulation of gene expression. However, it remains challenging to identify which parts of the nuclear genome are in proximity to nuclear speckles, due to physical separation between nuclear speckle cores and chromatin. We hypothesized that noncoding RNAs including small nuclear RNAs, 7SK and Malat1, which accumulate at the periphery of nuclear speckles (nsaRNA, nuclear speckle associated RNA), may extend to sufficient proximity to the nuclear genome. Leveraging a transcriptome-genome interaction assay (MARGI), we identified nsaRNA-interacting genomic sequences, which exhibited clustering patterns (nsaPeaks) in the genome, suggesting existence of relatively stable interaction sites for nsaRNAs in nuclear genome. Posttranscriptional pre-mRNAs, which are known to be clustered to nuclear speckles, exhibited proximity to nsaPeaks but rarely to other genomic regions. Furthermore, CDK9 proteins that localize to the vicinity of nuclear speckles produced ChIP-seq peaks that overlapped with nsaPeaks. Our combined DNA FISH and immunofluorescence analysis in 182 single cells revealed a 3-fold increase in odds for nuclear speckles to localize near an nsaPeak than its neighboring genomic sequence. These data suggest a model that nsaRNAs locate in sufficient proximity to nuclear genome and leave identifiable genomic footprints, thus revealing the parts of genome proximal to nuclear speckles.

bioinformatics

Homogenization of sub-genome secretome gene expression patterns in the allodiploid fungus Verticillium longisporum

Hybridization is an important evolutionary mechanism that can enable organisms to adapt to environmental challenges. It has previously been shown that the fungal allodiploid species Verticillium longisporum, causal agent of Verticillium stem striping in rape seed, has originated from at least three independent hybridization events between two haploid Verticillium species. To reveal the impact of genome duplication as a consequence of the hybridization, we studied the genome and transcriptome dynamics upon two independent V. longisporum hybridization events, represented by the hybrid lineages "A1/D1" and "A1/D3". We show that the V. longisporum genomes are characterized by extensive chromosomal rearrangements, including between parental chromosomal sets. V. longisporum hybrids display signs of evolutionary dynamics that are typically associated with the aftermath of allodiploidization, such as haploidization and a more relaxed gene evolution. Expression patterns of the two sub-genomes within the two hybrid lineages are more similar than those of the shared A1 parent between the two lineages, showing that expression patterns of the parental genomes homogenized within a lineage. However, as genes that display differential parental expression in planta do not typically display the same pattern in vitro, we conclude that sub-genome-specific responses occur in both lineages. Overall, our study uncovers the genomic and transcriptomic plasticity during evolution of the filamentous fungal hybrid V. longisporum and illustrate its adaptive potential. ImportanceVerticillium is a genus of plant-associated fungi that include a handful of plant pathogens that collectively affect a wide range of hosts. On several occasions, haploid Verticillium species hybridized into the stable allodiploid species Verticillium longisporum, which is, in contrast to haploid Verticillium species, a Brassicaceae specialist. Here, we studied the evolutionary genome and transcriptome dynamics of V. longisporum and the impact of the hybridization. V. longisporum genomes display a mosaic structure due do genomic rearrangements between the parental chromosome sets. Similar to other allopolyploid hybrids, V. longisporum displays an ongoing loss of heterozygosity and a more relaxed gene evolution. Also, differential parental gene expression is observed, with an enrichment for genes that encode secreted proteins. Intriguingly, the majority of these genes displays sub-genome-specific responses under differential growth conditions. In conclusion, hybridization has incited the genomic and transcriptomic plasticity that enables adaptation to environmental changes in a parental allele-specific fashion.

microbiology

The haplotype-resolved genome sequence of hexaploid Ipomoea batatas reveals its evolutionary history

Although the sweet potato, Ipomoea batatas, is the seventh most important crop in the world and the fourth most significant in China, its genome has not yet been sequenced. The reason, at least in part, is that the genome has proven very difficult to assemble, being hexaploid and highly polymorphic; it has a presumptive composition of two B1 and four B2 component genomes (B1B1B2B2B2B2). By using a novel haplotyping method based on de novo genome assembly, however, we have produced a half haplotype-resolved genome from [~]267Gb of paired-end sequence reads amounting to roughly 60-fold coverage. By phylogenetic tree analysis of homologous chromosomes, it was possible to estimate the time of two whole genome duplication events as occurring about 525,000 and 341,000 years ago. Our analysis also identified many clusters of genes for specialized compounds biosynthesis in this genome. This half haplotype-resolved hexaploid genome represents the first successful attempt to investigate the complexity of chromosome sequence composition directly in a polyploid genome, using direct sequencing of the polyploid organism itself rather than of any of its simplified proxy relatives. Adaptation and application of our approach should provide higher resolution in future genomic structure investigations, especially for similarly complex genomes.

Genomics

A Bacillus anthracis Genome Sequence from the Sverdlovsk 1979 Autopsy Specimens

Anthrax is a zoonotic disease that occurs naturally in wild and domestic animals but has been used by both state-sponsored programs and terrorists as a biological weapon. The 2001 anthrax letter attacks involved less than gram quantities of Bacillus anthracis spores while the earlier Soviet weapons program produced tons. A Soviet industrial production facility in Sverdlovsk proved deficient in 1979 when a plume of spores was accidentally released and resulted in one of the largest known human anthrax outbreak. In order to understand this outbreak and others, we have generated a B. anthracis population genetic database based upon whole genome analysis to identify all SNPs across a reference genome. Only ~12,000 SNPs were identified in this low diversity species and represents the breadth of its known global diversity. Phylogenetic analysis has defined three major clades (A, B and C) with B and C being relatively rare compared to A. The A clade has numerous subclades including a major polytomy named the Trans-Eurasian (TEA) group. The TEA radiation is a dominant evolutionary feature of B. anthracis, many contemporary populations, and must have resulted from large-scale dispersal of spores from a single source. Two autopsy specimens from the Sverdlovsk outbreak were deeply sequenced to produce draft B. anthracis genomes. This allowed the phylogenetic placement of the Sverdlovsk strain into a clade with two Asian live vaccine strains, including the Russian Tsiankovskii strain. The genome was examined for evidence of drug resistance manipulation or other genetic engineering, but none was found. Only 13 SNPs differentiated the virulent Sverdlovsk strain from its common ancestor with two vaccine strains. The Soviet Sverdlovsk strain genome is consistent with a wild type strain from Russia that had no evidence of genetic manipulation during its industrial production. This work provides insights into the world's largest biological weapons program and provides an extensive B. anthracis phylogenetic reference valuable for future anthrax investigations.\n\nImportanceThe 1979 Russian anthrax outbreak resulted from an industrial accident at the Soviet anthrax spore production facility in the city of Sverdlovsk. Deep genomic sequencing of two autopsy specimens generated a draft genome and phylogenetic placement of the Soviet Sverdlovsk anthrax strain. While it is known that Soviet scientists had genetically manipulated Bacillus anthracis, with the potential to evade vaccine prophylaxis and antibiotic therapeutics, there was no genomic evidence of this from the Sverdlovsk production strain genome. The whole genome SNP genotype of the Sverdlovsk strain was used to precisely identify it and its close relatives in the context of an extensive global B. anthracis strain collection. This genomic identity can now be used for forensic tracking of this weapons material on a global scale and for future anthrax investigations.

Genomics

Comparative Genomics Of Two Sequential Candida glabrata Clinical Isolates

Candida glabrata is an important fungal pathogen which develops rapidly antifungal resistance in treated patients. It is known that azole treatments lead to antifungal resistance in this fungal species and that multidrug efflux transporters are involved in this process. Specific mutations in the transcriptional regulator PDR1 result in upregulation of the transporters. In addition, we showed that the PDR1 mutations can contribute to enhance virulence in animal models. We were interested in this study to compare genomes of two specific C. glabrata related isolates, one of which was azole-susceptible (DSY562) while the other was azole-resistant (DSY565). DSY565 contained a PDR1 mutation (L280F) and was isolated after a time lapse of 50 days of azole therapy. We expected that genome comparisons between both isolates could reveal additional mutations reflecting host adaptation or even additional resistance mechanisms. The PacbBio technology used here yielded 14 major contigs (sizes 0.18 Mb-1.6 Mb) and mitochondrial genomes from both DSY562 and DSY565 isolates that were highly similar to each other. Comparisons of the clinical genomes with the published CBS138 genome indicated important genome rearrangements, but not between the clinical strains. Among unique features, several retrotransposons were identified in the genomes of the investigated clinical isolates. DSY562 and DSY565 contained each a large set of adhesin-like genes (101 and 107, respectively), which exceed by far the number of reported adhesins (66) in the CBS138 genome. Comparison between DSY562 and DSY565 yielded 17 non-synonymous SNPs (among which the expected PDR1 mutation) as well as small size indels in coding regions (11) but mainly in adhesin-like genes. The genomes were containing a DNA mismatch repair allele of MSH2 known to be involved in the so-called hypermutator phenotype of this yeast species and the number of accumulated mutations between both clinical isolates is consistent with the presence of a MSH2 defect. In conclusion, this study is the first to compare genomes of C. glabrata sequential clinical isolates using the PacBio technology as an approach. The genomes of these isolates taken in the same patient at two different time points were exhibiting limited variations, even if submitted to the host pressure.

genomics

Genomic architecture of codfishes featured by expansions of innate immune genes and short tandem repeats

BackgroundIncreased availability of genome assemblies for non-model organisms has resulted in invaluable biological and genomic insight into numerous vertebrates including teleosts. The sequencing and assembly of the Atlantic cod (Gadus morhua) genome and the genomes of many of its relatives (Gadiformes) demonstrated a shared loss 100 million years ago of the major histocompatibility complex (MHC) II genes. The recent publication of an improved version of the Atlantic cod genome assembly reported an extreme density of tandem repeats compared to other vertebrate genome assemblies. Highly contiguous genome assemblies are needed to further investigate the unusual immune system of the Gadiformes, and the high density of tandem repeats in this group.\n\nResultsHere, we have sequenced and assembled the genome of haddock (Melanogrammus aeglefinus) - a relative of Atlantic cod - using a combination of PacBio and Illumina reads. Comparative analyses uncover that the haddock genome contains an even higher density of tandem repeats outside and within protein coding sequences than Atlantic cod. Further, both species show an elevated number of tandem repeats in genes mainly involved in signal transduction compared to other teleosts. An in-depth characterization of the immune gene repertoire demonstrates a substantial expansion of MCHI in Atlantic cod compared to haddock. In contrast, the Toll-like receptors show a similar pattern of gene losses and expansions. For another gene family associated with the innate immune system, the NOD-like receptors (NLRs), we find a large expansion common to all teleosts, with possible lineage-specific expansions in zebrafish, stickleback and the codfishes.\n\nConclusionsThe generation of a highly contiguous genome assembly of haddock revealed that the high density of short tandem repeats as well as expanded immune gene families is not unique to Atlantic cod - but most likely a feature common to all codfishes. A shared expansion of NLR genes in teleosts suggests that the NLRs have a more substantial role in the innate immunity of teleosts than other vertebrates. Moreover, we find that high copy number genes combined with variable genome assembly qualities may impede complete characterization, i.e. the number of NLRs might be underestimates in the different teleost species.

genomics

Novel metrics for quantifying bacterial genome composition skews

BackgroundBacterial genomes have characteristic compositional skews, which are differences in nucleotide frequency between the leading and lagging DNA strands across a segment of a genome. It is thought that these strand asymmetries arise as a result of mutational biases and selective constraints, particularly for energy efficiency. Analysis of compositional skews in a diverse set of bacteria provides a comparative context in which mutational and selective environmental constraints can be studied. These analyses typically require finished and well-annotated genomic sequences.\n\nResultsWe present three novel metrics for examining genome composition skews; all three metrics can be computed for unfinished or partially-annotated genomes. The first two metrics, (dot-skew and cross-skew) depend on sequence and gene annotation of a single genome, while the third metric (residual skew) highlights unusual genomes by subtracting a GC content-based model of a library of genome sequences. We applied these metrics to all 7738 available bacterial genomes, including partial drafts, and identified outlier species. A number of these outliers (i.e., Borrelia, Ehrlichia, Kinetoplastibacterium, and Phytoplasma) display similar skew patterns despite only distant phylogenetic relationship. While unrelated, some of the outlier bacterial species share lifestyle characteristics, in particular intracellularity and biosynthetic dependence on their hosts.\n\nConclusionsOur novel metrics appear to reflect the effects of biosynthetic constraints and adaptations to life within one or more hosts on genome composition. We provide results for each analyzed genome, software and interactive visualizations at http://db.systemsbiology.net/gestalt/skew_metrics.

genomics

Comprehensive, Integrated, and Phased Whole-Genome Analysis of the Primary ENCODE Cell Line K562

K562 is widely used in biomedical research. It is one of three tier-one cell lines of ENCODE and also most commonly used for large-scale CRISPR/Cas9 screens. Although its functional genomic and epigenomic characteristics have been extensively studied, its genome sequence and genomic structural features have never been comprehensively analyzed. Such information is essential for the correct interpretation and understanding of the vast troves of existing functional genomics and epigenomics data for K562. We performed and integrated deep-coverage whole-genome (short-insert), mate-pair, and linked-read sequencing as well as karyotyping and array CGH analysis to identify a wide spectrum of genome characteristics in K562: copy numbers (CN) of aneuploid chromosome segments at high-resolution, SNVs and Indels (both corrected for CN in aneuploid regions), loss of heterozygosity, mega-base-scale phased haplotypes often spanning entire chromosome arms, structural variants (SVs) including small and large-scale complex SVs and non-reference retrotransposon insertions. Many SVs were phased, assembled, and experimentally validated. We identified multiple allele-specific deletions and duplications within the tumor suppressor gene FHIT. Taking aneuploidy into account, we re-analyzed K562 RNA-seq and whole-genome bisulfite sequencing data for allele-specific expression and allele-specific DNA methylation. We also show examples of how deeper insights into regulatory complexity are gained by integrating genomic variant information and structural context with functional genomics and epigenomics data. Furthermore, using K562 haplotype information, we produced an allele-specific CRISPR targeting map. This comprehensive whole-genome analysis serves as a resource for future studies that utilize K562 as well as a framework for the analysis of other cancer genomes.

genomics

Independent assessment and improvement of wheat genome assemblies using Fosill jumping libraries

BackgroundThe accurate sequencing and assembly of very large, often polyploid, genomes remain a challenging task, limiting long range sequence information and phased sequence variation for applications such as plant breeding. The 15 Gb hexaploid bread wheat genome has been particularly challenging to sequence, and several contending approaches recently generated accurate long range assemblies. Understanding errors in these assemblies is important for optimising future sequencing and assembly approaches and for comparative genomics.\n\nResultsHere we use a Fosill 38 Kb jumping library to assess medium and longer range order of different publicly available wheat genome assemblies. Modifications to the Fosill protocol generated longer Illumina sequences and enabled comprehensive genome coverage. Analyses of two independent BAC based chromosome-scale assemblies, two independent Illumina whole genome shotgun assemblies, and a hybrid long read (PacBio) and short read (Illumina) assembly were carried out. We revealed a variety of discrepancies using Fosill mate-pair mapping and validated several of each class. In addition, Fosill mate-pairs were used to scaffold a whole genome Illumina assembly, leading to a three-fold increase in N50 values.\n\nConclusionsOur analyses, using an independent means to validate different wheat genome assemblies, show that whole genome shotgun assemblies are significantly more accurate by all measures compared to BAC-based chromosome scale assemblies. Although current whole genome assemblies are reasonably accurate and useful, additional steps will be needed for the rapid, cost effective and complete sequencing and assembly of wheat genomes.

genomics