bioRxiv ScienceSearch

SEARCH · bioRxiv Science

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29Linked to original sources

Expanding an expanded genome: long-read sequencing of Trypanosoma cruzi

Although the genome of Trypanosoma cruzi, the causative agent of Chagas disease, was first made available in 2005, with additional strains reported later, the intrinsic genome complexity of this parasite (abundance of repetitive sequences and genes organized in tandem) has traditionally hindered high-quality genome assembly and annotation. This also limits diverse types of analyses that require high degree of precision. Long reads generated by third-generation sequencing technologies are particularly suitable to address the challenges associated with T. cruzi{acute}s genome since they permit directly determining the full sequence of large clusters of repetitive sequences without collapsing them. This, in turn, allows not only accurate estimation of gene copy numbers but also circumvents assembly fragmentation. Here, we present the analysis of the genome sequences of two T. cruzi clones: the hybrid TCC (DTU TcVI) and the non-hybrid Dm28c (DTU TcI), determined by PacBio SMRT technology. The improved assemblies herein obtained permitted us to accurately estimate gene copy numbers, abundance and distribution of repetitive sequences (including satellites and retroelements). We found that the genome of T. cruzi is composed of a \"core compartment\" and a \"disruptive compartment\" which exhibit opposite gene and GC content composition. New tandem and disperse repetitive sequences were identified, including some located inside coding sequences. Additionally, homologous chromosomes were separately assembled, allowing us to retrieve haplotypes as separate contigs instead of a unique mosaic sequence. Finally, manual annotation of surface multigene families MUC and trans-sialidases allows now a better overview of these complex groups of genes.

genomics

Long-read genome sequence and assembly of Leptopilina boulardi: a specialist Drosophila parasitoid

BackgroundLeptopilina boulardi is a specialist parasitoid belonging to the order Hymenoptera, which attacks the larval stages of Drosophila. The Leptopilina genus has enormous value in the biological control of pests as well as in understanding several aspects of host-parasitoid biology. However, none of the members of Figitidae family has their genomes sequenced. In order to improve the understanding of the parasitoid wasps by generating genomic resources, we sequenced the whole genome of L. boulardi.\n\nFindingsHere, we report a high-quality genome of L. boulardi, assembled from 70Gb of Illumina reads and 10.5Gb of PacBio reads, forming a total coverage of 230X. The 375Mb draft genome has an N50 of 275Kb with 6315 scaffolds >500bp, and encompasses >95% complete BUSCOs. The GC% of the genome is 28.26%, and RepeatMasker identified 868105 repeat elements covering 43.9% of the assembly. A total of 25259 protein-coding genes were predicted using a combination of ab-initio and RNA-Seq based methods, with an average gene size of 3.9Kb. 78.11% of the predicted genes could be annotated with at least one function.\n\nConclusionOur study provides a highly reliable assembly of this parasitoid wasp, which will be a valuable resource to researchers studying parasitoids. In particular, it can help delineate the host-parasitoid mechanisms that are part of the Drosophila - Leptopilina model system.

genomics

The draft genome of the invasive walking stick, Medauroidea extradendata, reveals extensive lineage-specific gene family expansions of cell wall degrading enzymes in Phasmatodea

Plant cell wall components are the most abundant macromolecules on Earth. The study of the breakdown of these molecules is thus a central question in biology. Surprisingly, plant cell wall breakdown by herbivores is relatively poorly understood, as nearly all early work focused on the mechanisms used by symbiotic microbes to breakdown plant cell walls in insects such as termites. Recently, however, it has been shown that many organisms make endogenous cellulases. Insects, and other arthropods, in particular have been shown to express a variety of plant cell wall degrading enzymes in many gene families with the ability to break down all the major components of the plant cell wall. Here we report the genome of a walking stick, Medauroidea extradentata, an obligate herbivore that makes uses of endogenously produced plant cell wall degrading enzymes. We present a draft of the 3.3Gbp genome along with an official gene set that contains a diversity of plant cell wall degrading enzymes. We show that at least one of the major families of plant cell wall degrading enzymes, the pectinases, have undergone a striking lineage-specific gene family expansion in the Phasmatodea. This genome will be a useful resource for comparative evolutionary studies with herbivores in many other clades and will help elucidate the mechanisms by which metazoans breakdown plant cell wall components.\n\nData availabilityThe Medauroidea extradentata genome assembly, Med v1.0, is available for download via NCBI (Bioproject: PRJNA369247). The genome, annotation files, and official gene set Mext_OGS_v1.0 are also available at the i5k NAL workspace (https://i5k.nal.usda.gov/medauroidea-extradentata) and at github (https://github.com/pbrec/medauroidea_genome_resources). The genomic raw reads are available via NCBI SRA: SRR6383867 and the raw transcriptomic reads are available at NCBI SRA: SRR6383868, SRR6383869.

genomics

MinION sequencing enables rapid whole genome assembly of Rickettsia typhi in a resource-limited setting

The infrastructure challenges and costs of next-generation sequencing have been largely overcome, for many sequencing applications, by Oxford Nanopore Technologies portable MinION sequencer. However the question remains open whether MinION-based bacterial whole-genome sequencing (WGS) is by itself sufficient for the accurate assessment of phylogenetic and epidemiological relationships between isolates and whether such tasks can be undertaken in resource-limited settings. To investigate this question, we sequenced the genome of an isolate of Rickettsia typhi, an important and neglected cause of fever across much of the tropics and subtropics, for which only three genomic sequences previously existed. We prepared and sequenced libraries on a MinION in Vientiane, Lao PDR using v9.5 chemistry and in parallel we sequenced the same isolate on the Illumina platform in a genomics laboratory in the UK. The MinION sequence reads yielded a single contiguous assembly, in which the addition of Illumina data revealed 226 base-substitution and 5,856 in/del errors. The combined assembly represents the first complete genome sequence of a human R. typhi isolate collected in the last 50 years and differed from the genomes of existing strains collected over a 90-year time period at very few sites, and with no re-arrangements. Filtering based on the known error profile of MinION data improved the accuracy of the Nanopore-only assembly. However, the frequency of false-positive errors remained greater than true sequence divergence from recorded sequences. While Nanopore-only sequencing cannot yet recover phylogenetic signal in R. typhi, such an approach may be applicable for more diverse organisms.

genomics

Genome evolution in Burkholderia spp

BackgroundThe genus Burkholderia consists of species that occupy remarkably diverse ecological niches. Its best known members are important pathogens, B. mallei and B. pseudomallei, which cause glanders and melioidosis, respectively. Burkholderia genomes are unusual due to their multichromosomal organization.\n\nResultsWe performed integrated genomic analysis of 127 Burkholderia strains. The pan-genome is open with the saturation to be reached between 86,000 and 88,000 genes. The reconstructed rearrangements indicate a strong avoidance of intra-replichore inversions that is likely caused by selection against the transfer of large groups of genes between the leading and the lagging strands. Translocated genes also tend to retain their position in the leading or the lagging strand, and this selection is stronger for large syntenies. Integrated reconstruction of chromosome rearrangements in the context of strains phylogeny reveals parallel rearrangements that may indicate inversion-based phase variation and integration of new genomic islands. In particular, we detected parallel inversions in the second chromosomes of B. pseudomallei with breakpoints formed by genes encoding membrane components of multidrug resistance complex, that may be linked to a phase variation mechanism. Two genomic islands, spreading horizontally between chromosomes, were detected in the B. cepacia group.\n\nConclusionsThis study demonstrates the power of integrated analysis of pan-genomes, chromosome rearrangements, and selection regimes. Non-random inversion patterns indicate selective pressure, inversions are particularly frequent in a recent pathogen B. mallei, and, together with periods of positive selection at other branches, may indicate adaptation to new niches. One such adaptation could be a possible phase variation mechanism in B. pseudomallei.

genomics

A chromosome-scale assembly of the sorghum genome using nanopore sequencing and optical mapping

The advent of long-read sequencing technologies has greatly facilitated assemblies of large eukaryotic genomes. In this paper, Oxford Nanopore sequences generated on a MinION sequencer were combined with BioNano Genomics Direct Label and Stain (DLS) optical maps to generate a chromosome-scale de novo assembly of the repeat-rich Sorghum bicolor Tx430 genome. The final hybrid assembly consists of 29 scaffolds, encompassing in most cases entire chromosome arms. It has a scaffold N50 value of 33.28Mbps and covers >90% of Sorghum bicolor expected genome length. A sequence accuracy of 99.67% was obtained in unique regions after aligning contigs against Illumina Tx430 data. Alignments showed that 99.4% of the 34,211 public gene models are present in the assembly, including 94.2% mapping end-to-end. Comparisons of the DLS optical maps against the public Sorghum Bicolor v3.0.1 BTx623 genome assembly suggest the presence of substantial genomic rearrangements whose origin remains to be determined.

genomics

Computational analysis of the Plasmodiophora brassicae genome: mitochondrial sequence description and metabolic pathway database design

Plasmodiophora brassicae is an obligate biotrophic pathogenic protist responsible for clubroot, a root gall disease of Brassicaceae species. In addition to the reference genome of the P. brassicae European e3 isolate and the draft genomes of Canadian or Chinese isolates, we present the genome of eH, a second European isolate. Refinement of the annotation of the eH genome led to the identification of the mitochondrial genome sequence, which was found to be bigger than that of Spongospora subterranea, another plant parasitic Plasmodiophorid phylogenetically related to P. brassicae. New pathways were also predicted, such as those for the synthesis of spermidine, a polyamine up-regulated in clubbed regions of roots. A P. brassicae pathway genome database was created to facilitate the functional study of metabolic pathways in transcriptomics approaches. These available tools can help in our understanding of the regulation of P. brassicae metabolism during infection and in response to diverse constraints.

genomics

Repeat elements organize 3D genome structure and mediate transcription in the filamentous fungus Epichloë festucae

Structural features of genomes, including the three-dimensional arrangement of DNA in the nucleus, are increasingly seen as key contributors to the regulation of gene expression. However, studies on how genome structure and nuclear organization influence transcription have so far been limited to a handful of model species. This narrow focus limits our ability to draw general conclusions about the ways in which three-dimensional structures are encoded, and to integrate information from three-dimensional data to address a broader gamut of biological questions. Here, we generate a complete and gapless genome sequence for the filamentous fungus, Epichloe festucae. Coupling it with RNAseq and HiC data, we investigate how the structure of the genome contributes to the suite of transcriptional changes that an Epichloe species needs to maintain symbiotic relationships with its grass host. Our results reveal a unique \"patchwork\" genome, in which repeat-rich blocks of DNA with discrete boundaries are interspersed by gene-rich sequences. In contrast to other species, the three-dimensional structure of the genome is anchored by these repeat blocks, which act to isolate transcription in neighbouring gene-rich regions. Genes that are differentially expressed in planta are enriched near the boundaries of these repeat-rich blocks, suggesting that their three-dimensional orientation partly encodes and regulates the symbiotic relationship formed by this organism.

genomics

A hybrid de novo genome assembly of the honeybee, Apis mellifera, with chromosome-length scaffolds

BackgroundThe ability to generate long sequencing reads and access long-range linkage information is revolutionizing the quality and completeness of genome assemblies. Here we use a hybrid approach that combines data from four genome sequencing and mapping technologies to generate a new genome assembly of the honeybee Apis mellifera. We first generated contigs based on PacBio sequencing libraries, which were then merged with linked-read 10x Chromium data followed by scaffolding using a BioNano optical genome map and a Hi-C chromatin interaction map, complemented by a genetic linkage map.\n\nResultsEach of the assembly steps reduced the number of gaps and incorporated a substantial amount of additional sequence into scaffolds. The new assembly (Amel_HAv3) is significantly more contiguous and complete than the previous one (Amel_4.5), based mainly on Sanger sequencing reads. N50 of contigs is 120-fold higher (5.381 Mbp compared to 0.053 Mbp) and we anchor >98% of the sequence to chromosomes. All of the 16 chromosomes are represented as single scaffolds with an average of three sequence gaps per chromosome. The improvements are largely due to the inclusion of repetitive sequence that was unplaced in previous assemblies. In particular, our assembly is highly contiguous across centromeres and telomeres and includes hundreds of AvaI and AluI repeats associated with these features.\n\nConclusionsThe improved assembly will be of utility for refining gene models, studying genome function, mapping functional genetic variation, identification of structural variants, and comparative genomics.

genomics

Integrating genomic resources to present full gene and promoter capture probe sets for bread wheat

BackgroundWhole genome shotgun re-sequencing of wheat is expensive because of its large, repetitive genome. Moreover, sequence data can fail to map uniquely to the reference genome making it difficult to unambiguously assign variation. Re-sequencing using target capture enables sequencing of large numbers of individuals at high coverage to reliably identify variants associated with important agronomic traits.\n\nResultsWe present and validate two gold standard capture probe sets for hexaploid bread wheat, a gene and a promoter capture, which are designed using recently developed genome sequence and annotation resources. The captures can be combined or used independently. We demonstrate that the capture probe sets effectively enrich the high confidence genes and promoters that were identified in the genome alongside a large proportion of the low confidence genes and promoters. Finally, we demonstrate successful sample multiplexing that allows generation of adequate sequence coverage for SNP calling while significantly reducing cost per sample for gene and promoter capture.\n\nConclusionsWe show that a capture design employing an island strategy can enable analysis of the large gene/promoter space of wheat with only 2x160 Mb probe sets. Furthermore, these assays extend the regions of the wheat genome that are amenable to analyses beyond its exome, providing tools for detailed characterization of these regulatory regions in large populations.

genomics

Genome-wide maps of distal gene regulatory regions active in the human placenta

Placental dysfunction is implicated in many pregnancy complications, including preeclampsia and preterm birth (PTB). While both these syndromes are influenced by environmental risk factors, they also have a substantial genetic component that is not well understood. Precisely controlled gene expression during development is crucial to proper placental function and often mediated through gene regulatory enhancers. However, we lack accurate maps of placental enhancer activity due to the challenges of assaying the placenta and the difficulty of comprehensively identifying enhancers. To address the gap in our knowledge of gene regulatory elements in the placenta, we used a two-step machine learning pipeline to synthesize existing functional genomics studies, transcription factor (TF) binding patterns, and evolutionary information to predict placental enhancers. The trained classifiers accurately distinguish enhancers from the genomic background and placental enhancers from enhancers active in other tissues. Genomic features collected from tissues and cell lines involved in pregnancy are the most predictive of placental regulatory activity. Applying the classifiers genome-wide enabled us to create a map of 33,010 predicted placental enhancers, including 4,562 high-confidence enhancer predictions. The genome-wide placental enhancers are significantly enriched nearby genes associated with placental development and birth disorders and for SNPs associated with gestational age. These genome-wide predicted placental enhancers provide candidate regions for further testing in vitro, will assist in guiding future studies of genetic associations with pregnancy phenotypes, and aid interpretation of potential mechanisms of action for variants found through genetic studies.

genomics

Comparative genomics-first approach to understand diversification of secondary metabolite biosynthetic pathways in symbiotic dinoflagellates

Symbiotic dinoflagellates of the genus Symbiodinium are photosynthetic and unicellular. They possess smaller nuclear genomes than other dinoflagellates and produce structurally specialized, biologically active, secondary metabolites. Polyketide biosynthetic genes of toxic dinoflagellates have been studied extensively using transcriptomic analyses; however, a comparative genomic approach to understand secondary metabolism has been hampered by their large genome sizes. Here, we use a combined genomic and metabolomics approach to investigate the structure and diversification of secondary metabolite genes to understand how chemical diversity arises in three decoded Symbiodinium genomes (A3, B1 and C). Our analyses identify 71 polyketide synthase and 41 non-ribosomal peptide synthetase genes from two newly decoded genomes of clades A3 and C. Additionally, phylogenetic analyses indicate that almost all of the gene families are derived from lineage-specific gene duplications in Symbiodinium clades, suggesting divergence for environmental adaptation. Few metabolic pathways are conserved among the three clades and we detect metabolic similarity only in the recently diverged clades, B1 and C. We establish that secondary metabolism protein architecture guides substrate specificity and that gene duplication and domain shuffling have resulted in diversification of secondary metabolism genes.

genomics

Use of Hyperspectral Reflectance-Derived Relationship Matrices for Genomic Prediction of Grain Yield in Wheat

Hyperspectral reflectance phenotyping and genomic selection are two emerging technologies that have the potential to increase plant breeding efficiency by improving prediction accuracy for grain yield. Hyperspectral cameras quantify canopy reflectance across a wide range of wavelengths that are associated with numerous biophysical and biochemical processes in plants. Genomic selection models utilize genome-wide marker or pedigree information to predict the genetic values of breeding lines. In this study, we propose a multi-kernel GBLUP approach to genomic selection that uses genomic marker-, pedigree-, and hyperspectral reflectance-derived relationship matrices to model the genetic main effects and genotype x environment (G x E) interactions across environments within a bread wheat (Triticum aestivum L.) breeding program. We utilized an airplane equipped with a hyperspectral camera to phenotype five differentially managed treatments of the yield trials conducted by the Bread Wheat Improvement Program, International Maize and Wheat Improvement Center (CIMMYT) at Ciudad Obregon, Mexico over four breeding cycles. We observed that single-kernel models using hyperspectral reflectance-derived relationship matrices performed similarly or superior to marker-and pedigree-based genomic selection models when predicting within and across environments. Multi-kernel models combining marker/pedigree information with hyperspectral reflectance phentoypes had the highest prediction accuracies; however, improvements in accuracy over marker-and pedigree-based models were marginal when correcting for days to heading. Our results demonstrates the potential of hyperspectral imaging in predicting grain yield within a multi-environment context, it also supports further studies on integration of hyperspectral reflectance phenotyping in breeding programs.

genomics

Genome Analyses of a New Mycoplasma Species From the scorpion Centruroides vittatus.

Arthropod Mycoplasma are little known endosymbionts in insects, primarily known as plant disease vectors. Mycoplasma in other arthropods such as arachnids are unknown. We report the first complete Mycoplasma genome sequenced, identified, and annotated from a scorpion, Centruroides vittatus, and designate it as Mycoplasma vittatus. We find the genome is at least a 683,827 bp single circular chromosome with a GC content of 43.7% and with 1,010 protein-coding genes. The putative virulence determinants include 20 genes associated with the virulence operon associated with protein synthesis (SSU ribosomal proteins) and nine genes with fluoroquinolone resistance. Comparative analysis revealed that the M. vittatus genome is smaller than other Mycoplasma genomes and exhibits a higher GC content. Phylogenetic analysis shows M. vittatus as part of the Hominis group of Mycoplasma. As arthropod genomes accumulate, further novel Mycoplasma genomes may be identified and characterized.

genomics

Synthetic and genomic regulatory elements reveal aspects of cis-regulatory grammar in Mouse Embryonic Stem Cells

In embryonic stem cells (ESCs), a core network of transcription factors establish and maintain the gene expression program necessary to grow indefinitely in cell culture and generate all three primary germ layers. To understand how interactions between four key pluripotency transcription factors (TFs), SOX2, POU5F1 (OCT4), KLF4, and ESRRB, contribute to cis-regulation in mouse ESCs, we assayed two massively parallel reporter assay (MPRA) libraries composed of different combinations of binding sites for these TFs. One library was an exhaustive set of synthetic cis-regulatory elements and the second was a set of genomic sequences with comparable configurations of binding sites. Comparisons between the libraries allowed us to determine the regulatory grammar requirements for these binding sites in constrained synthetic contexts versus genomic sequence contexts. We found that binding site quality is a common attribute for active elements in both the synthetic and genomic contexts. For synthetic regulatory elements, the level of expression is mostly determined by the number of binding sites but is tuned by a grammar that includes position effects. Surprisingly, this grammar appears to only play a small role in setting the output levels of genomic sequences. The relative activity of genomic sequences is best explained by the predicted affinity of binding sites, regardless of identity, and optimized spacing between sites. Our findings highlight the need for detailed examinations of complex sequence space when trying to understand cis-regulatory grammar in the genome.

genomics

Estimation of Rearrangement Break Rates Across the Genome

Genomic rearrangements provide an important source of novel functions by recombining genes and motifs throughout and between genomes. However, understanding how rearrangement functions to shape genomes is hard because reconstructing rearrangements is a combinatoric problem which often has many solutions. In lieu of reconstructing the history of rearrangements, we answer the question of where rearrangements are occurring in the genome by remaining agnostic to the types of rearrangement and solving the simpler problem of estimating the rate at which double-strand breaks occur at every site in a genome. We phrase this problem in graph theoretic terms and find that it is a special case of the minimum cover problem for an interval graph. We employ and modify existing algorithms for efficiently solving this problem. We implement this method as a Python program, named BRAG, and use it to estimate the break rates in the genome of the model Ascomycete mold, Neurospora crassa. We find evidence that rearrangements are more common in the subtelomeric regions of the chromosomes, which facilitates the evolution of novel genes.

genomics

Genomic, transcriptomic, and structural analysis of Pseudomonas virus PA5oct highlights the molecular complexity among Jumbo phages

Pseudomonas virus PA5oct has a large, linear, double-stranded DNA genome (287,182 bp) and is related to Escherichia phages 121Q/PBECO 4, Klebsiella phage vB_KleM-RaK2, Klebsiella phage K64-1, and Cronobacter phage vB_CsaM_GAP32. A protein-sharing network analysis highlights the conserved core genes within this clade. Combining genome, RNAseq and mass spectrometry analyses of its virion proteins allowed us to accurately identify genes and elucidate regulatory elements for this phage (ncRNAs, tRNAs and promoter elements). In total PA5oct encodes 462 CDS (compared to 345 in silico predicted genes using automated annotation pipelines), of which 25.32%, have been identified as virion-associated based on ESI-MS/MS. The RNAseq-based temporal genome organization suggests a gradual take-over by viral transcripts from 21%, 69%, and 92% at 5, 15 and 25 min after infection, respectively. Like many large phages, PA5oct is not organized into contiguous regions of temporal transcription. However, although the temporal regulation of the PA5oct genome expression reveals specific genome clusters expressed in early and late infection, many genes encoding experimentally observed structural proteins surprisingly appear to remain almost untranscribed throughout the infection cycle. Within the host, operons associated with elements of a cryptic Pf1-like prophage are upregulated, as are operons responsible for Psl exopolysaccharide (pslE-J) and periplasmic nitrate reductase (napA-F) production. The characterization described here represents a crucial step towards understanding the genomic complexity as well as molecular diversity of jumbo viruses.

genomics

Genome duplication and reorganization in Aquilegia

BackgroundWhole-genome duplications (WGD) have dominated the evolutionary history of plants. One consequence of WGD is a dramatic restructuring of the genome as it undergoes diploidization, a process under which deletions and rearrangements of various sizes scramble the genetic material, leading to a repacking of the genome and eventual return to diploidy. Here, we investigate the history of WGD in the columbine genus Aquilegia, a basal eudicot, and use it to illuminate the origins of the core eudicots.\n\nResultsWithin-genome synteny confirms that columbines are ancient tetraploids, and comparison with the grape genome reveals that this tetraploidy appears to be shared with the core eudicots. Thus, the ancient gamma hexaploidy found in all core eudicots must have involved a two-step process: first tetraploidy in the ancestry of all eudicots, then hexaploidy in the ancestry of core eudicots. Furthermore, the precise pattern of synteny sharing suggests that the latter involved allopolyploidization, and that core eudicots thus have a hybrid origin.\n\nConclusionsNovel analyses of synteny sharing together with the well-preserved structure of the columbine genome reveal that the gamma hexaploidy at the root of core eudicots is likely a result of hybridization between a tetraploid and a diploid species.

genomics