bioRxiv Science⌕ Search

SEARCH · bioRxiv Science

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,693 records · Page 94Linked to original sources

Signatures of host specialization and a recent transposable element burst in the dynamic one-speed genome of the fungal barley powdery mildew pathogen

Powdery mildews are biotrophic pathogenic fungi infecting a number of economically important plants. The grass powdery mildew, Blumeria graminis, has become a model organism to study host specialization of obligate biotrophic fungal pathogens. We resolved the large-scale genomic architecture of B. graminis forma specialis hordei (Bgh) to explore the potential influence of its genome organization on the co-evolutionary process with its host plant, barley (Hordeum vulgare). The near-chromosome level assemblies of the Bgh reference isolate DH14 and one of the most diversified isolates, RACE1, enabled a comparative analysis of these haploid genomes, which are highly enriched with transposable elements (TEs). We found largely retained genome synteny and gene repertoires, yet detected copy number variation (CNV) of secretion signal peptide-containing protein-coding genes (SPs) and locally disrupted synteny blocks. Genes coding for sequence-related SPs are often locally clustered, but neither the SP clusters nor TEs are enriched in specific genomic regions. Extended comparative analysis with different host-specific B. graminis formae speciales revealed the existence of a core suite of SPs, but also isolate-specific SP sets as well as congruence of SP CNV and phylogenetic relationship. We further detected evidence for a recent, lineage-specific expansion of TEs in the Bgh genome. The characteristics of the Bgh genome (largely retained synteny, CNV of SP genes, recently proliferated TEs and a lack of compartmentalization) are consistent with a \"one-speed\" genome that differs in its architecture and (co-)evolutionary pattern from the \"two-speed\" genomes reported for several other filamentous phytopathogens.

genomics↗

Nuclear and mitochondrial genomes of the hybrid fungal plant pathogen Verticillium longisporum display a mosaic structure

Allopolyploidization, genome duplication through interspecific hybridization, is an important evolutionary mechanism that can enable organisms to adapt to environmental changes or stresses. This increased adaptive potential of allopolyploids can be particularly relevant for plant pathogens in their quest for host immune response evasion. Allodiploidization likely caused the shift in host range of the fungal pathogen plant Verticillium longisporum, as V. longisporum mainly infects Brassicaceae plants in contrast to haploid Verticillium spp. In this study, we investigated the allodiploid genome structure of V. longisporum and its evolution in the hybridization aftermath. The nuclear genome of V. longisporum displays a mosaic structure, as numerous contigs consists of sections of both parental origins. V. longisporum encountered extensive genome rearrangements, whereas the contribution of gene conversion is negligible. Thus, the mosaic genome structure mainly resulted from genomic rearrangements between parental chromosome sets. Furthermore, a mosaic structure was also found in the mitochondrial genome, demonstrating its bi-parental inheritance. In conclusion, the nuclear and mitochondrial genomes of V. longisporum parents interacted dynamically in the hybridization aftermath. Conceivably, novel combinations of DNA sequence of different parental origin facilitated genome stability after hybridization and consecutive niche adaptation of V. longisporum.

genomics↗

Whole genome sequence of Mapuche-Huilliche Native Americans

BackgroundWhole human genome sequencing initiatives provide a compendium of genetic variants that help us understand population history and the basis of genetic diseases. Current data mostly focuses on Old World populations and information on the genomic structure of Native Americans, especially those from the Southern Cone is scant.\n\nResultsHere we present a high-quality complete genome sequence of 11 Mapuche-Huilliche individuals (HUI) from Southern Chile (85% genomic and 98% exonic coverage at > 30X), with 96-97% high confidence calls. We found approximately 3.1x106 single nucleotide variants (SNVs) per individual and identified 403,383 (6.9%) of novel SNVs that are not included in current sequencing databases. Analyses of large-scale genomic events detected 680 copy number variants (CNVs) and 4,514 structural variants (SVs), including 398 and 1,910 novel events, respectively. Global ancestry composition of HUI genomes revealed that the cohort represents a marginally admixed population from the Southern Cone, whose genetic component is derived from early Native American ancestors. In addition, we found that HUI genomes display highly divergent and novel variants with potential functional impact that converge in ontological categories essential in cell metabolic processes.\n\nConclusionsMapuche-Huilliche genomes contain a unique set of small- and large-scale genomic variants in functionally linked genes, which may contribute to susceptibility for the development of common complex diseases or traits in admixed Latinos and Native American populations. Our data represents an ancestral reference panel for population-based studies in Native and admixed Latin American populations.

genomics↗

Systematic Discovery of Conservation States for Single-Nucleotide Annotation of the Human Genome

Comparative genomics sequence data is an important source of information for interpreting genomes. Genome-wide annotations based on this data have largely focused on univariate scores or binary calls of evolutionary constraint. Here we present a complementary whole genome annotation approach, ConsHMM, which applies a multivariate hidden Markov model to learn de novo different conservation states based on the combinatorial and spatial patterns of which species align to and match a reference genome in a multiple species DNA sequence alignment. We applied ConsHMM to a 100-way vertebrate sequence alignment to annotate the human genome at single nucleotide resolution into 100 different conservation states. These states have distinct enrichments for other genomic information including gene annotations, chromatin states, and repeat families, which were used to characterize their biological significance. Conservation states have greater or complementary predictive information than standard constraint based measures for a variety of genome annotations. Bases in constrained elements have distinct heritability enrichments depending on the conservation state assignment, demonstrating their relevance to analyzing phenotypic associated variation. The conservation states also highlight differences in the conservation patterns of bases prioritized by a number of scores used for variant prioritization. The ConsHMM method and conservation state annotations provide a valuable resource for interpreting genomes and genetic variation.

genomics↗

Bacillus safensis FO-36b and Bacillus pumilus SAFR-032: A Whole Genome Comparison of Two Spacecraft Assembly Facility Isolates

BackgroundBacillus strains producing highly resistant spores have been isolated from cleanrooms and space craft assembly facilities. Organisms that can survive such conditions merit planetary protection concern and if that resistance can be transferred to other organisms, a health concern too. To further efforts to understand these resistances, the complete genome of Bacillus safensis strain FO-36b, which produces spore resistant to peroxide and radiation was determined. The genome was compared to the complete genome of B. pumilus SAFR-032, as well as draft genomes of B. safensis JPL-MERTA-8-2 and the type strain B. pumilus ATCC7061T. In addition, comparisons were made to 61 draft genomes that have been mostly identified as strains of B. pumilus or B. safensis.\n\nResultsThe FO-36b gene order is essentially the same as that in SAFR-032 and other B. pumilus strains [1]. The annotated genome has 3850 open reading frames and 40 noncoding RNAs and riboswitches. Of these, 307 are not shared by SAFR-032, and 65 are also not shared by either MERTA or ATCC7061T. The FO-36b genome was found to have ten unique reading frames and two phage-like regions, which have homology with the Bacillus bacteriophage SPP1 (NC_004166) and Brevibacillus phage Jimmer1 (NC_029104). Differing remnants of the Jimmer1 phage are found in essentially all safensis/pumilus strains. Seven unique genes are part of these phage elements. Comparison of gyrA sequences from FO-36b, SAFR-032, ATCC7061T, and 61 other draft genomes separate the various strains into three distinct clusters. Two of these are subgroups of B. pumilus while the other houses all the B. safensis strains.\n\nConclusionsIt is not immediately obvious that the presence or absence of any specific gene or combination of genes is responsible for the variations in resistance seen. It is quite possible that distinctions in gene regulation can change the level of expression of key proteins thereby changing the organisms resistance properties without gain or loss of a particular gene. What is clear is that phage elements contribute significantly to genome variability. The larger comparison of multiple strains indicates that many strains named as B. pumilus actually belong to the B. safensis group.

genomics↗

MinION re-sequencing of Giardia genomes and de novo assembly of a new Giardia isolate

BackgroundGenomes of the parasite Giardia duodenalis are relatively small for eukaryotic genomes, yet there are only six publicly available. Difficulties in assembling the tetraploid G. duodenalis genome from short read sequencing data likely contribute to this lack of genomic information. We sequenced three isolates of G. duodenalis (AWB, BGS, and beaver) on the Oxford Nanopore Technologies MinION whose long reads have the potential to address genomic areas that are problematic for short reads.\n\nResultsUsing a hybrid approach that combines MinION long reads and Illumina short reads to take advantage of the continuity of the long reads and the accuracy of the short reads we generated reference quality genomes for each isolate. The genomes for two of the isolates were evaluated against the available reference genomes for comparison. The third genome for which there is no previous data was then assembled. The long reads were used to find structural variants in each isolate to examine heterozygosity. Consistent with previous findings based on SNPs, Giardia BGS was found to be considerably more heterozygous than the other isolates that are from Assemblage A. We also find an enrichment of variant-specific surface proteins in some of the structural variant regions.\n\nConclusionsOur results show that the MinION can be used to generate reference quality genomes in Giardia and further be used to identify structural variant regions that are an important source of genetic variation not previously examined in these parasites.

genomics↗

A high-quality, long-read de novo genome assembly to aid conservation of Hawaii’s last remaining crow species

Genome-level data can provide researchers with unprecedented precision to examine the causes and genetic consequences of population declines, and to apply these results to conservation management. Here we present a high-quality, long-read, de novo genome assembly for one of the worlds most endangered bird species, the Alala. As the only remaining native crow species in Hawaii, the Alala survived solely in a captive breeding program from 2002 until 2016, at which point a long-term reintroduction program was initiated. The high-quality genome assembly was generated to lay the foundation for both comparative genomics studies, and the development of population-level genomic tools that will aid conservation and recovery efforts. We illustrate how the quality of this assembly places it amongst the very best avian genomes assembled to date, comparable to intensively studied model systems. We describe the genome architecture in terms of repetitive elements and runs of homozygosity, and we show that compared with more outbred species, the Alala genome is substantially more homozygous. We also provide annotations for a subset of immunity genes that are likely to be important for conservation applications, and we discuss how this genome is currently being used as a roadmap for downstream conservation applications.

genomics↗

A high-quality grapevine downy mildew genome assembly reveals rapidly evolving and lineage-specific putative host adaptation genes

Downy mildews are obligate biotrophic oomycete pathogens that cause devastating plant diseases on economically important crops. Plasmopara viticola is the causal agent of grapevine downy mildew, a major disease in vineyards worldwide. We sequenced the genome of Pl. viticola with PacBio long reads and obtained a new 92.94 Mb assembly with high continuity (359 scaffolds for a N50 of 706.5 kb) due to a better resolution of repeat regions. This assembly presented a high level of gene completeness, recovering 1,592 genes encoding secreted proteins involved in plant-pathogen interactions. Pl. viticola had a two-speed genome architecture, with secreted protein-encoding genes preferentially located in gene-sparse, repeat-rich regions and evolving rapidly, as indicated by pairwise dN/dS values. We also used short reads to assemble the genome of Plasmopara muralis, a closely related species infecting grape ivy (Parthenocissus tricuspidata). The lineage-specific proteins identified by comparative genomics analysis included a large proportion of RxLR cytoplasmic effectors and, more generally, genes with high dN/dS values. We identified 270 candidate genes under positive selection, including several genes encoding transporters and components of the RNA machinery potentially involved in host specialization. Finally, the Pl. viticola genome assembly generated here will allow the development of robust population genomics approaches for investigating the mechanisms involved in adaptation to biotic and abiotic selective pressures in this species.\n\nDATA AVAILABILITYRaw reads and genome assemblies have been deposited in GenBank (BioProjects PRJNA329579 for Pl. viticola and PRJNA448661 for Pl. muralis). Genome assemblies, gene annotations and analysis files (e.g. orthology relationships, full tables for GO enrichment analyses, pairwise dN/dS values and branch-site tests) have been deposited in Dataverse (Pl. viticola assembly and annotation: doi.org/10.15454/4NYHD6, Pl. muralis assembly and annotation: doi.org/10.15454/Q1QJYK, analysis files: doi.org/10.15454/8NZ8X9). Links to the data and information about the grapevine downy mildew genome project can be found at http://grapevine-downy-mildew-genome.com/.

genomics↗

SMRT long-read sequencing and Direct Label and Stain optical maps allow the generation of a high-quality genome assembly for the European barn swallow (Hirundo rustica rustica)

BackgroundThe barn swallow (Hirundo rustica) is a migratory bird that has been the focus of a large number of ecological, behavioural and genetic studies. To facilitate further population genetics and genomic studies, here we present a reference genome assembly for the European subspecies (H. r. rustica).\n\nFindingsAs part of the Genome10K (G10K) effort on generating high quality vertebrate genomes, we have assembled a highly contiguous genome assembly using Single Molecule Real-Time (SMRT) DNA sequencing and several Bionano optical map technologies. We compared and integrated optical maps derived both from the Nick, Label, Repair and Stain and from the Direct Label and Stain (DLS) technologies. As proposed by Bionano, the DLS more than doubled the scaffold N50 with respect to the nickase. The dual enzyme hybrid scaffold led to a further marginal increase in scaffold N50 and an overall increase of confidence in the scaffolds. After removal of haplotigs, the final assembly is approximately 1.21 Gbp in size, with a scaffold N50 value of over 25.95 Mbp.\n\nConclusionsThis high-quality genome assembly represents a valuable resource for further studies of population genetics and genomics in the barn swallow, and for studies concerning the evolution of avian genomes. It also represents one of the very first genomes assembled by combining SMRT long-read sequencing with the new Bionano DLS technology for scaffolding. The quality of this assembly demonstrates the potential of this methodology to substantially increase the contiguity of genome assemblies.

genomics↗

Coverage-versus-Length plots, a simple quality control step for de novo yeast genome sequence assemblies

Illumina sequencing has revolutionized yeast genomics, with prices for commercial draft genome sequencing now below $200. The popular SPAdes assembler makes it simple to generate a de novo genome assembly for any yeast species. However, whereas making genome assemblies has become routine, understanding what they contain is still challenging. Here, we show how graphing the information that SPAdes provides about the length and coverage of each scaffold can be used to investigate the nature of an assembly, and to diagnose possible problems. Scaffolds derived from mitochondrial DNA, ribosomal DNA, and yeast plasmids can be identified by their high coverage. Contaminating data, such as cross-contamination from other samples in a multiplex sequencing run, can be identified by its low coverage. Scaffolds derived from the bacteriophage PhiX174 and Lambda DNAs that are frequently used as molecular standards in Illumina protocols can also be detected. Assemblies of yeast genomes with high heterozygosity, such as interspecies hybrids, often contain two types of scaffold: regions of the genome where the two alleles assembled into two separate scaffolds and each has a coverage level C, and regions where the two alleles co-assembled (collapsed) into a single scaffold that has a coverage level 2C. Visualizing the data with Coverage-versus-Length (CVL) plots, which can be done using Microsoft Excel or Google Sheets, provides a simple method to understand the structure of a genome assembly and detect aberrant scaffolds or contigs. We provide a Python script that allows assemblies to be filtered to remove contaminants identified in CVL plots.\n\n100-word article summaryWe describe a simple new method, Coverage-versus-Length plots, for examining de novo genome sequence assemblies. These plots enable researchers to detect scaffolds that have unusually high or unusually low coverage, which allows contaminants, and scaffolds that come from atypical parts of the organisms DNA complement, to be detected. We show that contaminants are common in yeast genomes sequenced in multiplex Illumina runs. We provide instructions for making plots using Microsoft Excel or Google Sheets, and software for filtering assemblies to remove contaminants. Contaminants can be detected and removed, even without knowing their source.

genomics↗

Improving recovery of member genomes from enrichment reactor microbial communities using MinION--based long read metagenomics

New long read sequencing technologies offer huge potential for effective recovery of complete, closed genomes. While much progress has been made on cultured isolates, the ability of these methods to recover genomes of member taxa in complex microbial communities is less clear. Here we examine the ability of long read data to recover genomes from enrichment reactor metagenomes. Such modified communities offer a moderate level of complexity compared to the source communities and so are realistic, yet tractable, systems to use for this problem. We sampled an enrichment bioreactor designed to target anaerobic ammonium-oxidising bacteria (AnAOB) and sequenced genomic DNA using both short read (Illumina 301bp PE) and long read data (MinION Mk1B) from the same extraction aliquot. The community contained 23 members, of which 16 had genome bins defined from an assembly of the short read data. Two distinct AnAOB species from genus Candidatus Brocadia were present and had complete genomes, of which one was the most abundant member species in the community. We can recover a 4Mb genome, in 2 contigs, of long read assembled sequence that is unambiguously associated with the most abundant AnAOB member genome. We conclude that obtaining near closed, complete genomes of members of low-medium microbial communities using MinION long read sequence is feasible.

genomics↗

Chromatin profiling of the repetitive and non-repetitive genome of the human fungal pathogen Candida albicans

BackgroundEukaryotic genomes are packaged into chromatin structures with pivotal roles in regulating all DNA-associated processes. Post-translational modifications of histone proteins modulate chromatin structure leading to rapid, reversible regulation of gene expression and genome stability which are key steps in environmental adaptation. Candida albicans is the leading fungal pathogen in humans, and can rapidly adapt and thrive in diverse host niches. The contribution of chromatin to C. albicans biology is largely unexplored.\n\nResultsHere, we harnessed genome-wide sequencing approaches to generate the first comprehensive chromatin profiling of histone modifications (H3K4me3, H3K9Ac, H4K16Ac and {gamma}-H2A) across the C. albicans genome and relate it to gene expression. We demonstrate that gene-rich non-repetitive regions are packaged in canonical euchromatin associated with histone modifications that mirror their transcriptional activity. In contrast, repetitive regions are assembled into distinct chromatin states: subtelomeric regions and the rDNA locus are assembled into canonical heterochromatin, while Major Repeat Sequences and transposons are packaged in chromatin bearing features of euchromatin and heterochromatin. Genome-wide mapping of {gamma}H2A, a marker of genome instability, allowed the identification of potential recombination-prone genomic sites. Finally, we present the first quantitative chromatin profiling in C. albicans to delineate the role of the chromatin modifiers Sir2 and Set1 in controlling chromatin structure and gene expression.\n\nConclusionsThis study presents the first genome-wide chromatin profiling of histone modifications associated with the C. albicans genome. These epigenomic maps provide an invaluable resource to understand the contribution of chromatin to C. albicans biology.

genomics↗

Genome-wide recombination map construction from single individuals using linked-read sequencing

Meiotic recombination is a major molecular mechanism generating genomic diversity. Recombination rates vary across the genome, often involving localized crossover \"hotspots\" and \"coldspots\". Studying the molecular basis and mechanism underlying this variation within and among individuals has been challenging due to the high cost and effort required to construct individualized genome-wide maps of recombination crossovers. In this study we introduce a new method to detect recombination crossovers across the genome from sperm DNA using Illumina sequencing of linked-read libraries produced using 10X Genomics technology. We leverage the long range information provided by the linked short reads to phase and assign haplotype states to each DNA molecule. When applied to DNA from gametes of a diploid organism, the majority of linked-read molecules can be used to faithfully reconstruct an individuals two haplotypes present at each location in the genome. A valuable rare fraction of molecules that span meiotic crossovers between the two chromosome haplotypes can then be isolated from the broader population of nonrecombinant molecules. Our pipeline, called ReMIX, allows us to characterize the genomic location and intensity of meiotic crossovers in a single individual and faithfully detects previously described recombination hotspots discovered by studies using mapping panels in mice. With a median crossover resolution of the mouse and stickleback being 15kb and 23kb respectively, ReMIX provides a powerful, high-throughput, low-cost approach to quantify recombination variation across the genome opening up numerous opportunities to study recombination variation with high genomic resolution in multiple individuals. ReMIX source code is available at at https://github.com/adreau/ReMIX.

genomics↗

Physiological changes during cellular ageing in fission yeast drive non-random patterns of genome rearrangements

Aberrant repair of DNA double-strand breaks can recombine distant pairs of chromosomal breakpoints. Such chromosomal rearrangements are a hallmark of ageing and compromise the structure and function of genomes. Rearrangements are challenging to detect in non-dividing cell populations, because they reflect individually rare, heterogeneous events. The genomic distribution of de novo rearrangements in non-dividing cells, and their dynamics during ageing, remain therefore poorly characterized. Studies of genomic instability during ageing have focussed on mitochondrial DNA, small genetic variants, or proliferating cells. To gain a better understanding of genome rearrangements during cellular ageing, we focused on a single diagnostic measure - DNA breakpoint junctions - allowing us to interrogate the changing genomic landscape in non-dividing cells of fission yeast (Schizosaccharomyces pombe). Aberrant DNA junctions that accumulated with age were associated with microhomology sequences and R-loops. Global hotspots for age-associated breakpoint formation were evident near telomeric genes and linked to remote breakpoints on the same or different chromosomes, including the mitochondrial chromosome. An unexpected mechanism of genomic instability caused more local hotspots: age-associated reduction in an RNA-binding protein could trigger R-loop formation at target loci. This finding suggests that biological processes other than transcription or replication can drive genome rearrangements. Notably, we detected similar signatures of genome rearrangements that accumulated in old brain cells of humans. These findings provide insights into the unique patterns and potential mechanisms of genome rearrangements in non-dividing cells, which can be triggered by ageing-related changes in gene-regulatory proteins.

genomics↗

The polyploid genome of the mitotic parthenogenetic root knot nematode Meloidogyne enterolobii

Root-knot nematodes (genus Meloidogyne) are plant parasitic species that cause huge economic loss in the agricultural industry and affect the prosperity of communities in developing countries. Control methods against these plant pests are sparse and the current preferred method is deployment of plant cultivars bearing resistance genes against Meloidogyne species. However, some species such as M. enterolobii are not controlled by the resistance genes deployed in the most important crop plants cultivated in Europe. The recent identification of this species in Europe is thus a major concern. Like the other most damaging Meloidogyne species (e.g. M. incognita, M. arenaria and M. javanica), M. enterolobii reproduces by obligatory mitotic parthenogenesis. Genomic singularities such as a duplicated genome structure and a relatively high proportion of transposable elements have previously been described in the above mentioned mitotic parthenogenetic Meloidogyne.\n\nTo gain a better understanding of the genomic and evolutionary background we sequenced the genome of M. enterolobii using high coverage short and long read technologies. The information contained in the long reads helped produce a highly contiguous genome assembly of M. enterolobii, thus enabling us to perform high quality annotations of coding and non-coding genes, and transposable elements.\n\nThe genome assembly and annotation reveals a genome structure similar to the ones described in the other mitotic parthenogenetic Meloidogyne, described as recent hybrids. Most of the genome is present in 3 different copies that show high divergence. Because most of the genes belong to these duplicated regions only few gene losses took place, which suggest a recent polyploidization. The most likely hypothesis to reconcile high divergence between genome copies despite few gene losses and translocations is also a recent hybrid origin. Consistent with this hypothesis, we found an abundance of transposable elements at least as high as the one observed in the mitotic parthenogenetic nematodes M. incognita and M. javanica.

genomics↗

From single nuclei to whole genome assembly

A large proportion of Earth's biodiversity constitutes organisms that cannot be cultured, have cryptic life-cycles and/or live submerged within their substrates1-4. Genomic data are key to unravel both their identity and function5. The development of metagenomic methods6,7 and the advent of single cell sequencing8-10 have revolutionized the study of life and function of cryptic organisms by upending the need for large and pure biological material, and allowing generation of genomic data from complex or limited environmental samples. Genome assemblies from metagenomic data have so far been restricted to organisms with small genomes, such as bacteria11, archaea12 and certain eukaryotes13. On the other hand, single cell technologies have allowed the targeting of unicellular organisms, attaining a better resolution than metagenomics8,9,14-16, moreover, it has allowed the genomic study of cells from complex organisms one cell at a time17,18. However, single cell genomics are not easily applied to multicellular organisms formed by consortia of diverse taxa, and the generation of specific workflows for sequencing and data analysis is needed to expand genomic research to the entire tree of life, including sponges19, lichens3,20, intracellular parasites21,22, and plant endophytes23,24. Among the most important plant endophytes are the obligate mutualistic symbionts, arbuscular mycorrhizal (AM) fungi, that pose an additional challenge with their multinucleate coenocytic mycelia25. Here, the development of a novel single nuclei sequencing and assembly workflow is reported. This workflow allows, for the first time, the generation of reference genome assemblies from large scale, unbiased sorted, and sequenced AM fungal nuclei circumventing tedious, and often impossible, culturing efforts. This method opens infinite possibilities for studies of evolution and adaptation in these important plant symbionts and demonstrates that reference genomes can be generated from complex non-model organisms by isolating only a handful of their nuclei.

genomics↗

A potential genomic recombination site upstream of the rfb locus in Leptospira interrogans is associated with serogroup Serjoe and serovar Hardjo classification

Leptospirosis is a zoonotic disease caused by pathogenic spirochetes of the genus Leptospira. It has a global distribution and affects domestic animals, including cattle. In livestock production, Leptospira interrogans serogroup Sejroe serovar Hardjo is the major reproductive disease leading to economic losses. The whole-genome sequence of the first Brazilian clinical isolate classified as L. interrogans serogroup Sejroe serovar Hardjo strain Norma enabled the evaluation of its genomic features. Here, we investigated particularities of this isolate, obtained from a leptospirosis outbreak. Bioinformatic analysis using the L. interrogans serovar Hardjo str. Norma was applied as a reference for genomic evaluation and comparative analysis among L. interrogans and L. borgpetersenii serovars. Our data suggest the occurrence of genomic recombination in L. interrogans serovar Hardjo str. Norma encompassing 45 Kb located upstream of the rfb locus. A hallmark of genetic evolution was predicted through an orthologue analysis that identified that sugar enzymes associated with carbohydrate and lipid biosynthesis and metabolism composed this genetic module. Comparative genomics revealed a wide range of relatedness among the bacterial strains of serogroup Sejroe that are classified as L. interrogans and L. borgpetersenii species. Furthermore, identification of an IS3 family suggests a genetic recombination site in L. interrogans serovar Hardjo str. Norma that is distinct among L. interrogans serovars and may contribute to clarify the taxonomic classification of Leptospira spp.\n\nImpact StatementLeptospirosis remains an important neglected disease with worldwide distribution. This zoonotic disease impacts in the livestock production and the bovine infection is currently associated to species L. borgpetersenii and L. interrogans serovar Hardjo. L.interrogans serovar Hardjo infection is recognized as reproductive disease associated with abortion and economic lost. In this context, we studied a unique whole genome sequence of L. interrogans serovar Hardjo subtype Hardjo-prajitno isolated from bovine leptospirosis outbreaks in Brazilian dairy farm, one of the greatest country of milk production in world. We compared L. interrogans and L. borgpetersenii genomes with L. interrogans serogroup Sejroe serovar Hardjo subtype Hardjo-prajitno focusing on rfb locus and sugars biosynthesis. Leptospira spp. taxonomy and serology information are strictly associated with rfb locus and we found high correlation among bacterial strains classified in serogroup Sejroe. Although L. interrogans and L. borgpetersenii classified in serogroup Sejroe possess a greater genetic correlation, we uniquely identified that serovar Hardjo strains often possess identical loci carrying predicted sugar biossinthesis genes and mobile elements. The Sru (Sejroe specific Rfb Upstream locus) locus associated to rfb locus probably contribute to Leptospira spp. genetic information concerning serogroup and serovar degrees of taxonomic and serology in this microbiology field.\n\nData SummaryAll Leptospira spp. genome sequences used in this study were retrieved from National Center for Biotechnology Information (NCBI) (Table 1) with NCBI ID: NZ_CP006723.1, NZ_CP012603.1, NC_004342.2, NZ_CP011934.1, NZ_AKXA02000040.1, NZ_CP013147.1, NZ_CP012029.1, NZ_CP015048.1, NC_008508.1, GCA_000216175.3,GCA_000244115.3, NC005823.1, GCA_000244395.3, GCA_000346975.1.\n\nO_TBL View this table:\norg.highwire.dtl.DTLVardef@843b1eorg.highwire.dtl.DTLVardef@1453917org.highwire.dtl.DTLVardef@1a75127org.highwire.dtl.DTLVardef@1c0e185org.highwire.dtl.DTLVardef@160aa4_HPS_FORMAT_FIGEXP M_TBL O_FLOATNOTable 1:C_FLOATNO O_TABLECAPTIONLeptospira spp. strains used in this study.\n\nC_TABLECAPTION C_TBL HighlightsO_LIThe study identified potential molecular features associated with serovar Hardjo subtype Hardjo-prajitno, present in both L. interrogans and L. borgpetersenii, by comparative genomics.\nC_LIO_LIA new potential recombinant site found upstream of the rfb locus contains proteins that correlate with serogroup taxonomy in the Leptospira genus.\nC_LIO_LIThe proteins encoded in the recombinant locus are associated with the synthesis of serological surface determinants such as carbohydrates and lipopolysaccharides.\nC_LI

genomics↗

Improved genome assembly and annotation of the soybean aphid (Aphis glycines Matsumura)

Aphids are an economically important insect group due to their role as plant disease vectors. Despite this economic impact, genomic resources have only been generated for a small number of aphid species. The soybean aphid (Aphis glycines Matsumura) was the third aphid species to have its genome sequenced and the first to use long-read sequence data. However, version 1 of the soybean aphid genome assembly has low contiguity (contig N50 = 57 KB, scaffold N50 = 174 KB), poor representation of conserved genes and the presence of genomic scaffolds likely derived from parasitoid wasp contamination. Here, I use recently developed methods to reassemble the soybean aphid genome. The version 2 genome assembly is highly contiguous, containing half of the genome in only 40 scaffolds (contig N50 = 2.00 Mb, scaffold N50 = 2.51 Mb) and contains 11% more conserved single copy arthropod genes than version 1. To demonstrate the utility this improved assembly, I identify a region of conserved synteny between aphids and Drosophila containing members of the Osiris gene family that was split over multiple scaffolds in the original assembly. The improved genome assembly and annotation of A. glycines demonstrates the benefit of applying new methods to old data sets and will provide a useful resource for future comparative genome analysis of aphids.

genomics↗