bioRxiv Science⌕ Search

SEARCH · bioRxiv Science

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,567 records · Page 87Linked to original sources

Genomes of novel Myxococcota reveal severely curtailed machineries for predation and cellular differentiation

Cultured Myxococcota are predominantly aerobic soil inhabitants, characterized by their highly coordinated predation and cellular differentiation capacities. Little is currently known regarding yet-uncultured Myxococcota from anaerobic, non-soil habitats. We analyzed genomes representing one novel order (o__JAFGXQ01) and one novel family (f__JAFGIB01) in the Myxococcota from an anoxic freshwater spring in Oklahoma, USA. Compared to their soil counterparts, anaerobic Myxococcota possess smaller genomes, and a smaller number of genes encoding biosynthetic gene clusters (BGCs), peptidases, one- and two-component signal transduction systems, and transcriptional regulators. Detailed analysis of thirteen distinct pathways/processes crucial to predation and cellular differentiation revealed severely curtailed machineries, with the notable absence of homologs for key transcription factors (e.g. FruA and MrpC), outer membrane exchange receptor (TraA), and the majority of sporulation-specific and A-motility-specific genes. Further, machine-learning approaches based on a set of 634 genes informative of social lifestyle predicted a non-social behavior for Zodletone Myxococcota. Metabolically, Zodletone Myxococcota genomes lacked aerobic respiratory capacities, but encoded genes suggestive of fermentation, dissimilatory nitrite reduction, and dissimilatory sulfate-reduction (in f_JAFGIB01) for energy acquisition. We propose that predation and cellular differentiation represent a niche adaptation strategy that evolved circa 500 Mya in response to the rise of soil as a distinct habitat on earth. ImportanceThe Myxococcota is a phylogenetically coherent bacterial lineage that exhibits unique social traits. Cultured Myxococcoat are predominantly aerobic soil-dwelling microorganisms that are capable of predation and fruiting body formation. However, multiple yet-uncultured lineages within the Myxococcota has been encountered in a wide range of non-soil, predominantly anaerobic habitats; and the metabolic capabilities, physiological preferences, and capacity of social behavior of such lineages remains unclear. Here, we analyzed genomes recovered from a metagenomic analysis of an anoxic freshwater spring in Oklahoma, USA that represent novel, yet-uncultured, orders and families in the Myxococcota. The genomes appear to lack the characteristic hallmarks for social behavior encountered in Myxococcota genomes, and displayed a significantly smaller genome size and a smaller number of genes encoding biosynthetic gene clusters, peptidases, signal transduction systems, and transcriptional regulators. Such perceived lack of social capacity we confirmed through detailed comparative genomic analysis of thirteen pathways associated with Myxococcota social behavior, as well as the implementation of machine learning approaches to predict social behavior based on genome composition. Metabolically, these novel Myxococcota are predicted to be strict anaerobes, utilizing fermentation, nitrate rductio, and dissimilarity sulfate reduction for energy acquisition. Our result highlight the broad patterns of metabolic diversity within the yet-uncultured Myxococcota and suggest that the evolution of predation and fruiting body formation in the Myxococcoat has occurred in response to soil formation as a distinct habitat on earth.

microbiology↗

CoVizu: Rapid analysis and visualization of the global diversity of SARS-CoV-2 genomes

Phylogenetics has played a pivotal role in the genomic epidemiology of SARS-CoV-2, such as tracking the emergence and global spread of variants, and scientific communication. However, the rapid accumulation of genomic data from around the world -- with over two million genomes currently available in the GISAID database -- is testing the limits of standard phylogenetic methods. Here, we describe a new approach to rapidly analyze and visualize large numbers of SARS-CoV-2 genomes. Using Python, genomes are filtered for problematic sites, incomplete coverage, and excessive divergence from a strict molecular clock. All differences from the reference genome, including indels, are extracted using minimap2, and compactly stored as a set of features for each genome. For each Pango lineage (https://cov-lineages.org), we collapse genomes with identical features into variants, generate 100 bootstrap samples of the feature set union to generate weights, and compute the symmetric differences between the weighted feature sets for every pair of variants. The resulting distance matrices are used to generate neigihbor-joining trees in RapidNJ and converted into a majority-rule consensus tree for the lineage. Branches with support values below 50% or mean lengths below 0.5 differences are collapsed, and tip labels on affected branches are mapped to internal nodes as directly-sampled ancestral variants. Currently, we process about million genomes in approximately nine hours on 34 cores. The resulting trees are visualized using the JavaScript framework D3.js as beadplots, in which variants are represented by horizontal line segments, annotated with beads representing samples by collection date. Variants are linked by vertical edges to represent branches in the consensus tree. These visualizations are published at https://filogeneti.ca/CoVizu. All source code was released under an MIT license at https://github.com/PoonLab/covizu.

bioinformatics↗

Patterns of gene co-expression under water-deficit treatments and pan-genome occupancy in Brachypodium distachyon.

Natural populations are characterized by abundant genetic diversity driven by a range of different types of mutation. The tractability of sequencing complete genomes has allowed new insights into the variable composition of genomes, summarized as a species pan-genome. These analyses demonstrate that many genes are absent from the first reference genomes, whose analysis dominated the initial years of the genomic era. Our field now turns towards understanding the functional consequence of these highly variable genomes. Here, we analyzed weighted gene co-expression networks from leaf transcriptome data for drought response in the purple false brome Brachypodium distachyon and the differential expression of genes putatively involved in adaptation to this stressor. We specifically asked whether genes with variable "occupancy" in the pan-genome - genes which are either present in all studied genotypes or missing in some genotypes - show different distributions among co-expression modules. Co-expression analysis united genes expressed in drought-stressed plants into nine modules covering 72 hub genes (87 hub isoforms), and genes expressed under controlled water conditions into 13 modules, covering 190 hub genes (251 hub isoforms). We find that low occupancy pan-genes are under-represented among several modules, while other modules are over-enriched for low-occupancy pan-genes. We also provide new insight into the regulation of drought response in B. distachyon, specifically identifying one module with an apparent role in primary metabolism that is strongly responsive to drought. Our work shows the power of integrating pan-genomic analysis with transcriptomic data using factorial experiments to understand the functional genomics of environmental response.

plant biology↗

Trait-trait relationships and functional tradeoffs vary with genome size in prokaryotes

We report genomic traits that have been associated with the life history of prokaryotes and highlight conflicting findings concerning earlier observed trait correlations and tradeoffs. In order to address possible explanations for these contradictions we examined trait-trait variations of 11 genomic traits from ~ 18,000 sequenced genomes. The studied trait-trait variations suggested: (i) the predominance of two resistance and resilience-related orthogonal axes and (ii) at least in free living species with large effective population sizes whose evolution is little affected by genetic drift an overlap between a resilience axis and an axis of resource usage efficiency. These findings imply that resistance associated traits of prokaryotes are globally decoupled from resilience related traits and in the case of free-living communities also from resource use efficiencies associated traits. However, further inspection of pairwise scatterplots showed that resistance and resilience traits tended to be positively related for genomes up to roughly five million base pairs and negatively for larger genomes. This in turn may preclude a globally consistent assignment of prokaryote genomic traits to the competitor - stress-tolerator - ruderal (CSR) schema that sorts species depending on their location along disturbance and productivity gradients into three ecological strategies and may serve as an explanation for conflicting findings from earlier studies. All reviewed genomic traits featured significant phylogenetic signals and we propose that our trait table can be applied to extrapolate genomic traits from taxonomic marker genes. This will enable to empirically evaluate the assembly of these genomic traits in prokaryotic communities from different habitats and under different productivity and disturbance scenarios as predicted via the resistance-resilience framework formulated here.

ecology↗

The miR-430 locus with extreme promoter density is a transcription body organizer, which facilitates long range regulation in zygotic genome activation

In anamniote embryos the major wave of zygotic genome activation (ZGA) starts during the mid-blastula transition. This major wave of ZGA is facilitated by several mechanisms, including dilution of repressive maternal factors and accumulation of activating transcription factors during the fast cell division cycles preceding the mid-blastula transition. However, a set of genes escape global genome repression and are activated substantially earlier, during what is called, the minor wave of genome activation. While the mechanisms underlying the major wave of genome activation have been studied extensively, the minor wave of genome activation is little understood. In zebrafish the earliest expressed RNA polymerase II (Pol II) transcribed genes are activated in a pair of large transcription bodies depleted of chromatin, abundant in elongating Pol II and nascent RNAs (Hadzhiev et al., 2019; Hilbert et al., 2021). This transcription body includes the miR-430 gene cluster required for maternal mRNA clearance. Here we explored the genomic, chromatin organisation and cis-regulatory mechanisms of the minor wave of genome activation occurring in the transcription body. By long read genome sequencing we identified a remarkable cluster of miR-430 genes with over 300 promoters and spanning 0.6 Mb, which represent the highest promoter density of the genome. We demonstrate that the miR-430 gene cluster is required for the formation of the transcription body and acts as a transcription organiser for minor wave activation of a set of zinc finger genes scattered on the same chromosome arm, which share promoter features with the miR-430 cluster. These promoter features are shared among minor wave genes overall and include the TATA-box and sharp transcription start site profile. Single copy miR-430 promoter transgene reporter experiments indicate the importance of promoter-autonomous mechanisms regulating escape from global repression of the early embryo. These results together suggest that formation of the transcription body in the early embryo is the result of high promoter density coupled to a minor wave-specific core promoter code for transcribing key minor wave ZGA genes, which are required for the overhaul of the transcriptome during early embryonic development.

developmental biology↗

The OceanDNA MAG catalog contains over 50,000 prokaryotic genomes originated from various marine environments

Marine microorganisms are immensely diverse and play fundamental roles in global geochemical cycling. Recent metagenome-assembled genome studies, with special attention to large-scale projects such as Tara Oceans, have expanded the genomic repertoire of marine microorganisms. However, published marine metagenome data has not been fully explored yet. Here, we collected 2,057 marine metagenomes (>29 Tera bps of sequences) covering various marine environments and developed a new genome reconstruction pipeline. We reconstructed 52,325 qualified genomes composed of 8,466 prokaryotic species-level clusters spanning 59 phyla, including genomes from deep-sea deeper than 1,000 m (n=3,337), low-oxygen zones of <90 mol O2 per kg water (n=7,884), and polar regions (n=7,752). Novelty evaluation using a genome taxonomy database shows that 6,256 species (73.9%) are novel and include genomes of high taxonomic novelty such as new class candidates. These genomes collectively expanded the known phylogenetic diversity of marine prokaryotes by 34.2% and the species representatives cover 26.5 - 42.0% of prokaryote-enriched metagenomes. This genome resource, thoroughly leveraging accumulated metagenomic data, illuminates uncharacterized marine microbial dark matter lineages.

microbiology↗

A bacterial genome and culture collection of gut microbial in weanling piglet

The microbiota hosted in the pig gastrointestinal tract are important for productivity of livestock. However, the individual species and functional repertoires that make up the pig gut microbiome remain largely undefined. Here we comprehensively investigated the genomes and functions of the piglet gut microbiome using culture-based and metagenomics approaches. A collection included 266 cultured genomes and 482 metagenome-assembled genomes (MAGs) that were clustered to 428 species across 10 phyla was established. Among these clustered species, 333 genomes represent potential new species. Less matches between cultured genomes and MAGs revealed a substantial bias for the acquisition of reference genomes by the two strategies. Glycoside hydrolases was the dominant category of carbohydrate-active enzymes. 445 secondary metabolite biosynthetic genes were predicted from 292 genomes with bacteriocin being the most. Pan genome analysis of Limosilactobacillus reuteri uncover the biosynthesis of reuterin was strain-specific and the production was experimentally determined. These genomic resources will enable a comprehensive characterization of the microbiome composition and function of pig gut.

microbiology↗

Conserved genomic landscapes of differentiation across Populus speciation continuum

Speciation, the continuous process by which new species form, is often investigated by looking at the variation of nucleotide diversity and differentiation across the genome (hereafter genomic landscapes). A key challenge lies in how to determine the main evolutionary forces at play shaping these patterns. One promising strategy, albeit little used to date, is to comparatively investigate these genomic landscapes as a progression through time by using a series of species pairs along a divergence gradient. Here, we resequenced 201 whole-genomes from eight closely related Populus species, with pairs of species at different stages along the divergence gradient to learn more about speciation processes. Using population structure and ancestry analyses, we document extensive introgression between some species pairs, especially those with parapatric distributions. We further investigate genomic landscapes, focusing on within-species (nucleotide diversity and recombination rate) and among-species (relative and absolute divergence) summary statistics of diversity and divergence. We observe highly conserved patterns of genomic divergence across species pairs. Independent of the stage across the divergence gradient, we find support for signatures of linked selection (i.e., the interaction between natural selection and genetic linkage) in shaping these genomic landscapes, along with gene flow and standing genetic variation. We highlight the importance of investigating genomic patterns on multiple species across a divergence gradient and discuss prospects to better understand the evolutionary forces shaping the genomic landscapes of diversity and differentiation.

evolutionary biology↗

Short '1.2x genome' infectious clone initiates deltavirus replication in Boa constrictor cells

Human hepatitis D virus (HDV), discovered in 1977, represented the sole known deltavirus for decades. The dependence on hepatitis B virus (HBV) co-infection and its glycoproteins for infectious particle formation led to the assumption that deltaviruses are human-only pathogens. However, since 2018, several reports have described identification of HDV-like agents from various hosts but without co-infecting hepadnaviruses. Indeed, we demonstrated that Swiss snake colony virus 1 (SwSCV-1) uses arenaviruses as the helper for infectious particle formation, thus shaking the dogmatic alliance with hepadnaviruses for completing deltavirus life cycle. In vitro systems enabling helper virus-independent replication are key for studying the newly discovered deltaviruses. Others and we have successfully used constructs containing multimers of the deltavirus genome for the replication of various deltaviruses via transfection in cell culture. Here, we report the establishment of deltavirus infectious clones with 1.2x genome inserts bearing two copies of the genomic and antigenomic ribozymes. We used SwSCV-1 as the model to compare the ability of the previously reported "2x genome" and the "1.2x genome" plasmid constructs/infectious clones to initiate replication in cell culture. Using immunofluorescence, qRT-PCR, immuno- and northern blotting, we found the 2x and 1.2x genome clones to similarly initiate deltavirus replication in vitro and both induced a persistent infection of snake cells. We hypothesize that duplicating the ribozymes facilitates the cleavage of genome multimers into unit-length pieces during the initial round of replication. The 1.2x genome constructs enable easier introduction of modifications required for studying deltavirus replication and cellular interactions. IMPORTANCEHepatitis D virus (HDV) is a satellite virus infecting humans with strict association to hepatitis B virus (HBV) co-infection because HBV glycoproteins can mediate infectious HDV particle formation. For decades, HDV was the sole representative of deltaviruses, which had led to hypotheses suggesting that it evolved in humans, the only known natural host. Recent sequencing studies have led to the discovery of HDV-like sequences across a wide range of species, representing a paradigm shift in deltavirus evolution. Molecular biology tools such as infectious clones, which enable initiation of deltavirus infection without helper virus, are key to demonstrate that the recently found deltaviruses are capable of independent replication. Such tools will enable identification of the potential helper viruses. Here, we report a 1.2x genome copy strategy for designing plasmid-based infectious clones to study deltaviruses and to demonstrate that plasmid delivery into cultured snake cells sufficiently initiates replication of different deltaviruses.

molecular biology↗

Deep mining of the Sequence Read Archive reveals bipartite coronavirus genomes and inter-family Spike glycoprotein recombination

Genetic variation in RNA viruses is generated by point mutation and recombination as well as reassortment in the case of viruses with segmented genomes. While point mutation concerns only few sites per genome copy, recombination and reassortment can affect large genome regions, possibly facilitating the sudden emergence of novel traits. The contribution of recombination and reassortment to genomic plasticity and their rates remain poorly understood and might be underappreciated because of the lack of a comprehensive description of the virosphere. Here we employed a computational approach that directly queries primary sequencing data in a highly parallelized way and involves a targeted viral genome assembly strategy. By screening more than 213,000 data sets from the Sequence Read Archive repository and using two metrics that quantitatively assess assembly quality we discovered 25 novel nidoviruses from a wide range of vertebrate hosts. These include eight fish coronaviruses with bipartite genomes, a giant 36.1 kilobase coronavirus genome with a duplicated Spike glycoprotein (S) gene, and 16 additional so far undescribed vertebrate nidoviruses. Some of these novel virus genomes encode protein domains that have not been described for nidoviruses. We provide evidence for a possible inter-family homologous recombination event involving S between ancestral bipartite coronaviruses and unsegmented tobaniviruses and report a case example of an individual fish simultaneously infected with members from both virus families. Our results shed light on the evolution and genomic plasticity of coronaviruses and identify recombinants with a possibly improved ability to cross species barriers, which might elevate their pandemic potential.

microbiology↗

A new lineage of non-photosynthetic green algae with extreme organellar genomes

BackgroundThe plastid genomes of the green algal order Chlamydomonadales tend to expand their non-coding regions, but this phenomenon is poorly understood. Here we shed new light on organellar genome evolution in Chlamydomonadales by studying a previously unknown non-photosynthetic lineage. We established cultures of two new Polytoma-like flagellates, defined their basic characteristics and phylogenetic position, and obtained complete organellar genome sequences and a transcriptome assembly for one of them. ResultsWe discovered a novel deeply diverged chlamydomonadalean lineage that has no close photosynthetic relatives and represents an independent case of photosynthesis loss. To accommodate these organisms we establish the new genus Leontynka, with two species (L. pallida and L. elongata) distinguishable through both their morphological and molecular characteristics. Notable features of the colourless plastid of L. pallida deduced from the plastid genome (plastome) sequence and transcriptome assembly include the retention of ATP synthase, thylakoid-associated proteins, the carotenoid biosynthesis pathway, and a plastoquinone-based electron transport chain, the latter two modules having an obvious functional link to the eyespot present in Leontynka. Most strikingly, the ~362 kbp plastome of L. pallida is by far the largest among the non-photosynthetic eukaryotes investigated to date due to an extreme proliferation of sequence repeats. These repeats are also present in coding sequences, with one repeat type found in the exons of 11 out of 34 protein-coding genes, with up to 36 copies per gene, thus affecting the encoded proteins. The mitochondrial genome of L. pallida is likewise exceptionally large, with its >104 kbp surpassed only by the mitogenome of Haematococcus lacustris among all members of Chlamydomonadales hitherto studied. It is also bloated with repeats, though entirely different from those in the L. pallida plastome, which contrasts with the situation in H. lacustris where both the organellar genomes have accumulated related repeats. Furthermore, the L. pallida mitogenome exhibits an extremely high GC content in both coding and non-coding regions and, strikingly, a high number of predicted G-quadruplexes. ConclusionsWith its unprecedented combination of plastid and mitochondrial genome characteristics, Leontynka pushes the frontiers of organellar genome diversity and is an interesting model for studying organellar genome evolution.

evolutionary biology↗

Rapid genomic evolution in Brassica rapa with bumblebee selection in experimental evolution

BackgroundInsect pollinators shape rapid phenotypic evolution of traits related to floral attractiveness and plant reproductive success. However, the underlying genomic changes remain largely unknown despite their importance in predicting adaptive responses to natural or to artificial selection. Based on a nine-generation experimental evolution study with fast cycling Brassica rapa plants adapting to bumblebees, we investigate the genomic evolution associated with the previously observed parallel phenotypic evolution. In this current evolve and resequencing (E&R) study, we conduct a genomic scan of the allele frequency changes along the genome in bumblebee-pollinated and hand-pollinated plants and perform a genomic principal component analysis (PCA). ResultsWe highlight rapid genomic evolution associated with the observed phenotypic evolution mediated by bumblebees. Controlling for genetic drift, we observe significant changes in allelic frequencies at multiple loci. However, this pattern differs according to the replicate of bumblebee-pollinated plants, suggesting putative non-parallel genomic evolution. Finally, our study underlines an increase in genomic differentiation implying the putative involvement of multiple loci in short-term pollinator adaptation. ConclusionsOverall, our study enhances our understanding of the complex interactions between pollinator and plants, providing a steppingstone towards unravelling the genetic basis of plant genomic adaptation to biotic factors in the environment.

evolutionary biology↗

binny: an automated binning algorithm to recover high-quality genomes from complex metagenomic datasets

The reconstruction of genomes is a critical step in genome-resolved metagenomics and for multi-omic data integration from microbial communities. Here, we present binny, a binning tool that produces complete and pure metagenome-assembled genomes (MAG) from both contiguous and highly fragmented genomes. Based on established metrics, binny outperforms or is highly competitive with commonly-used and state- of-the-art binning methods and finds unique genomes that could not be detected by other methods. binny uses k-mer-composition and coverage by metagenomic reads for iterative, non-linear dimension reduction of genomic signatures, as well as subsequent automated contig clustering with cluster assessment using lineage-specific marker gene sets. When compared to seven widely used binning algorithms, binny provides substantial amounts of uniquely identified MAGs and almost always recovers the most near-complete (>95% pure, >90% complete) and high-quality (>90% pure, >70% complete) genomes from simulated data sets from the Critical Assessment of Metagenome Interpretation (CAMI) initiative, as well as substantially more high-quality draft genomes, as defined by the Minimum Information about a Metagenome-Assembled Genome (MIMAG) standard, from a real-world benchmark comprised of metagenomes from various environments than any other tested method.

bioinformatics↗

Putative host-derived insertions in the genome of circulating SARS-CoV-2 variants

Insertions in the SARS-CoV-2 genome have the potential to drive viral evolution, but the source of the insertions is often unknown. Recent proposals have suggested that human RNAs could be a source of some insertions, but the small size of many insertions makes this difficult to confirm. Through an analysis of available direct RNA sequencing data from SARS-CoV-2 infected cells, we show that viral-host chimeric RNAs are formed through what are likely stochastic RNA-dependent RNA polymerase template switching events. Through an analysis of the publicly available GISAID SARS-CoV-2 genome collection, we identified two genomic insertions in circulating SARS-CoV-2 variants that are identical to regions of the human 18S and 28S rRNAs. These results provide direct evidence of the formation of viral-host chimeric sequences and the integration of host genetic material into the SARS-CoV-2 genome, highlighting the potential importance of host-derived insertions in viral evolution. IMPORTANCEThroughout the COVID-19 pandemic, the sequencing of SARS-CoV-2 genomes has revealed the presence of insertions in multiple globally circulating lineages of SARS-CoV-2, including the Omicron variant. The human genome has been suggested to be the source of some of the larger insertions, but evidence for this kind of event occurring is still lacking. Here, we leverage direct RNA sequencing data and SARS-CoV-2 genomes to show host-viral chimeric RNAs are generated in infected cells and two large genomic insertions have likely been formed through the incorporation of host rRNA fragments into the SARS-CoV-2 genome. These host-derived insertions may increase the genetic diversity of SARS-CoV-2 and expand its strategies to acquire genetic materials, potentially enhancing its adaptability, virulence, and spread.

bioinformatics↗

The salmon louse genome may be much larger than sequencing suggests

The genome size of organisms impacts their evolution and biology and is often assumed to be characteristic of a species. Here we present the first published estimates of genome size of the ecologically and economically important ectoparasite, Lepeophtheirus salmonis (Copepoda, Caligidae). Four independent L. salmonis genome assemblies of the North Atlantic subspecies Lepeophtheirus salmonis salmonis, including two chromosome level assemblies, yield assemblies ranging from 665 - 790 Mbps. These genome assemblies are congruent in their findings, and appear very complete with Benchmarking Universal Single-Copy Orthologs analyses finding >92% of expected genes and transcriptome datasets routinely mapping >90% of reads. However, two cytometric techniques, flow cytometry and Feulgen image analysis densitometry, yield measurements of 1.3-1.6 Gb in the haploid genome. Interestingly, earlier cytometric measurements reported genome sizes of 939 and 567 Mbps in L. salmonis salmonis samples from Bay of Fundy and Norway, respectively. Available data thus suggest that the genome sizes of salmon lice are variable. Current understanding of eukaryotic genome dynamics suggests that the most likely explanation for such variability involves repetitive DNA, which for L. salmonis makes up {approx}60% of the genome assemblies.

genetics↗

Genomic library of Bordetella

BackgroundThe re-emergence of whooping cough and geographic disparities in vaccine escape or antimicrobial resistance dynamics, underline the importance of a unified definition of Bordetella pertussis strains. Understanding of the evolutionary adaptations of Bordetella pathogens to humans and animals requires comparative studies with environmental bordetellae. MethodsWe have set-up a unified library of Bordetella genomes by merging previously existing Oxford and Pasteur databases, importing genomes from public repositories, and developing harmonized genotyping schemes. We developed a genus-wide cgMLST genotyping scheme and incorporated a previous B. pertussis cgMLST scheme. Specific schemes were developed to define antigenic, virulence and macrolide resistance profiles. Genomic sequencing of 83 French B. bronchiseptica isolates and of B. tumulicola, B. muralis and B. tumbae type strains was performed. ResultsThe public library currently includes 2,581 Bordetella isolates and their provenance data, and 2,084 genomes. The "classical Bordetella" (B. bronchiseptica, B. parapertussis and B. pertussis), which form a single genomic species (B. bronchiseptica genomic species, BbGS), were overrepresented (n=2,382). The phylogenetic analysis of Bordetella genomes associated the three novel species B. tumulicola, B. muralis and B. tumbae in a clade with B. petrii and revealed 18 yet undescribed species. A sister lineage of the classical bordetellae, provisionally named Bbs lineage II, was uncovered and may represent a novel species (average nucleotide identity with BbGS strains: [~]95%). It comprised strain HT200 from India, two strains of genogroup 6 from the USA and six clinical isolates from France; this lineage lacked ptxP and its fim2 gene was divergent. Within B. pertussis, vaccine antigen sequence types marked important phylogenetic subdivisions, and macrolide resistance markers (23S_rRNA allele 13 and fhaB3) confirmed the current restriction of this phenotype in China with few exceptions. ConclusionsThe genomic platform provides an expandable resource for unified genotyping of Bordetella strains and will facilitate collective evolutionary and epidemiological understanding of the re-emergence of whooping cough and other Bordetella infections. Data summaryBordetella genomes list and accession numbers: Supplementary Table S4 Bordetella genus phylogeny dataset (92 isolates): https://bigsdb.pasteur.fr/cgi-bin/bigsdb/bigsdb.pl?db=pubmlst_bordetella_isolates&page=query&project_list=23&submit=1 B. bronchiseptica phylogeny dataset (213 isolates): https://bigsdb.pasteur.fr/cgi-bin/bigsdb/bigsdb.pl?db=pubmlst_bordetella_isolates&page=query&project_list=24&submit=1 B. pertussis phylogeny (124 isolates): https://bigsdb.pasteur.fr/cgi-bin/bigsdb/bigsdb.pl?db=pubmlst_bordetella_isolates&page=query&project_list=25&submit=1 iTOL interactive trees: https://itol.embl.de/shared/1l7Fw0AvKOoCF

microbiology↗

Regional mutational signature activities in cancer genomes

Cancer genomes harbor a catalog of somatic mutations. The type and genomic context of these mutations depend on their causes, and allow their attribution to particular mutational signatures. Previous work has shown that mutational signature activities change over the course of tumor development, but investigations of genomic region variability in mutational signatures have been limited. Here, we expand upon this work by constructing regional profiles of mutational signature activities over 2,203 whole genomes across 25 tumor types, using data aggregated by the Pan-Cancer Analysis of Whole Genomes (PCAWG) consortium. We present GenomeTrackSig as an extension to the TrackSig R package to construct regional signature profiles using optimal segmentation and the expectation-maximization (EM) algorithm. We find that 426 genomes from 20 tumor types display at least one change in mutational signature activities (changepoint), and 257 genomes contain at least one of 54 recurrent changepoints shared by seven or more genomes of the same tumor type. Five recurrent changepoint locations are shared by multiple tumor types. Within these regions, the particular signature changes are often consistent across samples of the same type and some, but not all, are characterized by signatures associated with subclonal expansion. The changepoints we found cannot strictly be explained by gene density, mutation density, or cell-of-origin chromatin state. We hypothesize that they reflect a confluence of factors including evolutionary timing of mutational processes, regional differences in somatic mutation rate, large-scale changes in chromatin state that may be tissue type-specific, and changes in chromatin accessibility during subclonal expansion. These results provide insight into the regional effects of DNA damage and repair processes, and may help us localize genomic and epigenomic changes that occur during cancer development.

cancer biology↗

An ancient endogenous DNA virus in the human genome

The genomes of eukaryotes preserve a striking diversity of ancient viruses in the form of endogenous viral elements (EVEs). Study of this genomic fossil record provides insights into the diversity, origin and evolution of viruses across geological timescales. In particular, Mavericks have emerged as one of the oldest groups of viruses infecting vertebrates ([&ge;]419 My). They have been found in the genomes of fish, amphibians and non-avian reptiles but had been overlooked in mammals. Thus, their evolutionary history and the causes of their demise in mammals remain puzzling questions. Here, we conduct a detailed evolutionary study of two Maverick-like integrations found on human chromosomes 7 and 8. We performed a comparative analysis of the integrations and determined their orthology across placental mammals (Eutheria) via the syntenic arrangement of neighbouring genes. The integrations were absent at the orthologous sites in the genomes of marsupials and monotremes. These observations allowed us to reconstruct a time-calibrated phylogeny and infer the age of their most recent common ancestor at 268.61 (199.70-344.54) My. In addition, we estimate the age of the individual integrations at ~105 My which represent the oldest non-retroviral EVEs found in the human genome. Our findings suggest that active Mavericks existed in the ancestors of modern mammals ~172 My ago (Jurassic Period) and potentially to the end of the Early Cretaceous. We hypothesise Mavericks could have gone extinct in mammals from the evolution of an antiviral defence system or from reduced opportunities for transmission in terrestrial hosts. ImportanceThe genomes of vertebrates preserve an enormous diversity of endogenous viral elements (remnants of ancient viruses that accumulate in host genomes over evolutionary time). Although retroviruses account for the vast majority of these elements, diverse DNA viruses have also been found and novel lineages are being described. Here we analyse two elements found in the human genome belonging to an ancient group of DNA viruses called Mavericks. We study their evolutionary history, finding that the elements are shared between humans and many different species of placental mammals. These observations suggest the elements inserted at least ~105 Mya in the most recent common ancestor of placentals. We further estimate the age of the viral ancestor around 268 My. Our results provide evidence for some of the oldest viral integrations in the human genome and insights into the ancient interactions of viruses with the ancestors of modern-day mammals.

evolutionary biology↗