bioRxiv Science⌕ Search

SEARCH · bioRxiv Science

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,801 records · Page 100Linked to original sources

Putative host-derived insertions in the genome of circulating SARS-CoV-2 variants

Insertions in the SARS-CoV-2 genome have the potential to drive viral evolution, but the source of the insertions is often unknown. Recent proposals have suggested that human RNAs could be a source of some insertions, but the small size of many insertions makes this difficult to confirm. Through an analysis of available direct RNA sequencing data from SARS-CoV-2 infected cells, we show that viral-host chimeric RNAs are formed through what are likely stochastic RNA-dependent RNA polymerase template switching events. Through an analysis of the publicly available GISAID SARS-CoV-2 genome collection, we identified two genomic insertions in circulating SARS-CoV-2 variants that are identical to regions of the human 18S and 28S rRNAs. These results provide direct evidence of the formation of viral-host chimeric sequences and the integration of host genetic material into the SARS-CoV-2 genome, highlighting the potential importance of host-derived insertions in viral evolution. IMPORTANCEThroughout the COVID-19 pandemic, the sequencing of SARS-CoV-2 genomes has revealed the presence of insertions in multiple globally circulating lineages of SARS-CoV-2, including the Omicron variant. The human genome has been suggested to be the source of some of the larger insertions, but evidence for this kind of event occurring is still lacking. Here, we leverage direct RNA sequencing data and SARS-CoV-2 genomes to show host-viral chimeric RNAs are generated in infected cells and two large genomic insertions have likely been formed through the incorporation of host rRNA fragments into the SARS-CoV-2 genome. These host-derived insertions may increase the genetic diversity of SARS-CoV-2 and expand its strategies to acquire genetic materials, potentially enhancing its adaptability, virulence, and spread.

bioinformatics↗

The salmon louse genome may be much larger than sequencing suggests

The genome size of organisms impacts their evolution and biology and is often assumed to be characteristic of a species. Here we present the first published estimates of genome size of the ecologically and economically important ectoparasite, Lepeophtheirus salmonis (Copepoda, Caligidae). Four independent L. salmonis genome assemblies of the North Atlantic subspecies Lepeophtheirus salmonis salmonis, including two chromosome level assemblies, yield assemblies ranging from 665 - 790 Mbps. These genome assemblies are congruent in their findings, and appear very complete with Benchmarking Universal Single-Copy Orthologs analyses finding >92% of expected genes and transcriptome datasets routinely mapping >90% of reads. However, two cytometric techniques, flow cytometry and Feulgen image analysis densitometry, yield measurements of 1.3-1.6 Gb in the haploid genome. Interestingly, earlier cytometric measurements reported genome sizes of 939 and 567 Mbps in L. salmonis salmonis samples from Bay of Fundy and Norway, respectively. Available data thus suggest that the genome sizes of salmon lice are variable. Current understanding of eukaryotic genome dynamics suggests that the most likely explanation for such variability involves repetitive DNA, which for L. salmonis makes up {approx}60% of the genome assemblies.

genetics↗

Genomic library of Bordetella

BackgroundThe re-emergence of whooping cough and geographic disparities in vaccine escape or antimicrobial resistance dynamics, underline the importance of a unified definition of Bordetella pertussis strains. Understanding of the evolutionary adaptations of Bordetella pathogens to humans and animals requires comparative studies with environmental bordetellae. MethodsWe have set-up a unified library of Bordetella genomes by merging previously existing Oxford and Pasteur databases, importing genomes from public repositories, and developing harmonized genotyping schemes. We developed a genus-wide cgMLST genotyping scheme and incorporated a previous B. pertussis cgMLST scheme. Specific schemes were developed to define antigenic, virulence and macrolide resistance profiles. Genomic sequencing of 83 French B. bronchiseptica isolates and of B. tumulicola, B. muralis and B. tumbae type strains was performed. ResultsThe public library currently includes 2,581 Bordetella isolates and their provenance data, and 2,084 genomes. The "classical Bordetella" (B. bronchiseptica, B. parapertussis and B. pertussis), which form a single genomic species (B. bronchiseptica genomic species, BbGS), were overrepresented (n=2,382). The phylogenetic analysis of Bordetella genomes associated the three novel species B. tumulicola, B. muralis and B. tumbae in a clade with B. petrii and revealed 18 yet undescribed species. A sister lineage of the classical bordetellae, provisionally named Bbs lineage II, was uncovered and may represent a novel species (average nucleotide identity with BbGS strains: [~]95%). It comprised strain HT200 from India, two strains of genogroup 6 from the USA and six clinical isolates from France; this lineage lacked ptxP and its fim2 gene was divergent. Within B. pertussis, vaccine antigen sequence types marked important phylogenetic subdivisions, and macrolide resistance markers (23S_rRNA allele 13 and fhaB3) confirmed the current restriction of this phenotype in China with few exceptions. ConclusionsThe genomic platform provides an expandable resource for unified genotyping of Bordetella strains and will facilitate collective evolutionary and epidemiological understanding of the re-emergence of whooping cough and other Bordetella infections. Data summaryBordetella genomes list and accession numbers: Supplementary Table S4 Bordetella genus phylogeny dataset (92 isolates): https://bigsdb.pasteur.fr/cgi-bin/bigsdb/bigsdb.pl?db=pubmlst_bordetella_isolates&page=query&project_list=23&submit=1 B. bronchiseptica phylogeny dataset (213 isolates): https://bigsdb.pasteur.fr/cgi-bin/bigsdb/bigsdb.pl?db=pubmlst_bordetella_isolates&page=query&project_list=24&submit=1 B. pertussis phylogeny (124 isolates): https://bigsdb.pasteur.fr/cgi-bin/bigsdb/bigsdb.pl?db=pubmlst_bordetella_isolates&page=query&project_list=25&submit=1 iTOL interactive trees: https://itol.embl.de/shared/1l7Fw0AvKOoCF

microbiology↗

Regional mutational signature activities in cancer genomes

Cancer genomes harbor a catalog of somatic mutations. The type and genomic context of these mutations depend on their causes, and allow their attribution to particular mutational signatures. Previous work has shown that mutational signature activities change over the course of tumor development, but investigations of genomic region variability in mutational signatures have been limited. Here, we expand upon this work by constructing regional profiles of mutational signature activities over 2,203 whole genomes across 25 tumor types, using data aggregated by the Pan-Cancer Analysis of Whole Genomes (PCAWG) consortium. We present GenomeTrackSig as an extension to the TrackSig R package to construct regional signature profiles using optimal segmentation and the expectation-maximization (EM) algorithm. We find that 426 genomes from 20 tumor types display at least one change in mutational signature activities (changepoint), and 257 genomes contain at least one of 54 recurrent changepoints shared by seven or more genomes of the same tumor type. Five recurrent changepoint locations are shared by multiple tumor types. Within these regions, the particular signature changes are often consistent across samples of the same type and some, but not all, are characterized by signatures associated with subclonal expansion. The changepoints we found cannot strictly be explained by gene density, mutation density, or cell-of-origin chromatin state. We hypothesize that they reflect a confluence of factors including evolutionary timing of mutational processes, regional differences in somatic mutation rate, large-scale changes in chromatin state that may be tissue type-specific, and changes in chromatin accessibility during subclonal expansion. These results provide insight into the regional effects of DNA damage and repair processes, and may help us localize genomic and epigenomic changes that occur during cancer development.

cancer biology↗

An ancient endogenous DNA virus in the human genome

The genomes of eukaryotes preserve a striking diversity of ancient viruses in the form of endogenous viral elements (EVEs). Study of this genomic fossil record provides insights into the diversity, origin and evolution of viruses across geological timescales. In particular, Mavericks have emerged as one of the oldest groups of viruses infecting vertebrates ([≥]419 My). They have been found in the genomes of fish, amphibians and non-avian reptiles but had been overlooked in mammals. Thus, their evolutionary history and the causes of their demise in mammals remain puzzling questions. Here, we conduct a detailed evolutionary study of two Maverick-like integrations found on human chromosomes 7 and 8. We performed a comparative analysis of the integrations and determined their orthology across placental mammals (Eutheria) via the syntenic arrangement of neighbouring genes. The integrations were absent at the orthologous sites in the genomes of marsupials and monotremes. These observations allowed us to reconstruct a time-calibrated phylogeny and infer the age of their most recent common ancestor at 268.61 (199.70-344.54) My. In addition, we estimate the age of the individual integrations at ~105 My which represent the oldest non-retroviral EVEs found in the human genome. Our findings suggest that active Mavericks existed in the ancestors of modern mammals ~172 My ago (Jurassic Period) and potentially to the end of the Early Cretaceous. We hypothesise Mavericks could have gone extinct in mammals from the evolution of an antiviral defence system or from reduced opportunities for transmission in terrestrial hosts. ImportanceThe genomes of vertebrates preserve an enormous diversity of endogenous viral elements (remnants of ancient viruses that accumulate in host genomes over evolutionary time). Although retroviruses account for the vast majority of these elements, diverse DNA viruses have also been found and novel lineages are being described. Here we analyse two elements found in the human genome belonging to an ancient group of DNA viruses called Mavericks. We study their evolutionary history, finding that the elements are shared between humans and many different species of placental mammals. These observations suggest the elements inserted at least ~105 Mya in the most recent common ancestor of placentals. We further estimate the age of the viral ancestor around 268 My. Our results provide evidence for some of the oldest viral integrations in the human genome and insights into the ancient interactions of viruses with the ancestors of modern-day mammals.

evolutionary biology↗

Mitochondrial genomes in Perkinsus decode conserved frameshifts in all genes

Mitochondrial genomes of apicomplexans, dino-flagellates and chrompodellids, that collectively make up the Myzozoa, are uncommonly reduced in coding capacity and display divergent gene configuration and expression mechanisms. They encode only three proteins -- COB, COX1, COX3 -- contain rRNAs fragmented to [~]100-200 base pair elements, and employ extensive recombination, RNA trans-splicing, and RNA-editing for genome maintenance and expression. The early-diverging Perkinsozoa is the final major myzozoan lineage whose mitochondrial genomes remain poorly characterized. Previous reports of Perkinsus cox1 and cob partial gene sequences have indicated independent acquisition of non-canonical features, namely the occurrence of multiple frameshifts in both genes. To determine ancestral myzozoan mitochondrial genome features, as well as any novel ones in Perkinsozoa, we sequenced and assembled four Perkinsus species mitochondrial genomes. These data show a simple ancestral genome with the common reduced coding capacity, but one already prone to rearrangement. Moreover, we identified 75 frameshifts across the four species that are present in all genes, that are highly conserved in gene location, and that occur as four distinct types. A decoding mechanism apparently employs unused codons at the frameshift sites that advance translation either +1 or +2 frames to the next used codon. The locations of the frameshifts are seemingly positioned to regulate protein folding of the nascent protein as it emerges from the ribosome. COX3 is distinct in containing only one frameshift and showing strong selection against residues that are otherwise frequently encoded at the frameshift positions in COX3 and COB. All genes also lack cysteine codons implying a further constraint on these genomes with reduction to only 19 different amino acids. Furthermore, mitochondrion-encoded rRNA fragment complements are incomplete in Perkinsus spp. but some are found in the nuclear DNA, suggesting these may be imported into the organelle as for tRNAs. Perkinsus demonstrates additional remarkable trajectories of organelle genome evolution including pervasive integration of frameshift translation into genome expression.

molecular biology↗

Shaping the Genome via Lengthwise Compaction, Phase Separation, and Lamina Adhesion

The link between genomic structure and biological function is yet to be consolidated, it is, however, clear that physical manipulation of the genome, driven by the activity of a variety of proteins, is a crucial step. To understand the consequences of the physical forces underlying genome organization, we build a coarse-grained polymer model of the genome, featuring three fundamentally distinct classes of interactions: lengthwise compaction, i.e., compaction of chromosomes along its contour, self-adhesion among epigenetically similar genomic segments, and adhesion of chromosome segments to the nuclear envelope or lamina. We postulate that these three types of interactions sufficiently represent the concerted action of the different proteins organizing the genome architecture and show that an interplay among these interactions can recapitulate the architectural variants observed across the tree of life. The model elucidates how an interplay of forces arising from the three classes of genomic interactions can drive drastic, yet predictable, changes in the global genome architecture, and makes testable predictions. We posit that precise control over these interactions in vivo is key to the regulation of genome architecture.

biophysics↗

GenErode: a bioinformatics pipeline to investigate genome erosion in endangered and extinct species

BackgroundMany wild species have suffered drastic population size declines over the past centuries, which have led to genomic erosion processes characterized by reduced genetic diversity, increased inbreeding, and accumulation of harmful mutations. Yet, genomic erosion estimates of modern-day populations often lack concordance with dwindling population sizes and conservation status of threatened species. One way to directly quantify the genomic consequences of population declines is to compare genome-wide data from pre-decline museum samples and modern samples. However, doing so requires computational data processing and analysis tools specifically adapted to comparative analyses of degraded, ancient or historical, DNA data with modern DNA data as well as personnel trained to perform such analyses. ResultsHere, we present a highly flexible, scalable, and modular pipeline to compare patterns of genomic erosion using samples from disparate time periods. The GenErode pipeline uses state-of-the-art bioinformatics tools to simultaneously process whole-genome re-sequencing data from ancient/historical and modern samples, and to produce comparable estimates of several genomic erosion indices. No programming knowledge is required to run the pipeline and all bioinformatic steps are well-documented, making the pipeline accessible to users with different backgrounds. GenErode is written in Snakemake and Python3 and uses Conda and Singularity containers to achieve reproducibility on high-performance compute clusters. The source code is freely available on GitHub (https://github.com/NBISweden/GenErode). ConclusionsGenErode is a user-friendly and reproducible pipeline that enables the standardization of genomic erosion indices from temporally sampled whole genome re-sequencing data.

bioinformatics↗

The contribution of Kaposis sarcoma-associated herpesvirus ORF7 and its zinc-finger motif to viral genome cleavage and capsid formation

Kaposis sarcoma-associated herpesvirus (KSHV) is the causative agent of endothelial and B cell malignancies. During KSHV lytic infection, lytic-related proteins are synthesized, viral genomes are replicated as a tandemly repeated form, and subsequently, capsids are assembled. The herpesvirus terminase complex is proposed to package an appropriate genome unit into an immature capsid, by cleavage of terminal repeats (TRs) flanking tandemly linked viral genomes. Although the mechanism of capsid formation in - and {beta}-herpesviruses are well-studied, in KSHV, it remains largely unknown. It has been proposed that KSHV ORF7 is a terminase subunit, and ORF7 harbors a zinc-finger motif, which is conserved among other herpesviral terminases. However, the biological significance of ORF7 is unknown. We previously reported that KSHV ORF17 is essential for the cleavage of inner scaffold proteins in capsid maturation, and ORF17 knockout (KO) induced capsid formation arrest between the procapsid and B-capsid stages. However, it remains unknown if ORF7-mediated viral DNA cleavage occurs before or after ORF17-mediated scaffold collapse. We analyzed the role of ORF7 during capsid formation using ORF7-KO-, ORF7&17-double-KO (DKO)-, and ORF7-zinc-finger motif mutant-KSHVs. We found that ORF7 acted after ORF17 in the capsid formation process, and ORF7-KO-KSHV produced incomplete capsids harboring non-spherical internal structures, which resembled soccer balls. This soccer ball-like capsid was formed after ORF17-mediated B-capsid formation. Moreover, ORF7-KO- and zinc-finger motif KO-KSHV failed to appropriately cleave the TR on replicated genome and had a defect in virion production. Thus, our data revealed that ORF7 contributes to terminase-mediated viral genome cleavage and capsid formation. IMPORTANCEIn herpesviral capsid formation, the viral terminase complex cleaves the TR sites on newly synthesized tandemly repeating genomes and inserts an appropriate genomic unit into an immature capsid. Herpes simplex virus 1 (HSV-1) UL28 is a subunit of the terminase complex that cleaves the replicated viral genome. However, the physiological importance of the UL28 homolog, KSHV ORF7, remains poorly understood. Here, using several ORF7-deficent KSHVs, we found that ORF7 acted after ORF17-mediated scaffold collapse in the capsid maturation process. Moreover, ORF7 and its zinc-finger motif were essential for both cleavage of TR sites on the KSHV genome and virus production. ORF7-deficient KSHVs produced incomplete capsids that resembled a soccer ball. To our knowledge, this is the first report showing ORF7-KO-induced soccer ball-like capsids production and ORF7 function in the KSHV capsid assembly process. Our findings provide insights into the role of ORF7 in KSHV capsid formation.

microbiology↗

Patterns of lineage-specific genome evolution in the brood parasitic black-headed duck (Heteronetta atricapilla)

Occurring independently at seven separate origins across the avian tree of life, obligate brood parasitism is a unique suite of traits observed in only approximately 1% of all bird species. Obligate brood parasites exhibit varied physiological, morphological, and behavioural traits across lineages, but common among all obligate brood parasites is that the females lay their eggs in the nest of other species. Unique among these species is the black-headed duck (Heteronetta atricapilla), a generalist brood parasite that is the only obligate brood parasite among waterfowl. This provides an opportunity to assess evolutionary changes in traits associated with brood parasitism, notably the loss of parental care behaviours, with an unspecialized brood parasite. We generated new high-quality genome assemblies and genome annotations of the black-headed duck and three related non-parasitic species (freckled duck, African pygmy-goose, and ruddy duck). With these assemblies and existing public genome assemblies, we produced a whole genome alignment across Galloanserae to identify conserved non-coding regions with atypical accelerations in the black-headed duck and coding genes with evidence of positive selection, as well as to resolve uncertainties in the duck phylogeny. To complement these data, we sequenced a population sample of black-headed ducks, allowing us to conduct McDonald-Kreitman tests of lineage-specific selection. We resolve the existing polytomy between our focal taxa with concordance from coding and non-coding sequences, and we observe stronger signals of evolution in non-coding regions of the genome than in coding regions. Collectively, the new high-quality genomes, comparative genome alignment, and population genomics provide a detailed picture of genome evolution in the only brood parasitic duck.

evolutionary biology↗

Reovirus efficiently reassorts genome segments during coinfection and superinfection

Reassortment, or genome segment exchange, increases diversity among viruses with segmented genomes. Previous studies on the limitations of reassortment have largely focused on parental incompatibilities that restrict generation of viable progeny. However, less is known about whether factors intrinsic to virus replication influence reassortment. Mammalian orthoreovirus (reovirus) encapsidates a segmented, double- stranded RNA genome, replicates within cytoplasmic factories, and is susceptible to host antiviral responses. We sought to elucidate the influence of infection multiplicity, timing, and compartmentalized replication on reovirus reassortment in the absence of parental incompatibilities. We used an established post-PCR genotyping method to quantify reassortment frequency between wild-type and genetically-barcoded type 3 reoviruses. Consistent with published findings, we found that reassortment increased with infection multiplicity until reaching a peak of efficient genome segment exchange during simultaneous coinfection. However, reassortment frequency exhibited a substantial decease with increasing time to superinfection, which strongly correlated with viral transcript abundance. We hypothesized that physical sequestration of viral transcripts within distinct virus factories or superinfection exclusion also could influence reassortment frequency during superinfection. Imaging revealed that transcripts from both wild-type and barcoded viruses frequently co-occupied factories with superinfection time delays up to 16 hours. Additionally, primary infection dampened superinfecting virus transcription with a 24 hour, but not shorter, time delay to superinfection. Thus, in the absence of parental incompatibilities and with short times to superinfection, reovirus reassortment proceeds efficiently and is largely unaffected by compartmentalization of replication and superinfection exclusion. However, reassortment may be limited by superinfection exclusion with greater time delays to superinfection. IMPORTANCEReassortment, or genome segment exchange between viruses, can generate novel virus genotypes and pandemic virus strains. For viruses to reassort their genome segments, they must replicate within the same physical space by coinfecting the same host cell. Even after entry into the host cell, many viruses with segmented genomes synthesize new virus transcripts and assemble and package their genomes within cytoplasmic replication compartments. Additionally, some viruses can interfere with subsequent infection of the same host or cell. However, spatial and temporal influences on reassortment are only beginning to be explored. We found that infection multiplicity and transcript abundance are important drivers of reassortment during coinfection and superinfection, respectively, for reovirus, which has a segmented, double-stranded RNA genome. We also provide evidence that compartmentalization of transcription and packaging is unlikely to influence reassortment, but the length of time between primary and subsequent reovirus infection can alter reassortment frequency.

microbiology↗

Genome-wide analysis of heat stress-stimulated transposon mobility in the human fungal pathogen Cryptococcus deneoformans

We recently reported transposon mutagenesis as a significant driver of spontaneous mutations in the human fungal pathogen Cryptococcus deneoformans during murine infection. Mutations caused by transposable element (TE) insertion into reporter genes were dramatically elevated at high temperature (37{degrees} versus 30{degrees}) in vitro, suggesting that heat stress stimulates TE mobility in the Cryptococcus genome. To explore the genome-wide impact of TE mobilization, we generated transposon accumulation lines by in vitro passage of C. deneoformans strain XL280 for multiple generations at both 30{degrees} and at the host-relevant temperature of 37{degrees}. Utilizing whole-genome sequencing, we identified native TE copies and mapped multiple de novo TE insertions in these lines. Movements of the T1 DNA transposon occurred at both temperatures with a strong bias for insertion between gene-coding regions. By contrast, the Tcn12 retrotransposon integrated primarily within genes and movement occurred exclusively at 37{degrees}. In addition, we observed a dramatic amplification in copy number of the Cnl1 (C. neoformans LINE-1) retrotransposon in sub-telomeric regions under heat-stress conditions. Comparing TE mutations to other sequence variations detected in passaged lines, the increase in genomic changes at elevated temperature was primarily due to mobilization of the retroelements Tcn12 and Cnl1. Finally, we found multiple TE movements (T1, Tcn12 and Cnl1) in the genomes of single C. deneoformans isolates recovered from infected mice, providing evidence that mobile elements are likely to facilitate microevolution and rapid adaptation during infection. Significance StatementRising global temperatures and climate change are predicted to increase fungal diseases in plants and mammals. However, the impact of heat stress on genetic changes in environmental fungi is largely unexplored. Environmental stressors can stimulate the movement of mobile DNA elements (transposons) within the genome to alter the genetic landscape. This report provides a genome-wide assessment of heat stress-induced transposon mobilization in the human fungal pathogen Cryptococcus. Transposon copies accumulated in genomes more rapidly following growth at the higher, host-relevant temperature. Additionally, movements of multiple elements were detected in the genomes of cryptococci recovered from infected mice. These findings suggest that heat stress-stimulated transposon mobility contributes to rapid adaptive changes in fungi both in the environment and during infection.

microbiology↗

Genomic analysis of a synthetic reversed sequence reveals default chromatin states in yeast and mammalian cells

Up to 93% of the human genome may show evidence of transcription, yet annotated transcripts account for less than 5%. It is unclear what makes up this major discrepancy, and to what extent the excess transcription has a definable biological function, or is just a pervasive byproduct of non-specific RNA polymerase binding and transcription initiation. Understanding the default state of the genome would be informative in determining whether the observed pervasive activity has a definable function. The genome of any modern organism has undergone billions of years of evolution, making it unclear whether any observed genomic activity, or lack thereof, has been selected for. We sought to address this question by introducing a completely novel 100-kb locus into the genomes of two eukaryotic organisms, S. cerevisiae and M. musculus, and characterizing its genomic activity based on chromatin accessibility and transcription. The locus was designed by reversing (but not complementing) the sequence of the human HPRT1 locus, including [~]30-kb of both upstream and downstream regulatory regions, allowing retention of sequence features like repeat frequency and GC content but ablating coding information and transcription factor binding sites. We also compared this reversed locus with a synthetic version of the normal human HPRT1 locus in both organismal contexts. Despite neither the synthetic HPRT1 locus nor its reverse version coding for any promoters evolved for gene expression in yeast, we observed widespread transcriptional activity of both loci. This activity was observed both when the loci were present as episomes and when chromosomally integrated, although it did not correspond to any of the known HPRT1 functional regulatory elements. In contrast, when integrated in the mouse genome, the synthetic HPRT1 locus showed transcriptional activity corresponding precisely to the HPRT1 coding sequence, while the reverse locus displayed no activity at all. Together, these results show that genomic sequences with no coding information are active in yeast, but relatively inactive in mouse, indicating a potentially major difference in "default genomic states" between these two divergent eukaryotes.

molecular biology↗

The compact genome of the sponge Oopsacas minuta (Hexactinellida) is lacking key metazoan core genes

BackgroundBilaterian animals today represent 99% of animal biodiversity. Elucidating how bilaterian hallmarks emerged is a central question of animal evo-devo and evolutionary genomics. Studies of non-bilaterian genomes have suggested that the ancestral animal already possessed a diversified developmental toolkit, including some pathways required for bilaterian body plans. Comparing genomes within the early branching metazoan Porifera phylum is key to identify which changes and innovations contributed to the successful transition towards bilaterians. ResultsHere, we report the first whole genome comprehensive analysis of a glass sponge, Oopsacas minuta, a member of the Hexactinellida. Studying this class of sponge is evolutionary relevant because it differs from the three other Porifera classes in terms of development, tissue organization, ecology and physiology. Although O. minuta does not exhibit drastic body simplifications, its genome is among the smallest animal genomes sequenced so far, surprisingly lacking several metazoan core genes (including Wnt and several key transcription factors). Our study also provided the complete genome of the symbiotic organism dominating the associated microbial community: a new Thaumarchaeota species. ConclusionsThe genome of the glass sponge O. minuta differs from all other available sponge genomes by its compactness and smaller number of predicted proteins. The unexpected losses of numerous genes considered as ancestral and pivotal for metazoan morphogenetic processes most likely reflect the peculiar syncytial organization in this group. Our work further documents the importance of convergence during animal evolution, with multiple emergences of sponge skeleton, electrical signaling and multiciliated cells.

evolutionary biology↗

Evolutionary dynamics of genome size and content during the adaptive radiation of Heliconiini butterflies

Heliconius butterflies, a speciose genus of Mullerian mimics, represent a classic example of an adaptive radiation that includes a range of derived dietary, life history, physiological and neural traits. However, key lineages within the genus, and across the broader Heliconiini tribe, lack genomic resources, limiting our understanding of how adaptive and neutral processes shaped genome evolution during their radiation. We have generated highly contiguous genome assemblies for nine new Heliconiini, 29 additional reference-assembled genomes, and improve 10 existing assemblies. Altogether, we provide a major new dataset of annotated genomes for a total of 63 species, including 58 species within the Heliconiini tribe. We use this extensive dataset to generate a robust and dated heliconiine phylogeny, describe major patterns of introgression, explore the evolution of genome architecture, and the genomic basis of key innovations in this enigmatic group, including an assessment of the evolution of putative regulatory regions at the Heliconius stem. Our work illustrates how the increased resolution provided by such dense genomic sampling improves our power to generate and test gene-phenotype hypotheses, and precisely characterize how genomes evolve.

evolutionary biology↗

Improved Genome Editing by an Engineered CRISPR-Cas12a

CRISPR-Cas12a is an RNA-guided, programmable genome editing enzyme found within bacterial adaptive immune pathways. Unlike CRISPR-Cas9, Cas12a uses only a single catalytic site to both cleave target double-stranded DNA (dsDNA) (cis-activity) and indiscriminately degrade single-stranded DNA (ssDNA) (trans-activity). To investigate how the relative potency of cis- versus trans-DNase activity affects Cas12a-mediated genome editing, we first used structure-guided engineering to generate variants of Lachnospiraceae bacterium Cas12a (LbCas12a) that selectively disrupt trans-activity. The resulting engineered mutant with the biggest differential between cis- and trans-DNase activity in vitro showed minimal genome editing activity in human cells, motivating a second set of experiments using directed evolution to generate additional mutants with robust genome editing activity. Notably, these engineered and evolved mutants had enhanced ability to induce homology-directed repair (HDR) editing by 2-18-fold depending on the genomic locus. Finally, we found that a site-specific reversion mutation produced improved Cas12a (iCas12a) variants with superior genome editing efficiency at genomic sites that are difficult to edit using wild-type Cas12a. This strategy of coupled rational engineering and directed evolution establishes a pipeline for creating improved genome editing tools by combining structural insights with randomization and selection. The availability of experimental and predicted structures of other CRISPR-Cas enzymes will enable this strategy to be applied to improve the efficacy of other genome editing proteins.

biochemistry↗

Using ARCADE (ARChaeplastida Annotation DatabasE) to understand the evolution of genome size in land plants

The abundance of plant genomic information caused by the decrease of sequencing costs contrasts with the lack of databases that integrate genome annotation, taxonomy and phenotypes to produce statistically sound, biologically meaningful knowledge. Here we present ARCADE (ARChaeplastida Annotation DatabasE), a database of 171 high-quality archaeplastidian non-redundant proteomes gathered from six primary genomic databases, together with proteome quality metrics anda growing number of associated metadata. As a case study to demonstrate the usefulness of ARCADE, we used it to investigate the expansion and contraction of protein domains associated with the evolution of genome size (hereafter GS). GS varies greatly among land plants and the synthesis of large genomes can be costly to cells. Although GS has been studied extensively for decades, the molecular mechanisms involved in the adaptations of plants to the increase in GS are still poorly understood. We used the annotation and phylogenetic information available in ARCADE, together with estimated GS values available for 83 land plant species, to search for associations between the abundance of protein domain families in these species and GS variation through phylogenetic-aware methods. Additionally, we estimated the GS for the ancestral nodes of the extant land plant species. GS seems to be decreasing along the course of evolution, except for a few branches that might have undergone independent GS increases. We found 7 Pfam correlated with the variation in GS in land plants, mainly related to nucleotide metabolism, DNA repair and genome organization. We found larger genomes to have a greater frequency of the Histone 2A superfamily, responsible for diverse functions, including the nucleosome formation and silencing of transposable elements. These molecular functions we found correlated to GS variation suggests they may be associated with preserving genome stability in larger genomes, and might indicate the evolution of mechanisms to cope with the variation in GS in land plants. ARCADE is available at https://bit.ly/ARCADE_OSF.

plant biology↗

VarSCAT: A computational tool for sequence context annotations of genomic variants

The sequence contexts of genomic variants play important roles in understanding biological significances of variants and potential sequencing related variant calling issues. However, methods for assessing the diverse sequence contexts of genomic variants such as tandem repeats and unambiguous annotations have been limited. Herein, we describe the Variant Sequence Context Annotation Tool (VarSCAT) for annotating the sequence contexts of genomic variants, including breakpoint ambiguities, flanking sequences, variant nomenclatures, adjacent variants, and tandem repeats with user customizable options. Our analysis demonstrate that VarSCAT is more versatile and customizable than current methods or strategies for annotating variants in short tandem repeat (STR) regions. Variant sequence context annotations of high-confidence human variant sets with VarSCAT revealed that more than 75% of all human individual germline and clinically relevant insertions and deletions (indels) have breakpoint ambiguities. Moreover, we illustrate that more than 80% of human individual germline small variants in STR regions are indels and that the sizes of these indels correlated with STR motif sizes. VarSCAT is available at https://github.com/elolab/VarSCAT. Author SummaryThe sequence contexts have significant impacts on the biological and technical aspects of genomic variants. The sequence contexts, such as tandem repeats or nearby indels, may increase the mutation rate of a region compared to other genome regions. Besides, variants in specific sequence contexts like STRs may also have distinguished biological consequences, which can lead to certain human diseases and thus they may be used as biomarkers for disease diagnosis and treatments. Moreover, potential ambiguous variant representations such as equivalent or redundant indels are also related with their sequence contexts, which may complicate variant harmonization from different sources. Our previous study demonstrated that more than half of false positive indel calls detected through next generation sequencing data are related with STRs. Thus, the sequence contexts of genomic variants are important and cannot be ignored. However, the current methods or strategies for assessing the sequence contexts of genomic variants are limited and not feasible to use. Here, we developed a computational tool VarSCAT for sequence contexts annotation of genomic variants. Our tool provides diverse sequence contexts annotations providing users information to further explore the variants of their interests. By applying VarSCAT to high confidence human variant sets, we demonstrate the influence of sequence context of genomic variants and emphasize the importance of sequence context assessment.

bioinformatics↗

Refine your search to explore more results.