bioRxiv Science⌕ Search

Biology subjects

Wahba, L.

Publications and source records attributed to Wahba, L..

3 recordsLinked to original sources

Recurrent expansion and rapid evolution of the Drosophilid RNAi pathway in testis

Multiple classes of selfish genetic elements, including transposable elements, meiotic drivers, and viruses, are suppressed by small interfering RNAs (siRNAs) that guide RNA interference (RNAi). These are interlocked within continually evolving genetic conflicts that are characterized by extremely rapid dynamics of change in both sequence and copy number. Indeed, it was shown that across Drosophila, the core RNAi machinery evolves under positive selection, and that the RNAi effector AGO2 has additional copies in some species. This contrasts with the core microRNA (miRNA) machinery in Drosophila, which evolves under negative selection and maintains one-to-one orthologs not only amongst flies, but even to mammals. Here, we analyze >300 long read genome assemblies of Drosophila species to generalize these attributes. Not only do we find recurrent expansion of AGO2 across [~]100 species, including lineages with ongoing amplification of AGO2, we also find dozens of species with extra copies of the other core RNAi factors, r2d2 and dicer-2. In many cases, these additional RNAi factor copies co-exist, or are nested, within a lineage bearing an ancestral expansion of AGO2. Transcriptome data provide evidence that core RNAi factors, including certain lineage-specific copies, are biased for testis expression. Finally, we use small RNA data to annotate hairpin RNAs (hpRNAs) in D. pseudoobscura, one of the species with prominent amplification of RNAi factors. We find rampant de novo hpRNA loci in this species, whose siRNAs are predominantly expressed in testis. All together, these findings highlight evolutionary plasticity of the fly RNAi pathway and affirm that it is preferentially deployed in the germline of the heterogametic sex. When considered alongside abundant genetic data for recurrent sex ratio meiotic drive against the Y chromosome, these genomic data strongly imply that a fundamental role of endogenous RNAi is to control sex chromosome conflicts.

evolutionary biology↗

CGC1, a new reference genome for Caenorhabditis elegans

The original 100.3 Mb reference genome for Caenorhabditis elegans, generated from the wild-type laboratory strain N2, has been crucial for analysis of C. elegans since 1998 and has been considered complete since 2005. Unexpectedly, this long-standing reference was shown to be incomplete in 2019 by a genome assembly from the N2-derived strain VC2010. Moreover, genetically divergent versions of N2 have arisen over decades of research and hindered reproducibility of C. elegans genetics and genomics. Here we provide a 106.4 Mb gap-free, telomere-to-telomere genome assembly of C. elegans, generated from CGC1, an isogenic derivative of the N2 strain. We used improved long-read sequencing and manual assembly of 43 recalcitrant genomic regions to overcome deficiencies of prior N2 and VC2010 assemblies, and to assemble tandem repeat loci including a 772-kb sequence for the 45S rRNA genes. While many differences from earlier assemblies came from repeat regions, unique additions to the genome were also found. Of 19,972 protein-coding genes in the N2 assembly, 19,790 (99.1%) encode products that are unchanged in the CGC1 assembly. The CGC1 assembly also may encode 183 new protein-coding and 163 new ncRNA genes. CGC1 thus provides both a completely defined reference genome and corresponding isogenic wild-type strain for C. elegans, allowing unique opportunities for model and systems biology.

genomics↗

Identification of a pangolin niche for a 2019-nCoV-like coronavirus through an extensive meta-metagenomic search

In numerous instances, tracking the biological significance of a nucleic acid sequence can be augmented through the identification of environmental niches in which the sequence of interest is present. Many metagenomic datasets are now available, with deep sequencing of samples from diverse biological niches. While any individual metagenomic dataset can be readily queried using web-based tools, meta-searches through all such datasets are less accessible. In this brief communication, we demonstrate such a meta-meta-genomic approach, examining close matches to the Wuhan coronavirus 2019-nCoV in all high-throughput sequencing datasets in the NCBI Sequence Read Archive accessible with the keyword "virome". In addition to the homology to bat coronaviruses observed in descriptions of the 2019-nCoV sequence (F. Wu et al. 2020, Nature, doi.org/10.1038/s41586-020-2008-3; P. Zhou et al. 2020, Nature, doi.org/10.1038/s41586-020-2012-7), we note a strong homology to numerous sequence reads in a metavirome dataset generated from the lungs of deceased Pangolins reported by Liu et al. (Viruses 11:11, 2019, http://doi.org/10.3390/v11110979). Our observations are relevant to discussions of the derivation of 2019-nCoV and illustrate the utility and limitations of meta-metagenomic search tools in effective and rapid characterization of potentially significant nucleic acid sequences. ImportanceMeta-metagenomic searches allow for high-speed, low-cost identification of potentially significant biological niches for sequences of interest.

bioinformatics↗