bioRxiv ScienceSearch

Biology subjects

Tan, M. H.

Publications and source records attributed to Tan, M. H..

4 recordsLinked to original sources

More from less: Genome skimming for nuclear markers for animal phylogenomics, a case study using decapod crustaceans

Low coverage genome sequencing is rapid and cost-effective for recovering complete mitochondrial genomes for animal phylogenomics. The recovery of high copy number nuclear genes, including histone H3, 18S and 28S ribosomal RNAs, is also possible using this approach. In this study, we explore the potential of the genome skimming (GS) to recover additional nuclear genes from shallow sequencing projects. Using an in silico baited approach, we recover three additional core histone genes (H2A, H2B and H4) from our existing collection of low coverage decapod crustacean dataset (99 species, 69 genera, 38 families, 10 infraorders). Phylogenetic analyses based on various combinations of mitochondrial and nuclear genes for the entire decapod dataset and 40 species of crayfish (Infraorder Astacidea) found that the evolutionary rates for different classes of genes varied widely. The highlight being a very high level of congruence found between trees from the six nuclear genes and those derived from the mitogenome sequences for freshwater crayfish. These findings indicate that nuclear genes recovered from the same genome skimming datasets designed to obtain mitogenomes can be used to support more robust and comprehensive phylogenetic analyses. Further, a search for additional intron-less nuclear genes identified several high copy number genes across the decapod dataset and recovery of NaK, PEPCK and GAPDH gene fragments is possible at slightly elevated coverage, suggesting the potential and utility of GS in recovering even more nuclear genetic information for phylogenetic studies from these inexpensive and increasingly abundant datasets.

evolutionary biology

First draft genome of the Labyrinthula genus, an opportunistic seagrass pathogen, reveals novel insight into marine protist phylogeny, ecology and CAZyme cell-wall degradation

Labyrinthula spp. are saprobic, marine protists that also act as opportunistic pathogens and are the causative agents of seagrass wasting disease (SWD). Despite the threat of local- and large-scale SWD outbreaks, there are currently gaps in our understanding of the drivers of SWD, particularly surrounding Labyrinthula virulence and ecology. Given these uncertainties, we investigated Labyrinthula from a novel genomic perspective by presenting the first draft genome and predicted proteome of a pathogenic isolate of Labyrinthula SR_Ha_C, generated from a hybrid assembly of Nanopore and Illumina sequences. Phylogenetic and cross-phyla comparisons revealed insights into the evolutionary history of Stramenopiles. Genome annotation showed evidence of glideosome-type machinery and an apicoplast protein typically found in protist pathogens and parasites. Proteins involved in Labyrinthulas actin-myosin mode of transport, as well as carbohydrate degradation were also prevalent. Further, CAZyme functional predictions revealed a repertoire of enzymes involved in breakdown of cell-wall and carbohydrate storage compounds common to seagrasses. The relatively low number of CAZymes annotated from the genome of Labyrinthula SR_Ha_C compared to other Labyrinthulea species may reflect the conservative annotation parameters, a specialised substrate affinity and the scarcity of characterised protist enzymes. Inherently, there is high probability for finding both unique and novel enzymes from Labyrinthula spp. This study provides resources for further exploration of Labyrinthula ecology and evolution, and will hopefully be the catalyst for new hypothesis-driven SWD research revealing more details of molecular interactions between Labyrinthula species and its host substrate.

genomics

A CRISPR-based SARS-CoV-2 diagnostic assay that is robust against viral evolution and RNA editing

Extensive testing is essential to break the transmission of the new coronavirus SARS-CoV-2, which causes the ongoing COVID-19 pandemic. Recently, CRISPR-based diagnostics have emerged as attractive alternatives to quantitative real-time PCR due to their faster turnaround time and their potential to be used in point-of-care testing scenarios. However, existing CRISPR-based assays for COVID-19 have not considered viral genome mutations and RNA editing in human cells. Here, we present the VaNGuard (Variant Nucleotide Guard) test that is not only specific and sensitive for SARS-CoV-2, but can also detect the virus when its genome or transcriptome has evolved or has been edited by deaminases in infected human cells. We show that an engineered AsCas12a enzyme is more tolerant of mismatches than wildtype LbCas12a and that multiplexed Cas12a targeting can overcome the presence of single nucleotide variations. Our assay can be completed in 30 minutes with a dipstick for a rapid point-of-care test.Competing Interest StatementThe authors have declared no competing interest.View Full Text

genomics

Direct RNA sequencing reveals structural differences between transcript isoforms

The ability to correctly assign structure information to an individual transcript in a continuous and phased manner is critical to understanding RNA function. RNA structure play important roles in every step of an RNAs lifecycle, however current short-read high throughput RNA structure mapping strategies are long, complex and cannot assign unique structures to individual gene-linked isoforms in shared sequences. To address these limitations, we present an approach that combines structure probing with SHAPE-like compound NAI-N3, nanopore direct RNA sequencing, and one-class support vector machines to detect secondary structures on near full-length RNAs (PORE-cupine). PORE-cupine provides rapid, direct, accurate and robust structure information along known RNAs and recapitulates global structural features in human embryonic stem cells. The majority of gene-linked isoforms showed structural differences in shared sequences both local and distal to the alternative splice site, highlighting the importance of long-read sequencing for phasing of structures. Structural differences between gene-linked isoforms are associated with differential translation efficiencies globally, highlighting the role of structure as a pervasive mechanism for regulating isoform-specific gene expression inside cells.

biochemistry