bioRxiv ScienceSearch

Biology subjects

Lovell, N.

Publications and source records attributed to Lovell, N..

3 recordsLinked to original sources

Evidence for ZAP-independent CpG reduction in SARS-CoV-2 genome, and pangolin coronavirus origin of 5'UTR

SARS-CoV-2, the causative agent of COVID-19, has an RNA genome, which is, overall, closely related to the bat coronavirus sequence RaTG13. However, the ACE2-binding domain of this virus is more similar to a coronavirus isolated from a Guangdong pangolin. In addition to this unique feature, the genome of SARS-CoV-2 (and its closely related coronaviruses) has a low CpG content. This has been postulated to be the signature of an evolutionary pressure exerted by the host antiviral protein ZAP. Here, we analyzed the sequences of a wide range of viruses using both alignment-based and alignment free approaches to investigate the origin of SARS-CoV-2 genome. Our analyses revealed a high level of similarity between the 5UTR of SARS-CoV-2 and that of the Guangdong pangolin coronavirus. This suggests bat and pangolin coronaviruses might have recombined at least twice (in the 5UTR and ACE2 binding regions) to seed the formation of SARS-CoV-2. An alternative hypothesis is that the lineage preceding SARS-CoV-2 is a yet to be sampled bat coronavirus whose ACE2 binding domain and 5UTR are distinct from other known bat coronaviruses. Additionally, we performed a detailed analysis of viral genome compositions as well as expression and RNA binding data of ZAP to show that the low CpG abundance in SARS-CoV-2 is not related to an evolutionary pressure from ZAP.

genomics

Integrative analysis of mutated genes and mutational processes reveals seven colorectal cancer subtypes

Colorectal cancer (CRC) is one of the leading causes of cancer-related deaths in the world. It has been reported that [~]10%-15% of individuals with colorectal cancer experience a causative mutation in the known susceptibility genes, highlighting the importance of identifying mutations for early detection in high risk individuals. Through extensive sequencing projects such as the International Cancer Genome Consortium (ICGC), a large number of somatic point mutations have been identified that can be used to identify cancer-associated genes, as well as the signature of mutational processes defined by the tri-nucleotide sequence context (motif) of mutated sites. Mutation is the hallmark of cancer genome, and many studies have reported cancer subtyping based on the type of frequently mutated genes, or the proportion of mutational processes, however, none of these cancer subtyping methods consider these features simultaneously. This highlights the need for a better and more inclusive subtype classification approach to enable biomarker discovery and thus inform drug development for CRC. In this study, we developed a statistical pipeline based on a novel concept gene-motif, which merges mutated gene information with tri-nucleotide motif of mutated sites, to identify cancer subtypes, in this case CRCs. Our analysis identified for the first time, 3,131 gene-motif combinations that were significantly mutated in 536 ICGC colorectal cancer samples compared to other cancer types, identifying seven CRC subtypes with distinguishable phenotypes and biomarkers. Interestingly, we identified several genes that were mutated in multiple subtypes but with unique sequence contexts. Taken together, our results highlight the importance of considering both the mutation type and mutated genes in identification of cancer subtypes and cancer biomarkers.

cancer biology

Evidence for enhancer noncoding RNAs (enhancer-ncRNAs) with gene regulatory functions relevant to neurodevelopmental disorders

Noncoding RNAs (ncRNAs) comprise a significant proportion of the mammalian genome, but their biological significance in neurodevelopment disorders is poorly understood. In this study, we identified 908 brain-enriched noncoding RNAs comprising at least one nervous system-related eQTL polymorphism that is associated with protein coding genes and also overlap with chromatin states characterised as enhancers. We referred to such noncoding RNAs with putative enhancer activity as brain enhancer-ncRNAs. By integrating GWAS SNPs and Copy Number Variation (CNV) data from neurodevelopment disorders, we found that 265 enhancer-ncRNAs were either mutated (CNV deletion or duplication) or contain at least one GWAS SNPs in the context of such conditions. Of these, the eQTL-associated gene for 82 enhancer-ncRNAs did not overlap with either GWAS SNPs or CNVs suggesting in such contexts that mutations to neurodevelopment gene enhancers disrupt ncRNA interaction. Taken together, we identified 49 novel NDD-associated ncRNAs that influence genomic enhancers during neurodevelopment, suggesting enhancer mutations may be relevant to the functions for such ncRNAs in neurodevelopmental disorders.

neuroscience