bioRxiv Science⌕ Search

Biology subjects

Petrov, A. I.

Publications and source records attributed to Petrov, A. I..

2 recordsLinked to original sources

Comprehensive Survey of Conserved RNA Secondary Structures in Full-Genome Alignment of Hepatitis C Virus

Hepatitis C virus (HCV) is a plus-stranded RNA virus that often chronically infects liver hepatocytes and causes liver cirrhosis and cancer. These viruses replicate their genomes employing error-prone replicases. Thereby, they routinely generate a large "cloud" of RNA genomes which - by trial and error - comprehensively explore the sequence space available for functional RNA genomes that maintain the ability for efficient replication and immune escape. In this context, it is important to identify which RNA secondary structures in the sequence space of the HCV genome are conserved, likely due to functional requirements. Here, we provide the first genome-wide multiple sequence alignment (MSA) with the prediction of RNA secondary structures throughout all representative full-length HCV genomes. We selected 57 representative genomes by clustering all complete HCV genomes from the BV-BRC database based on k-mer distributions and dimension reduction and adding RefSeq sequences. We include annotations of previously recognized features for easy comparison to other studies. Our results indicate that mainly the core coding region, the C-terminal NS5A region, and the NS5B region contain secondary structure elements that are conserved beyond coding sequence requirements, indicating functionality on the RNA level. In contrast, the genome regions in between contain less highly conserved structures. The results provide a complete description of all conserved RNA secondary structures and make clear that functionally important RNA secondary structures are present in certain HCV genome regions but are largely absent from other regions. Full-genome alignments of all branches of Hepacivirus C are provided in the supplement.

genomics↗

R2DT: computational framework for template-based RNA secondary structure visualisation across non-coding RNA types

Non-coding RNAs (ncRNA) are essential for all life, and the functions of many ncRNAs depend on their secondary (2D) and tertiary (3D) structure. Despite proliferation of 2D visualisation software, there is a lack of methods for automatically generating 2D representations in consistent, reproducible, and recognisable layouts, making them difficult to construct, compare and analyse. Here we present R2DT, a comprehensive method for visualising a wide range of RNA structures in standardised layouts. R2DT is based on a library of 3,632 templates representing the majority of known structured RNAs, from small RNAs to the large subunit ribosomal RNA. R2DT has been applied to ncRNA sequences from the RNAcentral database and produced >13 million diagrams, creating the worlds largest RNA 2D structure dataset. The software is freely available at https://github.com/rnacentral/R2DT and a web server is found at https://rnacentral.org/r2dt.

bioinformatics↗