bioRxiv · 10.1101/2021.03.08.434363
Genome-Wide Covariation in SARS-CoV-2
Abstract
The SARS-CoV-2 virus causing the global pandemic is a coronavirus with a genome of about 30Kbase length [Song et al., 2019]. The design of vaccines and choice of therapies depends on the structure and mutational stability of encoded proteins in the open reading frames(ORFs) of this genome. In this study, we computed, using Expectation Reflection, the genome-wide covariation of the SARS-CoV-2 genome based on an alignment of {approx} 130000 SARS-CoV-2 complete genome sequences obtained from GISAID[Shu & McCauley, 2017]. We used this covariation to compute the Direct Information between pairs of positions across the whole genome, investigating potentially important relationships within the genome, both within each encoded protein and between encoded proteins. We then computed the covariation within each clade of the virus. The covariation detected recapitulates all clade determinants and each clade exhibits distinct covarying pairs.
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Cresswell-Clay, E. C., Periwal, V.. 2021-03-08. Genome-Wide Covariation in SARS-CoV-2. https://doi.org/10.1101/2021.03.08.434363
Cite the original work for its findings. Save a collection to share your selection of sources.