bioRxiv ScienceSearch

Biology subjects

Gupta, R. M.

Publications and source records attributed to Gupta, R. M..

2 recordsLinked to original sources

Phylogenomic analysis of SARS-CoV-2 genomes from western India reveals unique linked mutations

India has become the third worst-hit nation by the COVID-19 pandemic caused by the SARS-CoV-2 virus. Here, we investigated the molecular, phylogenomic, and evolutionary dynamics of SARS-CoV-2 in western India, the most affected region of the country. A total of 90 genomes were sequenced. Four nucleotide variants, namely C241T, C3037T, C14408T (Pro4715Leu), and A23403G (Asp614Gly), located at 5UTR, Orf1a, Orf1b, and Spike protein regions of the genome, respectively, were predominant and ubiquitous (90%). Phylogenetic analysis of the genomes revealed four distinct clusters, formed owing to different variants. The major cluster (cluster 4) is distinguished by mutations C313T, C5700A, G28881A are unique patterns and observed in 45% of samples. We thus report a newly emerging pattern of linked mutations. The predominance of these linked mutations suggests that they are likely a part of the viral fitness landscape. A novel and distinct pattern of mutations in the viral strains of each of the districts was observed. The Satara district viral strains showed mutations primarily at the 3' end of the genome, while Nashik district viral strains displayed mutations at the 5' end of the genome. Characterization of Pune strains showed that a novel variant has overtaken the other strains. Examination of the frequency of three mutations i.e., C313T, C5700A, G28881A in symptomatic versus asymptomatic patients indicated an increased occurrence in symptomatic cases, which is more prominent in females. The age-wise specific pattern of mutation is observed. Mutations C18877T, G20326A, G24794T, G25563T, G26152T, and C26735T are found in more than 30% study samples in the age group of 10-25. Intriguingly, these mutations are not detected in the higher age range 61-80. These findings portray the prevalence of unique linked mutations in SARS-CoV-2 in western India and their prevalence in symptomatic patients. ImportanceElucidation of the SARS-CoV-2 mutational landscape within a specific geographical location, and its relationship with age and symptoms, is essential to understand its local transmission dynamics and control. Here we present the first comprehensive study on genome and mutation pattern analysis of SARS-CoV-2 from the western part of India, the worst affected region by the pandemic. Our analysis revealed three unique linked mutations, which are prevalent in most of the sequences studied. These may serve as a molecular marker to track the spread of this viral variant to different places.

genomics

Deep learning enables genetic analysis of the human thoracic aorta

The aorta is the largest blood vessel in the body, and enlargement or aneurysm of the aorta can predispose to dissection, an important cause of sudden death. While rare syndromes have been identified that predispose to aortic aneurysm, the common genetic basis for the size of the aorta remains largely unknown. By leveraging a deep learning architecture that was originally developed to recognize natural images, we trained a model to evaluate the dimensions of the ascending and descending thoracic aorta in cardiac magnetic resonance imaging. After manual annotation of just 116 samples, we applied this model to 3,840,140 images from the UK Biobank. We then conducted a genome-wide association study in 33,420 individuals, revealing 68 loci associated with ascending and 35 with descending thoracic aortic diameter, of which 10 loci overlapped. Integration of common variation with transcriptome-wide analyses, rare-variant burden tests, and single nucleus RNA sequencing prioritized SVIL, a gene highly expressed in vascular smooth muscle, that was significantly associated with the diameter of the ascending and descending aorta. A polygenic score for ascending aortic diameter was associated with a diagnosis of thoracic aortic aneurysm in the remaining 391,251 UK Biobank participants who did not undergo imaging (HR = 1.44 per standard deviation; P = 3.7{middle dot}10-12). Defining the genetic basis of the diameter of the aorta may enable the identification of asymptomatic individuals at risk for aneurysm or dissection and facilitate the prioritization of potential therapeutic targets for the prevention or treatment of aortic aneurysm. Finally, our results illustrate the potential for rapidly defining novel quantitative traits derived from a deep learning model, an approach that can be more broadly applied to biomedical imaging data.

genomics