bioRxiv ScienceSearch

Biology subjects

Hoff, K.

Publications and source records attributed to Hoff, K..

4 recordsLinked to original sources

SARS-CoV-2 Variant Identification Using a Genome Tiling Array and Genotyping Probes

With over three million deaths worldwide attributed to the respiratory disease COVID-19 caused by the novel coronavirus SARS-CoV-2, it is essential that continued efforts be made to track the evolution and spread of the virus globally. We previously presented a rapid and cost-effective method to sequence the entire SARS-CoV-2 genome with 95% coverage and 99.9% accuracy. This method is advantageous for identifying and tracking variants in the SARS-CoV-2 genome when compared to traditional short read sequencing methods which can be time consuming and costly. Herein we present the addition of genotyping probes to our DNA chip which target known SARS-CoV-2 variants. The incorporation of the genotyping probe sets along with the advent of a moving average filter have improved our sequencing coverage and accuracy of the SARS-CoV-2 genome.

bioinformatics

Identification of Fibronectin 1 as a candidate genetic modifier in a Col4a mutant mouse model of Gould syndrome

Collagen type IV alpha 1 and alpha 2 (COL4A1 and COL4A2) are major components of almost all basement membranes. COL4A1 and COL4A2 mutations cause a multisystem disorder called Gould syndrome which can affect any organ but typically involves the cerebral vasculature, eyes, kidneys and skeletal muscles. The manifestations of Gould syndrome are highly variable and animal studies suggest that allelic heterogeneity and genetic context contribute to the clinical variability. We previously characterized a mouse model of Gould syndrome caused by a Col4a1 mutation in which the severities of ocular anterior segment dysgenesis (ASD), myopathy, and intracerebral hemorrhage (ICH) were dependent on genetic background. Here, we performed a genetic modifier screen to provide insight into the mechanisms contributing to Gould syndrome pathogenesis and identified a single locus (modifier of Gould syndrome 1; MoGS1) on Chromosome 1 that suppressed ASD. A separate screen showed that the same locus ameliorated myopathy. Interestingly, MoGS1 had no effect on ICH, suggesting that this phenotype may be mechanistically distinct. We refined the MoGS1 locus to a 4.3 Mb interval containing 18 protein coding genes, including Fn1 which encodes the extracellular matrix component fibronectin 1. Molecular analysis showed that the MoGS1 locus increased Fn1 expression raising the possibility that suppression is achieved through a compensatory extracellular mechanism. Furthermore, we show evidence of increased integrin linked kinase levels and focal adhesion kinase phosphorylation in Col4a1 mutant mice that is partially restored by the MoGS1 locus implicating the involvement of integrin signaling. Taken together, our results suggest that tissue-specific mechanistic heterogeneity contributes to the variable expressivity of Gould syndrome and that perturbations in integrin signaling may play a role in ocular and muscular manifestations.

genetics

BRAKER2: Automatic Eukaryotic Genome Annotation with GeneMark-EP+ and AUGUSTUS Supported by a Protein Database

Full automation of gene prediction has become an important bioinformatics task since the advent of next generation sequencing. The eukaryotic genome annotation pipeline BRAKER1 had combined self-training GeneMark-ET with AUGUSTUS to generate genes coordinates with support of transcriptomic data. Here, we introduce BRAKER2, a pipeline with GeneMark-EP+ and AUGUSTUS externally supported by cross-species protein sequences aligned to the genome. Among the challenges addressed in the development of the new pipeline was generation of reliable hints to the locations of protein-coding exon boundaries from likely homologous but evolutionarily distant proteins. Under equal conditions, the gene prediction accuracy of BRAKER2 was shown to be higher than the one of MAKER2, yet another genome annotation pipeline. Also, in comparison with BRAKER1 supported by a large volume of transcript data, BRAKER2 could produce a better gene prediction accuracy if the evolutionary distances to the reference species in the protein database were rather small. All over, our tests demonstrated that fully automatic BRAKER2 is a fast and accurate method for structural annotation of novel eukaryotic genomes.

bioinformatics

Integrative analysis of genomic variants reveals new associations of candidate haploinsufficient genes with congenital heart disease

Congenital Heart Disease (CHD) affects approximately 7-9 children per 1000 live births. Numerous genetic studies have established a role for rare genomic variants at the copy number variation (CNV) and single nucleotide variant level. In particular, the role of de novo mutations (DNM) has been highlighted in syndromic and non-syndromic CHD. To identify novel haploinsufficient CHD disease genes we performed an integrative analysis of CNVs and DNMs identified in probands with CHD including cases with sporadic thoracic aortic aneurysm (TAA). We assembled CNV data from 7,958 cases and 14,082 controls and performed a gene-wise analysis of the burden of rare genomic deletions in cases versus controls. In addition, we performed mutation rate testing for DNMs identified in 2,489 parent-offspring trios. Our combined analysis revealed 21 genes which were significantly affected by rare genomic deletions and/or constrained non-synonymous de novo mutations in probands. Fourteen of these genes have previously been associated with CHD while the remaining genes (FEZ1, MYO16, ARID1B, NALCN, WAC, KDM5B and WHSC1) have only been associated in singletons and small cases series, or show new associations with CHD. In addition, a systems level analysis revealed shared contribution of CNV deletions and DNMs in CHD probands, affecting protein-protein interaction networks involved in Notch signaling pathway, heart morphogenesis, DNA repair and cilia/centrosome function. Taken together, this approach highlights the importance of re-analyzing existing datasets to strengthen disease association and identify novel disease genes.

genetics