bioRxiv Science⌕ Search

Biology subjects

Loboda, A. A.

Publications and source records attributed to Loboda, A. A..

3 recordsLinked to original sources

Discordant genotype calls across technology platforms elucidate variants with systematic errors in next-generation sequencing

Large-scale next-generation sequencing datasets have been transformative for informing clinical variant interpretation and as reference panels for statistical and population genetic efforts. While such resources are often treated as ground truth, we find that in widely used reference datasets such as the Genome Aggregation Database (gnomAD), some variants pass gold standard filters yet are systematically different in their genotype calls across genotype discovery approaches. The inclusion of such discordant sites in study designs involving multiple genotype discovery strategies could bias results and lead to false-positive hits in association studies due to technological artifacts rather than a true relationship to the phenotype. Here, we describe this phenomenon of discordant genotype calls across genotype discovery approaches, characterize the error mode of wrong calls, provide a blacklist of discordant sites identified in gnomAD that should be treated with caution in analyses, and present a metric and machine learning classifier trained on gnomAD data to identify likely discordant variants in other datasets. We find that different genotype discovery approaches have different sets of variants at which this problem occurs but that there are characteristic variant features that can be used to predict discordant behavior. Discordant sites are largely shared across ancestry groups, though different populations are powered for discovery of different variants. We find that the most common error mode is that of a variant being heterozygous for one approach and homozygous for the other, with heterozygous in the genomes and homozygous reference in the exomes making up the majority of miscalls.

genomics↗

The 22q11.2 region regulates presynaptic gene-products linked to schizophrenia

To study how the 22q11.2 deletion predisposes to psychiatric disease, we generated induced pluripotent stem cells from deletion carriers and controls, as well as utilized CRISPR/Cas9 to introduce the heterozygous deletion into a control cell line. Upon differentiation into neural progenitor cells, we found the deletion acted in trans to alter the abundance of transcripts associated with risk for neurodevelopmental disorders including Autism Spectrum Disorder. In more differentiated excitatory neurons, altered transcripts encoded presynaptic factors and were associated with genetic risk for schizophrenia, including common (per-SNP heritability p ({tau}c)= 4.2 x 10-6) and rare, loss of function variants (p = 1.29x10-12). These findings suggest a potential relationship between cellular states, developmental windows and susceptibility to psychiatric conditions with different ages of onset. To understand how the deletion contributed to these observed changes in gene expression, we developed and applied PPItools, which identifies the minimal protein-protein interaction network that best explains an observed set of gene expression alterations. We found that many of the genes in the 22q11.2 interval interact in presynaptic, proteasome, and JUN/FOS transcriptional pathways that underlie the broader alterations in psychiatric risk gene expression we identified. Our findings suggest that the 22q11.2 deletion impacts genes and pathways that may converge with risk loci implicated by psychiatric genetic studies to influence disease manifestation in each deletion carrier.

neuroscience↗

A platform for case-control matching enables association studies without genotype sharing

Acquiring a sufficiently powered cohort of control samples can be time consuming or, sometimes, impossible. Accordingly, an ability to leverage control samples that were already collected and sequenced elsewhere could dramatically improve power in all genetic association studies. However, since majority of the genotyped and sequenced human DNA samples to date are subject to strict data sharing regulations, large-scale sharing of, in particular, control samples is extremely challenging. Using insights from image recognition, we developed a method allowing selection of the best-matching controls in an external pool of samples that is compliant with personal genotype data protection restrictions. Our approach uses singular value decomposition of the matrix of case genotypes to rank controls in another study by similarity to cases. We demonstrate that this recovers an accurate case-control association analysis for both ultra-rare and common variants and implement and provide online access to a library of ~17,000 controls that enables association studies for case cohorts lacking control subjects.

genetics↗