bioRxiv ScienceSearch

Biology subjects

Freimer, N. B.

Publications and source records attributed to Freimer, N. B..

5 recordsLinked to original sources

ForestQC: quality control on genetic variants from next-generation sequencing data using random forest

Next-generation sequencing technology (NGS) enables discovery of nearly all genetic variants present in a genome. A subset of these variants, however, may have poor sequencing quality due to limitations in sequencing technology or in variant calling algorithms. In genetic studies that analyze a large number of sequenced individuals, it is critical to detect and remove those variants with poor quality as they may cause spurious findings. In this paper, we present a statistical approach for performing quality control on variants identified from NGS data by combining a traditional filtering approach and a machine learning approach. Our method uses information on sequencing quality such as sequencing depth, genotyping quality, and GC contents to predict whether a certain variant is likely to contain errors. To evaluate our method, we applied it to two whole-genome sequencing datasets where one dataset consists of related individuals from families while the other consists of unrelated individuals. Results indicate that our method outperforms widely used methods for performing quality control on variants such as VQSR of GATK by considerably improving the quality of variants to be included in the analysis. Our approach is also very efficient, and hence can be applied to large sequencing datasets. We conclude that combining a machine learning algorithm trained with sequencing quality information and the filtering approach is an effective approach to perform quality control on genetic variants from sequencing data.\n\nAuthor SummaryGenetic disorders can be caused by many types of genetic mutations, including common and rare single nucleotide variants, structural variants, insertions and deletions. Nowadays, next generation sequencing (NGS) technology allows us to identify various genetic variants that are associated with diseases. However, variants detected by NGS might have poor sequencing quality due to biases and errors in sequencing technologies and analysis tools. Therefore, it is critical to remove variants with low quality, which could cause spurious findings in follow-up analyses. Previously, people applied either hard filters or machine learning models for variant quality control (QC), which failed to filter out those variants accurately. Here, we developed a statistical tool, ForestQC, for variant QC by combining a filtering approach and a machine learning approach. We applied ForestQC to one family-based whole genome sequencing (WGS) dataset and one general case-control WGS dataset, to evaluate our method. Results show that ForestQC outperforms widely used methods for variant QC by considerably improving the quality of variants. Also, ForestQC is very efficient and scalable to large-scale sequencing datasets. Our study indicates that combining filtering approaches and machine learning approaches enables effective variant QC.

bioinformatics

Contribution of common and rare variants to bipolar disorder susceptibility in extended pedigrees from population isolates

Current evidence from case/control studies indicates that genetic risk for psychiatric disorders derives primarily from numerous common variants, each with a small phenotypic impact. The literature describing apparent segregation of bipolar disorder (BP) in numerous multigenerational pedigrees suggests that, in such families, large-effect inherited variants might play a greater role. To evaluate this hypothesis, we conducted genetic analyses in 26 Colombian (CO) and Costa Rican (CR) pedigrees ascertained for BP1, the most severe and heritable form of BP. In these pedigrees, we performed microarray SNP genotyping of 856 individuals and high-coverage whole-genome sequencing of 454 individuals. Compared to their unaffected relatives, BP1 individuals had higher polygenic risk scores estimated from SNPs associated with BP discovered in independent genome-wide association studies, and also displayed a higher burden of rare deleterious single nucleotide variants (SNVs) and rare copy number variants (CNVs) in genes likely to be relevant to BP1. Parametric and non-parametric linkage analyses identified 15 BP1 linkage peaks, encompassing about 100 genes, although we observed no significant segregation pattern for any particular rare SNVs and CNVs. These results suggest that even in extended pedigrees, genetic risk for BP appears to derive mainly from small to moderate effect rare and common variants.

genetics

Coronary artery disease risk and lipidomic profiles are similar in familial and population-ascertained hyperlipidemias

Aims: To characterize and compare coronary artery disease (CAD) risk and detailed lipidomic profiles of individuals with familial and population-ascertained hyperlipidemias.\n\nMethods and Results: We determined incident CAD risk for 760 members of 66 hyperlipidemic families ([≥] 2 first degree relatives with the same hyperlipidemia) and 19,644 Finnish FINRISK population study participants. We also quantified 151 lipid species in plasma or serum samples from 550 members of 73 hyperlipidemic pedigrees and 897 FINRISK participants using a mass spectrometric shotgun lipidomics platform. Hyperlipidemias (LDL-C or triacylglycerides over 90th population percentile) were associated with increased CAD risk (high LDL-C: HR 1.74, 95% CI 1.48-2.04; high triacylglycerides: HR 1.38, 95% CI 1.09-1.74) and the risk estimates were very similar between the family and population samples. High LDL-C was associated with altered levels of 105 lipid species in families (p-value range 0.033-7.3*10-20 at 5% false discovery rate) and 51 species in the population samples (p-value range 0.017-6.8*10-21). Hypertriglyceridemia was associated with altered levels of 117 lipid species in families (p-value range 0.035-1.8*10-49) and 119 species in the population sample (p-value range 0.038-2.3*10-56). The lipidomics profiles of hyperlipidemias were highly similar in families and population samples.\n\nConclusion: We identified distinct lipidomic profiles associated with high LDL-C and triacylglyceride levels. CAD risk, lipidomic profiles and genetic profiles are highly similar between familial and population-ascertained hyperlipidemias, providing evidence of similar and overlapping underlying mechanisms. Our results do not support different screening and treatment for such hyperlipidemias.

epidemiology

Whole Genome Sequencing in Psychiatric Disorders: the WGSPD Consortium

As technology advances, whole genome sequencing (WGS) is likely to supersede other genotyping technologies. The rate of this change depends on its relative cost and utility. Variants identified uniquely through WGS may reveal novel biological pathways underlying complex disorders and provide high-resolution insight into when, where, and in which cell type these pathways are affected. Alternatively, cheaper and less computationally intensive approaches may yield equivalent insights. Understanding the role of rare variants in the noncoding gene-regulating genome, through pilot WGS projects, will be critical to determine which of these two extremes best represents reality. With large cohorts, well-defined risk loci, and a compelling need to understand the underlying biology, psychiatric disorders have a role to play in this preliminary WGS assessment. The WGSPD consortium will integrate data for 18,000 individuals with psychiatric disorders, beginning with autism spectrum disorder, schizophrenia, bipolar disorder, and major depressive disorder, along with over 150,000 controls.

genomics

Genetic variation and gene expression across multiple tissues and developmental stages in a non-human primate

By analyzing multi-tissue gene expression and genome-wide genetic variation data in samples from a vervet monkey pedigree, we generated a transcriptome resource and produced the first catalogue of expression quantitative trait loci (eQTLs) in a non-human primate model. This catalogue contains more genome-wide significant eQTLs, per sample, than comparable human resources, and reveals sex and age-related expression patterns. Findings include a master regulatory locus that likely plays a role in immune function, and a locus regulating hippocampal long non-coding RNAs (lncRNAs), whose expression correlates with hippocampal volume. This resource will facilitate genetic investigation of quantitative traits, including brain and behavioral phenotypes relevant to neuropsychiatric disorders.

genetics