bioRxiv ScienceSearch

Biology subjects

Ripatti, S.

Publications and source records attributed to Ripatti, S..

10 recordsLinked to original sources

Coronary artery disease risk and lipidomic profiles are similar in familial and population-ascertained hyperlipidemias

Aims: To characterize and compare coronary artery disease (CAD) risk and detailed lipidomic profiles of individuals with familial and population-ascertained hyperlipidemias.\n\nMethods and Results: We determined incident CAD risk for 760 members of 66 hyperlipidemic families ([≥] 2 first degree relatives with the same hyperlipidemia) and 19,644 Finnish FINRISK population study participants. We also quantified 151 lipid species in plasma or serum samples from 550 members of 73 hyperlipidemic pedigrees and 897 FINRISK participants using a mass spectrometric shotgun lipidomics platform. Hyperlipidemias (LDL-C or triacylglycerides over 90th population percentile) were associated with increased CAD risk (high LDL-C: HR 1.74, 95% CI 1.48-2.04; high triacylglycerides: HR 1.38, 95% CI 1.09-1.74) and the risk estimates were very similar between the family and population samples. High LDL-C was associated with altered levels of 105 lipid species in families (p-value range 0.033-7.3*10-20 at 5% false discovery rate) and 51 species in the population samples (p-value range 0.017-6.8*10-21). Hypertriglyceridemia was associated with altered levels of 117 lipid species in families (p-value range 0.035-1.8*10-49) and 119 species in the population sample (p-value range 0.038-2.3*10-56). The lipidomics profiles of hyperlipidemias were highly similar in families and population samples.\n\nConclusion: We identified distinct lipidomic profiles associated with high LDL-C and triacylglyceride levels. CAD risk, lipidomic profiles and genetic profiles are highly similar between familial and population-ascertained hyperlipidemias, providing evidence of similar and overlapping underlying mechanisms. Our results do not support different screening and treatment for such hyperlipidemias.

epidemiology

Refining fine-mapping: effect sizes and regional heritability

Recent statistical approaches have shown that the set of all available genetic variants explains considerably more phenotypic variance of complex traits and diseases than the individual variants that are robustly associated with these phenotypes. However, rapidly increasing sample sizes constantly improve detection and prioritization of individual variants driving the associations between genomic regions and phenotypes. Therefore, it is useful to routinely estimate how much phenotypic variance the detected variants explain for each region by taking into account the correlation structure of variants and the uncertainty in their causal status. Here we extend the FINEMAP software to estimate the effect sizes and regional heritability under the probabilistic model that assumes a handful of causal variants per each region. Using the UK Biobank data to simulate GWAS regions with only a few causal variants, we demonstrate that FINEMAP provides higher precision and enables more detailed decomposition of regional heritability into individual variants than the variance component model implemented in BOLT or the fixed-effect model implemented in HESS. Using data from 51 serum biomarkers and four lipid traits from the FINRISK study, we estimate that FINEMAP captures on average 24% more regional heritability than the variant with the lowest P-value alone and 20% less than BOLT. Our simulations suggest how an upward bias of BOLT and a downward bias of FINEMAP could together explain the observed difference between the methods. We conclude that FINEMAP enables computationally efficient estimation of effect sizes and regional heritability in the era of biobank scale data.

genetics

Deep coverage whole genome sequences and plasma lipoprotein(a) in individuals of European and African ancestries

Lipoprotein(a), Lp(a), is a modified low-density lipoprotein particle where apolipoprotein(a) (protein product of the LPA gene) is covalently attached to apolipoprotein B. Lp(a) is a highly heritable, causal risk factor for cardiovascular diseases and varies in concentrations across ancestries. To comprehensively delineate the inherited basis for plasma Lp(a), we performed deep-coverage whole genome sequencing in 8,392 individuals of European and African American ancestries. Through whole genome variant discovery and direct genotyping of all structural variants overlapping LPA, we quantified the 5.5kb kringle IV-2 copy number (KIV2-CN), a known LPA structural polymorphism, and developed a model for its imputation. Through common variant analysis, we discovered a novel locus (SORT1) associated with Lp(a)-cholesterol, and also genetic modifiers of KIV2-CN. Furthermore, in contrast to previous GWAS studies, we explain most of the heritability of Lp(a), observing Lp(a) to be 85% heritable among African Americans and 75% among Europeans, yet with notable inter-ethnic heterogeneity. Through analyses of aggregates of rare coding and non-coding variants with Lp(a)-cholesterol, we found the only genome-wide significant signal to be at a non-coding SLC22A3 intronic window also previously described to be associated with Lp(a); however, this association was mitigated by adjustment with KIV2-CN. Finally, using an additional imputation dataset (N=27,344), we performed Mendelian randomization of LPA variant classes, finding that genetically regulated Lp(a) is more strongly associated with incident cardiovascular diseases than directly measured Lp(a), and is significantly associated with measures of subclinical atherosclerosis in African Americans.

genomics

Deep-coverage whole genome sequences and blood lipids among 16,324 individuals

Deep-coverage whole genome sequencing at the population level is now feasible and offers potential advantages for locus discovery, particularly in the analysis rare mutations in non-coding regions. Here, we performed whole genome sequencing in 16,324 participants from four ancestries at mean depth >29X and analyzed correlations of genotypes with four quantitative traits - plasma levels of total cholesterol, low-density lipoprotein cholesterol (LDL-C), high-density lipoprotein cholesterol, and triglycerides. We conducted a discovery analysis including common or rare variants in coding as well as non-coding regions and developed a framework to interpret genome sequence for dyslipidemia risk. Common variant association yielded loci previously described with the exception of a few variants not captured earlier by arrays or imputation. In coding sequence, rare variant association yielded known Mendelian dyslipidemia genes and, in non-coding sequence, we detected no rare variant association signals after application of four approaches to aggregate variants in non-coding regions. We developed a new, genome-wide polygenic score for LDL-C and observed that a high polygenic score conferred similar effect size to a monogenic mutation (~30 mg/dl higher LDL-C for each); however, among those with extremely high LDL-C, a high polygenic score was considerably more prevalent than a monogenic mutation (23% versus 2% of participants, respectively).

genomics

Phenome-wide association studies (PheWAS) across large "real-world data" population cohorts support drug target validation

Phenome-wide association studies (PheWAS), which assess whether a genetic variant is associated with multiple phenotypes across a phenotypic spectrum, have been proposed as a possible aid to drug development through elucidating mechanisms of action, identifying alternative indications, or predicting adverse drug events (ADEs). Here, we evaluate whether PheWAS can inform target validation during drug development. We selected 25 single nucleotide polymorphisms (SNPs) linked through genome-wide association studies (GWAS) to 19 candidate drug targets for common disease therapeutic indications. We independently interrogated these SNPs through PheWAS in four large \"real-world data\" cohorts (23andMe, UK Biobank, FINRISK, CHOP) for association with a total of 1,892 binary endpoints. We then conducted meta-analyses for 145 harmonized disease endpoints in up to 697,815 individuals and joined results with summary statistics from 57 published GWAS. Our analyses replicate 70% of known GWAS associations and identify 10 novel associations with study-wide significance after multiple test correction (P<1.8x10-6; out of 72 novel associations with FDR<0.1). By leveraging directionality and point estimate of the effect sizes, we describe new associations that may predict ADEs, e.g., acne, high cholesterol, gout and gallstones for rs738409 (p.I148M) in PNPLA3; or asthma for rs1990760 (p.T946A) in IFIH1. We further propose how quantitative estimates of genetic safety/efficacy profiles can be used to help prioritize candidate targets for a specific indication. Our results demonstrate PheWAS as a powerful addition to the toolkit for drug discovery.\n\nOne Sentence SummaryMatching genetics with phenotypes in 800,000 individuals predicts efficacy and on-target safety of future drugs.

genetics

Haplotype sharing provides insights into fine-scale population history and disease in Finland

Finland provides unique opportunities to investigate population and medical genomics because of its adoption of unified national electronic health records, detailed historical and birth records, and serial population bottlenecks. We assemble a comprehensive view of recent population history ([&le;]100 generations), the timespan during which most rare disease-causing alleles arose, by comparing pairwise haplotype sharing from 43,254 Finns to geographically and linguistically adjacent countries with different population histories, including 16,060 Swedes, Estonians, Russians, and Hungarians. We find much more extensive sharing in Finns, with at least one [&ge;] 5 cM tract on average between pairs of unrelated individuals. By coupling haplotype sharing with fine-scale birth records from over 25,000 individuals, we find that while haplotype sharing broadly decays with geographical distance, there are pockets of excess haplotype sharing; individuals from northeast Finland share several-fold more of their genome in identity-by-descent (IBD) segments than individuals from southwest regions containing the major cities of Helsinki and Turku. We estimate recent effective population size changes over time across regions of Finland and find significant differences between the Early and Late Settlement Regions as expected; however, our results indicate more continuous gene flow than previously indicated as Finns migrated towards the northernmost Lapland region. Lastly, we show that haplotype sharing is locally enriched among pairs of individuals sharing rare alleles by an order of magnitude, especially among pairs sharing rare disease causing variants. Our work provides a general framework for using haplotype sharing to reconstruct an integrative view of recent population history and gain insight into the evolutionary origins of rare variants contributing to disease.

genetics

Quantifying the impact of rare and ultra-rare coding variation across the phenotypic spectrum

There is a limited understanding about the impact of rare protein truncating variants across multiple phenotypes. We explore the impact of this class of variants on 13 quantitative traits and 10 diseases using whole-exome sequencing data from 100,296 individuals. Protein truncating variants in genes intolerant to this class of mutations increased risk of autism, schizophrenia, bipolar disorder, intellectual disability, ADHD. In individuals without these disorders, there was an association with shorter height, lower education, increased hospitalization and reduced age. Gene sets implicated from GWAS did not show a significant protein truncating variants-burden beyond what captured by established Mendelian genes. In conclusion, we provide the most thorough investigation to date of the impact of rare deleterious coding variants on complex traits, suggesting widespread pleiotropic risk.\n\nMain abbreviations

genetics

An interaction map of circulating metabolites, immune gene networks and their genetic regulation

The interaction between metabolism and the immune system plays a central role in many cardiometabolic diseases. We integrated blood transcriptomic, metabolomic, and genomic profiles from two population-based cohorts, including a subset with 7-year follow-up sampling. We identified topologically robust gene networks enriched for diverse immune functions including cytotoxicity, viral response, B cell, platelet, neutrophil, and mast cell/basophil activity. These immune gene modules showed complex patterns of association with 158 circulating metabolites, including lipoprotein subclasses, lipids, fatty acids, amino acids, and CRP. Genome-wide scans for module expression quantitative trait loci (mQTLs) revealed five modules with mQTLs of both cis and trans effects. The strongest mQTL was in ARHGEF3 (rs1354034) and affected a module enriched for platelet function. Mast cell/basophil and neutrophil function modules maintained their metabolite associations during 7-year follow-up, while our strongest mQTL in ARHGEF3 also displayed clear temporal stability. This study provides a detailed map of natural variation at the blood immuno-metabolic interface and its genetic basis, and facilitates subsequent studies to explain inter-individual variation in cardiometabolic disease.

genomics

biMM: Efficient estimation of genetic variances andcovariances for cohorts with high-dimensional phenotype measurements

Genetic research utilizes a decomposition of trait variances and covariances into genetic and environmental parts. Our software package biMM is a computationally efficient implementation of a bivariate linear mixed model for settings where hundreds of traits have been measured on partially overlapping sets of individuals.\n\nAvailabilityImplementation in R freely available at www.iki.fi/mpirinen.

genetics

The rate of false polymorphisms introduced when imputing genotypes from global imputation panels

Previous studies1,2 have shown that large multi-population imputation reference panels increases the number of well-imputed variants. However, to our knowledge, no previous studies have evaluated the rate of introduced variation in monomorphic sites of the study population when using imputation panels with admixed populations. In this study we evaluate the rate of false positive variants introduced by the imputation of Finnish genotype data using global reference panels (Haplotype Reference Consortium1; HRC, and the 1000Genomes project Phase I3; 1000G) and compare the results to a Finnish population-specific reference panel combining whole genome and exome sequenced samples. In sites that were monomorphic in our test set, we observed high false positive rates for the global reference panels (4.0% for 1000G and 2.6% for HRC) compared to the Finnish panel (0.26%). This rate was even higher (7.4%) when using a combination panel of 1000G and Finnish whole genome sequences with cross-panel imputation.

genetics