bioRxiv ScienceSearch

Biology subjects

Samuli Ripatti

Publications and source records attributed to Samuli Ripatti.

9 recordsLinked to original sources

Genetic loci associated with coronary artery disease harbor evidence of selection and antagonistic pleiotropy

Traditional genome-wide scans for positive selection have mainly uncovered selective sweeps associated with monogenic traits. While selection on quantitative traits is much more common, very few signals have been detected because of their polygenic nature. We searched for positive selection signals underlying coronary artery disease (CAD) in worldwide populations, using novel approaches to quantify relationships between polygenic selection signals and CAD genetic risk. We identified new candidate adaptive loci that appear to have been directly modified by disease pressures given their significant associations with CAD genetic risk. These candidates were all uniquely and consistently associated with many different male and female reproductive traits suggesting selection may have also targeted these because of their direct effects on fitness. This suggests the presence of widespread antagonistic-pleiotropic tradeoffs on CAD loci, which provides a novel explanation for the maintenance and high prevalence of CAD in modern humans. Lastly, we found that positive selection more often targeted CAD gene regulatory variants using HapMap3 lymphoblastoid cell lines, which further highlights the unique biological significance of candidate adaptive loci underlying CAD. Our study provides a novel approach for detecting selection on polygenic traits and evidence that modern human genomes have evolved in response to CAD-induced selection pressures and other early-life traits sharing pleiotropic links with CAD.\n\nAuthor SummaryHow genetic variation contributes to disease is complex, especially for those such as coronary artery disease (CAD) that develop over the lifetime of individuals. One of the fundamental questions about CAD -- whose progression begins in young adults with arterial plaque accumulation leading to life-threatening outcomes later in life -- is why natural selection has not removed or reduced this costly disease. It is the leading cause of death worldwide and has been present in human populations for thousands of years, implying considerable pressures that natural selection should have operated on. Our study provides new evidence that genes underlying CAD have recently been modified by natural selection and that these same genes uniquely and extensively contribute to human reproduction, which suggests that natural selection may have maintained genetic variation contributing to CAD because of its beneficial effects on fitness. This study provides novel evidence that CAD has been maintained in modern humans as a byproduct of the fitness advantages those genes provide early in human lifecycles.

Genomics

Whole genome view of the consequences of a population bottleneck using 2926 genome sequences from Finland and United Kingdom

Isolated populations with enrichment of variants due to recent population bottlenecks provide a powerful resource for identifying disease-associated genetic variants and genes. As a model of an isolate population, we sequenced the genomes of 1463 Finnish individuals as part of the Sequencing Initiative Suomi (SISu) Project. We compared the genomic profiles of the 1463 Finns to a sample of 1463 British individuals that were sequenced in parallel as part of the UK10K Project. Whereas there were no major differences in the allele frequency of common variants, a significant depletion of variants in the rare frequency spectrum was observed in Finns when comparing the two populations. On the other hand, we observed >2.1 million variants that were twice as frequent among Finns compared to Britons and 800,000 variants that were more than 10 times more frequent in Finns. Furthermore, in Finns we observed a relative proportional enrichment of variants in the minor allele frequency range between 2 - 5% (p < 2.2x10-16). When stratified by their functional annotations, loss-of-function (LoF) variants showed the highest proportional enrichment in Finns (p = 0.0291). In the noncoding part of the genome, variants in conserved regions (p = 0.002) and promoters (p = 0.01) were also significantly enriched in the Finnish samples. These functional categories represent the highest a priori power for downstream association studies of rare variants using population isolates.

Genetics

Genome-wide association study identifies 17 new loci influencing concentrations of circulating cytokines and growth factors

Circulating cytokines and growth factors are regulators of inflammation and have been implicated in autoimmune and metabolic diseases. In this genome-wide association study (GWAS) up to n=8,293 Finns we identified 27 loci with genome-wide association (P-value<1.2x10-9) for one or more cytokines, including 17 unidentified in previous GWASes. Fifteen of the associated SNPs had expression quantitative trait loci in whole blood. We provide strong genetic instruments to clarify the causal roles of cytokine signaling and upstream inflammation in immune-related and other chronic diseases. We further link known autoimmune disease variants including Crohn's disease, multiple sclerosis and ulcerative colitis with new inflammatory markers, which elucidate the molecular mechanisms underpinning these diseases and suggest potential drug targets.

Genomics

Genomic prediction of coronary heart disease

BackgroundGenetics plays an important role in coronary heart disease (CHD) but the clinical utility of a genomic risk score (GRS) relative to clinical risk scores, such as the Framingham Risk Score (FRS), is unclear.\n\nMethodsWe generated a GRS of 49,310 SNPs based on a CARDIoGRAMplusC4D Consortium meta-analysis of CHD, then independently tested this using five prospective population cohorts (three FINRISK cohorts, combined n=12,676, 757 incident CHD events; two Framingham Heart Study cohorts (FHS), combined n=3,406, 587 incident CHD events).\n\nResultsThe GRS was strongly associated with time to CHD event (FINRISK HR=1.74, 95% CI 1.61-1.86 per S.D. of GRS; Framingham HR=1.28, 95% CI 1.18-1.38), and was largely unchanged by adjustment for clinical risk scores or individual risk factors, including family history. Integration of the GRS with clinical risk scores (FRS and ACC/AHA13 score) improved prediction of CHD events within 10 years (meta-analysis C-index: +1.5-1.6%, P<0.001), particularly for individuals [&ge;]60 years old (meta-analysis C-index: +4.6-5.1%, P<0.001). Men in the top 20% of the GRS had 3-fold higher risk of CHD by age 75 in FINRISK and 2-fold in FHS, and attaining 10% cumulative CHD risk 18y earlier in FINRISK and 12y earlier in FHS than those in the bottom 20%. Furthermore, high genomic risk was partially compensated for by low systolic blood pressure, low cholesterol level, and non-smoking.\n\nConclusionsA GRS based on a large number of SNPs substantially improves CHD risk prediction and encodes decades of variation in CHD risk not captured by traditional clinical risk scores.

Genomics

Mergeomics: integration of diverse genomics resources to identify pathogenic perturbations to biological systems

Mergeomics is a computational pipeline (http://mergeomics.research.idre.ucla.edu/Download/Package/) that integrates multidimensional omics-disease associations, functional genomics, canonical pathways and gene-gene interaction networks to generate mechanistic hypotheses. It first identifies biological pathways and tissue-specific gene subnetworks that are perturbed by disease-associated molecular entities. The disease-associated subnetworks are then projected onto tissue-specific gene-gene interaction networks to identify local hubs as potential key drivers of pathological perturbations. The pipeline is modular and can be applied across species and platform boundaries, and uniquely conducts pathway/network level meta-analysis of multiple genomic studies of various data types. Application of Mergeomics to cholesterol datasets revealed novel regulators of cholesterol metabolism.

Systems Biology

FINEMAP: Efficient variable selection using summary data from genome-wide association studies

MotivationThe goal of fine-mapping in genomic regions associated with complex diseases and traits is to identify causal variants that point to molecular mechanisms behind the associations. Recent fine-mapping methods using summary data from genome-wide association studies rely on exhaustive search through all possible causal configurations, which is computationally expensive.\n\nResultsWe introduce FINEMAP, a software package to efficiently explore a set of the most important causal configurations of the region via a shotgun stochastic search algorithm. We show that FINEMAP produces accurate results in a fraction of processing time of existing approaches and is therefore a promising tool for analyzing growing amounts of data produced in genome-wide association studies.\n\nAvailabilityFINEMAP v1.0 is freely available for Mac OS X and Linux at http://www.christianbenner.com.\n\nContact: christian.benner@helsinki.fi, matti.pirinen@helsinki.fi

Genetics

metaCCA: Summary statistics-based multivariate meta-analysis of genome-wide association studies using canonical correlation analysis

A dominant approach to genetic association studies is to perform univariate tests between genotype-phenotype pairs. However, analysing related traits together increases statistical power, and certain complex associations become detectable only when several variants are tested jointly. Currently, modest sample sizes of individual cohorts and restricted availability of individual-level genotype-phenotype data across the cohorts limit conducting multivariate tests.\n\nWe introduce metaCCA, a computational framework for summary statistics-based analysis of a single or multiple studies that allows multivariate representation of both genotype and phenotype. It extends the statistical technique of canonical correlation analysis to the setting where original individual-level records are not available, and employs a covariance shrinkage algorithm to achieve robustness.\n\nMultivariate meta-analysis of two Finnish studies of nuclear magnetic resonance metabolomics by metaCCA, using standard univariate output from the program SNPTEST, shows an excellent agreement with the pooled individual-level analysis of original data. Motivated by strong multivariate signals in the lipid genes tested, we envision that multivariate association testing using metaCCA has a great potential to provide novel insights from already published summary statistics from high-throughput phenotyping technologies.\n\nCode is available at https://github.com/aalto-ics-kepaco.

Bioinformatics

Towards a molecular systems model of coronary artery disease

Coronary artery disease (CAD) is a complex disease driven by myriad interactions of genetics and environmental factors. Traditionally, studies have analyzed only one disease factor at a time, providing useful but limited understanding of the underlying etiology. Recent advances in cost-effective and high-throughput technologies, such as single nucleotide polymorphism (SNP) genotyping, exome/genome sequencing, gene expression microarrays and metabolomics assays have enabled the collection of millions of data points in many thousands of individuals. In order to make sense of such omics data, effective analytical methods are needed. We review and highlight some of the main results in this area, focusing on integrative approaches that consider multiple modalities simultaneously. Such analyses have the potential to uncover the genetic basis of CAD, produce genomic risk scores (GRS) for disease prediction, disentangle the complex interactions underlying disease, and predict response to treatment.

Systems Biology

Cell specific eQTL analysis without sorting cells

Expression quantitative trait locus (eQTL) mapping on tissue, organ or whole organism data can detect associations that are generic across cell types. We describe a new method to focus upon specific cell types without first needing to sort cells. We applied the method to whole blood data from 5,683 samples and demonstrate that SNPs associated with Crohn's disease preferentially affect gene expression within neutrophils.

Genetics