bioRxiv ScienceSearch

Biology subjects

Rivas, M. A.

Publications and source records attributed to Rivas, M. A..

8 recordsLinked to original sources

Deep learning facilitates rapid cohort identification using human and veterinary clinical narratives

ObjectiveCurrently, dedicated tagging staff spend considerable effort assigning clinical codes to patient summaries for public health purposes, and machine-learning automated tagging is bottlenecked by availability of electronic medical records. Veterinary medical records, a largely untapped data source that could benefit both human and non-human patients, could fill the gap. Materials and MethodsIn this retrospective study, we trained long short-term memory (LSTM) recurrent neural networks (RNNs) on 52,722 human and 89,591 veterinary records. We established relevant baselines by training Decision Trees (DT) and Random Forests (RF) on the same data. We finally investigated the effect of merging data across clinical settings and probed model portability. ResultsWe show that the LSTM-RNNs accurately classify veterinary/human text narratives into top-level categories with an average weighted macro F1, score of 0.735/0.675 respectively. The evaluation metric for the LSTM was 7 and 8% higher than that of the DT and RF models respectively. We generally did not find evidence of model portability albeit moderate performance increases in select categories. DiscussionWe see a strong positive correlation between number of training samples and classification performance, which is promising for future efforts. The use of LSTM-RNN models represents a scalable structure that could prove useful in cohort selection, which could in turn better address emerging public health concerns. ConclusionDigitization of human and veterinary health information will continue to be a reality. Our approach is a step forward for these two domains to learn from, and inform, one another.

epidemiology

Significant shared heritability underlies suicide attempt and clinically predicted probability of attempting suicide

Suicide accounts for nearly 800,000 deaths per year worldwide with rates of both deaths and attempts rising. Family studies have estimated substantial heritability of suicidal behavior; however, collecting the sample sizes necessary for successful genetic studies has remained a challenge. We utilized two different approaches in independent datasets to characterize the contribution of common genetic variation to suicide attempt. The first is a patient reported suicide attempt phenotype from genotyped samples in the UK Biobank (337,199 participants, 2,433 cases). The second leveraged electronic health record (EHR) data from the Vanderbilt University Medical Center (VUMC, 2.8 million patients, 3,250 cases) and machine learning to derive probabilities of attempting suicide in 24,546 genotyped patients. We identified significant and comparable heritability estimates of suicide attempt from both the patient reported phenotype in the UK Biobank (h2SNP = 0.035, p = 7.12x10-4) and the clinically predicted phenotype from VUMC (h2SNP = 0.046, p = 1.51x10-2). A significant genetic overlap was demonstrated between the two measures of suicide attempt in these independent samples through polygenic risk score analysis (t = 4.02, p = 5.75x10-5) and genetic correlation (rg = 1.073, SE = 0.36, p = 0.003). Finally, we show significant but incomplete genetic correlation of suicide attempt with insomnia (rg = 0.34 - 0.81) as well as several psychiatric disorders (rg = 0.26 - 0.79). This work demonstrates the contribution of common genetic variation to suicide attempt. It points to a genetic underpinning to clinically predicted risk of attempting suicide that is similar to the genetic profile from a patient reported outcome. Lastly, it presents an approach for using EHR data and clinical prediction to generate quantitative measures from binary phenotypes that improved power for our genetic study.

genomics

Bayesian model comparison for rare variant association studies of multiple phenotypes

Whole genome sequencing studies applied to large populations or biobanks with extensive phenotyping raise new analytic challenges. The need to consider many variants at a locus or group of genes simultaneously and the potential to study many correlated phenotypes with shared genetic architecture provide opportunities for discovery and inference that are not addressed by the traditional one variant, one phenotype association study. Here, we introduce a Bayesian model comparison approach that we refer to as MRP (Multiple Rare-variants and Phenotypes) for rare-variant association studies that considers correlation, scale, and direction of genetic effects across a group of genetic variants, phenotypes, and studies. The approach requires only summary statistic data. To demonstrate the efficacy of MRP, we apply our method to exome sequencing data (N = 184,698) across 2,019 traits from the UK Biobank, aggregating signals in genes. MRP demonstrates an ability to recover previously-verified signals such as associations between PCSK9 and LDL cholesterol levels. We additionally find MRP effective in conducting meta-analyses in exome data. Notable non-biomarker findings include associations between MC1R and red hair color and skin color, IL17RA and monocyte count, IQGAP2 and mean platelet volume, and JAK2 and platelet count and crit (mass). Finally, we apply MRP in a multi-phenotype setting; after clustering the 35 biomarker phenotypes based on genetic correlation estimates into four clusters, we find that joint analysis of these phenotypes results in substantial power gains for gene-trait associations, such as in TNFRSF13B in one of the clusters containing diabetes and lipid-related traits. Overall, we show that the MRP model comparison approach is able to improve upon useful features from widely-used meta-analysis approaches for rare variant association analyses and prioritize protective modifiers of disease risk.

genetics

Large-scale phenome-wide association study of PCSK9 loss-of-function variants demonstrates protection against ischemic stroke

PCSK9 inhibitors are a potent new therapy for hypercholesterolemia and have been shown to decrease risk of coronary heart disease. Although short-term clinical trial results have not demonstrated major adverse effects, long-term data will not be available for some time. Genetic studies in large well-phenotyped biobanks offer a unique opportunity to predict drug effects and provide context for the evaluation of future clinical trial outcomes. We tested association of the PCSK9 loss-of-function variant rsll591147 (R46L) in a hypothesis-driven 11 phenotype set and a hypothesis-generating 278 phenotype set in 337,536 individuals of British ancestry in the United Kingdom Biobank (UKB), with independent discovery (n = 225K) and replication (n = 112K). In addition to the known association with lipid levels (OR 0.63) and coronary heart disease (OR 0.73), the T allele of rs11591147 showed a protective effect on ischemic stroke (OR 0.61, p = 0.002) but not hemorrhagic stroke in the hypothesis-driven screen. We did not observe an association with type 2 diabetes, cataracts, heart failure, atrial fibrillation, and cognitive dysfunction. In the phenome-wide screen, the variant was associated with a reduction in metabolic disorders, ischemic heart disease, coronary artery bypass graft operations, percutaneous coronary interventions and history of angina. A single variant analysis of UKB data using TreeWAS, a Bayesian analysis framework to study genetic associations leveraging phenotype correlations, also showed evidence of association with cerebral infarction and vascular occlusion. This result represents the first genetic evidence in a large cohort for the protective effect of PCSK9 inhibition on ischemic stroke, and corroborates exploratory evidence from clinical trials. PCSK9 inhibition was not associated with variables other than those related to low density lipoprotein cholesterol and atherosclerosis, suggesting that other effects are either small or absent.

genetics

Vulnerabilities of transcriptome-wide association studies

Transcriptome-wide association studies (TWAS) integrate GWAS and gene expression datasets to find gene-trait associations. In this Perspective, we explore properties of TWAS as a potential approach to prioritize causal genes, using simulations and case studies of literature-curated candidate causal genes for schizophrenia, LDL cholesterol and Crohns disease. We explore risk loci where TWAS accurately prioritizes the likely causal gene, as well as loci where TWAS prioritizes multiple genes, some of which are unlikely to be causal, because they share the same variants as eQTLs. We illustrate that TWAS is especially prone to spurious prioritization when using expression data from tissues or cell types that are less related to the trait, due to substantial variation in both expression levels and eQTL strengths across cell types. Nonetheless, TWAS prioritizes candidate causal genes at GWAS loci more accurately than simple baselines based on proximity to lead GWAS variant and expression in trait-related tissue. We discuss current strategies and future opportunities for improving the performance of TWAS for causal gene prioritization. Our results showcase the strengths and limitations of using expression variation across individuals to determine causal genes at GWAS loci and provide guidelines and best practices when using TWAS to prioritize candidate causal genes.

genetics

Medical relevance of protein-truncating variants across 337,208 individuals in the UK Biobank study

Protein-truncating variants can have profound effects on gene function and are critical for clinical genome interpretation and generating therapeutic hypotheses, but their relevance to medical phenotypes has not been systematically assessed. We characterized the effect of 18,228 protein-truncating variants across 135 phenotypes from the UK Biobank and found 27 associations between medical phenotypes and protein-truncating variants in genes outside the major histocompatibility complex. We performed phenome-wide analyses and directly measured the effect of homozygous carriers, commonly referred to as \"human knockouts,\" across medical phenotypes for genes implicated to be protective against disease or associated with at least one phenotype in our study and found several genes with strong pleiotropic or non-additive effects. Our results illustrate the importance of protein-truncating variants in a variety of diseases.

genetics

Base-Specific Mutational Intolerance Near Splice-Sites Clarifies Role Of Non-Essential Splice Nucleotides

Variation in RNA splicing (i.e., alternative splicing) plays an important role in many diseases. Variants near 5' and 3' splice sites often affect splicing, but the effects of these variants on splicing and disease have not been fully characterized beyond the 2 \"essential\" splice nucleotides flanking each exon. Here we provide quantitative measurements of tolerance to mutational disruptions by position and reference allele-alternative allele combination. We show that certain reference alleles are particularly sensitive to mutations, regardless of the alternative alleles into which they are mutated. Using public RNA-seq data, we demonstrate that individuals carrying such variants have significantly lower levels of the correctly spliced transcript compared to individuals without them, and confirm that these specific substitutions are highly enriched for known Mendelian mutations. Our results propose a more refined definition of the \"splice region\" and offer a new way to prioritize and provide functional interpretation of variants identified in diagnostic sequencing and association studies.

genomics

biMM: Efficient estimation of genetic variances andcovariances for cohorts with high-dimensional phenotype measurements

Genetic research utilizes a decomposition of trait variances and covariances into genetic and environmental parts. Our software package biMM is a computationally efficient implementation of a bivariate linear mixed model for settings where hundreds of traits have been measured on partially overlapping sets of individuals.\n\nAvailabilityImplementation in R freely available at www.iki.fi/mpirinen.

genetics