bioRxiv ScienceSearch

Biology subjects

Inouye, M.

Publications and source records attributed to Inouye, M..

11 recordsLinked to original sources

Elevated alpha 1 antitrypsin is a major component of GlycA-associated risk for future morbidity and mortality

Integration of electronic health records with systems-level biomolecular data has led to the discovery that GlycA, a complex nuclear magnetic resonance (NMR) spectroscopy biomarker, predicts long-term risk of disease onset and death from myriad causes. To determine the molecular underpinnings of the disease risk of the heterogeneous GlycA signal, we used machine learning to build imputation models for GlycAs constituent glycoproteins, then estimated glycoprotein levels in 11,861 adults across two population-based cohorts with long-term follow-up. While alpha-1-acid glycoprotein had the strongest correlation with GlycA, our analysis revealed that alpha-1 antitrypsin (AAT) was the most predictive of morbidity and mortality for the widest range of diseases, including heart failure (HR=1.60 per s.d., P=1x10-10), influenza and pneumonia (HR=1.37, P=6x10-10), and liver diseases (HR=1.81, P=1x10-6). Despite emerging evidence of AAT's role in suppressing inflammation, transcriptional analyses revealed elevated expression of diverse inflammatory immune pathways with elevated AAT levels, suggesting AAT is elevating to compensate for low-grade chronic inflammation. This study clarifies the molecular underpinnings of the GlycA biomarker and its associated disease risk, and indicates a previously unrecognised association between elevated AAT and severe disease onset and mortality.

systems biology

The landscape of incident disease risk for the biomarker GlycA and its mortality stratification in angiography patients

Integration of systems-level biomolecular information with electronic health records has led to the discovery of robust blood-based biomarkers predictive of future health and disease. Of recent intense interest is the GlycA biomarker, a complex nuclear magnetic resonance (NMR) spectroscopy signal reflective of acute and chronic inflammation, which predicts long term risk of diverse outcomes including cardiovascular disease, type 2 diabetes, and all-cause mortality. To systematically explore the specificity of the disease burden indicated by GlycA we analysed the risk for 468 common incident hospitalization and mortality outcomes occurring during an 8-year follow-up of 11,861 adults from Finland. Our analyses of GlycA replicated known associations, identified associations with specific cardiovascular disease outcomes, and uncovered new associations with risk of alcoholic liver disease (meta-analysed hazard ratio 2.94 per 1-SD, P=5x10-6), chronic renal failure (HR=2.47, P=3x10-6), glomerular diseases (HR=1.95, P=1x10-6), chronic obstructive pulmonary disease (HR=1.58, P=3x10-5), inflammatory polyarthropathies (HR=1.46, P=4x10-8), and hypertension (HR=1.21, P=5x10-5). We further evaluated GlycA as a biomarker in secondary prevention of 12-year cardiovascular mortality in 900 angiography patients with suspected coronary artery disease. We observed hazard ratios of 4.87 and 5.00 for 12-year mortality in angiography patients in the fourth and fifth quintiles by GlycA levels demonstrating the prognostic potential of GlycA for identification of high mortality-risk individuals. Both GlycA and C-reactive protein had shared as well as independent contributions to mortality hazard, emphasising the importance of chronic inflammation in secondary prevention of cardiovascular disease.

epidemiology

FastSpar: Rapid and scalable correlation estimation for compositional data

A common goal of microbiome studies is the elucidation of community composition and member interactions using counts of taxonomic units extracted from sequence data. Inference of interaction networks from sparse and compositional data requires specialised statistical approaches. A popular solution is SparCC, however its performance limits the calculation of interaction networks for very high-dimensional datasets. Here we introduce FastSpar, an efficient and parallelisable implementation of the SparCC algorithm which rapidly infers correlation networks and calculates p-values using an unbiased estimator. We further demonstrate that FastSpar reduces network inference wall time by 2-3 orders of magnitude compared to SparCC. FastSpar source code, precompiled binaries, and platform packages are freely available on GitHub: github.com/scwatts/FastSpar

bioinformatics

Genomic risk prediction of coronary artery disease in nearly 500,000 adults: implications for early screening and primary prevention

BackgroundCoronary artery disease (CAD) has substantial heritability and a polygenic architecture; however, genomic risk scores have not yet leveraged the totality of genetic information available nor been externally tested at population-scale to show potential utility in primary prevention.\n\nMethodsUsing a meta-analytic approach to combine large-scale genome-wide and targeted genetic association data, we developed a new genomic risk score for CAD (metaGRS), consisting of 1.7 million genetic variants. We externally tested metaGRS, individually and in combination with available conventional risk factors, in 22,242 CAD cases and 460,387 non-cases from UK Biobank.\n\nFindingsIn UK Biobank, a standard deviation increase in metaGRS had a hazard ratio (HR) of 1.71 (95% CI 1.68-1.73) for CAD, greater than any other externally tested genetic risk score. Individuals in the top 20% of the metaGRS distribution had a HR of 4.17 (95% CI 3.97-4.38) compared with those in the bottom 20%. The metaGRS had higher C-index (C=0.623, 95% CI 0.615-0.631) for incident CAD than any of four conventional factors (smoking, diabetes, hypertension, and body mass index), and addition of the metaGRS to a model of conventional risk factors increased C-index by 3.7%. In individuals on lipid-lowering or anti-hypertensive medications at recruitment, metaGRS hazard for incident CAD was significantly but only partially attenuated with HR of 2.83 (95% CI 2.61- 3.07) between the top and bottom 20% of the metaGRS distribution.\n\nInterpretationRecent genetic association studies have yielded enough information to meaningfully stratify individuals using the metaGRS for CAD risk in both early and later life, thus enabling targeted primary intervention in combination with conventional risk factors. The metaGRS effect was partially attenuated by lipid and blood pressure-lowering medication, however other prevention strategies will be required to fully benefit from earlier genomic risk stratification.\n\nFundingNational Health and Medical Research Council of Australia, British Heart Foundation, Australian Heart Foundation.

genetics

Non-parametric mixture models identify trajectories of childhood immune development relevant to asthma and allergy

Events in early life contribute to subsequent risk of asthma; however, the causes and trajectories of childhood wheeze are heterogeneous and do not always result in asthma. Similarly, not all atopic individuals develop wheeze, and vice versa. The reasons for these differences are unclear. Using unsupervised model-based cluster analysis, we identified latent clusters within a prospective birth cohort with deep immunological and respiratory phenotyping. We characterised each cluster in terms of immunological profile and disease risk, and replicated our results in external cohorts from the UK and USA. We discovered three distinct trajectories, one of which is a high-risk \"atopic\" cluster with increased propensity for allergic diseases throughout childhood. Atopy contributes varyingly to later wheeze depending on cluster membership. Our findings demonstrate the utility of unsupervised analysis in elucidating heterogeneity in asthma pathogenesis and provide a foundation for improving management and prevention of childhood asthma.

systems biology

Dynamics of the upper airway microbiome in the pathogenesis of asthma-associated persistent wheeze in preschool children

Repeated cycles of infection-associated lower airway inflammation drives the pathogenesis of persistent wheezing disease in children. Tracking these events across a birth cohort during their first five years, we demonstrate that >80% of infectious events indeed involve viral pathogens, but are accompanied by a shift in the nasopharyngeal microbiome (NPM) towards dominance by a small range of pathogenic bacterial genera. Unexpectedly, this change in NPM frequently precedes the appearance of viral pathogens and acute symptoms. In non-sensitized children these events are associated only with \"transient wheeze\" that resolves after age three. In contrast, in children developing early allergic sensitization, they are associated with ensuing development of persistent wheeze, which is the hallmark of the asthma phenotype. This suggests underlying pathogenic interactions between allergic sensitization and antibacterial mechanisms.

genetics

Power, false discovery rate and Winner’s Curse in eQTL studies

Investigation of the genetic architecture of gene expression traits has aided interpretation of disease and trait-associated genetic variants, however key aspects of expression quantitative trait (eQTL) study design and analysis remain understudied. We used extensive, empirically-driven simulations to explore eQTL study design and the performance of various analysis strategies. Across multiple testing correction methods, false discoveries of genes with eQTLs (eGenes) were substantially inflated when false discovery rate (FDR) control was applied to all tests, and only appropriately controlled using hierarchical procedures. All multiple testing correction procedures had low power and inflated FDR for eGenes whose causal SNPs had small allele frequencies using small sample sizes (e.g. frequency <10% in 100 samples), indicating that even moderately low frequency eQTL SNPs (eSNPs) in these studies are enriched for false discoveries. In scenarios with [&ge;]80% power, the top eSNP was the true simulated eSNP 90% of the time, but substantially less frequently for very common eSNPs (minor allele frequencies >25%). Overestimation of eQTL effect sizes, so-called \"Winners Curse\", was common in low and moderate power settings. To address this, we developed a bootstrap method (BootstrapQTL) which led to more accurate effect size estimation. These insights provide a foundation for future eQTL studies, especially those with sampling constraints and subtly different conditions.

genetics

Genome-wide analysis of genetic risk factors for rheumatic heart disease in Aboriginal Australians provides support for pathogenic molecular mimicry

BackgroundRheumatic heart disease (RHD) following Group A Streptococcus (GAS) infections is heritable and prevalent in Indigenous populations. Molecular mimicry between human and GAS proteins triggers pro-inflammatory cardiac valve-reactive T-cells.\n\nMethodsGenome-wide genetic analysis was undertaken in 1263 Aboriginal Australians (398 RHD cases; 865 controls). Single nucleotide polymorphisms (SNPs) were genotyped using Illumina HumanCoreExome BeadChips. Direct typing and imputation was used to fine-map the human leukocyte antigen (HLA) region. Epitope binding affinities were mapped for human cross-reactive GAS proteins, including M5 and M6.\n\nResultsThe strongest genetic association was intronic to HLA-DQA1 (rs9272622; P=1.86x10-7). Conditional analyses showed rs9272622 and/or DQA1*AA16 account for the HLA signal. HLA-DQA1*0101_DQB1*0503 (OR 1.44, 95%CI 1.09-1.90, P=9.56x10-3) and HLA-DQA1*0103_DQB1*0601 (OR 1.27, 95%CI 1.07-1.52, P=7.15x10-3) were risk haplotypes; HLA_DQA1*0301-DQB1*0402 (OR 0.30, 95%CI 0.14-0.65, P=2.36x10-3) was protective. Human myosin cross-reactive N-terminal and B repeat epitopes of GAS M5/M6 bind with higher affinity to DQA1/DQB1 alpha/beta dimers for the two risk haplotypes than the protective haplotype.\n\nConclusionsVariation at HLA_DQA1-DQB1 is the major genetic risk factor for RHD in Aboriginal Australians studied here. Cross-reactive epitopes bind with higher affinity to alpha/beta dimers formed by risk haplotypes, supporting molecular mimicry as the key mechanism of RHD pathogenesis.

genetics

FlashPCA2: principal component analysis of biobank-scale genotype datasets

MotivationPrincipal component analysis (PCA) is a crucial step in quality control of genomic data and a common approach for understanding population genetic structure. With the advent of large genotyping studies involving hundreds of thousands of individuals, standard approaches are no longer computationally feasible. We present FlashPCA2, a tool that can perform PCA on 1 million individuals faster than competing approaches, while requiring substantially less memory.\n\nAvailabilityhttps://github.com/gabraham/ashpca\n\nContactgad.abraham@unimelb.edu.au

genomics

Genomic analysis of Mycobacterium tuberculosis reveals complex etiology of tuberculosis in Vietnam including frequent introduction and transmission of Beijing lineage and positive selection for EsxW Beijing variant

Introduction Introduction Accession codes AUTHOR CONTRIBUTIONS COMPETING FINANCIAL INTERESTS Online Methods References Tuberculosis (TB) is a leading cause of death from infectious disease and the global burden is now higher than at any point in history 1,2 Despite coordinated efforts to control TB transmission, the factors contributing to its successful spread remain poorly understood. Vietnam is identified as one of 30 high burden countries for TB and MDR-TB with an incidence of 137 TB cases per 100,000 individuals in 2015 2 Recent phylogenomic analyses of the causative agent Mycobacterium tuberculosis (Mtb) in other high-prevalence regions have provided insights into the complex processes underlying TB transmission ...

genomics

An interaction map of circulating metabolites, immune gene networks and their genetic regulation

The interaction between metabolism and the immune system plays a central role in many cardiometabolic diseases. We integrated blood transcriptomic, metabolomic, and genomic profiles from two population-based cohorts, including a subset with 7-year follow-up sampling. We identified topologically robust gene networks enriched for diverse immune functions including cytotoxicity, viral response, B cell, platelet, neutrophil, and mast cell/basophil activity. These immune gene modules showed complex patterns of association with 158 circulating metabolites, including lipoprotein subclasses, lipids, fatty acids, amino acids, and CRP. Genome-wide scans for module expression quantitative trait loci (mQTLs) revealed five modules with mQTLs of both cis and trans effects. The strongest mQTL was in ARHGEF3 (rs1354034) and affected a module enriched for platelet function. Mast cell/basophil and neutrophil function modules maintained their metabolite associations during 7-year follow-up, while our strongest mQTL in ARHGEF3 also displayed clear temporal stability. This study provides a detailed map of natural variation at the blood immuno-metabolic interface and its genetic basis, and facilitates subsequent studies to explain inter-individual variation in cardiometabolic disease.

genomics