bioRxiv Science⌕ Search

SEARCH · bioRxiv Science

Results for “Molecular Biology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,675 records · Page 93Linked to original sources

Spatial visualization of A-to-I Editing in cells using Endonuclease V Immunostaining Assay (EndoVIA)

Adenosine-to-Inosine (A-to-I) editing is one of the most widespread post-transcriptional RNA modifications and is catalyzed by adenosine deaminases acting on RNA (ADARs). Varying across tissue types, A-to-I editing is essential for numerous biological functions and dysregulation leads to autoimmune and neurological disorders, as well as cancer. Recent evidence has also revealed a link between RNA localization and A-to-I editing, yet understanding of the mechanisms underlying this relationship and its biological impact remains limited. Current methods rely primarily on in vitro characterization of extracted RNA that ultimately erases subcellular localization and cell-to-cell heterogeneity. To address these challenges, we have repurposed Endonuclease V (EndoV), a magnesium dependent ribonuclease that cleaves inosine bases in edited RNA, to selectively bind and detect A-to-I edited RNA in cells. The work herein introduces Endonuclease V Immunostaining Assay (EndoVIA), a workflow that provides spatial visualization of edited transcripts, enables rapid quantification of overall inosine abundance, and maps the landscape of A-to-I editing within the transcriptome at the nanoscopic level.

molecular biology↗

Arabidopsis SWR1-associated protein methyl-CpG-binding domain 9 is required for histone H2A.Z deposition.

Deposition of the histone variant H2A.Z by the SWI2/SNF2-Related 1 chromatin remodeling complex (SWR1-C) is important for gene regulation in eukaryotes, but the composition of the Arabidopsis SWR1-C has not been thoroughly characterized. Here identify interacting partners of a conserved Arabidopsis SWR1 subunit, ACTIN-RELATED PROTEIN 6 (ARP6). We isolated nine predicted components, and identified additional interactors implicated in histone acetylation and chromatin biology. One of the novel interacting partners, methyl-CpG-binding domain 9 (MBD9), also strongly interacted with the Imitation SWItch (ISWI) chromatin remodeling complex. MBD9 was required for deposition of H2A.Z at a distinct subset of ARP6-dependent loci. MBD9 was preferentially bound to nucleosome-depleted regions at the 5 ends of genes containing high levels of activating histone marks. These data suggest that MBD9 is a SWR1-C interacting protein required for H2A.Z deposition at a subset of actively transcribing genes.

molecular biology↗

CRISPR/Cas9-mediated germline mutagenesis in the subsocial parasitoid wasp, Sclerodermus guani

The ectoparasitoid wasp Sclerodermus guani (Hymenoptera: Bethylidae), as a subsocial insect, is widely applied in biological control against beetle vectors of pine wood nematodes. Despite significant advances in behavioral research, functional genetics in S. guani remains underdeveloped due to the absence of efficient gene manipulation tools. In this study, we employed CRISPR-mediated mutagenesis to achieve germline gene knockout targeting the eye pigment associated gene kynurenine 3-monooxygenase (KMO). Phylogenetic analysis revealed that S. guani KMO shares a close relationship with its homolog in Prorops nasuta (Hymenoptera: Bethylidae). Two single-guide RNAs (sgRNAs), coupled with Cas9 protein with and without nuclear localization signal (NLS) were tested. Both sgRNAs induced specific in vitro DNA cleavage and in vivo heritable indels at the target genomic loci. Homozygous null mutant females and males exhibit a white-eye phenotype, which was identified during pupal stage. Optimal editing efficiency in vivo was achieved using the Cas9-NLS variant. Given the complication of germline gene editing in eusocial Hymenopterans, the application of CRISPR in the subsocial parasitoid wasp S. guani provides an accessible research platform for the molecular evolution of insect sociality.

molecular biology↗

Loss of CDC50A function drives Aβ/p3 production via increased β/α-secretase processing of APP

The Amyloid Precursor Protein (APP) undergoes extensive proteolytic processing to produce several biologically active metabolites which affect Alzheimers disease (AD) pathogenesis. Sequential cleavage of APP by {beta}- and {gamma}-secretases results in A{beta}, while cleavage by - and {gamma}-secretases produces the smaller p3 peptide. Here we report that in cells in which the P4-ATPase flippase subunit CDC50A has been knocked out, large increases in the products of {beta}- and -secretase cleavage of APP (sAPP{beta}/{beta}CTF and sAPP/CTF, respectively) and the downstream metabolites A{beta} and p3 are seen. These data indicate that APP cleavage by {beta}/-secretase are increased and suggest that phospholipid asymmetry plays an important role in APP metabolism and A{beta} production.

molecular biology↗

Csp1, A Cold-Shock Protein Homolog in Xylella fastidiosa Influences Pili Formation, Stress Response, and Gene Expression

Bacterial cold shock-domain proteins (CSPs) are conserved nucleic acid binding chaperones that play important roles in stress adaptation and pathogenesis. Csp1 is a temperature-independent cold shock protein homolog in Xylella fastidiosa, a bacterial plant pathogen of grapevine and other economically important crops. Csp1 contributes to stress tolerance and virulence in X. fastidiosa. However, besides general single stranded nucleic acid binding activity, little is known about the specific function(s) of this protein. To further investigate the role(s) of Csp1, we compared phenotypic differences between wild type and a csp1 deletion mutant ({Delta}csp1). We observed decreases in cellular aggregation and surface attachment with the {Delta}csp1 strain compared to the wild type. Transmission electron microscopy imaging revealed that {Delta}csp1 had reduced pili compared to the wild type and complemented strains. The {Delta}csp1 strain also showed reduced survival after long term growth, in vitro. Since Csp1 binds DNA and RNA, its influence on gene expression was also investigated. Long-read Nanopore RNA-Seq analysis of wild type and {Delta}csp1 revealed changes in expression of several genes important for attachment and biofilm formation in {Delta}csp1. One gene of intertest, pilA1, encodes a type IV pili subunit protein and was up regulated in {Delta}csp1. Deleting pilA1 increased surface attachment in vitro and reduced virulence in grapevines. X. fastidiosa virulence depends on bacterial attachment to host tissue and movement within and between xylem vessels. Our results show Csp1 may play a role in both virulence and stress tolerance by influencing expression of genes important for biofilm formation. ImportanceXylella fastidiosa is a major threat to the worldwide agriculture industry (1, 2). Despite its global importance, many aspects of X. fastidiosa biology and pathogenicity are poorly understood. There are currently few effective solutions to suppress X. fastidiosa disease development or eliminate bacteria from infected plants(3). Recently, disease epidemics due to X. fastidiosa have greatly expanded(2, 4, 5), exacerbating the need for better disease prevention and control strategies. Our studies show that Csp1 is involved in X. fastidiosa virulence and stress tolerance. Understanding how Csp1 influences pathogenesis and bacteria survival can aide in developing novel pathogen and disease control strategies. We also streamlined a bioinformatics protocol to process and analyze long read Nanopore bacterial RNA-Seq data, which has previously not been reported for X. fastidiosa.

molecular biology↗

Generalizable prediction of liquid-liquid phase separation from protein sequence

Liquid-liquid phase separation (LLPS) is emerging as a fundamental process supporting multiple facets of biological systems. This phenomenon enables the dynamic compartmentalization of biomolecules contributing to a wide range of cellular functions, though in many instances its precise role and evolution remain unclear. Protein phase separation naturally occurs within cells and is prevalent across all species. Despite a recent surge in protein LLPS discovery, current predictive models lack generalizability and fail to identify the full spectrum of phase-separating proteins. To address this shortcoming, we developed Phaseek, a hybrid model integrating contextual sequence encoding with statistical graph representations to score LLPS propensity of amino acid sequences. Phaseek accurately identifies phase-separating proteins across diverse biological contexts, predicting key functional regions and the effects of point mutations. Proteome-wide predictions for 18 species highlight important physicochemical features. Gene Ontology enrichments recapitulate known processes (e.g., nucleic acid binding, nuclear localization, chromatin organization) and suggest novel areas of investigation. Phylogenetic analysis of orthologs further suggests that LLPS is evolutionarily conserved beyond sequence similarity. In addition, we used Phaseek to design de novo phase-separating peptides and achieved a 70% success rate in vivo. Provided with a user-friendly implementation, Phaseek serves as a multipurpose LLPS predictor for advancing both fundamental and applied LLPS research.

molecular biology↗

Identification of HIV Tat and NF-κB binding proteins associated with semen-derived extracellular vesicles

Semen-derived extracellular vesicles (SEVs) have been shown to inhibit transactivation of the long terminal repeat (LTR) in human immunodeficiency virus type 1 (HIV-1, or HIV) and, hence, viral replication by blocking the interaction of the viruss transcriptional activator Tat and host transcription factors NF-{kappa}B and Sp1. The ability of SEVs to regulate the activities of transcription factors suggests that SEVs may contain transcription activators and repressors. Here, we identified host proteins in human SEVs that interacted with the Tat and NF-{kappa}B subunit p65. Integrative network and pathway enrichment analyses of these complexes revealed associations with an array of biological functions regulating genome transcription. In particular, several proteins in SEVs could bind to both Tat and NF-{kappa}B: the scaffolding and cell signaling regulatory protein AKAP9, the G protein signaling regulator ARHGEF28, the small nuclear RNA processor INTS1, the epigenetic reader BRD2, and the transcription elongation inhibitor NELFB. NF-{kappa}B p65-bound NELFB also interacted with HEXIM1, another transcription elongation inhibitor, suggesting that SEVs may inhibit HIV propagation through networks of transcriptional regulation and repression. One Sentence SummaryProteins in vesicles shed from human semen may repress HIV by targeting transcription factors.

molecular biology↗

A high-throughput microbial glycomics platform for prebiotic development

The mammalian intestine contains diverse carbohydrate pools that govern the gut microbiome composition. Structurally distinct polysaccharides, also called glycans, are differentially consumed by gut microbial subsets and direct their abundance by controlling gene expression and metabolite production. Therefore, identifying gut microbial accessible carbohydrates (MACs) is necessary to develop new prebiotics that beneficially manipulate the gut microbiome. However, no methods exist to efficiently examine MACs in biologically-derived mixtures. Here, we present a high-throughput platform to detect MACs from various plant, animal, and microbial sources using a genome-wide library of engineered Bacteroides thetaiotaomicron (Bt) strains that harness their endogenous glycan detection machinery. We demonstrate that this platform exhibits specific and sensitive responses to glycan mixtures and use bacterially-encoded proteins to characterize a previously unknown MAC from yeast. Expanding this technology across gut Bacteroides species will generate a broadly applicable approach to characterize heterogeneous glycan mixtures and identify prebiotic substrates.

molecular biology↗

Impact of Environmental Variables on the Seasonal Dynamics and Relative Abundance of Endosymbionts in Glossina Species in Northern Nigeria

This study explores how environmental factors influence the seasonal patterns and endosymbiont abundance in Glossina species across four ecological zones in northern Nigeria. Tsetse flies--the vectors of African trypanosomiasis--host both obligate and facultative endosymbionts, including Wigglesworthia glossinidia, Sodalis glossinidius, and Wolbachia pipientis, which affect their physiology and vector capacity. Over two years, 7,632 tsetse flies were collected and examined for species distribution, symbiont prevalence, and local climate variables (temperature, humidity, and vegetation index). Glossina tachinoides was most prevalent (55.78%), followed by G. morsitans submorsitans (29.36%) and G. palpalis palpalis (14.86%), each showing site-specific distribution. Endosymbiont prevalence rose markedly during the wet season (e.g., Yankari: 73.41% to 94.83%, p < 0.0001), especially for Sodalis, which declined under dry conditions. Strong negative correlations with temperature (r = -0.99, p = 0.0001) and positive correlations with humidity (r = +0.99, p = 0.0005) were observed. These patterns reflect the vulnerability of tsetse-symbiont systems to climate stress and underscore challenges for control strategies. Broad-area methods may suit G. tachinoides, while targeted trapping is more suitable for G. palpalis. The resilience of Wolbachia suggests its utility for paratransgenic control. The study emphasizes the importance of integrated, climate-aware surveillance and intervention strategies to mitigate trypanosomiasis risk. Author SummaryTsetse flies transmit African trypanosomiasis, a disease affecting livestock and rural livelihoods in sub-Saharan Africa. These flies harbor bacterial endosymbionts that influence their biology and vector competence. This study investigated how seasonal changes in temperature, humidity, and vegetation affect the abundance of tsetse flies and their symbionts in northern Nigeria. We found that certain symbionts, especially Sodalis glossinidius, decline sharply during hot-dry periods, while Wolbachia remains stable. Our findings highlight the need for adaptive, climate-informed vector control strategies and provide insights for integrating microbial monitoring into trypanosomiasis surveillance systems.

molecular biology↗

Expression of a Malassezia codon optimized mCherry fluorescent protein in a bicistronic vector

The use of fluorescent proteins allows a multitude of approaches from live imaging and fixed cells to labelling of whole organisms, making it a foundation of diverse experiments. Tagging a protein of interest or specific cell type allows visualization and studies of cell localization, cellular dynamics, physiology, and structural characteristics. In specific instances fluorescent fusion proteins may not be properly functional as a result of structural changes that hinder protein function, or when overexpressed may be cytotoxic and disrupt normal biological processes. In our study, we describe application of a bicistronic vector incorporating a Picornavirus 2A peptide sequence between a NAT antibiotic selection marker and mCherry. This allows expression of multiple genes from a single open reading frame and production of discrete protein products through a cleavage event within the 2A peptide. We demonstrate integration of this bicistronic vector into a model Malassezia species, the haploid strain M. furfur CBS 14141, with both active selection, high fluorescence, and proven proteolytic cleavage. Potential applications of this technology can include protein functional studies, Malassezia cellular localization, and co-expression of genes required for targeted mutagenesis.

molecular biology↗

Proteomic Analysis of Urine from Youths Indulging in Gaming

Video game addiction manifests as an escalating enthusiasm and uncontrolled use of digital games, yet there are no objective indicators for gaming addiction. This study employed mass spectrometry proteomics to analyze the proteomic differences in the urine of adolescents addicted to gaming compared to those who do not play video games. The study included 10 adolescents addicted to gaming and 9 non-gaming adolescents as a control group. The results showed that there were 125 significantly different proteins between the two groups. Among these, 11 proteins have been reported to change in the body after the intake of psychotropic drugs and are associated with addiction: Calmodulin, ATP synthase subunit alpha, ATP synthase subunit beta, Acid ceramidase, Tomoregulin-2, Calcitonin, Apolipoprotein E, Glyceraldehyde-3-phosphate dehydrogenase, Heat shock protein beta-1, CD63 antigen, Ephrin type-B receptor 4, Tomoregulin-2. Additionally, several proteins were found to interact with pathways related to addiction: Dickkopf-related protein 3, Nicastrin, Leucine-rich repeat neuronal protein 4, Cerebellin-4. Enriched biological pathways discovered include those related to nitric oxide synthase, amphetamine addiction, and numerous calcium ion pathways, all of which are associated with addiction. Moreover, through the analysis of differentially expressed proteins, we speculated about some proteins not yet fully studied, which might play a significant role in the mechanisms of addiction: Protein kinase C and casein kinase substrate in neurons protein, Cysteine-rich motor neuron 1 protein, Bone morphogenetic protein receptor type-2, Immunoglobulin superfamily member 8. In the analysis of urinary proteins in adolescents addicted to online gaming, we identified several proteins that have previously been reported in studies of drug addiction.

molecular biology↗

Mathematical Programming and Graph Neural Networks illuminate functional heterogeneity of pathways in disease

We employ a computational framework that integrates mathematical programming and graph neural networks to elucidate functional phenotypic heterogeneity in disease by classifying entire pathways under various conditions of interest. Our approach combines two distinct, yet seamlessly integrated, modeling schemes. First, we leverage Prior Knowledge Networks (PKNs) to reconstruct gene networks from genomic and transcriptomic data. We demonstrate how this can be achieved through mathematical programming optimization and provide examples using comprehensive established databases. We then tailor Graph Neural Networks (GNNs) to classify each network as a single data point at graph-level, using various node embeddings and edge attributes. These networks may vary in their biological or molecular annotations, which serves as a labeling scheme for their supervised classification. We apply the framework to the human DNA damage and repair pathway using the TP53 regulon in a pancancer study across cell-lines and tumor samples to classify Gene Regulatory Networks (GRNs) across different TP53 mutation types. This approach allows us to identify mutations with distinguishable functional profiles which can be related to specific phenotypes, thus providing a data-driven pipeline for genotype-to-phenotype translation. This scalable approach enables the classification of diverse conditions within the multi-factorial nature of diseases and disentangles their polygenic complexity by revealing new functional patterns through a causal representation.

systems biology↗

Cryo-EM structure of the vault from human brain reveals symmetry mismatch at its caps

The vault protein is expressed in most eukaryotic cells, where it is assembled on polyribosomes into large hollow barrel-shaped complexes. Despite its widespread and abundant presence in cells, the biological function of the vault remains unclear. In this study, we describe the cryo-EM structure of vault particles that were imaged as a contamination of a preparation to extract tau filaments from brain tissue of an individual with progressive supranuclear palsy (PSP). We identify a mechanism of symmetry mismatch at the caps of the vault, from 39-fold to 13-fold symmetry, where two out of three monomers are sequentially excluded from the cap, resulting in a narrow, greasy pore at the tip of the vault. Our structure offers valuable insights for engineering carboxy-terminal modifications of the major vault protein (MVP) for potential therapeutic applications.

molecular biology↗

Structural basis for terminal loop recognition and processing of pri-miRNA-18a by hnRNP A1

Post-transcriptional mechanisms play a predominant role in the control of microRNA (miRNA) production. Recognition of the terminal loop of precursor miRNAs by RNA-binding proteins (RBPs) influences their processing; however, the mechanistic and structural basis for how levels of individual or subsets of miRNAs are regulated is mostly unexplored. We previously described a role for hnRNP A1, an RBP implicated in many aspects of RNA processing, as an auxiliary factor that promotes the Microprocessor-mediated processing of pri-mir-18a. Here, we reveal the mechanistic basis for this stimulatory role of hnRNP A1 by combining integrative structural biology with biochemical and functional assays. We demonstrate that hnRNP A1 forms a 1:1 complex with pri-mir-18a that involves binding of both RNA recognition motifs (RRMs) to cognate RNA sequence motifs in the conserved terminal loop of pri-mir-18a. Terminal loop binding induces an allosteric destabilization of base-pairing in the pri-mir-18a stem that promotes its down-stream processing. Our results highlight terminal loop RNA recognition by RNA-binding proteins as a general principle of miRNA biogenesis and regulation.

molecular biology↗

RNA-seq gene expression profiling of the bladder in a mouse model of urogenital schistosomiasis

Background: Parasitic flatworms of the Schistosoma genus cause schistosomiasis, which affects over 230 million people. Schistosoma haematobium causes the urogenital form of schistosomiasis (UGS), which can lead to hematuria, fibrosis, and increased risk of secondary infections by bacteria or viruses. UGS is also linked to bladder cancer. To understand the bladder pathology during S. haematobium infection, our group previously developed a mouse model that involves the injection of S. haematobium eggs into the bladder wall. Using this model, we studied changes in epigenetics profile, as well as changes in gene and protein expression in the host bladder tissues. In the current study, we expand upon this work by examining the expression level of both host and parasite genes using RNA sequencing (RNA-seq) in the mouse bladder wall injection model of S. haematobium infection. Methods: We used a mouse model of S. haematobium infection in which parasite eggs or vehicle control were injected into the bladder walls of female BALB/c mice. RNA-seq was performed on the RNA isolated from the bladders four days after bladder wall injection. Results/Conclusions: RNA-seq analysis of egg- and vehicle control-injected bladders revealed the differential expression of 1025 mouse genes in the egg-injected bladders, including genes associated with cellular infiltration, immune cell chemotaxis, cytokine signaling, and inflammation We also observed the upregulation of immune checkpoint-related genes, which suggests that while the infection causes an inflammatory response, it also dampens the response to avoid excessive inflammation-related damage to the host. Identifying these changes in host signaling and immune responses improves our understanding of the infection and how it may contribute to the development of bladder cancer. Analysis of the differential gene expression of the parasite eggs between bladder-injected versus uninjected eggs revealed 119 S. haematobium genes associated with transcription, intracellular signaling, and metabolism. The analysis of the parasite genes also revealed fewer transcript reads compared to that found in the analysis of mouse genes, highlighting the challenges of studying parasite egg biology in the mouse model of S. haematobium infection. Author summaryMore than 230 million people worldwide are estimated to carry infection with parasites belonging to the Schistosoma genus, which cause morbidity associated with parasite egg deposition. Praziquantel, the drug of choice to treat the infection, does not prevent reinfection, and its decades-long history as the main treatment raises concerns for drug resistance. Of the schistosome species, Schistosoma haematobium causes urogenital disease and has a strong association with bladder cancer. The possibility for drug resistance and the gap in knowledge with respect to the mechanisms driving S. haematobium-related bladder cancer highlight the need to better understand the biology of the infection to aid in the development of new therapeutic strategies. In this study, we used a mouse model of S. haematobium infection that delivers parasite eggs directly to the host mouse bladder wall, and we examined the changes in the gene expression profile of the host and the parasite by RNA-sequencing. The results corroborated previous findings with respect to the hosts inflammatory responses against the parasite eggs, as well as revealed alterations in other immune response genes that deepen our understanding of the mechanisms involved in urogenital schistosomiasis pathogenesis.

molecular biology↗

Identification of Disease-relevant, Sex-based Proteomic Differences in iPSC-derived Vascular Smooth Muscle

The prevalence of cardiovascular disease varies with sex, and the impact of intrinsic sex-based differences on vasculature is not well understood. Animal models can provide important insight into some aspects of human biology, however not all discoveries in animal systems translate well to humans. To explore the impact of chromosomal sex on proteomic phenotypes, we used iPSC-derived vascular smooth muscle cells from healthy donors of both sexes to identify sex-based proteomic differences and their possible effects on cardiovascular pathophysiology. Our analysis confirmed that differentiated cells have a proteomic profile more similar to healthy primary aortic smooth muscle than iPSCs. We also identified sex-based differences in iPSC- derived vascular smooth muscle in pathways related to ATP binding, glycogen metabolic process, and cadherin binding as well as multiple proteins relevant to cardiovascular pathophysiology and disease. Additionally, we explored the role of autosomal and sex chromosomes in protein regulation, identifying that proteins on autosomal chromosomes also show sex-based regulation that may affect the protein expression of proteins from autosomal chromosomes. This work supports the biological relevance of iPSC-derived vascular smooth muscle cells as a model for disease, and further exploration of the pathways identified here can lead to the discovery of sex-specific pharmacological targets for cardiovascular disease. SignificanceIn this work, we have differentiated 4 male and 4 female iPSC lines into vascular smooth muscle cells, giving us the ability to identify statistically-significant sex-specific proteomic markers that are relevant to cardiovascular disease risk (such as PCK2, MTOR, IGFBP2, PTGR2, and SULTE1).

molecular biology↗

Maximizing Quantitative Phosphoproteomics of Kinase Signaling Expands the Mec1 and Tel1 Networks

Global phosphoproteome analysis is crucial for comprehensive and unbiased investigation of kinase-mediated signaling. However, since each phosphopeptide represents a unique entity for defining identity, site-localization, and quantitative changes, phosphoproteomics often suffers from lack of redundancy and statistical power for generating high confidence datasets. Here we developed a phosphoproteomic approach in which data consistency among experiments using reciprocal stable isotope labeling defines a central filtering rule for achieving reliability in phosphopeptide identification and quantitation. We find that most experimental error or biological variation in phosphopeptide quantitation does not revert in quantitation once light and heavy media are swapped between two experimental conditions. Exclusion of non-reverting data-points from the dataset not only reduces quantitation error and variation, but also drastically reduces false positive identifications. Application of our approach in combination with extensive fractionation of phosphopeptides by HILIC identifies new substrates of the Mec1 and Tel1 kinases, expanding our understanding of the DNA damage signaling network regulated by these kinases. Overall, the proposed quantitative phosphoproteomic approach should be generally applicable for investigating kinase signaling networks with high confidence and depth.

molecular biology↗

Meta-analysis of Gene Expression Microarray Datasets in Chronic Obstructive Pulmonary Disease

Chronic obstructive pulmonary disease (COPD) was classified by the Centers for Disease Control and Prevention in 2014 as the 3rd leading cause of death in the United States (US). The main cause of COPD is exposure to tobacco smoke and air pollutants. Problems associated with COPD include under-diagnosis of the disease and an increase in the number of smokers worldwide. The goal of our study is to identify disease variability in the gene expression profiles of COPD subjects compared to controls. We used pre-existing, publicly available microarray expression datasets to conduct a meta-analysis. Our inclusion criteria for microarray datasets selected for smoking status, age and sex of blood donors reported. Our datasets used Affymetrix, Agilent microarray platforms (7 datasets, 1,262 samples). We re-analyzed the curated raw microarray expression data using R packages, and used Box-Cox power transformations to normalize datasets. To identify significant differentially expressed genes we ran an analysis of variance with a linear model with disease state, age, sex, smoking status and study as effects that also included binary interactions. We found 1,513 statistically significant (Benjamini-Hochberg-adjusted p-value <0.05) differentially expressed genes with respect to disease state (COPD or control). We further filtered these genes for biological effect using results from a Tukey test post-hoc analysis (Benjamini-Hochberg-adjusted p-value <0.05 and 10% two-tailed quantiles of mean differences between COPD and control), to identify 304 genes. Through analysis of disease, sex, age, and also smoking status and disease interactions we identified differentially expressed genes involved in a variety of immune responses and cell processes in COPD. We also trained a logistic regression model using the 304 genes as features, which enabled prediction of disease status with 84% accuracy. Our results give potential for improving the diagnosis of COPD through blood and highlight novel gene expression disease signatures.

molecular biology↗