bioRxiv ScienceSearch

Biology subjects

Wu, P.

Publications and source records attributed to Wu, P..

9 recordsLinked to original sources

Learning from Longitudinal Data in Electronic Health Record and Genetic Data to Improve Cardiovascular Event Prediction

BackgroundCurrent approaches to predicting Cardiovascular disease rely on conventional risk factors and cross-sectional data. In this study, we asked whether: i) machine learning and deep learning models with longitudinal EHR information can improve the prediction of 10-year CVD risk, and ii) incorporating genetic data can add values to predictability.\n\nMethodsWe conducted two experiments. In the first experiment, we modeled longitudinal EHR data with aggregated features and temporal features. We applied logistic regression (LR), random forests (RF) and gradient boosting trees (GBT) and Convolutional Neural Networks (CNN) and Recurrent Neural Networks, using Long Short-Term Memory (LSTM) units. In the second experiment, we proposed a late-fusion framework to incorporate genetic features.\n\nResultsOur study cohort included 109, 490 individuals (9,824 were cases and 99, 666 were controls) from Vanderbilt University Medical Centers (VUMC) de-identified EHRs. American College of Cardiology and the American Heart Association (ACC/AHA) Pooled Cohort Risk Equations had areas under receiver operating characteristic curves (AUROC) of 0.732 and areas under receiver under precision and recall curves (AUPRC) of 0.187. LSTM, CNN and GBT with temporal features achieved best results, which had AUROC of 0.789, 0.790, and 0.791, and AUPRC of 0.282, 0.280 and 0.285, respectively. The late fusion approach achieved a significant improvement for the prediction performance.\n\nConclusionsMachine learning and deep learning with longitudinal features improved the 10-year CVD risk prediction. Incorporating genetic features further enhanced 10-year CVD prediction performance, underscoring the importance of integrating relevant genetic data whenever available in the context of routine care.

epidemiology

Using Topic Modeling via Non-negative Matrix Factorization to Identify Relationships between Genetic Variants and Disease Phenotypes: A Case Study of Lipoprotein(a) (LPA)

Genome-wide and phenome-wide association studies are commonly used to identify important relationships between genetic variants and phenotypes. Most of these studies have treated diseases as independent variables and suffered from heavy multiple adjustment burdens due to the large number of genetic variants and disease phenotypes. In this study, we propose using topic modeling via non-negative matrix factorization (NMF) for identifying associations between disease phenotypes and genetic variants. Topic modeling is an unsupervised machine learning approach that can be used to learn the semantic patterns from electronic health record data. We chose rs10455872 in LPA as the predictor since it has been shown to be associated with increased risk of hyperlipidemia and cardiovascular diseases (CVD). Using data of 12,759 individuals from the biobank at Vanderbilt University Medical Center, we trained a topic model using NMF from 1,853 distinct phecodes extracted from the cohorts electronic health records and generated six topics. We quantified their associations with rs10455872 in LPA. Topics indicating CVD had positive correlations with rs10455872 (P < 0.001), replicating a previous finding. We also identified a negative correlation between LPA and a topic representing lung cancer (P < 0.001). Our results demonstrate the applicability of topic modeling in exploring the relationship between the genome and clinical diseases.\n\nAuthor summaryIdentifying the clinical associations of genetic variants remains crucial in understanding how the human genome modulates disease risk. Traditional phenome-wide association studies consider each disease phenotype as an independent variable, however, diseases often present as complex clusters of comorbid conditions. In this study, we propose using topic modeling to model electronic health record data as a mixture of topics (e.g., disease clusters or relevant comorbidities) and testing associations between topics and genetic variants. Our results demonstrated the feasibility of using topic modeling to replicate and discover novel associations between the human genome and clinical diseases.

genetics

Bacterial Glycosyltransferase-mediated Cell-surface Chemoenzymatic Glycan Editing: Methods and Applications

AbstractChemoenzymatic glycan editing that modifies glycan structures directly on the cell surface has emerged as a complementary tool to metabolic oligosaccharide engineering. In this article, we report the discovery that three bacterial enzymes--Pasteurella multocida 2-3-sialyltransferase M144D mutant (Pm2,3ST-M144D), Photobacterium damsel 2-6-sialyltransferase (Pd2,6ST) and Helicobacter mustelae 1-2-fucosyltransferase (Hm1,2FT)--can serve as highly efficient tools for cell-surface glycan editing. Among these three enzymes, the two sialyltransferases were also found to be tolerant to large substituents introduced to the C-5 position of the cytidine monophosphate N-acetylneuraminic acid donor, including biotin and fluorescent dyes. Combining these enzymes with our previously discovered Helicobacter pylori 1-3-FT, we developed a live cell-based assay to probe host-cell glycan-mediated influenza A virus (IAV) infection including both wild-type and mutant strains of human H1N1 and H3N2 influenza subtypes. At high SiaNAc2-6-Gal levels, the ability of a viral strain to induce the host cell death is positively correlated with the SiaNAc2-6-Gal binding affinity of its haemagglutinin. Surprisingly, the creation of sLeX on the host cell surface via in situ 1-3-Fuc editing also exacerbated the killing induced by several wild-type IAV strains as well as a mutant known as HK68-MTA. Structural alignment of HAs from the wild-type HK68 and HK68-MTA revealed the formation of a putative hydrogen bond between Trp222 of HA-HK68-MTA and the C-4 hydroxyl group of the 1-3-linked fucose of sLeX. This interaction is likely to be responsible for the better binding affinity of HA-HK68-MTA to sLeX and accordingly the enhanced host-cell killing compared with the wild-type HK68.

biochemistry

The Biological Evaluation of Fusidic Acid and Its Hydrogenation Derivative as Antimicrobial and Anti-inflammatory Agents

Fusidic acid (WU-FA-00) is the only commercially available antimicrobial from the fusidane family that has a narrow spectrum of activity against Gram-positive bacteria. Herein, the hydrogenation derivative (WU-FA-01) of fusidic acid was prepared, and both compounds were examined against a panel of six bacterial strains. In addition, their anti-inflammation properties were evaluated using a 12-O-tetradecanoylphorbol-13-acetate (TPA)-induced mouse ear edema model. The results of the antimicrobial assay revealed that both WU-FA-00 and WU-FA-01 displayed a high level of antimicrobial activity against Gram-positive strains. Moreover, killing kinetic studies were performed, and the results were in accordance with the MIC and MBC results. We also demonstrated that the topical application of WU-FA-00 and WU-FA-01 effectively decreased TPA-induced ear edema in a dose-dependent manner. This inhibitory effect was associated with the inhibition of TPA-induced up-regulation of pro-inflammation cytokines IL-1{beta}, TNF- and COX-2. WU-FA-01 significantly suppressed the expression levels of p65, I{kappa}B-, and p-I{kappa}B- in the TPA-induced mouse ear model. Overall, our results showed that WU-FA-00 and WU-FA-01 not only had effective antimicrobial activities in vitro, especially to the Gram-positive bacteria, but also possessed strong anti-inflammatory effects in vivo. These results provide a scientific basis for developing fusidic acid derivatives as antimicrobial and anti-inflammatory agents.

pharmacology and toxicology

Single-step Enzymatic Glycoengineering for the Construction of Antibody-cell Conjugates

Employing live cells as therapeutics is a direction of future drug discovery. An easy and robust method to modify the surfaces of cells directly to incorporate novel functionalities is highly desirable. However, many current methods for cell-surface engineering interfere with cells endogenous properties. Here we report an enzymatic approach that enables the transfer of biomacromolecules, such as a full length IgG antibody, to the glycocalyx on the surfaces of live cells when the antibody is conjugated to the enzymes natural donor substrate GDP-fucose. This method is fast and biocompatible with little interference to cells endogenous functions. We applied this method to construct two antibody-cell conjugates (ACCs) using different immune cells, and the modified cells exhibited specific tumor targeting and resistance to inhibitory signals produced by tumor cells, respectively. Remarkably, Herceptin-NK-92MI conjugates exhibits enhanced activities to induce the lysis of HER2+ cancer cells both ex vivo and in a murine tumor model, indicating its potential for further development as a clinical candidate.

bioengineering

Integration and analysis of CPTAC proteomics data in the context of cancer genomics in the cBioPortal

The Clinical Proteomic Tumor Analysis Consortium (CPTAC) has produced extensive mass spectrometry based proteomics data for selected breast, colon and ovarian tumors from The Cancer Genome Atlas (TCGA). We have incorporated the CPTAC proteomics data into the cBioPotal to support easy exploration and integrative analysis of these proteomic datasets in the context of the clinical and genomics data from the same tumors. cBioPortal is an open source platform for exploring, visualizing, and analyzing multi-dimensional cancer genomics and clinical data. The public instance of the cBioPortal (http://cbioportal.org/) hosts more than 100 cancer genomics studies including all of the data from TCGA. Its biologist-friendly interface provides many rich analysis features, including a graphical summary of gene-level data across multiple platforms, correlation analysis between genes or other data types, survival analysis, and network visualization. Here, we present the integration of the CPTAC mass spectrometry based proteomics data into the cBioPortal, consisting of 77 breast, 95 colorectal, and 174 ovarian tumors that already have been profiled by TCGA for mutations, copy number alterations, gene expression, and DNA methylation. As a result, the CPTAC data can now be easily explored and analyzed in the cBioPortal in the context of clinical and genomics data. By integrating CPTAC data into cBioPortal, limitations of TCGA proteomics array data can be overcome while also providing a user-friendly web interface, a web API and an R client to query the mass spectrometry data together with genomic, epigenomic, and clinical data.

cancer biology

Radial glial lineage progression and differential intermediate progenitor amplification underlie striatal compartments and circuit organization

The circuitry of the striatum is characterized by two organizational plans: the division into striosome and matrix compartments, thought to mediate evaluation and action, and the direct and indirect pathways, thought to promote or suppress behavior. The developmental origins of and relationships between these organizations are unknown, leaving a conceptual gap in understanding the cortico-basal ganglia system. Through genetic fate mapping, we demonstrate that striosome-matrix compartmentalization arises from a lineage program embedded in lateral ganglionic eminence radial glial progenitors mediating neurogenesis through two distinct types of intermediate progenitors (IPs). The early phase of this program produces striosomal spiny projection neurons (SPNs) through fate-restricted apical IPs (aIPSs) with limited capacity; the late phase produces matrix SPNs through fate-restricted basal IPs (bIPMs) with expanded capacity. Remarkably, direct and indirect pathway SPNs arise within both aIPS and bIPM pools, suggesting that striosome-matrix architecture is the fundamental organizational plan of basal ganglia circuitry organization.

neuroscience

Resequencing the Escherichia coli genome by GenoCare single molecule sequencing platform

Next generation sequencing (NGS) has revolutionized life sciences research. Recently, a new class of third-generation sequencing platforms has arrived to meet increasing demands in the clinic, capable of directly measuring DNA and RNA sequences at the single-molecule level without amplification. Here, we use the new GenoCare single molecule sequencing platform from Direct Genomics to resequence the E. coli genome and show comparable performance to the Illumina MiSeq system. Our platform detects single-molecule fluorescence by total internal reflection microscopy, with sequencing-by-synthesis chemistry. With a consensus sequence of 99.71% nucleotide identity to that of the Illumina MiSeq systems, GenoCare was determined to be a reliable platform for single-molecule sequencing, with strong potential for clinical applications.

genomics

Single Molecule Sequencing Of M13 Virus Genome Without Amplification

Third generation sequencing is a direct measurement of DNA/RNA sequences at the single molecule level without amplification. In this study, we report sequencing of the genome of the M13 virus by a new single molecule sequencing platform. Our platform detects single molecule fluorescence by the total internal reflection microscope technique, with sequencing-by-synthesis chemistry. We sequenced the genome of M13 to a depth of 316x and 100% coverage. The consensus sequence accuracy is 100%. We demonstrated that single molecule sequencing has no significant GC bias.

genomics