bioRxiv ScienceSearch

Biology subjects

Glicksberg, B. S.

Publications and source records attributed to Glicksberg, B. S..

4 recordsLinked to original sources

Prioritizing Small Molecule as Candidates for Drug Repositioning using Machine Learning

Drug repositioning, i.e. identifying new uses for existing drugs and research compounds, is a cost-effective drug discovery strategy that is continuing to grow in popularity. Prioritizing and identifying drugs capable of being repositioned may improve the productivity and success rate of the drug discovery cycle, especially if the drug has already proven to be safe in humans. In previous work, we have shown that drugs that have been successfully repositioned have different chemical properties than those that have not. Hence, there is an opportunity to use machine learning to prioritize drug-like molecules as candidates for future repositioning studies. We have developed a feature engineering and machine learning that leverages data from publicly available drug discovery resources: RepurposeDB and DrugBank. ChemVec is the chemoinformatics-based feature engineering strategy designed to compile molecular features representing the chemical space of all drug molecules in the study. ChemVec was trained through a variety of supervised classification algorithms (Naive Bayes, Random Forest, Support Vector Machines and an ensemble model combining the three algorithms). Models were created using various combinations of datasets as Connectivity Map based model, DrugBank Approved compounds based model, and DrugBank full set of compounds; of which RandomForest trained using Connectivity Map based data performed the best (AUC=0.674). Briefly, our study represents a novel approach to evaluate a small molecule for drug repositioning opportunity and may further improve discovery of pleiotropic drugs, or those to treat multiple indications.

bioinformatics

Learning and Mapping Lyme Disease Patient Trajectories from Electronic Medical Data for Stratification of Disease Risk and Therapeutic Response

BackgroundLyme disease (LD) is an epidemic, tick-borne illness with approximately 329,000 incidences diagnosed each year in United States. Long-term use of antibiotics is associated with serious complications, including post-treatment Lyme disease syndrome (PTLDS). The landscape of comorbidities and health trajectories associated with LD and associated treatments is not fully understood. Consequently, there is an urgent need to improve clinical management of LD based on a more precise understanding of disease and patient stratification.\n\nMethodsWe used a precision medicine machine-learning approach based on high-dimensional electronic medical records (EMRs) to characterize the heterogeneous comorbidities in a LD population and develop systematic predictive models for identifying medications that influence the risk of subsequent comorbidities.\n\nFindingsWe identified 3, 16, and 17 comorbidities at broad disease categories associated with LD within 2, 5, and 10 years of diagnosis, respectively. At higher resolution of ICD-9 levels, we pinpointed specific co-morbid diseases on a timescale that matched the symptoms associated with PTLDS. We identified 7, 30, and 35 medications that influenced the risks of the reported comorbidities within 2, 5, and 10 years, respectively. These medications included six previously associated with the identified comorbidities and 29 new findings. For instance, the first-line antibiotic doxycycline exhibited a consistently protective effect for typical symptoms of LD, including backache Not Otherwise Specified (NOS) and chronic rhinitis, but consistently increased the risk of cataract NOS, tear film insufficiency NOS, and nocturia.\n\nInterpretationOur approach and findings suggest new hypotheses for precision medicine treatments regimens and drug repurposing opportunities tailored to the phenotypic profiles of LD patients.\n\nFundingThe Steven & Alexandra Cohen Foundation

bioinformatics

Genetic Identification Of A Common Collagen Disease In Puerto Ricans Via Identity-By-Descent Mapping In A Health System

Achieving confidence in the causality of a disease locus is a complex task that often requires supporting data from both statistical genetics and clinical genomics. Here we describe a combined approach to identify and characterize a genetic disorder that leverages distantly related patients in a health system and population-scale mapping. We utilize genomic data to uncover components of distant pedigrees, in the absence of recorded pedigree information, in the multi-ethnic BioMe biobank in New York City. By linking to medical records, we discover a locus associated with genetic relatedness that also underlies extreme short stature. We link the gene, COL27A1, with a little-known genetic disease, previously thought to be rare and recessive. We demonstrate that disease manifests in both heterozygotes and homozygotes, indicating a common collagen disorder impacting up to 2% of individuals of Puerto Rican ancestry, leading to a better understanding of the continuum of complex and Mendelian disease.

genomics

Co-localization of Conditional eQTL and GWAS Signatures in Schizophrenia

Causal genes and variants within genome-wide association study (GWAS) loci can be identified by integrating GWAS statistics with expression quantitative trait loci (eQTL) and determining which SNPs underlie both GWAS and eQTL signals. Most analyses, however, consider only the marginal eQTL signal, rather than dissecting this signal into multiple independent eQTL for each gene. Here we show that analyzing conditional eQTL signatures, which could be important under specific cellular or temporal contexts, leads to improved fine mapping of GWAS associations. Using genotypes and gene expression levels from post-mortem human brain samples (N=467) reported by the CommonMind Consortium (CMC), we find that conditional eQTL are widespread; 63% of genes with primary eQTL also have conditional eQTL. In addition, genomic features associated with conditional eQTL are consistent with context specific (i.e. tissue, cell type, or developmental time point specific) regulation of gene expression. Integrating the Psychiatric Genomics Consortium schizophrenia (SCZ) GWAS and CMC conditional eQTL data reveals forty loci with strong evidence for co-localization (posterior probability >0.8), including six loci with co-localization of conditional eQTL. Our co-localization analyses support previously reported genes and identify novel genes for schizophrenia risk, and provide specific hypotheses for their functional follow-up.

genetics