bioRxiv ScienceSearch

Biology subjects

Keiser, M. J.

Publications and source records attributed to Keiser, M. J..

5 recordsLinked to original sources

Predicting Cellular Drug Sensitivity using Conditional Modulation of Gene Expression

Selecting drugs most effective against a tumors specific transcriptional signature is an important challenge in precision medicine. To assess oncogenic therapy options, cancer cell lines are dosed with drugs that can differentially impact cellular viability. Here we show that basal gene expression patterns can be conditioned by learned small molecule structure to better predict cellular drug sensitivity, achieving an R2 of 0.7190{+/-}0.0098 (a 5.61% gain). We find that 1) transforming gene expression values by learned small molecule representations outperforms raw feature concatenation, 2) small molecule structural features meaningfully contribute to learned representations, and 3) an affine transformation best integrates these representations. We analyze conditioning parameters to determine how small molecule representations modulate gene expression embeddings. This ongoing work formalizes in silico cellular screening as a conditional task in precision oncology applications that can improve drug selection for cancer treatment.

bioinformatics

Deep learning from multiple experts improves identification of amyloid neuropathologies

Pathologists can label pathologies differently, making it challenging to yield consistent assessments in the absence of one ground truth. To address this problem, we present a DL approach that draws on a cohort of experts, weighs each contribution, and is robust to noisy labels. We collected 100,495 annotations on 20,099 candidate amyloid beta neuropathologies (cerebral amyloid angiopathy (CAA), and cored and diffuse plaques) from three institutions, independently annotated by five experts. DL methods trained on a consensus-of-two strategy yielded 12.6-26% improvements by area under the precision recall curve (AUPRC) when compared to those that learned individualized annotations. This strategy surpassed individual-expert models, even when unfairly assessed on benchmarks favoring them. Moreover, ensembling over individual models was robust to hidden random annotators. In blind prospective tests of 52,555 subsequent expert-annotated images, the models labeled pathologies like their human counterparts (consensus model AUPRC=0.74 cored; 0.69 CAA). This study demonstrates a means to combine multiple ground truths into a common-ground DL model that yields consistent diagnoses informed by multiple and potentially variable expert opinion.

pathology

Trans-channel fluorescence learning improves high-content screening for Alzheimer's disease therapeutics

In microscopy-based drug screens, fluorescent markers carry critical information on how compounds affect different biological processes. However, practical considerations may hinder the use of certain fluorescent markers. Here, we present a deep learning method for overcoming this limitation. We accurately generated predicted fluorescent signals from other related markers and validated this new machine learning (ML) method on two biologically distinct datasets. We used the ML method to improve the selection of biologically active compounds for Alzheimers disease (AD) from high-content high-throughput screening (HCS). The ML method identified novel compounds that effectively blocked tau aggregation, which would have been missed by traditional screening approaches unguided by ML. The method improved triaging efficiency of compound rankings over conventional rankings by raw image channels. We reproduced this ML pipeline on a biologically independent cancer-based dataset, demonstrating its generalizability. The approach is disease-agnostic and applicable across diverse fluorescence microscopy datasets.

pharmacology and toxicology

Adding stochastic negative examples into machine learning improves molecular bioactivity prediction

Multitask deep neural networks learn to predict ligand-target binding by example, yet public pharmacological datasets are sparse, imbalanced, and approximate. We constructed two hold-out benchmarks to approximate temporal and drug-screening test scenarios whose characteristics differ from a random split of conventional training datasets. We developed a pharmacological dataset augmentation procedure, Stochastic Negative Addition (SNA), that randomly assigns untested molecule-target pairs as transient negative examples during training. Under the SNA procedure, ligand drug-screening benchmark performance increases from R2 = 0.1926 {+/-} 0.0186 to 0.4269{+/-}0.0272 (121.7%). This gain was accompanied by a modest decrease in the temporal benchmark (13.42%). SNA increases in drug-screening performance were consistent for classification and regression tasks and outperformed scrambled controls. Our results highlight where data and feature uncertainty may be problematic, but also show how leveraging uncertainty into training improves predictions of drug-target relationships.

bioinformatics

Interpretable classification of Alzheimer’s disease pathologies with a convolutional neural network pipeline

Neuropathologists assess vast brain areas to identify diverse and subtly-differentiated morphologies. Standard semi-quantitative scoring approaches, however, are coarse-grained and can lack precise neuroanatomic localization. We report a proof-of-concept deep learning pipeline identifying specific neuropathologies--amyloid plaques and cerebral amyloid angiopathy--in immunohistochemical-stained archival slides. Using automated segmentation of stained objects and a cloud-based interface, we annotated >70,000 plaque candidates from 43 whole slide images (WSIs) to train and evaluate convolutional neural networks. Networks achieved strong plaque classification (0.993 and 0.744 areas under the receiver operating characteristic and precision recall curve, respectively) on a 10 WSI hold-out set. Prediction confidence maps visualized morphology distributions from the full-WSI level down to 20x magnification. Resulting plaque-burden scores correlated well with established semi-quantitative scores. Finally, saliency mapping demonstrated that networks learned patterns agreeing with accepted pathologic features. This scalable means to augment a neuropathologists ability may suggest a route to neuropathologic deep phenotyping.

pathology