bioRxiv Science⌕ Search

Biology subjects

Amini, A. P.

Publications and source records attributed to Amini, A. P..

3 recordsLinked to original sources

Benchmarking Uncertainty Quantification for Protein Engineering

Machine learning sequence-function models for proteins could enable significant ad vances in protein engineering, especially when paired with state-of-the-art methods to select new sequences for property optimization and/or model improvement. Such methods (Bayesian optimization and active learning) require calibrated estimations of model uncertainty. While studies have benchmarked a variety of deep learning uncertainty quantification (UQ) methods on standard and molecular machine-learning datasets, it is not clear if these results extend to protein datasets. In this work, we implemented a panel of deep learning UQ methods on regression tasks from the Fitness Landscape Inference for Proteins (FLIP) benchmark. We compared results across different degrees of distributional shift using metrics that assess each UQ methods accuracy, calibration, coverage, width, and rank correlation. Additionally, we compared these metrics using one-hot encoding and pretrained language model representations, and we tested the UQ methods in a retrospective active learning setting. These benchmarks enable us to provide recommendations for more effective design of biological sequences using machine learning.

bioengineering↗

A nanoparticle priming agent reduces cellular uptake of cell-free DNA and enhances the sensitivity of liquid biopsies

Liquid biopsies are enabling minimally invasive monitoring and molecular profiling of diseases across medicine, but their sensitivity remains limited by the scarcity of cell-free DNA (cfDNA) in blood. Here, we report an intravenous priming agent that is given prior to a blood draw to increase the abundance of cfDNA in circulation. Our priming agent consists of nanoparticles that act on the cells responsible for cfDNA clearance to slow down cfDNA uptake. In tumor-bearing mice, this agent increases the recovery of circulating tumor DNA (ctDNA) by up to 60-fold and improves the sensitivity of a ctDNA diagnostic assay from 0% to 75% at low tumor burden. We envision that this priming approach will significantly improve the performance of liquid biopsies across a wide range of clinical applications in oncology and beyond.

genomics↗

Deep self-supervised learning for biosynthetic gene cluster detection and product classification

Natural products are chemical compounds that form the basis of many therapeutics used in the pharmaceutical industry. In microbes, natural products are synthesized by groups of colocalized genes called biosynthetic gene clusters (BGCs). With advances in high-throughput sequencing, there has been an increase of complete microbial isolate genomes and metagenomes, from which a vast number of BGCs are undiscovered. Here, we introduce a self-supervised learning approach designed to identify and characterize BGCs from such data. To do this, we represent BGCs as chains of functional protein domains and train a masked language model on these domains. We assess the ability of our approach to detect BGCs and characterize BGC properties in bacterial genomes. We also demonstrate that our model can learn meaningful representations of BGCs and their constituent domains, detect BGCs in microbial genomes, and predict BGC product classes. These results highlight self-supervised neural networks as a promising framework for improving BGC prediction and classification. Author summaryBiosynthetic gene clusters (BGCs) encode for natural products of diverse chemical structures and function, but they are often difficult to discover and characterize. Many bioinformatic and deep learning approaches have leveraged the abundance of genomic data to recognize BGCs in bacterial genomes. However, the characterization of BGC properties remains the main bottleneck in identifying novel BGCs and their natural products. In this paper, we present a self-supervised masked language model that learns meaningful representations of BGCs with improved downstream detection and classification.

bioinformatics↗