bioRxiv Science⌕ Search

Biology subjects

Bacallado, S.

Publications and source records attributed to Bacallado, S..

2 recordsLinked to original sources

Calibrated prediction of scarce adverse drug reaction labels with conditional neural processes

Adverse drug reactions (ADRs) are a major source of concern in the development of novel pharmaceuticals. ADRs may be identified in the late stages of development or even after commercialization, which may lead to failure or discontinuation after spending enormous resources on candidate molecules. Thus, predicting ADRs early in the process could help reduce costs by avoiding future failures. However, due to the low number of drugs approved, the amount of historical datapoints on ADRs is limited, which makes their prediction challenging for traditional chemoinformatics methods. Interestingly, each approved drug may have been annotated for hundreds of ADRs, which opens the door to framing ADR prediction as a multi-task or meta-learning problem. In this work, we adopt a meta-learning approach to ADR prediction by applying conditional neural processes (CNPs) to the publicly available Side Effect Resource (SIDER). Our results suggest that CNPs are competitive against single-task baselines even when trained on sparse datasets with missing labels. Furthermore, we find that their predictions are well-calibrated. Finally, we evaluate their performance on ADRs associated to different physiological systems and confirm good predictions across organ classes. Our findings suggest that meta-learning strategies may be beneficial for data-limited clinical endpoints like ADRs.

bioinformatics↗

Protein Language Models Uncover Carbohydrate-Active Enzyme Function in Metagenomic

In metagenomics, the pool of uncharacterized microbial enzymes presents a challenge for functional annotation. Among these, carbohydrate-active enzymes (CAZymes) stand out due to their pivotal roles in various biological processes related to host health and nutrition. Here, we present CAZyLingua, the first tool that harnesses protein language model embeddings to build a deep learning framework that facilitates the annotation of CAZymes in metagenomic datasets. Our benchmarking results showed on average a higher F1 score (reflecting an average of precision and recall) on the annotated genomes of Bacteroides thetaiotaomicron, Eggerthella lenta and Ruminococcus gnavus compared to the traditional sequence homology-based method in dbCAN2. We applied our tool to a paired mother/infant longitudinal dataset and revealed unannotated CAZymes linked to microbial development during infancy. When applied to metagenomic datasets derived from patients affected by fibrosis-prone diseases such as Crohns disease and IgG4-related disease, CAZyLingua uncovered CAZymes associated with disease and healthy states. In each of these metagenomic catalogs, CAZyLingua discovered new annotations that were previously overlooked by traditional sequence homology tools. Overall, the deep learning model CAZyLingua can be applied in combination with existing tools to unravel intricate CAZyme evolutionary profiles and patterns, contributing to a more comprehensive understanding of microbial metabolic dynamics.

bioinformatics↗