bioRxiv Science⌕ Search

Biology subjects

Didi, K.

Publications and source records attributed to Didi, K..

3 recordsLinked to original sources

GPU-accelerated homology search with MMseqs2

Rapidly growing protein databases demand faster sensitive sequence similarity detection. We present GPU-accelerated search utilizing intra-query parallelization delivering 6x faster single-protein searches compared to state-of-the-art CPU methods on 2x64 cores--speeds previously requiring large protein batches. It is most cost effective, including in large-batches at 0.45x MMseqs2-CPU speed (8 GPUs delivering 2.4x). It accelerates ColabFold structure prediction 31.8x compared to AlphaFold2 and Foldseek search 4-27x. MMseqs2-GPU is open-source at mmseqs.com.

bioinformatics↗

BioModelsML: Building a FAIR and reproducible collection of machine learning models in life sciences and medicine for easy reuse

Machine learning (ML) models are widely used in life sciences and medicine; however, they are scattered across various platforms and there are several challenges that hinder their accessibility, reproducibility and reuse. In this manuscript, we present the formalisation and pilot implementation of community protocol to enable FAIReR (Findable, Accessible, Interoperable, Reusable, and Reproducible) sharing of ML models. The protocol consists of eight steps, including sharing model training code, dataset information, reproduced figures, model evaluation metrics, trained models, Dockerfiles, model metadata, and FAIR dissemination. Applying these measures we aim to build and share a comprehensive public collection of FAIR ML models in the BioModels repository through incentivized community curation. In a pilot implementation, we curated diverse ML models to demonstrate the feasibility of our approach and we discussed the current challenges. Building a FAIReR collection of ML models will directly enhance the reproducibility and reusability of ML models, minimising the effort needed to reimplement models, maximising the impact on the application and significantly accelerating the advancement in the field of life science and medicine.

bioinformatics↗

AbNatiV: VQ-VAE-based assessment of antibody and nanobody nativeness for engineering, selection, and computational design

Monoclonal antibodies have emerged as key therapeutics, and nanobodies are rapidly gaining momentum following the approval of the first nanobody drug in 2019. Nonetheless, the development of these biologics as therapeutics remains a challenge. Despite the availability of established in vitro directed evolution technologies that are relatively fast and cheap to deploy, the gold standard for generating therapeutic antibodies remains discovery from animal immunization or patients. Immune-system derived antibodies tend to have favourable properties in vivo, including long half-life, low reactivity with self-antigens, and low toxicity. Here, we present AbNatiV, a deep-learning tool for assessing the nativeness of antibodies and nanobodies, i.e., their likelihood of belonging to the distribution of immune-system derived human antibodies or camelid nanobodies. AbNatiV is a multi-purpose tool that accurately predicts the nativeness of Fv sequences from any source, including synthetic libraries and computational design. It provides an interpretable score that predicts the likelihood of immunogenicity, and a residue-level profile that can guide the engineering of antibodies and nanobodies indistinguishable from immune-system-derived ones. We further introduce an automated humanisation pipeline, which we applied to two nanobodies. Wet-lab experiments show that AbNatiV-humanized nanobodies retain binding and stability at par or better than their wild type, unlike nanobodies humanised relying on conventional structural and residue-frequency analysis. We make AbNatiV available as downloadable software and as a webserver.

bioengineering↗