bioRxiv Science⌕ Search

Biology subjects

Saquand, E.

Publications and source records attributed to Saquand, E..

2 recordsLinked to original sources

So ManyFolds, So Little Time: Efficient Protein Structure Prediction With pLMs and MSAs

In recent years, machine learning approaches for de novo protein structure prediction have made significant progress, culminating in AlphaFold which approaches experimental accuracies in certain settings and heralds the possibility of rapid in silico protein modelling and design. However, such applications can be challenging in practice due to the significant compute required for training and inference of such models, and their strong reliance on the evolutionary information contained in multiple sequence alignments (MSAs), which may not be available for certain targets of interest. Here, we first present a streamlined AlphaFold architecture and training pipeline that still provides good performance with significantly reduced computational burden. Aligned with recent approaches such as OmegaFold and ESMFold, our model is initially trained to predict structure from sequences alone by leveraging embeddings from the pretrained ESM-2 protein language model (pLM). We then compare this approach to an equivalent model trained on MSA-profile information only, and find that the latter still provides a performance boost - suggesting that even state-of-the-art pLMs cannot yet easily replace the evolutionary information of homologous sequences. Finally, we train a model that can make predictions from either the combination, or only one, of pLM and MSA inputs. Ultimately, we obtain accuracies in any of these three input modes similar to models trained uniquely in that setting, whilst also demonstrating that these modalities are complimentary, each regularly outperforming the other.

bioinformatics↗

CanSig: Discovering de novo shared transcriptional programs in single cancer cells

Single-cell RNA sequencing (scRNA-seq) facilitates the discovery of gene signatures that define cell states across patients, which could be used in patient stratification and drug discovery. However, the lack of standardization in computational methodologies to analyse these data impedes the reproducibility of signature detection. To address this, we developed CanSig, a comprehensive benchmarking tool that evaluates methods for identifying transcriptional signatures in cancer. CanSig integrates metrics for batch correction and biological signal conservation with a gene signature correlation metric to score according to rediscovery, cross-dataset reproducibility, and clinical relevance. We applied CanSig to ten methods and to ten scRNA-seq datasets from four human cancer types--glioblastoma, breast cancer, lung adenocarcinoma, and cutaneous squamous cell carcinoma-- representing 116 patients and 105,000 malignant cells. Our results identify BBKNN as a leading method. We showed that the signatures identified with these methods correlate with clinically relevant outcomes, including patient survival and lymph node metastasis. Thus, CanSig establishes a standardized framework for reproducible cancer transcriptomics analysis.

bioinformatics↗