bioRxiv Science⌕ Search

Biology subjects

Khamisi, L.

Publications and source records attributed to Khamisi, L..

2 recordsLinked to original sources

pLM-SAV: A Δ-Embedding Approach for Predicting Pathogenic Single Amino Acid Variants

Predicting whether single amino acid variants (SAVs) in proteins lead to pathogenic outcomes is a critical challenge in molecular biology and precision medicine. Experimental determination of all possible mutation effects is infeasible, and while state-of-the-art tools such as AlphaMissense show promise, their diagnostic performance is insufficient and they are often difficult to run locally. We developed pLM-SAV, a simple yet effective predictor that leverages protein language models (pLMs). {Delta}-embeddings, computed as the difference between wild-type and mutant sequence embeddings, are used as input for a convolutional neural network. To prevent data leakage, we trained our model on a well-characterized, labeled set of Eff10k and evaluated it on a non-homologous subset of ClinVar data. Our results demonstrate that this approach performs exceptionally well on the Eff10k test folds and reasonably on ClinVar test sets. Notably, pLM-SAV excels in resolving ambiguous predictions by AlphaMissense. We also found that an ensemble method, REVEL, outperforms both AlphaMissense and pLM-SAV, thus, we integrated these REVEL- enhanced predictions into our widely used AlphaMissense web application. Our results demonstrate that an SAV predictor trained on labeled data can achieve high predictive performance. Unlike previous methods such as VESPA, pLM-SAV uses no handcrafted features or substitution matrices, relying solely on pLM-derived representations. We anticipate that incorporating delta-embeddings into other mutation effect predictors or mutant structure prediction methods will further enhance their accuracy and utility in diverse biological contexts. Availability and ImplementationFreely available at https://doi.org/10.5281/zenodo.15502498 and https://alphamissense.hegelab.org.

bioinformatics↗

DNA-dependent phase separation by human SSB2 (NABP1/OBFC2A) protein points to adaptations to eukaryotic genome repair processes

Single-stranded DNA binding proteins (SSBs) are ubiquitous across all domains of life and play essential roles via stabilizing and protecting single-stranded (ss) DNA as well as organizing multiprotein complexes during DNA replication, recombination, and repair. Two mammalian SSB paralogs (hSSB1 and hSSB2 in humans) were recently identified and shown to be involved in various genome maintenance processes. Following our recent discovery of the liquid-liquid phase separation (LLPS) propensity of E. coli (Ec) SSB, here we show that hSSB2 also forms LLPS condensates under physiologically relevant ionic conditions. Similar to that seen for EcSSB, we demonstrate the essential contribution of hSSB2s C-terminal intrinsically disordered region (IDR) to condensate formation, and the selective enrichment of various genome metabolic proteins in hSSB2 condensates. However, in contrast to EcSSB-driven LLPS that is inhibited by ssDNA binding, hSSB2 phase separation requires single-stranded nucleic acid binding, and is especially facilitated by ssDNA. Our results reveal an evolutionarily conserved role for SSB-mediated LLPS in the spatiotemporal organization of genome maintenance complexes. At the same time, differential LLPS features of EcSSB and hSSB2 point to functional adaptations to prokaryotic versus eukaryotic genome metabolic contexts.

biochemistry↗