bioRxiv Science⌕ Search

Biology subjects

Shemesh, S.

Publications and source records attributed to Shemesh, S..

2 recordsLinked to original sources

DeltaMut: An Integrative Database of AlphaFold2-Derived Missense Variant Structures

The widespread use of next-generation sequencing has led to a surge in the number of identified variants with uncertain effects on protein function. These variants pose a significant challenge in diagnostics and hinder patient treatment strategies. Numerous variant effect predictors (VEPs) are available to assess variant impact, but they primarily rely on sequence-derived information. The recent development of AlphaFold2 has raised questions about whether information retrieved from wild-type or predicted structures of missense variants can improve the predictive power of these algorithms. While the AlphaFold Protein Structure Database serves as a valuable resource for wild-type protein structures, a large-scale collection of missense variant structures is not available, limiting current efforts to wild-type conformations and a handful of modeled variants. To address this limitation, we developed DeltaMut, a comprehensive database containing over 77,000 protein structures, including 65,000 pathogenic and neutral missense variants. All structural models were generated using ParaFold, a high-performance computing-optimized implementation of AlphaFold2. The large-scale and systematic generation of variant protein structures distinguish DeltaMut as a unique resource for both expansive statistical studies and detailed, case-specific investigations of variant-induced structural changes. Furthermore, the DeltaMut database is freely accessible without registration. HighlightsO_LIDeltaMut is currently the largest database of AlphaFold2-predicted variant structures. C_LIO_LIContains 77,713 structures covering 12,101 wild-type and 65,612 variant proteins. C_LIO_LI70.6% of predicted structures have high or very high confidence (pLDDT [≥] 70). C_LIO_LIFreely accessible web server with visualization and download of variant models. C_LI

bioinformatics↗

Optimal release of gene drives in population connectivity networks

Gene drives, genetic constructs that can spread deleterious alleles in wild populations, have the potential to address some of the major pressing challenges of the Anthropocene such as invasive species, spread of disease vectors, and agricultural pests. However, responsible and effective deployment of gene drive requires taking into account the complex nature of real-world population connectivity networks. In particular, it is unclear how the topological position of the deployment site affects the spread process and its final outcome. Here we develop a framework for modeling gene drive spread in population connectivity networks, and study the eco-evolutionary dynamics of gene drive spread under complex population structures. We investigated the relationship between the position of the deployment site in the topology of the network and whether the gene drive is eventually lost, fixed, or maintained at an intermediate frequency. We identified network centrality measures of deployment sites that are highly correlated with the outcome of deployment for different gene drive designs and across diverse network topologies. We also show that there is a trade-off between the time-to-fixation and the final outcome, implying that multiple centrality measures of the deployment site would need to be considered when aiming to achieve rapid and successful population control using gene drives.

ecology↗