bioRxiv Science⌕ Search

Biology subjects

Idanwekhai, K.

Publications and source records attributed to Idanwekhai, K..

2 recordsLinked to original sources

Adaptive Machine Learning Framework enables Unprecedented Yield and Purity of Adeno-Associated Viral Vectors for Gene Therapy

Adeno-associated viral (AAV) vectors for gene therapy are becoming integral to modern medicine, providing therapeutic options for diseases once deemed incurable. Currently, optimizing viral vector purification is a critical bottleneck in the gene therapy industry, impacting product efficacy and safety as well as accessibility and cost to patients. Traditional optimization methods are resource-intensive and often fail to adjust the purification process parameters to maximize the resulting product yield and quality. To address this challenge, we developed a machine learning framework that leverages Bayesian optimization to systematically refine affinity chromatography parameters (sample load, flow rate, and the formulation of chromatographic media) to improve AAV purification. The efficiency of this closed-loop workflow in iteratively optimizing the vectors yield, purity, and transduction efficiency was demonstrated by purifying clinically-relevant serotypes AAV2, AAV5, and AAV9 from HEK293 cell lysates using the affinity adsorbent AAVidity. We show that three cycles of Bayesian optimization elevated yields from a baseline of 70% to 99%, while reducing host-cell impurities by 230-to-400-fold across all serotypes. The optimized parameters consistently produced vectors with high purity and preserved high transduction activity, essential for therapeutic efficacy and safety, demonstrating serotype versatility - a key challenge in AAV manufacturing. By streamlining parameter optimization and enhancing productivity, our adaptive machine learning framework accelerates process development and reduces costs, advancing the accessibility and clinical translation of AAV-based gene therapies.

molecular biology↗

Machine learning of three-dimensional protein structures to predict the functional impacts of genome variation

Research in the human genome sciences generates a substantial amount of genetic data for hundreds of thousands of individuals, which concomitantly increases the number of variants with unknown significance (VUS). Bioinformatic analyses can successfully reveal rare variants and variants with clear associations to disease-related phenotypes. These studies have made a significant impact on how clinical genetic screens are interpreted and how patients are stratified for treatment. There are few, if any, comparable computational methods for variants to biological activity predictions. To address this gap, we developed a machine learning method that uses protein three-dimensional structures from AlphaFold to predict how a variant will influence changes to a genes downstream biological pathways. We trained state-of-the-art machine learning classifiers to predict which protein regions will most likely impact transcriptional activities of two proto-oncogenes, nuclear factor erythroid 2 (NFE2)-related factor 2 (Nrf2) and c-MYC. We have identified classifiers that attain accuracies higher than 80%, which have allowed us to identify a set of key protein regions that lead to significant perturbations in c-MYC or Nrf2 transcriptional pathway activities. SignificanceThe vast majority of mutations are either unspecified and/or their downstream biological implications are poorly understood. We have created a method that utilizes protein structure to cluster mutations from population-scale repositories to predict downstream functional impacts. The broader impacts of this approach include advanced filtering of mutations that are likely to impact genome function.

bioinformatics↗