bioRxiv Science⌕ Search

Biology subjects

van Lent, P.

Publications and source records attributed to van Lent, P..

4 recordsLinked to original sources

Comparing metabolic engineering scenarios using simulated design-build-test-learn-cycles

Design-Build-Test-Learn (DBTL) cycles are a widely employed engineering framework in metabolic engineering. Nonetheless, their performance depends on a wide range of experimental and algorithmic design choices, whose combined effects on the successful optimization of microbial strains remain an open question. In this study, we performed in-silico DBTL cycles based on metabolic kinetic models to quantitatively assess how key process parameters affect strain optimization outcomes across four distinct metabolic pathway models. This includes parameters governing DNA library design, experimental budget limitations, and machine learning configuration. The results show that screening capacity is a dominant driver of optimization success, whereas DNA sequencing capacity has surprisingly little impact, despite its importance for model training. Selecting top-producing strains for sequencing consistently outperforms stratified sampling, highlighting a trade-off between predictive accuracy and optimization efficiency. DNA library structure strongly affects performance: increasing the number of editable positions generally improves outcomes, while expanding the set of gene targets can hinder optimization due to increased dimensionality or sparse sampling. Together, these findings offer actionable guidance for designing more effective DBTL workflows and underscore the value of simulation frameworks for exploring metabolic engineering strategies prior to experimental implementation.

bioengineering↗

Machine Learning-Assisted Pathway Optimization in Large Combinatorial Design Spaces: a p-Coumaric Acid Case Study

Combinatorial pathway optimization is an important tool for industrial metabolic engineering to improve titer, yield, or productivity of strains. Machine learning has been increasingly applied on many aspects of the Design-Build-Test-Learn (DBTL) cycle, an engineering framework that aims to navigate through the large landscape of theoretically possible designs using an iterative approach. While machine learning-assisted recommendation strategies have been successfully used to optimize strains, they have so far been limited to relatively small design spaces with few targeted elements. This small design space may limit key strengths of these approaches, such as strong predictive capabilities of supervised machine learning and exploration-exploitation schemes widely used in reinforcement learning and Bayesian optimization. In this work, two DBTL cycles are performed on Saccharomyces cerevisiae for p-coumaric acid production. We first perform a large library transformation on eighteen genes with twenty promoters, which expands the size of the combinatorial design space significantly (approximately 170 million configurations), followed by a smaller model-guided recommendation round. We use a machine learning-assisted recommendation strategy, based on the gradient bandit algorithm, parametrized to balance explo- ration and exploitation. We show that our recommendation strategy has a better performance than strain recommendation strategy using greedy strategies, such as feature importance-based methods. While balancing between exploration and exploitation has been shown to be impor- tant in many applications, we provide the first direct experimental illustration of this effect by recommending strains for scenarios with increasing exploitative-ness. A clear effect of the exploration-exploitation scenario on the p-coumaric acid production distribution of strains is observed, where a balanced scenario shows a higher variation in production over an exploratory or exploitative scenario. Interestingly, using an alternative top-producing parent strain with this balanced exploration-exploitation scheme gives the highest p-coumaric acid production, suggest- ing that model predictions outside of the training data distribution can still be used to perform successful strain recommendation. Overall, these results suggest that using machine learning- assisted strategies with balanced exploration-exploitation can be used to efficiently explore large combinatorial design spaces. The best engineered strain shows an increase in p-coumaric acid production of 137% over the parent strains and a 0.07g/g yield on glucose.

bioengineering↗

A Metagenomic Study of Antibiotic Resistance Across Diverse Soil Types and Geographical Locations

BackgroundSoil naturally harbours antibiotic resistant bacteria and is considered to be a reservoir for antibiotic resistance. The overuse of antibiotics across human, animal, and environmental sectors has intensified this issue leading to an increased acquisition of antibiotic resistant genes by bacteria in soil. Various biogeographical factors, such as soil pH, temperature, and pollutants, play a role in the spread and emergence of antibiotic resistance in soil. In this study, we utilised publicly available metagenomic datasets from four different soil types (rhizosphere, urban, natural, and rural areas) sampled from nine distinct geographic locations to explore the patterns of antibiotic resistance in soils from different regions. ResultsBradyrhizobium was predominant in vegetation soil types regardless of soil pH and temperature. ESKAPE pathogen Pseudomonas aeruginosa was prevalent in rural soil samples. Antibiotic resistance gene families such as 16s rRNA with mutations conferring resistance to aminoglycoside antibiotics, OXA {beta}-lactamase, ANT(3), and the RND and MFS efflux pump gene were identified in all soil types, with their abundances influenced by anthropogenic activities, vegetation, and climate in different geographical locations. Plasmids were more abundant in rural soils and were linked to aminoglycoside resistance. Integrons and integrative elements identified were associated with commonly used and naturally occurring antibiotics, showing similar abundances across different soil types and geographical locations. ConclusionAntimicrobial resistance in soil may be driven by anthropogenic activities and biogeographical factors, increasing the risk of bacteria developing resistance and leading to higher morbidity and mortality rates in humans and animals.

microbiology↗

When do longer reads matter? A benchmark of long read de novo assembly tools for eukaryotic genomes

BackgroundAssembly algorithm choice should be a deliberate, well-justified decision when researchers create genome assemblies for eukaryotic organisms from third-generation sequencing technologies. While third-generation sequencing by Oxford Nanopore Technologies (ONT) and Pacific Biosciences (PacBio) have overcome the disadvantages of short read lengths specific to next-generation sequencing (NGS), third-generation sequencers are known to produce more error-prone reads, thereby generating a new set of challenges for assembly algorithms and pipelines. Since the introduction of third-generation sequencing technologies, many tools have been developed that aim to take advantage of the longer reads, and researchers need to choose the correct assembler for their projects. ResultsWe benchmarked state-of-the-art long-read de novo assemblers, to help readers make a balanced choice for the assembly of eukaryotes. To this end, we used 13 real and 72 simulated datasets from different eukaryotic genomes, with different read length distributions, imitating PacBio CLR, PacBio HiFi, and ONT sequencing to evaluate the assemblers. We include five commonly used long read assemblers in our benchmark: Canu, Flye, Miniasm, Raven and Redbean. Evaluation categories address the following metrics: reference-based metrics, assembly statistics, misassembly count, BUSCO completeness, runtime, and RAM usage. Additionally, we investigated the effect of increased read length on the quality of the assemblies, and report that read length can, but does not always, positively impact assembly quality. ConclusionsOur benchmark concludes that there is no assembler that performs the best in all the evaluation categories. However, our results shows that overall Flye is the best-performing assembler, both on real and simulated data. Next, the benchmarking using longer reads shows that the increased read length improves assembly quality, but the extent to which that can be achieved depends on the size and complexity of the reference genome.

bioinformatics↗