bioRxiv Science⌕ Search

Biology subjects

Shahreen, N.

Publications and source records attributed to Shahreen, N..

6 recordsLinked to original sources

Minimal Gene Signatures Enable High-Accuracy Prediction of Antibiotic Resistance in Pseudomonas aeruginosa

Antimicrobial resistance (AMR) in Pseudomonas aeruginosa poses a critical global health challenge, with current diagnostics relying on slow, culture-based methods. Here, we present a ML framework leveraging transcriptomic data to predict antibiotic resistance with high accuracy. We applied a genetic algorithm to 414 clinical isolates to identify minimal, highly predictive gene sets ([~]35-40 genes) distinguishing resistant from susceptible strains for meropenem, ciprofloxacin, tobramycin, and ceftazidime. Automated ML classifiers trained on these sets achieved accuracies of 96-99% on test data (F1 scores: 0.93-0.99), surpassing clinical deployment thresholds. Multiple distinct, non-overlapping gene subsets exhibited comparable performance, indicating that resistance acquisition broadly impacts the expression of diverse regulatory and metabolic genes. Comparison with known resistance markers from CARD and operon annotations revealed a substantial number of previously unannotated clusters, highlighting significant knowledge gaps in current AMR understanding. Mapping these genes onto independently modulated gene sets (iModulons) revealed transcriptional adaptations across diverse genetic regions. Overall, this study presents a streamlined machine-learning workflow for transcriptomic data and offers a pathway toward rapid diagnostics and personalized treatment strategies against AMR.

systems biology↗

Artificial Neural Network Reveals the Role of Transport Proteins in Rhodopseudomonas palustris CGA009 During Lignin Breakdown Product Catabolism

Rhodopseudomonas palustris, a versatile bacterium with diverse biotechnological applications, can effectively breakdown lignin, a complex and abundant polymer in plant biomass. This study investigates the metabolic response of R. palustris when catabolizing various lignin breakdown products (LBPs), including the monolignols p-coumaryl alcohol, coniferyl alcohol, sinapyl alcohol, p-coumarate, sodium ferulate, and kraft lignin. Transcriptomics and proteomics data were generated for those specific LBP breakdown conditions and used as features to train machine learning models, with growth rates as the target. Three models--Artificial Neural Networks (ANN), Random Forest (RF), and Support Vector Machine (SV)--were compared, with ANN achieving the highest predictive accuracy for both transcriptomics (94%) and proteomics (96%) datasets. Permutation feature importance analysis of the ANN models identified the top twenty genes and proteins influencing growth rates. Combining results from both transcriptomics and proteomics, eight key transport proteins were found to significantly influence the growth of R. palustris on LBPs. Re-training the ANN using only these eight transport proteins achieved predictive accuracies of 86% and 76% for proteomics and transcriptomics, respectively. This work highlights the potential of ANN-based models to predict growth-associated genes and proteins, shedding light on the metabolic behavior of R. palustris in lignin degradation under aerobic and anaerobic conditions. ImportanceThis study is significant as it addresses the biotechnological potential of Rhodopseudomonas palustris in lignin degradation, a key challenge in converting plant biomass into commercially important products. By training machine learning models with transcriptomics and proteomics data, particularly Artificial Neural Networks (ANN), the work achieves high predictive accuracy for growth rates on various lignin breakdown products (LBPs). Identifying top genes and proteins influencing growth, especially eight key transport proteins, offers insights into the metabolic niche of R. palustris. The ability to predict growth rates using just these few proteins highlights the efficiency of ANN models in distilling complex biological systems into manageable predictive frameworks. This approach not only enhances our understanding of lignin derivative catabolism but also paves the way for optimizing R. palustris for sustainable bioprocessing applications, such as bioplastic production, under varying environmental conditions.

systems biology↗

Robust Prediction of Enzyme Variant Kinetics with RealKcat

Predicting enzyme kinetics directly from sequence remains a central challenge in computational biology, particularly in resolving the effects of mutations at catalytically essential residues. Existing models frequently overlook the functional consequences of such perturbations, often defaulting to wild-type predictions even in cases of substantial activity loss, thereby limiting their reliability for enzyme design and mechanistic inference. Here, we introduce RealKcat, a machine learning framework trained on KinHub-27k, a rigorously curated dataset of 27,176 experimentally reported enzyme-substrate entries consolidated from BRENDA, SABIO-RK, and UniProt and verified across 2,158 primary sources. To ensure biochemical realism, kinetic parameters were collapsed into order-of-magnitude bins, enabling predictions that are tolerant to experimental noise yet sensitive to functional shifts. RealKcat integrates ESM embeddings for enzyme sequences with ChemBERTa embeddings of affiliated substrate, producing a unified feature space of the chemical conversion that supports robust multi-class classification of both catalytic turnover (kCat) and substrate affinity (KM). Across cross-validation, hold-out, out-of-distribution, and few-shot evaluations--including a dense mutational landscape of alkaline phosphatase (PafA)--RealKcat consistently capturead the direction and magnitude of mutation-induced changes, while preserving discrimination in both wild-type and mutant contexts. Importantly, structural descriptors were deliberately excluded, as naive integration of structural features has been shown to impair model generalization, underscoring the primacy of rigorous dataset curation, biologically informed task formulation, and balanced evaluation metrics. RealKcat establishes a scalable and mutation-sensitive framework for enzyme kinetics prediction, offering a biologically grounded platform for enzyme engineering, metabolic modeling, and therapeutic design. Significance StatementEnzymes catalyze biochemical reactions that sustain life, and accurate measurement of their efficiency--expressed through turnover number (kCat) and substrate affinity (KM)--is fundamental to biotechnology, synthetic biology, and even pharmaceutical innovation. Yet experimental assays remain prohibitive, time-intensive, and sensitive to conditions such as pH, temperature, and ionic strength of the assay buffer, while existing computational approaches often lack sensitivity to catalytic-site mutations and are constrained by inconsistencies in public databases. RealKcat addresses these gaps by introducing a rigorously curated dataset (KinHub-27k) derived from manual review of 2,158 articles and augmented with 5,278 synthetic catalytic variants generated through alanine substitution at annotated catalytic residues. Leveraging protein and substrate embeddings and a classification scheme based on order-of-magnitude kinetic bins, RealKcat achieves state-of-the-art functional e-accuracy and, critically, demonstrates sensitivity to catalytic perturbations. By adopting e-accuracy--a performance metric that evaluates predictions within {+/-}1 order of magnitude, aligning with the practical utility of enzyme kinetics--RealKcat provides biologically meaningful assessments that conventional metrics often obscure. This work establishes a robust, mutation-aware predictive platform that advances computational enzyme design and extends applicability to biomanufacturing, metabolic engineering, and precision medicine.

systems biology↗

Enzyme-constrained Metabolic Model of Treponema pallidum Identified Glycerol-3-phosphate Dehydrogenase as an Alternate Electron Sink

Treponema pallidum, the causative agent of syphilis, poses a significant global health threat. Its strict intracellular lifestyle and challenges in in vitro cultivation have impeded detailed metabolic characterization. In this study, we present iTP251, the first genome-scale metabolic model of T. pallidum, reconstructed and extensively curated to capture its unique metabolic features. These refinements included the curation of key reactions such as pyrophosphate-dependent phosphorylation and pathways for nucleotide synthesis, amino acid synthesis, and cofactor metabolism. The model demonstrated high predictive accuracy, validated by a MEMOTE score of 92%. To further enhance its predictive capabilities, we developed ec-iTP251, an enzyme-constrained version of iTP251, incorporating enzyme turnover rate and molecular weight information for all reactions having gene-protein-reaction associations. Ec-iTP251 provides detailed insights into protein allocation across carbon sources, showing strong agreement with proteomics data (Pearsons correlation of 0.88) in the central carbon pathway. Moreover, the thermodynamic analysis revealed that lactate uptake serves as an additional ATP-generating strategy to utilize unused proteomes, albeit at the cost of reducing the driving force of the central carbon pathway by 27%. Subsequent analysis identified glycerol-3-phosphate dehydrogenase as an alternative electron sink, compensating for the absence of a conventional electron transport chain while maintaining cellular redox balance. These findings highlight T. pallidums metabolic adaptations for survival and redox balance in intracellular environments, providing a foundation for future research into its unique bioenergetics. IMPORTANCEThis study advances our understanding of Treponema pallidum, the syphilis-causing pathogen, through the reconstruction of iTP251, the first genome-scale metabolic model for this organism, and its enzyme-constrained version, ec-iTP251. The work addresses challenges of studying T. pallidum due to its strict intracellular nature and difficulties in in vitro cultivation. Validated with strong agreement to proteomics data, the model demonstrates high predictive reliability. Key insights include unique metabolic adaptations such as lactate uptake for ATP production and alternative redox-balancing mechanisms. These findings provide a robust framework for future studies aimed at unraveling the pathogens survival strategies and identifying potential metabolic vulnerabilities.

systems biology↗

A thermodynamic bottleneck in the TCA cycle contributes to acetate overflow in Staphylococcus aureus

During aerobic growth, S. aureus relies on acetate overflow metabolism, a process where glucose is incompletely oxidized to acetate, for its bioenergetic needs. Acetate is not immediately captured as a carbon source and is excreted as waste by cells. The underlying factors governing acetate overflow in S. aureus have not been identified. Here, we show that acetate overflow is favored due to a thermodynamic bottleneck in the TCA cycle, specifically involving the oxidation of succinate to fumarate by succinate dehydrogenase. This bottleneck reduces flux through the TCA cycle, making it more efficient for S. aureus to generate ATP via acetate overflow metabolism. Additionally, the protein allocation cost of maintaining ATP flux through the restricted TCA cycle is greater than that of acetate overflow metabolism. Finally, we show that the TCA cycle bottleneck provides S. aureus the flexibility to redirect carbon towards maintaining redox balance through lactate overflow when oxygen becomes limiting, albeit at the expense of ATP production through acetate overflow. Overall, our findings suggest that overflow metabolism offers S. aureus distinct bioenergetic advantages over a thermodynamically constrained TCA cycle, potentially supporting its commensal-pathogen lifestyle.

systems biology↗

Optimal Protein Allocation Controls the Inhibition of GltA and AcnB in Neisseria gonorrhoeae

Neisseria gonorrhea (Ngo) is a major concern for global public health due to its severe implications for reproductive health. Understanding its metabolic phenotype is crucial for comprehending its pathogenicity. Despite Ngos ability to encode TCA cycle proteins, GltA and AcnB, their activities are notably restricted. To investigate this phenomenon, we used the iNgo_557 metabolic model and incorporated a constraint on total cellular protein content. Our results indicate that low cellular protein content severely limits GltA and AcnB activity, leading to a shift towards acetate overflow for ATP production, which is more efficient in terms of protein usage. Surprisingly, increasing cellular protein content alleviates this restriction on GltA and AcnB and delays the onset of acetate overflow, highlighting protein allocation as a critical determinant in understanding Ngos metabolic phenotype. These findings underscore the significance of Ngos metabolic adaptation in light of optimal protein allocation, providing a blueprint to understand Ngos metabolic landscape.

systems biology↗