bioRxiv Science⌕ Search

Biology subjects

Shiraishi, T.

Publications and source records attributed to Shiraishi, T..

4 recordsLinked to original sources

Predicting protein complexes in biosynthetic gene clusters

Biosynthetic gene clusters (BGCs) are contiguous genomic regions that encode diverse proteins responsible for natural product biosynthesis. These proteins collectively produce various secondary metabolites with complex chemical structure, including antibiotics and mycotoxins, yet the complete biosynthetic pathways have been experimentally resolved for only a limited number of compounds. Protein-protein interactions within BGCs have recently been recognized as key determinants of intermediate transfer, enzymatic regulation, and structural stability. However, many BGCs still contain proteins of unknown function that cannot be predicted by conventional sequence-based bioinformatics tools, hindering a comprehensive understanding of their biosynthetic pathways. To address this challenge, we built a high-throughput complex prediction pipeline by replacing AlphaFold3s multiple sequence alignment generation with a faster MMSeqs2. We systematically screened 487,828 protein pairs derived from 2,437 BGCs registered in the Minimum Information about a Biosynthetic Gene cluster (MIBiG) database and predicted 15,438 heteromeric interactions with an ipTM [≥] 0.6. Among them, 1,390 protein pairs exhibited structural homology with an RMSD [≤] 2.0 [A]. These predictions highlight interesting molecular mechanisms involving proteins previously annotated as "uncharacterized" or "potentially dysfunctional". Our analysis further showed that the ipSAE metric can distinguish correct heterocomplex pairs when multiple functionally homologous proteins are present within a BGC. Overall, our computational analysis revealed molecular interaction networks among proteins encoded by each BGC and identified enzyme complexes that are likely functional only when assembled. These predicted complexes may represent previously unrecognized links in their biosynthetic pathways. The complete results are available in a reusable format at https://doi.org/10.5281/zenodo.17451667 to support future experimental validation.

bioinformatics↗

Biosynthesis of Kaitocephalin: A Neuroprotective Natural Product Featuring a Peptide-Like yet Non-Peptidic Scaffold

Kaitocephalin (KCP, 1) is a neuroprotective natural product that acts as an antagonist of ionotropic glutamate receptors, making it a highly promising lead for drug discovery. It possesses a unique scaffold composed of three amino acids connected via C-C bonds, which appears peptide-like but is formed without peptide bonds. In this study, we identified the KCP biosynthetic gene cluster (kpb cluster) in the producing fungus Eupenicillium shearii through integrated genomic and transcriptomic analyses. LC-MS/MS profiling and chemical derivatization of E. shearii extracts led to the discovery of four novel pathway-related metabolites (2-5). In vitro enzymatic assays with 2(S)-dechlorokaito lactate (4) as a substrate enabled functional characterization of KpbI, KpbM, and KpbB involved in KCP formation. Among them, the dioxygenase KpbI was found to catalyze an unprecedented two-step oxidation to form the D-serine moiety. In addition, isotope tracing experiments provided new insights into the origin of the L-proline moiety. These findings establish a foundation for future studies aimed at elucidating the complete biosynthetic mechanism of KCP. Table of Contents graphical abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=50 SRC="FIGDIR/small/683206v1_ufig1.gif" ALT="Figure 1"> View larger version (12K): org.highwire.dtl.DTLVardef@189618aorg.highwire.dtl.DTLVardef@62d5deorg.highwire.dtl.DTLVardef@c71341org.highwire.dtl.DTLVardef@1c14e59_HPS_FORMAT_FIGEXP M_FIG C_FIG

biochemistry↗

A novel transformer-based platform for the prediction and design of biosynthetic gene clusters for (un)natural products

Biosynthetic gene clusters (BGCs), comprising sets of functionally related genes responsible for synthesizing complex natural products, are a rich source of bioactive compounds with pharmaceutical potential. Here, we present a transformer-based framework that models functional domains as linguistic units to capture and predict their positional relationships within genomes. Using a RoBERTa architecture, we trained models on four progressively broader datasets: bacterial BGCs, Actinomycetes genomes, bacterial genomes, and bacterial plus fungal genomes. Evaluation using 2,492 experimentally-validated BGCs from the MIBiG database showed that more than 60% of true domains were ranked first and over 80% within the top 10 candidates. Our models also achieved classification accuracies exceeding 70% for major compound classes including polyketides (PKs) and terpenes. To explore model-guided BGC design, we compared predictions from the BGC-trained and genome-trained models using the BGC for the bacterial diterpenoid cyclooctatin as a case study. The genome-trained model uniquely predicted several domains absent from both the original BGC and the prediction by the BGC-trained model. Heterologous expression of one of those predicted domains in Streptomyces albus, together with the biosynthetic genes for cyclooctatin, yielded an unknown cyclooctatin derivative. This framework not only provides a novel BGC prediction method using machine learning but also facilitates rational design of artificial BGCs. Future integration of transcriptomic, protein structural, and phylogenetic data will enhance the models predictive and generative capabilities, supporting accelerated discovery and engineering of natural products. Author SummaryBGCs encode diverse natural products, including antibiotics and anticancer agents. Identifying and designing BGCs in microbial genomes is crucial for discovering new bioactive compounds. In this study, we developed a transformer-based deep learning model that treats protein domains as language-like tokens and learns how they are arranged in genomes. By training on both known BGCs and whole genomes, the model successfully predicts biologically plausible combinations of domains, including those absent in known BGCs. We experimentally validated one such prediction by expressing a newly identified gene alongside known cyclooctatin biosynthetic genes, confirming the production of an unknown cyclooctatin derivative. Our results demonstrate how language models can uncover hidden biosynthetic potential and offer a promising new AI tool for natural product discovery and synthetic biology.

bioinformatics↗

Cancer/Testis Antigens Differentially Expressed In Indolent And Aggressive Prostate Cancer: Potential New Biomarkers And Targets For Immunotherapies

Current clinical tests for prostate cancer (PCa), such as the PSA test, are not fully capable of discerning patients that are highly likely to develop metastatic prostate cancer (MPCa). Hence, more accurate prediction tools are needed to provide treatment strategies that are focused on the different risk groups. Cancer/testis antigens (CTAs) are expressed during embryonic development and present aberrant expression in cancer making them ideal tumor specific biomarkers. Here, the potential use of a panel of CTAs as a biomarker for PCa detection as well as metastasis prediction is explored. We initially identified eight CTAs (CEP55, NUF2, PAGE4, PBK, RQCD1, SPAG4, SSX2 and TTK) that are differentially expressed in MPCa when compared to local disease and used this panel to compare the gene and protein expression profiles in paired PCa and normal adjacent prostate tissue. We identified differential expression of all eight CTAs at the protein level when comparing 80 paired samples of PCa and the adjacent non-cancer tissue. Using multiple logistic regression we also show that a panel of these CTAs present high accuracy to discriminate normal from tumor samples. In summary, this study provides evidence that a panel of CTAs, differentially expressed in aggressive PCa, is a potential biomarker for diagnosis and prognosis to be used in combination with the current clinically available tools and is also a potential target for immunotherapy development.

cancer biology↗