bioRxiv Science⌕ Search

Biology subjects

Oikawa, H.

Publications and source records attributed to Oikawa, H..

3 recordsLinked to original sources

TRACER navigates rearrangement-driven sesterterpene chemical space via multimodal enzyme-product representation learning

Skeletal rearrangement drives the immense structural complexity of terpene, yet predicting it remains a formidable challenge due to sequence-function decoupling in terpene synthases. Here, we established TRACER (terpene rearrangement annotation via co-attentive enzyme-product representation), a multimodal framework mapping the latent associations between sequence-derived enzyme representations and product chemotypes. Retrospective validation proved TRACERs exceptional precision in predicting compound classes and discriminating skeletal rearrangement (SR) from non-skeletal rearrangement (NSR) pathways. TRACER-guided genome mining characterized two bifunctional synthases, FsPS and AcPS, uncovering four unprecedented carbon skeletons. Density functional theory calculations deciphered these cyclization cascades, pinpointing a critical 5/6/11 tricyclic intermediate as the key branching node for scaffold diversification. Mutagenesis and molecular dynamics simulations suggested that E305 in FsPS enables rearrangement by maintaining active-site water exclusion, whereas its alanine mutation causes premature carbocation quenching. Collectively, this work establishes a predictive paradigm for the rational discovery and mechanistic elucidation of complex terpene architectures.

synthetic biology↗

Predicting protein complexes in biosynthetic gene clusters

Biosynthetic gene clusters (BGCs) are contiguous genomic regions that encode diverse proteins responsible for natural product biosynthesis. These proteins collectively produce various secondary metabolites with complex chemical structure, including antibiotics and mycotoxins, yet the complete biosynthetic pathways have been experimentally resolved for only a limited number of compounds. Protein-protein interactions within BGCs have recently been recognized as key determinants of intermediate transfer, enzymatic regulation, and structural stability. However, many BGCs still contain proteins of unknown function that cannot be predicted by conventional sequence-based bioinformatics tools, hindering a comprehensive understanding of their biosynthetic pathways. To address this challenge, we built a high-throughput complex prediction pipeline by replacing AlphaFold3s multiple sequence alignment generation with a faster MMSeqs2. We systematically screened 487,828 protein pairs derived from 2,437 BGCs registered in the Minimum Information about a Biosynthetic Gene cluster (MIBiG) database and predicted 15,438 heteromeric interactions with an ipTM [≥] 0.6. Among them, 1,390 protein pairs exhibited structural homology with an RMSD [≤] 2.0 [A]. These predictions highlight interesting molecular mechanisms involving proteins previously annotated as "uncharacterized" or "potentially dysfunctional". Our analysis further showed that the ipSAE metric can distinguish correct heterocomplex pairs when multiple functionally homologous proteins are present within a BGC. Overall, our computational analysis revealed molecular interaction networks among proteins encoded by each BGC and identified enzyme complexes that are likely functional only when assembled. These predicted complexes may represent previously unrecognized links in their biosynthetic pathways. The complete results are available in a reusable format at https://doi.org/10.5281/zenodo.17451667 to support future experimental validation.

bioinformatics↗

A single dimer of the SARS-CoV-2 N protein can associate with multiple fragments of single-stranded and stem-loop RNAs

The nucleocapsid (N) protein of SARS-CoV-2 associates with the viral genomic RNA (gRNA) and forms the ribonucleoprotein (RNP) granules. However, the detailed molecular structures of RNP and their formation mechanism are largely unknown. We used circular dichroism (CD) spectroscopy, fluorescence correlation spectroscopy (FCS) and single-molecule Forster resonance energy transfer (sm-FRET) spectroscopy to understand the interaction between the N protein and different structural units of RNA. We chose polyadenylate chains with 40, 30 and 20 bases possessing a single-stranded structure and three stem loops with 50, 41 and 29 bases selected from gRNA, and labeled their 5 and 3 ends by Alexa488 and Alexa647, respectively, for the FCS and sm-FRET measurements. We found that the N protein started to bind to the single-stranded RNAs at the concentrations between 10 and 100 nM. The binding of the N protein to one of the stem loops occurred at the concentration less than 10 nM without melting the stem loop. For all the samples, the binding of multiple molecules of the RNA fragments to a single dimer of the N protein was observed. These results demonstrate that the N protein acts as a non-specific binder to both single-stranded and stem-loop structures of RNA, and that the N protein might contract a long RNA chain by bridging its multiple segments. We propose that the RNP granules might be folded by the association of the numerous stem loops of gRNA triggered by the assembly of the N protein.

biophysics↗