bioRxiv Science⌕ Search

bioRxiv · 10.1101/2025.11.26.690826

ColiFormer: A Transformer-Based Codon Optimization Model Balancing Multiple Objectives for Enhanced E. coli Gene Expression

Abstract

Codon optimization has become a key strategy to improving heterologous gene expression in Escherichia coli. However, many existing methods focus primarily on maximizing the codon adaptation index (CAI) while overlooking the importance of codon harmonization, which is crucial for maintaining proper protein characteristics. In this study, we present ColiFormer, a transformer-based codon optimization framework that was fine-tuned on 3,676 high-expression E. coli genes curated from the NCBI database. Built upon the CodonTransformer BigBird architecture, ColiFormer employs self-attention mechanisms and a mathematical optimization method (the augmented Lagrangian approach) to balance multiple biological objectives simultaneously, including CAI, GC content, tRNA adaptation index (tAI), RNA stability, and minimization of negative cis-regulatory elements. Performance was evaluated on 37,053 native E. coli genes and 80 recombinant protein targets commonly used in industrial studies, and the results were compared with six established codon optimization approaches. Across all datasets, ColiFormer demonstrated significant improvements in CAI and tAI, maintained GC content within optimal ranges, and reduced the incidence of negative cis-regulatory elements, all with lower runtime costs than most alternative methods. These results, based on in silico metrics, indicate that ColiFormer consistently enhances recombinant protein expression and achieves superior performance across multiple benchmarks compared to other established techniques. The tool is available as an open-source software package, along with the benchmark sequences dataset used in this study.

Source connections

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Emam, O., Baddam, S., Elfikky, A. A., Cavarretta, F., Sanad, Y., Luka, G., Farag, I.. 2025-12-01. ColiFormer: A Transformer-Based Codon Optimization Model Balancing Multiple Objectives for Enhanced E. coli Gene Expression. https://doi.org/10.1101/2025.11.26.690826

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Modular core network constructed from Escherichia coli transcriptome datasets using a hypergraph-based pan-network approach

We have developed a method to integrate transcriptomic coexpression networks across diverse experimental conditions within a single species. Our framework expands previous pannetwork approaches into a hypergraph based pannetwork. It first identifies coexpressed gene clusters within each individual dataset and extracts frequently co-expressed gene sets across multiple datasets using a frequent itemset mining algorithm. This process yields a hypergraph where each hyperedge is assigned a frequency (universality, U). We then extracted a subnetwork comprising high U hyperedges as the core network and formed modules within it. We applied our method to 106 Escherichia coli transcriptome datasets from the GEO database. The modularity of the core network peaked at a universality cutoff of 15, which was subsequently used to define it. Approximately 70% of the resulting core modules correspond to operons, and conversely, approximately 70% of all operons are covered by these core modules. We visualized the core network via an inter-modular network and analyzed core module dataset relationships using a modularity profile matrix. Based on these analyses, we successfully visualized the dynamic reorganization of the coexpression network in response to environmental changes in bacteria.

bioinformatics↗

Phase Separation Potential of Marsupial RSX RNA Reveals Convergent Evolution of X-Chromosome Inactivation Mechanisms

Background: X-chromosome inactivation (XCI) evolved independently in eutherian and marsupial mammals, where it is orchestrated by the unrelated long non-coding RNAs Xist and RSX, respectively. Xist organizes a repressive nuclear compartment through multivalent RNA-protein interactions, but whether RSX exploits similar biophysical principles remains unknown. A recent paper has identified bona fide RSX interacting proteins. Results: We integrated proteome-scale RNA-protein interaction prediction, experimental validation, phase-separation propensity analysis, functional annotation and comparative RNA-structure modelling to characterize the RSX interaction landscape. Using catRAPID, we ranked 1,168 RNA-binding proteins from the native Monodelphis domestica proteome. Predictions were significantly enriched for experimentally identified RSX interactors, with 4.85-fold enrichment among the top 50 candidates (P approximately 1.6 x 10-5), increasing to approximately eightfold for proteins shared by the experimental RSX and Xist interactomes (P approximately 2 x 10-6). Among 30 high-confidence RSX interactors, 13 were experimentally supported, 17 were previously unrecognized candidates and 17 exhibited high phase-separation propensity. The network was enriched in ribonucleoprotein granules and nuclear bodies and converged on m6A regulators and SR-family splicing factors. Comparative modelling detected no conserved secondary or tertiary architecture between RSX and Xist. Conclusions: RSX and Xist appear to have converged not through RNA sequence or global structure, but through recruitment of related, condensation-prone protein networks. These findings identify interaction-network and biophysical convergence as a potential principle of lncRNA-mediated chromosome regulation and provide testable candidates for determining whether RSX establishes a condensate-like compartment on the marsupial inactive X.

bioinformatics↗

Systematic discovery of protein kinase-like domains reveals diverse evolutionary strategies in the human oral microbiome

The protein kinase-like (PKL) superfamily regulates metabolism, biofilm formation, host interactions, and antimicrobial resistance in bacteria, although much of its diversity remains unknown. We used a structure-based pipeline that combined protein structure prediction with structural similarity searches to search 5.1 million protein sequences in the Human Oral Microbiome Database (HOMD). Together with 30 families that had been previously characterized, we identified 20 PKL families new to this genome collection: 15 entirely novel and 5 previously reported as preliminary findings. Our taxonomic analysis revealed that the families were distributed either across multiple bacterial phyla or restricted to a single species; hits spanning domains suggest either ancient origins or horizontal transfer. Three families were examined in detail and show different evolutionary pathways: SEAE1 from Segetibacter aerophilus, which is a putative lipid kinase; POGI1 from Porphyromonas gingivalis, a putative ethanolamine kinase with the PKL fold that is limited to pathogens; and SrfA-N from Haemophilus parainfluenzae, a non-catalytic scaffold that has been convergently co-opted as a tripartite toxin platform. Taken together, these examples demonstrate enzymatic specialization, adaptation specific to pathogens and structural co-option. The families identified are candidates for further mechanistic study.

bioinformatics↗