bioRxiv Science⌕ Search

Biology subjects

Oler, E.

Publications and source records attributed to Oler, E..

3 recordsLinked to original sources

TransXplorer: An automated translational discovery platform for RNA-seq data

RNA-seq experiments routinely identify thousands of differentially expressed genes, but translating these into biological insights and therapeutic hypotheses often requires integrating multiple tools. Existing web platforms such as iDEP, NetworkAnalyst, and GEPIA2 address individual steps, differential expression, network visualization, or TCGA queries, but lack a unified environment spanning raw data processing to clinical and pharmacological interpretation. TransXplorer (https://transxplorer.org) is a freely available web platform that addresses this limitation by integrating the complete RNA-seq analytical workflow. It supports processing from raw FASTQ files using HISAT2 or Salmon, as well as direct GEO dataset import with automated metadata handling. Differential expression analysis is implemented via DESeq2, edgeR, and limma-voom, followed by functional enrichment across more than 1,800 species using Bioconductor resources. Batch effects are automatically detected and corrected using a composite of PVCA, kBET, and Silhouette metrics without requiring predefined batch annotations. Downstream analyses include co-expression network construction (WGCNA), protein-protein interaction mapping (STRING), cell-type deconvolution, and transcription factor inference using integrated DoRothEA and TFLink resources. The platform further links gene signatures to drug candidates through DGIdb and OpenTargets and enables survival and tumour-normal comparisons across TCGA cohorts. Application to cardiac endothelial differentiation (GSE151427) and kidney renal papillary cell carcinoma (TCGA-KIRP) datasets demonstrates accurate batch correction, biologically consistent pathway enrichment, recovery of expected cell-type proportions, and identification of clinically relevant genes and drug candidates. TransXplorer is freely available without a login.

bioinformatics↗

BioTransformer4.0 a comprehensive computational tool for small molecule metabolism prediction

BioTransformer 4.0, the successor to BioTransformer 3.0, is a freely available in silico metabolism prediction tool. It integrates both knowledge-based and machine learning approaches to predict metabolites for small molecules using one of seven modules: abiotic, environmental, CYP450, phase II, enzyme commission-based, human gut microbial, and all human metabolism. It also provides a customizable sequence prediction module that allows users to simulate multi-step metabolic transformations by chaining among the first six different modules. BioTransformer 4.0 can make predictions more efficiently and accurately than the previous version, as it includes more than 130 new reaction rules, and also an optional validation module to improve the efficiency by restricting the number of predicted metabolites, due to their similarity among real human metabolites. We evaluated its performance by running the six-step all-human metabolism prediction on the DrugBank dataset of 2,457 known biotransformations, and the PhytoHub dataset of 633 known biotransformations - achieving recall values of 87.2% (resp., 91.6%) for the DrugBank (resp., PhytoHub) datasets.

bioinformatics↗

Language model-guided anticipation and discovery of unknown metabolites

Despite decades of study, large parts of the mammalian metabolome remain unexplored. Mass spectrometry-based metabolomics routinely detects thousands of small molecule-associated peaks within human tissues and biofluids, but typically only a small fraction of these can be identified, and structure elucidation of novel metabolites remains a low-throughput endeavor. Biochemical large language models have transformed the interpretation of DNA, RNA, and protein sequences, but have not yet had a comparable impact on understanding small molecule metabolism. Here, we present an approach that leverages chemical language models to discover previously uncharacterized metabolites. We introduce DeepMet, a chemical language model that learns the latent biosynthetic logic embedded within the structures of known metabolites and exploits this understanding to anticipate the existence of as-of-yet undiscovered metabolites. Prospective chemical synthesis of metabolites predicted to exist by DeepMet directs their targeted discovery. Integrating DeepMet with tandem mass spectrometry (MS/MS) data enables automated metabolite discovery within complex tissues. We harness DeepMet to discover several dozen structurally diverse mammalian metabolites. Our work demonstrates the potential for language models to accelerate the mapping of the metabolome.

bioinformatics↗