bioRxiv ScienceSearch

EXPLORE THE ARCHIVE

Bioinformatics

Find computational methods and tools for biological research.

33 recordsLinked to original sources

Automatic bioinformatic software named entity recognition from literature

Bioinformatics software and databases are essential components of modern life science research, yet their mentions in the scientific literature are often inconsistent and difficult to systematically identify at scale. The lack of a comprehensive and up-to-date catalog of bioinformatics resources hinders efforts toward automated biomedical knowledge extraction and streamlined data analysis. Here we present SNAIL, a hybrid named entity recognition framework designed to automatically identify bioinformatics software and database (SW/DB) names from biomedical texts. SNAIL integrates complementary lexical and semantic modeling strategies. The lexical component captures orthographic patterns and contextual cues characteristic of SW/DB names, while the semantic component leverages contextual embeddings generated by transformer-based language models such as SciBERT, combined with an explicit token-masking strategy to enhance entity-focused representations. A large training corpus was constructed automatically through a hybrid pipeline that integrates citation-hinted extraction with large language model-assisted distillation. Evaluation on two independent benchmark datasets and real-world research articles demonstrates that SNAIL substantially outperforms existing approaches, including domain-specific methods such as bioNerDS2 and general-purpose large language models such as ChatGPT, Gemini, Grok and Claude. Applying SNAIL to large-scale literature analysis further reveals distinct journal-level preferences across bioinformatics subfields. These results demonstrate that SNAIL provides an accurate and scalable solution for identifying bioinformatics resources in scientific texts and enables systematic meta-analysis of tool usage and research trends.

bioinformatics

XpBrew and PanXpresso - automatic RNA-seq processing workflow and comprehensive collection of gene expression data

Rapid developments in sequencing technologies have reduced the costs of transcriptomic experiments and resulted in a plethora of publicly available RNA-seq datasets. This is a valuable resource that can be harnessed to obtain novel biological insights through data upcycling. In this wake, we introduce XpBrew, an end-to-end Python workflow that was applied to generate PanXpresso, a comprehensive collection of gene expression datasets covering the taxonomic breadth of plants, animals, fungi, bacteria and archaea. XpBrew (https://github.com/PuckerLab/XpBrew) and PanXpresso (https://doi.org/10.60507/FK2/OBIGQH) are freely available.

bioinformatics

PathFold: Predicting the Entire Protein Folding Pathway from Protein Sequence Alone

Recent advances in protein structure prediction, exemplified by AlphaFold, have largely addressed the determination of static structures, one aspect of the protein folding problem. However, predicting folding pathways, by which proteins reach their native states, remains a significant challenge. Here, we present PathFold, a deep learning framework that predicts protein folding pathways directly from sequence information. PathFold leverages an AlphaFold-based module to extract structural information from the sequence and generates a progressive folding trajectory from an extended conformation using a diffusion model. By modeling the full trajectory, it enables prediction of folding intermediates and transition pathways, analogous to those observed in steered molecular dynamics (SMD) simulations. The predicted pathways reveal well-defined intermediates and sequential folding events, and show agreement with experimental folding data, including measured {Phi}-values.

bioinformatics

GNMCADS: Sampling For Protein Conformation Diversity With Gaussian Network Model Guided Condition Annealed Diffusion Sampler

Proteins are dynamic molecules existing in diverse conformational states underlying their biological functions. Although recent approaches have enabled diverse conformational sampling by emulating molecular dynamics simulations, perturbing evolutionary information, or steering internal mechanisms of structure prediction models, predicting conformations resulting from major domain motions or motions that occur over long timescales still remains a challenge. To this end, we introduce GNMCADS, a conformational sampling strategy that enhances the diversity of protein diffusion models by selectively annealing the conditioning signal guided by the intrinsic dynamical organization of the sampled protein. Further, we implement GNMCADS in the diffusion module of AlphaFold3, enabling the generation of diverse protein conformations. When benchmarked across 92 proteins that include 54 class A GPCRs, 15 transporters, and 23 proteins with major domain movements, GNMCADS exhibits improved sampling diversity compared to other current conformational sampling methods.

bioinformatics

Calibration-free compression brings Evo 2 to its full million-token context on a single GPU

Evo 2 is the largest openly available genomic foundation model, but its forty billion parameter configuration cannot be loaded onto a single 80 GB accelerator, placing genome-scale analysis beyond most laboratories. We present TurboQuant-Bio, an open toolkit that compresses Evo 2s weights and attention cache to four bits without calibration data, and serves both through fused kernels. Compression is near-lossless across perplexity spanning the tree of life, genomic classification, splice-site prediction, gene completion and clinically relevant variant-effect prediction. It brings Evo 2 40B onto one 80 GB GPU and Evo 2 7B to its full million-token context within a 40 GB memory budget, an eightfold gain in reachable context. We further show that the released chunked-prefill path is silently incorrect, returning plausible but uncorrelated likelihoods, and derive the block-wise continuation that repairs it: a complete 580-kilobase bacterial genome is now scored in one context in 22 minutes rather than 13.7 hours.

bioinformatics

Scaling recipes for single-cell RNA sequencing foundation models: when do scaling laws hold?

Deep learning models exhibit empirical scaling laws whereby performance changes predictably with model size, dataset size, and training compute. Although these relationships are well established in domains such as language and image modelling, their applicability to biological data remains unclear. Here, we investigate scaling behaviour in foundation models trained on large collec tions of single-cell transcriptomes. We show that pre-training loss decreases systematically with model capacity and training compute, exhibiting a power law dependence on model size. The strength and regularity of these trends differ between model formulations. We identify and quantify empirical relationships linking the optimal learning rate and depth-to-width ratio to model size and depth or compute. These results demonstrate that scaling principles extend to transcriptomic modelling. More broadly, they provide a quantitative framework for estimating the expected returns from additional resources and selecting suit able hyperparameters and architectures, thereby supporting the development of increasingly capable foundation models for omics data.

bioinformatics

dnoise: Fast Native Data Reduction for Bruker timsTOF

Bruker timsTOF acquisitions produce dense native .d files whose storage, transfer, and archival become substantial at high throughput. We present dnoise, an open-source Rust tool that removes points directly from timsTOF frames and writes a native-compatible .d directory. dnoise retains ions that form coherent streaks across the ion-mobility dimension and applies acquisition-aware gates to signal that cannot be selected for fragmentation. On a three-species benchmark spanning ddaPASEF and diaPASEF at 5- and 15-minute gradients, default MS1-only denoising reduced the frame binary by 35 to 53%. Label-free quantification accuracy was preserved in both modes. ddaPASEF peptide-spectrum-match, peptide, and protein-group counts were unchanged, as expected with the searched MS/MS spectra untouched, and diaPASEF precursor and protein-group counts changed only slightly. Every tested processing run completed in 69 seconds or less on the benchmark workstation. Optional MS/MS denoising produced greater reduction but sacrificed several percent of identifications. Thus, a substantial fraction of native timsTOF frame data can be removed with little analytical change.

bioinformatics

Designing antimicrobials with programmable mechanism and safety

Antimicrobial peptides (AMPs) are a promising solution to antimicrobial resistance, yet generative models for their design cannot control the physicochemical properties and motifs that shape activity and selectivity. Here, we present OmegAMP, a conditional diffusion framework controlling net charge, mean hydrophobicity, and sequence length, supporting de novo, analog, and motif-guided design. Across 204 wet-lab characterized peptides, de novo generation yielded antimicrobials with broad activity against multidrug-resistant Gram-negative isolates. Analog generation converted six inactive prototypes into antimicrobials, with the prototype determining each analog's membrane-disruption mode and mammalian-cell safety. Motif-guided analog generation preserved lipopolysaccharide engagement of active prototypes, and a redesigned non-antimicrobial leucine zipper acquired antimicrobial activity while retaining DNA-perturbing character in vitro. In murine skin and thigh infection models, leads reduced bacterial burden, with a motif-guided DNA-perturbing lead matching the fluoroquinolone control systemically. OmegAMP opens a programmable route to new peptide antibiotics whose mechanism and safety follow from the chosen prototype.

bioinformatics

Integrated Transcriptomic and CRISPR Dependency Analysis Prioritizes a CDK1-AURKB Mitotic Vulnerability Axis in Diffuse Intrinsic Pontine Glioma

Diffuse intrinsic pontine glioma (DIPG), now classified within diffuse midline glioma, H3K27-altered, remains a lethal pediatric brainstem tumor with limited therapeutic options. Here, we integrated public DIPG transcriptomic datasets, protein-protein interaction modeling, functional enrichment, immune deconvolution, survival analysis, and DepMap CRISPR dependency data to nominate candidate mitotic vulnerabilities. Differential expression analysis comparing 27 DIPG tumors with 6 brainstem low-grade glioma comparator samples identified a proliferative transcriptional program enriched for chromosome segregation, nuclear division, and cell-cycle pathways. Network analysis prioritized a compact mitotic hub module containing CDK1, AURKB, TOP2A, CDC20, CDCA8, and related G2/M regulators. CIBERSORT analysis of an independent DIPG cohort inferred low cytotoxic T-cell signal, consistent with an immune-cold phenotype, although immune-cell fractions require orthogonal validation. Survival analysis showed that neither inferred immune scores nor a composite mitotic hub score significantly stratified overall survival. DepMap CRISPR gene-effect data nominated CDK1, AURKB, TOP2A, and BIRC5 as candidate dependencies across brain tumor models. These findings provide a computational framework for prioritizing mitotic vulnerabilities in DIPG and support experimental validation in disease-relevant models.

bioinformatics

Mural-VISTA: A tool for mural cell-vessel interaction assessment and multiscale single-cell topo-morphological analysis

Three-dimensional (3D) mural cell morphology is heterogeneous and coupled to vessel geometry, however, measurements from two-dimensional (2D) maximum intensity projections (MIP) obscure overlapping processes and cell-vessel contacts. Accordingly, we developed Mural-VISTA, a semi-automated Python workflow for mural cell-vessel interaction and single-cell topo-morphology analysis of reconstructed surface meshes. This workflow integrates mesh pretreatment, interactive centerline extraction, hierarchical segmentation of cell soma, main axis and secondary processes (branches), and extraction of 36 multiscale (cell process segment level, process level, and whole cell level) topo-morphological and vessel-referenced metrics. Mural-VISTA identified morphological changes in pericytes and vascular smooth muscle cells (vSMCs) with altered RhoA activity. Constitutive active RhoA (RhoA CA) over-expression reduced branch complexity and increased process alignment in both cell types, while increased whole-cell and branch solidity only in vSMCs. Dominant negative RhoA (RhoA DN) over-expression increased branch abundance and reduced branch solidity in pericytes but not vSMCs, suggesting cell-type specific effect of reduced RhoA activity. In conclusion, Mural-VISTA enables quantitative 3D profiling of mural cell architecture and its spatial relationship with the vessel.

bioinformatics

Basophilic Erythroblast Emerges as the Key Turning Point in Polycythemia Vera

Abstract Polycythemia vera (PV) is a rare, chronic myeloproliferative neoplasm driven by the JAK2V617F mutation and characterized by uncontrolled erythroid proliferation. Although the mutation arises in hematopoietic stem cells, the differentiation stage at which its transcriptional consequences first become biologically meaningful has remained undefined. Using a multi-layer transcriptomics integration approach that combined differential gene expression, NicheNet ligand-receptor analysis, pseudotime trajectory inference, and CNV profiling on scRNA seq data, alongside bulk transcriptome validation, we identified basophilic erythroblasts as the critical transition point at which JAK2V617F shifts from a genomically present but transcriptionally silent state to an actively trajectory-altering and treatment-responsive disease driver. Differential expression revealed a qualitatively distinct disease signature at this stage, including ERFE-mediated iron dysregulation, MAP2K2-driven RAS/MAPK co-activation, and epigenetic reprogramming. NicheNet showed the establishment of a TGF{beta} superfamily and chemokine-driven niche-remodeling axis, and pseudotime analysis demonstrated that basophilic erythroblasts are the first erythroid population to exhibit condition-dependent trajectory divergence, whereas earlier progenitors showed none despite carrying the mutation. Interferon- treatment showed its broadest counterresponse at this stage but declined sharply thereafter, identifying basophilic erythroblasts as both the principal therapeutic target and the point of maximum vulnerability in PV.

bioinformatics

AVOCODO: An open-source multimodal annotation platform for developmental EEG

Behavioral annotation of synchronized video recordings is an essential step in developmental electroencephalography (EEG) research, supporting both the identification of behavior-related artifacts and the investigation of brain-behavior relationships. Existing annotation workflows, however, are often fragmented: proprietary EEG software provides limited flexibility for behavioral coding, whereas dedicated behavioral annotation platforms typically lack native integration with EEG data. We developed AVOCODO (Audio/VideO CODing Optimization), an open-source MATLAB-based software platform that integrates synchronized behavioral annotation directly into the EEG workflow. AVOCODO reads native EGI MFF recordings, synchronizes embedded video with EEG, visualizes the audio spectrogram to facilitate precise annotation of vocalizations, and writes user-defined behavioral events directly back into the original MFF recording as native EEG event markers while simultaneously exporting annotations as CSV files. The software supports fully customizable behavioral coding schemes, optional EEG visualization for quality control, and reloading of previously annotated recordings for review and inter-rater verification. Since its initial development in 2024, AVOCODO has been applied internally across five developmental EEG studies involving approximately 500 pediatric participants and more than 3,000 EEG recordings. By bridging behavioral annotation and EEG preprocessing within a unified open-source workflow, AVOCODO has the potential to improve the efficiency, reproducibility, and scalability of behavioral annotation in developmental EEG research.

bioinformatics

EXTRARNAS: A Framework for Extracting RNA Structures with Multiple Tools

Accurate annotation of RNA base-pairing interactions is essential for structural analysis, benchmarking, and data-driven RNA structure prediction. Several tools can extract RNA interactions from three-dimensional coordinates, but their outputs are heterogeneous and may disagree, particularly for non-canonical base pairs. We present EXTRARNAS, a Java-based framework for automated, reproducible, and user-friendly large-scale extraction of RNA structural annotations with multiple tools. EXTRARNAS processes batches of RNA structures specified by PDB identifier and chain, or provided as local PDB files, executes annotation tools through a Docker-based environment, and parses tool-specific outputs using ANTLR4-based grammars. For each structure-tool pair, the framework generates standard BPSEQ files for canonical cis Watson-Crick interactions and introduces BPSEQE, a standardized text format for representing the extended secondary structure, preserving canonical, non-canonical, and multiple interactions per nucleotide. The current prototype supports RNAView, MC-Annotate, and RNAPolis Annotator. We demonstrate EXTRARNAS on eight RNA structures containing triple-helix motifs, comparing extracted canonical pairs against curated BPSEQ references and evaluating the recovery of manually validated Hoogsteen interactions. The results show consistent differences among tools, especially for non-canonical interactions, highlighting the need for standardized representations such as BPSEQE to support reproducible comparison and future consensus-based annotation.

bioinformatics

CyChat: a conversational Cytoscape app for no-code, reproducible network analysis

Network-based analyses of molecular interactions are useful for interpreting high-throughput omics data and identifying therapeutic targets. Cytoscape is the standard platform for these tasks, but users face a trade-off between accessible graphical workflows that are difficult to document and reproducible automation in Python or R that requires programming expertise. General-purpose coding assistants can generate Cytoscape Automation scripts, but remain external to Cytoscape. We present CyChat, a Cytoscape Desktop app that integrates a chat interface and a large language model (LLM) agent into the application. CyChat translates natural language into executable Cytoscape Automation workflows, runs generated Python code, and exports chat sessions with executed code as standalone Jupyter notebooks. To reduce setup barriers, CyChat includes an embedded Python runtime and supports both cloud-based and locally hosted LLMs. CyChat was evaluated across ten Cytoscape workflows using seven LLM providers, each represented by one LLM. The strongest configuration achieves a pass rate above 99%. In a qualitative evaluation based on a published network visualization, CyChat completes the task in 1.5-5 minutes, compared with 15-20 minutes for manual GUI workflows by computational biologists. CyChat is available through the Cytoscape App Store at https://apps.cytoscape.org/apps/cychat.

bioinformatics

Transcriptomic profile of a rat jaw opener (anterior digastric) and a jaw closer (superficial masseter).

Mammalian skeletal muscle research predominantly focuses on locomotor muscles, and feeding related muscles remain less extensively characterized despite their role in mastication, mandibular stabilization, and swallowing. In this study, we investigated the transcriptomic specialization of three functionally and developmentally unique rat muscles: the anterior digastric (AD), a jaw opening muscle; the superficial masseter (SM), a jaw closing muscle; and the Sternohyoid (SH), a non-mandibular muscle involved in swallowing. Differential gene expression and weighted gene co-expression network analysis were used to characterize the transcription level features associated with their distinct roles. Our results indicated that all three muscles predominantly expressed fast-twitch contractile isoforms. However, the AD showed lower overall expression of several contractile gene families, including myosin heavy chain, myosin light chain, and tropomyosin isoforms, while exhibiting elevated expression of slow/oxidative myosin isoforms like Myh7 and Myh2. Network analysis revealed that modules correlated with AD are strongly enriched for fatty acid catabolism, mitochondrial energy production, and vascular/extracellular matrix remodeling. Additionally, AD and SM shared a distinct gene set compared to SH, highlighting their common developmental origin from the first branchial arch. Our findings show that the rat feeding related muscles possess unique transcriptomic profiles shaped by their contractile functions, developmental origins, and metabolic functions.

bioinformatics

AmPair: automating housekeeping-gene primer design for species-level metataxonomics

Amplicon sequencing of the 16S rRNA gene is the most widely used approach for profiling bacterial communities, but its taxonomic resolution is typically limited to the genus level. Many species carry multiple divergent 16S rRNA alleles that overlap across species boundaries, an ambiguity that even full-length, long-read sequencing cannot fully resolve. Shotgun metagenomics achieves species-level resolution but remains costly, particularly when only a single genus is of interest. Amplicon sequencing of rapidly evolving, protein-coding housekeeping genes offers a cost-effective alternative, yet no tool exists to identify suitable primer sets for a given target taxon. Here we present AmPair, a Snakemake pipeline that, given a target genus and one or more candidate housekeeping genes, designs and ranks primer pairs binding conserved regions while flanking a variable region capable of species-level discrimination, and validates them in silico across all available genomes. Using the genus Bacillus and the housekeeping gene tuf as a case study, the primer set recommended by AmPair amplified 99% of 2,392 genomes; only 0.04% carried multiple alleles and none showed inter-species allele overlap, compared with 91.41% and 69.49%, respectively, for the standard 16S rRNA V1-V9 region. Applied to a Bacillus community profiled by Nanopore sequencing, the same primers resolved closely related species. AmPair thus offers a generalizable and accessible route to species-level community profiling.

bioinformatics

Rclade: automated taxonomic collapsing and geological-timescale annotation of time-calibrated phylogenetic trees in R

Background: Reproducible taxonomic collapsing and geological-timescale annotation of time-calibrated phylogenetic trees in R often require coordination among several packages and repeated code for label parsing, clade validation, plotting, and export. Workflow-managed analyses additionally benefit from non-interactive configuration, predictable diagnostics, and machine-readable exit status. Results: We present Rclade, an R package that consolidates the multi-package coordination required for taxonomic collapsing into a streamlined, single-function interface. Rclade provides (1) custom ggproto objects (GeomPolygonStraight/GeomSegmentStraight) that bypass coord_munch() interpolation to achieve straight-edge rendering of collapsed triangles in circular layouts; (2) automatic detection and parsing of four taxonomic-label formats (GTDB, Silva, NCBI, embedded) plus user-supplied custom regex, with explicit input-validation contracts and parsing-accuracy evaluation on real and derived test sets; and (3) workflow embeddability through YAML configuration, library-mode APIs, and standard Unix exit codes. Benchmarks on synthetic and real datasets (200-10,000 synthetic tips and real reference trees up to 10,122 tips; 5 replicates at every scale under a unified fully rendered measurement protocol) show that the full-pipeline overhead is modest for interactive use (median {approx}0.87 s in-session rendering and {approx}8.4 s process-level wall-clock at 10,000 tips). Conclusions: Rclade is a convenience layer over the ggtree/deeptime ecosystem that reduces boilerplate while adding targeted technical improvements for circular-layout rendering and format heterogeneity management.

bioinformatics

Model-based evaluation of Targeted-Antibacterial-Plasmids (TAPs) transfer kinetics and resensitization of pOXA-48 carbapenem-resistant Escherichia coli

Background Targeted-Antibacterial Plasmids (TAPs) are engineered mobile genetic elements that use bacterial conjugation to deliver selective CRISPR/Cas9 antibacterial activity against a specific target strain. Yet, the efficiency of TAPs is typically evaluated at a single time point, whereas the success of TAP-mediated resensitization critically depends on the dynamics of plasmid transfer and the complex interactions between bacterial subpopulations. This is the first study to evaluate the efficiency of a conjugation-based antibacterial approach at the subpopulation level, using an analytical framework analogous to that used for conventional antibiotics. Here, we investigate which process limits resensitization by TAPF-dCas9-OXA48: plasmid delivery, dCas9 activity, or the emergence of refractory and escape populations. Methods We fitted a mechanistic model of five interacting subpopulations (donors, recipients, transconjugants, escapers, and recusants) to 44 longitudinal conjugation experiments and used the fitted model to explore a range of biologically relevant scenarios. Results Using longitudinal conjugation data spanning 24 h, we show that up to 24% of recipients become recusants within 24h, refractory to further conjugation via entry exclusion, while secondary transconjugant emergence stays below 0.01%. Overall resensitization efficiency reaches up to 80%. Conclusion Plasmid transfer, rather than dCas9 repression, therefore appears to be the main bottleneck limiting the efficiency of TAPF-dCas9-OXA48 efficiency. These results identify plasmid delivery as a key engineering target for improving the performance of future TAPs.

bioinformatics