bioRxiv Science⌕ Search

Biology subjects

Jandalala, I.

Publications and source records attributed to Jandalala, I..

2 recordsLinked to original sources

GOFlowLLM - Curating miRNA literature with Large Language Models and flowcharts

The exponential growth of non-coding RNA research--with over 230,000 papers published since 2000--has created an urgent knowledge management crisis in molecular biology. Despite their crucial regulatory roles, microRNAs (miRNAs) face a significant curation bottleneck, with only 1,400 articles manually curated to the Gene Ontology (GO) knowledgebase over a decade. We present GOFlowLLM, an automated curation pipeline powered by reasoning-enabled Large Language Models (LLMs) that follows established GO curation flowcharts to extract and structure miRNA-mediated gene silencing data at scale. When evaluated on existing curation, GOFlowLLM selects the correct GO term in 90% of cases. Curators also agree with 95% of the systems reasoning steps and 90% of the evidence selected. Applied to 6,996 previously uncurated articles, our system identified 2,538 new candidate GO annotations on 1,785 articles in just 58 hours--potentially doubling the available miRNA GO curation. Manual review of a subset of these annotations shows that curators agreed with the selected term in 87% of cases, the models reasoning in 92% of cases, and the extracted evidence in 93%. GOFlowLLM demonstrates how LLMs can significantly accelerate biocuration while maintaining high-quality standards by following expert-designed reasoning frameworks. The integration of reasoning traces in our system provides transparent justification for annotations that can be reviewed by human curators, addressing one of the key challenges in adopting AI for scientific curation, potentially transforming how we manage the growing corpus of scientific knowledge in molecular biology. GoFlowLLM is available on github: https://github.com/RNAcentral/GO_Flow_LLM.

bioinformatics↗

RNAcentral in 2026: Genes and literature integration

RNAcentral was founded in 2014 to serve as a comprehensive database of non-coding RNA sequences. It began by providing a single unified interface to more specialised resources, and now contains 45 million sequences. It has grown beyond providing a single interface to many specialised resources and now provides several services and analyses. These include secondary structure prediction with R2DT, sequence search, and analysis with Rfam. Since its last publication in 2021, RNAcentral has developed two major features. First, literature integration with the development of LitScan and LitSumm. LitScan automatically identifies and links relevant publications to RNA entries, while LitSumm uses natural language processing to generate functional summaries from the literature. Together, these tools address the critical challenge of connecting sequence data with scattered functional knowledge across thousands of publications. Secondly, RNAcentral has created gene level entries. Gene level entries represent a large structural change to RNAcentral. While RNAcentral previously organized data exclusively at the sequence level, we now group related transcripts into gene-centric views. This allows researchers to explore all isoforms, splice variants, and related sequences for a gene in a unified interface, better reflecting biological organization and facilitating comparative analyses. RNAcentral is freely available at: https://rnacentral.org.

bioinformatics↗