bioRxiv Science⌕ Search

Biology subjects

Chin, K. Y.

Publications and source records attributed to Chin, K. Y..

2 recordsLinked to original sources

Assessing the Performance of LLMs in Multimodal Information Extraction for Biological Research: A Case Study on LLPS

Advances in experimental techniques have expanded the volume of biological data. This has increased the demand for structured information extraction from papers, with large language models (LLMs) considered promising. However, challenges remain, including limited validation in biology and unclear applicability to multimodal tasks that integrate text with domain-specific figures, such as microscopic images and scatter plots. Here, we developed a multimodal LLM (MLLM)-based workflow to extract the experimental conditions and phase status from the text and figures of experimental papers on liquid-liquid phase separation and validated the effect of various inputs, prompts, and MLLM types. As a result, the Gemini 2.5 Pro-based extraction achieved an F1-score of 0.847 by processing each figure as a processing unit and inputting domain-specific prompts reflecting manual extraction guidance. This study demonstrates the potential and limitations of MLLMs for extracting biological information and provides insights for advancing multimodal approaches in biology.

bioinformatics↗

Predicting condensate formation of protein and RNA under various environmental conditions

MotivationLiquid-liquid phase separation (LLPS) by biomolecules plays a central role in various biological phenomena and has garnered significant attention. The behavior of LLPS is strongly influenced by the characteristics of the RNAs and environmental factors such as pH and temperature, as well as the properties of the proteins. Recently, several databases of biomolecules associated with LLPS have been established, and prediction models of LLPS-related phenomena have been explored, leveraging these databases. However, a prediction model that concurrently considers proteins, RNAs, and experimental conditions has not been developed due to the limited information available from individual experiments in public databases. ResultsTo address this challenge, we have built a new dataset called RNAPSEC, which serves each individual experiment as a data point. This dataset was accomplished by manually collecting data from public literature. Utilizing RNAPSEC, we developed two distinct models that consider a protein, RNA, and experimental conditions. The first model can predict the LLPS behavior of a protein and RNA under specific conditions. The second model can predict the required conditions for a given protein and RNA to undergo LLPS. RNAPSEC and these prediction models are expected to accelerate our understanding of the roles of proteins, RNAs, and environmental factors in LLPS. AvailabilityThe codes for the prediction models and RNAPSEC are available at https://github.com/ycu-iil/RNAPSEC. Contactterayama@yokohama-cu.ac.jp

bioinformatics↗