bioRxiv Science⌕ Search

Biology subjects

Moriya, Y.

Publications and source records attributed to Moriya, Y..

3 recordsLinked to original sources

Extraction of biological terms using large language models enhances the usability of metadata in the BioSample database

BioSample is a comprehensive repository of experimental sample metadata, playing a crucial role in providing a comprehensive archive and enabling experiment searches regardless of type. However, the difficulty in comprehensively defining the rules for describing metadata and limited user awareness of best practices for metadata have resulted in substantial variability depending on the submitter. This inconsistency poses significant challenges to the findability and reusability of the data. Given the vast scale of BioSample, which hosts over 40 million records, manual curation is impractical. Rule-based automatic ontology mapping methods have been proposed to address this issue, but their effectiveness is limited by the heterogeneity of BioSample metadata. Recently, large language models (LLMs) have gained attention in natural language processing and have been expected as promising tools for automating metadata curation. In this study, we evaluated the performance of LLMs in extracting cell line names from BioSample descriptions using a gold-standard dataset derived from ChIP-Atlas, a secondary database of epigenomics experiment data, which manually curates samples. Our results demonstrated that LLM-assisted methods outperformed traditional approaches, achieving higher accuracy and coverage. We further extended this approach to extraction of information about experimentally manipulated genes from metadata where manual curation had not yet been applied in ChIP-Atlas. This also yielded successful results for the usage of the database, which facilitates more precise filtering of data and prevents misinterpretation caused by inclusion of unintended data. These findings underscore the potential of LLMs to improve the findability and reusability of experimental data in general, significantly reducing user workload and enabling more effective scientific data management.

bioinformatics↗

Global adaptive evolution involved in neuroticism and educational behaviors through the spread of anatomically modern humans

A cumulative process of cultural evolution occurred globally alongside the spread of anatomically modern humans (AMHs). This process of evolution was likely accompanied by global changes in mental and behavioral phenotypes, resulting in cultural differences between AMHs and Neanderthals. Globally, selective increases of frequencies were detected in alleles, which contributed to the suppression of neuroticism and advanced learning. This finding indicates that neuroticism was suppressed and learning behaviors developed through the global spread of AMHs. It suggests that global adaptation to psychosocial stress occurred, which contributed to cumulative cultural evolution by increasing cooperative learning. A comparison of phenotypic trends at the population level indicates that such globally occurring adaptive evolution, entailing neuroticism suppression and increased learning behaviors has been more significant in AMHs than in Neanderthals. However, adaptive evolution toward greater intelligence was more significant in Neanderthals than in AMHs. These findings suggest that mental traits involved in cultural development evolved differently in AMHs and Neanderthals, and that learning behaviors such as cooperative learning, rather than intelligence, played an important role in AMH evolution.

evolutionary biology↗

Enteropathway: the metabolic pathway database for the human gut microbiota

The human gut microbiota produces diverse, extensive metabolites which have the potential to affect host physiology. Despite significant efforts to identify metabolic pathways for producing these microbial metabolites, a comprehensive metabolic pathway database for the human gut microbiota is still lacking. Here, we present Enteropathway, a metabolic pathway database that integrates 3,121 compounds, 3,460 reactions, and 837 modules that were obtained from 835 manually curated scientific literature. Notably, 757 modules of these modules are new entries and cannot be found in any other databases. The database is accessible from a web application (https://enteropathway.org) that offers a metabolic diagram for graphical visualization of metabolic pathways, a customization interface, and an enrichment analysis feature for highlighting enriched modules on the metabolic diagram. Overall, Enteropathway is a comprehensive reference database and a tool for visual and statistical analysis in human gut microbiota studies and was designed to help researchers pinpoint new insights into the complex interplay between microbiota and host metabolism.

bioinformatics↗