bioRxiv ScienceSearch

bioRxiv · 10.1101/2020.12.07.415059

Communicating Regulatory High Throughput Sequencing Data Using BioCompute Objects

Abstract

For regulatory submissions of next generation sequencing (NGS) data it is vital for the analysis workflow to be robust, reproducible, and understandable. This project demonstrates that the use of the IEEE 2791-2020 Standard, (BioCompute objects [BCO]) enables complete and concise communication of NGS data analysis results. One arm of a clinical trial was replicated using synthetically generated data made to resemble real biological data. Two separate, independent analyses were then carried out using BCOs as the tool for communication of analysis: one to simulate a pharmaceutical regulatory submission to the FDA, and another to simulate the FDA review. The two results were compared and tabulated for concordance analysis: of the 118 simulated patient samples generated, the final results of 117 (99.15%) were in agreement. This high concordance rate demonstrates the ability of a BCO, when a verification kit is included, to effectively capture and clearly communicate NGS analyses within regulatory submissions. BCO promotes transparency and induces reproducibility, thereby reinforcing trust in the regulatory submission process.

Source connections

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

King, C. H. S., Keeney, J. G., Guimera, N., Das, S., Mazumder, R., Fochtman, B., Telawar, S., Patel, J., Walderhaug, M. O., Donaldson, E. F.. 2020-12-09. Communicating Regulatory High Throughput Sequencing Data Using BioCompute Objects. https://doi.org/10.1101/2020.12.07.415059

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Evaluating Large Language Models as Tools to Navigate Researchers in Rapidly Evolving Research Landscapes: A Case Study in Cancer Drug Response Prediction

Large Language Models (LLMs) have emerged as promising tools for assisting researchers in automating and accelerating the synthesis of literature reviews. However, their reliability is a significant concern due to issues like factual inaccuracies and hallucinations. The key question is whether LLMs can reliably provide comprehensive, up-to-date overviews and analyses. This study evaluates the performance of three leading LLMs (OpenAI's ChatGPT, Google's Gemini, and DeepSeek) on the complex task of generating a comprehensive survey paper on deep learning for cancer Drug Response Prediction (DRP). By testing both standard and Deep Research (DR) / Deep Think (DT) modes of LLMs with prompts of varying detail, this paper assesses key academic dimensions, including reference management, content quality, and analytical depth. Key findings reveal that while DR modes of LLMs significantly improve reliability by eliminating hallucinations, performance variations exist across models and prompts. A trade-off between reference quantity and integration quality was observed, and even the best-performing models lacked the analytical depth of human experts, often requiring extensive human supervision. The study concludes that LLMs currently serve as powerful assistive tools but still cannot replace the critical validation and synthesis provided by human researchers. Choosing the best LLM to use depends on the task in hand, while several strategies can be implemented to improve the produced output.

scientific communication and education

Research Funding for Male Reproductive Health and Infertility in the UK and USA

TitleResearch Funding for Male Reproductive Health and Infertility in the UK and USA [2016 - 2019] Study questionWhat is the research funding for male reproductive health and infertility in the UK and US between 2016 to 2019? Summary answerThe average funding for a research project in male reproductive health and infertility was not significantly different to that for female-based projects ({pound}653,733 in the UK and $779,707 in the US). However, only 0.07% and 0.83% of government funds from NIHR (UK) and NICHD (USA) was granted for male reproductive health, respectively. What is known alreadyThere is a marked paucity of data on research funding for male reproductive health. Study design, size, durationExamined government databases over a total 4-year period from January 2016 to December 2019. Participants/materials, setting, methodsInformation on the funding provided to male-based and female-based research was collected using public accessed web-databases from the UKRI-GTR, the NIHRs Open Data Summary, and the US NIH RePORT. Funded projects that began research activity between January 2016 to December 2019 were recorded, along with their grant and project details. Strict inclusion-exclusion criteria were followed for both UK and US data with a primary research focus of male infertility, reproductive health and disorders, and contraception development. Funding support was divided into three research groups: male-based, female-based, and not-specified research. Between the 4-year period, the UK is divided into 5 funding periods, starting from 2015/16 to 2019/20, and the US is divided into 5 fiscal years, from 2016 to 2020. Main results and the role of chanceBetween January 2016 to December 2019, UK agencies awarded a total of {pound}11,767,190 to 18 projects for male-based research and {pound}29,850,945 to 40 projects for female-based research. There was no statistically significant difference in funding average between the two research groups (P=0.56, W=392). The US NIH funded 76 projects totaling $59,257,746 for male-based research and 99 projects totaling $83,272,898 for female-based research. There was no statistically significant difference in funding average between the two groups (P=0.83, W=3834). Limitations, reasons for cautionThe findings of this study cannot be used to generalize and reflect global funding trends towards infertility and reproductive health as the data collected followed a narrow funding timeframe from government agencies and only two countries. Other funding sources such as charities, industry and major philanthropic organizations were not evaluated. Wider implications of the findingsThis is the first study examining funding granted by main government research agencies from the UK and US for male reproductive health. This study should stimulate further discussion of the challenges of tackling male infertility and reproductive health disorders and formulate appropriate investment strategies. Study funding/competing interest(s)CLRB is Editor for RBMO and has received lecturing fees from Merck, Pharmasure, and Ferring. His laboratory is funded by Bill and Melinda Gates Foundation, CSO, Genus. No other authors declare a conflict of interest.

scientific communication and education

Time Evolution of the Stroke Symptom-Herb Networks Based on TCM Prescriptions

Traditional Chinese Medicine (TCM) has its origins in distant antiquity and has piled up over a long time with much knowledge about diseases, especially stroke. Different combinations of symptom variables yield different combinations of herbs to form a myriad of prescriptions, and they have undergone repeated confirmation and are worthy objects of excavation and analysis. Herbal studies on stroke have developed from genomics to transcriptomics, proteomics and metabolomics, yet more thought is needed on putting time in a wider perspective of the symptoms and herbs for stroke in Chinese medicine. Due to this, we studied the dynamic structure of TCM prescriptions on stroke, using over 270 TCM prescription books containing 2231 prescriptions related to stroke recorded from 341 to 2000 CE. We labeled the functions of the prescriptions with the symptoms based on subject terms in MESH Neurologic Manifestations, then standardized the herbs in the prescriptions, and finally connected the co-occurring symptoms and herbs in the prescriptions to build an undirected complex network. The Stroke Symptom-Herb Networks (SSHNs) can be seen from its network characteristics that it is not a random network and has small-world characteristics. It has experienced two peaks in its nearly 1700-year history, during the Song dynasty, the Ming and Qing dynasties. From 600 years onwards, the core herb cluster has been initially formed. The comparison of sub-network similarities allowed us to identify several symptoms with similar herb clusters. We divided the community based on modularity, and by analyzing the community evolution, we found a more fixed historical evolutionary trend with Hemiplegia and Sialorrhea nodes and their associated symptom and herb nodes. In the time series analysis, we found many symptom-herb combinations that were consistently closely related to historic time depends on assessing the similarity between the symptoms and the herbs. The complex network provides a distinctive perspective for understanding the symptom-herb relationships embedded in TCM prescriptions in remote antiquity.

scientific communication and education