bioRxiv Science⌕ Search

Biology subjects

Keeney, J.

Publications and source records attributed to Keeney, J..

2 recordsLinked to original sources

BioCompute Objects to communicate a viral detection pipeline with potential for use in a regulatory environment

The volume of nucleic acid sequence data has exploded in recent years, and with it, the challenge of finding and transforming relevant data into meaningful information. Processing the abundance of data can require a dynamic ecosystem of customized tools. As analysis pipelines become more complex, there is an increased difficulty in communicating analysis details in a way that is understandable yet of sufficient detail to make informed decisions about results or repeat the analysis. This may be of particular interest to institutions and private companies that need to communicate complex computations in a regulatory environment. To meet this need for standard reporting, the open source BioCompute framework was developed as a standardized mechanism for communicating the details of an analysis in a concise and organized way, and other tools and interfaces were subsequently developed according to the standard. The goal of BioCompute is to streamline the process of communicating computational analyses. Reports that conform to the BioCompute standard are called BioCompute Objects (BCOs). Here, a comprehensive suite of BCOs is presented, representing interconnected elements of a computation that is modeled after those that might be found in a regulatory submission, but which can be shared publicly. Because BCOs are human and machine readable, they can be displayed in customized ways to further improve their utility, and an example of a collapsible format is shown. The work presented here serves as a real world implementation that imitates actual submissions, providing concrete examples. As an example, a pipeline designed to identify viral contaminants in biological manufacturing, such as for vaccines, is developed and rigorously tested to establish a rate of false positive detection, and is described in a BCO report. That pipeline relies on a specially curated database for alignment, and a set of synthetic reads for testing, both of which are also descriptively packaged in their own BCOs. All of the sufficiently complex processes associated with this analysis are therefore represented as BCOs that can be cross-referenced, demonstrating the modularity of BCOs, their ability to organize tremendous complexity, and their use in a lifelike regulatory environment.

bioinformatics↗

Strengthening the BioCompute Standard by Crowdsourcing on PrecisionFDA

BackgroundThe field of bioinformatics has grown at such a rapid pace that a gap in standardization exists when reporting an analysis. In response, the BioCompute project was created to standardize the type and method of information communicated when describing a bioinformatic analysis. Once the project became established, its goals shifted to broadening awareness and usage of BioCompute, and soliciting feedback from a larger audience. To address these goals, the BioCompute project collaborated with precisionFDA on a crowdsourced challenge that ran from May 2019 to October 2019. This challenge had a beginner track where participants submitted BCOs based on a pipeline of their choosing, and an advanced track where participants submitted applications supporting the creation of a BCO and verification of BCO conformance to specifications. ResultsIn total, there were 28 submissions to the beginner track (including submissions from a bioinformatics masters class at George Washington University) and three submissions to the advanced track. Three top performers were selected from the beginner track, while a single top performer was selected for the advanced track. In the beginner track, top performers differentiated themselves by submitting BCOs that included more than the minimally compliant content. Advanced track submissions were very impressive. They included a complete web application, a command line tool that produced a static result, and a dockerized container that automatically created the BCO as the tool was run. The ability to harmonize the correct function, a simple user experience, and the aesthetics of the tool interface differentiated the tools. ConclusionsDespite being new to the concept, most beginner track scores were high, indicating that most users understood the fundamental concepts of the BCO specification. Novice bioinformatics students were an ideal cohort for this Challenge because of their lack of familiarity with BioCompute, broad diversity of research interests, and motivation to submit high-quality work. This challenge was successful in introducing the BCO to a wider audience, obtaining feedback from that audience, and resulting in a tool novices may use for BCO creation and conformance. In addition, the BCO specification itself was improved based on feedback illustrating the utility of a "wisdom of the crowd" approach to standards development.

bioinformatics↗