bioRxiv ScienceSearch

Biology subjects

Mungall, C.

Publications and source records attributed to Mungall, C..

3 recordsLinked to original sources

Identifiers for the 21st century:How to design, provision, and reuse persistent identifiers to maximize utility and impact of life science data

In many disciplines, data is highly decentralized across thousands of online databases (repositories, registries, and knowledgebases). Wringing value from such databases depends on the discipline of data science and on the humble bricks and mortar that make integration possible; identifiers are a core component of this integration infrastructure. Drawing on our experience and on work by other groups, we outline ten lessons we have learned about the identifier qualities and best practices that facilitate large-scale data integration. Specifically, we propose actions that identifier practitioners (database providers) should take in the design, provision and reuse of identifiers; we also outline important considerations for those referencing identifiers in various circumstances, including by authors and data generators. While the importance and relevance of each lesson will vary by context, there is a need for increased awareness about how to avoid and manage common identifier problems, especially those related to persistence and web-accessibility/resolvability. We focus strongly on web-based identifiers in the life sciences; however, the principles are broadly relevant to other disciplines.

bioinformatics

WTFgenes: What's The Function of these genes? Static sites for model-based gene set analysis

A common technique for interpreting experimentally-identified lists of genes is to look for enrichment of genes associated to particular ontology terms. The most common technique uses the hypergeometric distribution; more recently, a model-based approach was proposed. These approaches must typically be run using downloaded software, or on a server. We develop a collapsed likelihood for model-based gene set analysis and present WTFgenes, an implementation of both hypergeometric and model-based approaches, that can be published as a static site with computation run in JavaScript on the user's web browser client. Apart from hosting files, zero server resources are required: the site can (for example) be served directly from Amazon S3 or GitHub Pages. A C++11 implementation yielding identical results runs roughly twice as fast as the JavaScript version. WTFgenes is available from https://github.com/evoldoers/wtfgenes under the BSD3 license. A demonstration for the Gene Ontology is usable at https://evoldoers.github.io/wtfgo. Contact: Ian Holmes ihholmes+wtfgenes@gmail.com.

bioinformatics

The Monarch Initiative: Insights across species reveal human disease mechanisms

The principles of genetics apply across the whole tree of life: on a cellular level, we share mechanisms with species from which we diverged millions or even billions of years ago. We can exploit this common ancestry at the level of sequences, but also in terms of observable outcomes (phenotypes), to learn more about health and disease for humans and all other species. Applying the range of available knowledge to solve challenging disease problems requires unified data relating genomics, phenotypes, and disease; it also requires computational tools that leverage these multimodal data to inform interpretations by geneticists and to suggest experiments. However, the distribution and heterogeneity of databases is a major impediment: databases tend to focus either on a single data type across species, or on single species across data types. Although each database provides rich, high-quality information, no single one provides unified data that is comprehensive across species, biological scales, and data types. Without a big-picture view of the data, many questions in genetics are difficult or impossible to answer. The Monarch Initiative (https://monarchinitiative.org) is an international consortium dedicated to providing computational tools that leverage a computational representation of phenotypic data for genotype-phenotype analysis, genomic diagnostics, and precision medicine on the basis of a large-scale platform of multimodal data that is deeply integrated across species and covering broad areas of disease.

evolutionary biology