bioRxiv ScienceSearch

Biology subjects

Izzo, M.

Publications and source records attributed to Izzo, M..

3 recordsLinked to original sources

PhenoMeNal: Processing and analysis of Metabolomics data in the Cloud

BackgroundMetabolomics is the comprehensive study of a multitude of small molecules to gain insight into an organisms metabolism. The research field is dynamic and expanding with applications across biomedical, biotechnological and many other applied biological domains. Its computationally-intensive nature has driven requirements for open data formats, data repositories and data analysis tools. However, the rapid progress has resulted in a mosaic of independent - and sometimes incompatible - analysis methods that are difficult to connect into a useful and complete data analysis solution.\n\nFindingsThe PhenoMeNal (Phenome and Metabolome aNalysis) e-infrastructure provides a complete, workflow-oriented, interoperable metabolomics data analysis solution for a modern infrastructure-as-a-service (IaaS) cloud platform. PhenoMeNal seamlessly integrates a wide array of existing open source tools which are tested and packaged as Docker containers through the projects continuous integration process and deployed based on a kubernetes orchestration framework. It also provides a number of standardized, automated and published analysis workflows in the user interfaces Galaxy, Jupyter, Luigi and Pachyderm.\n\nConclusionsPhenoMeNal constitutes a keystone solution in cloud infrastructures available for metabolomics. It provides scientists with a ready-to-use, workflow-driven, reproducible and shareable data analysis platform harmonizing the software installation and configuration through user-friendly web interfaces. The deployed cloud environments can be dynamically scaled to enable large-scale analyses which are interfaced through standard data formats, versioned, and have been tested for reproducibility and interoperability. The flexible implementation of PhenoMeNal allows easy adaptation of the infrastructure to other application areas and omics research domains.

bioinformatics

FAIRsharing: working with and for the community to describe and link data standards, repositories and policies

In this modern, data-driven age, governments, funders and publishers expect greater transparency and reuse of research data, as well as greater access to and preservation of the data that supports research findings. Community-developed standards, such as those for the identification1 and reporting2 of data, underpin reproducible and reusable research, aid scholarly publishing, and drive both the discovery and evolution of scientific practice. The number of these standardization efforts, driven by large organizations or at the grass root level, has been on the rise since the early 2000s. Thousands of community-developed standards are available (across all disciplines), many of which have been created and/or implemented by several thousand data repositories. Nevertheless, their uptake by the research community, however, has been slow and uneven. This is mainly because investigators lack incentives to follow and adopt standards. The situation is exacerbated if standards are not promptly implemented by databases, repositories and other research tools, or endorsed by infrastructures. Furthermore, the fragmentation of community efforts results in the development of arbitrarily different, incompatible standards. In turn, this leads to standards becoming rapidly obsolete in fast-evolving research areas.\n\nAs with any other digital object, standards, databases and repositories are dynamic in nature, with a life cycle that encompasses formulation, development and maintenance; their status in this cycle may vary depending on the level of activity of the developing group or community. There is an urgent need for a service that enhances the information available on the evolving constellation of heterogeneous standards, databases and repositories, guides users in the selection of these resources, and that works with developers and maintainers of these resources to foster collaboration and promote harmonization. Such an informative and educational service is vital to reduce the knowledge gap among those involved in producing, managing, serving, curating, preserving, publishing or regulating data. A diverse set of stakeholders-representing academia, industry, funding agencies, standards organizations, infrastructure providers and scholarly publishers-- both national and domain-specific as well global and general organizations-- have come together as a community, representing the core adopters, advisory board members, and/or key collaborators of the FAIRsharing resource. Here, we introduce its mission and community network. We present an evaluation of the standards landscape, focusing on those for reporting data and metadata - the most diverse and numerous of the standards - and their implementation by databases and repositories. We report on the ongoing challenge to recommend resources, and we discuss the importance of making standards invisible to the end users. We report on the ongoing challenge to recommend resources, and we discuss the importance of making standards invisible to the end users. We present guidelines that highlight the role each stakeholder group must play to maximize the visibility and adoption of standards, databases and repositories.

scientific communication and education

BioSharing: Harnessing Metadata Standards For The Data Commons

The use of community-driven metadata standards, such as minimal information guidelines, terminologies, formats/models, is essential to ensure that data and other digital research outputs are Findable, Accessible, Interoperable, and Reusable, according to the FAIR principles. As with other types of digital assets, metadata standards also need be FAIR. Their discoverability and accessibility is ensured by BioSharing, the most comprehensive resource of metadata standards, interlinked to data repositories and policies, available in the life, environmental and biomedical sciences. With its growing content, endorsements, and collaborative network, BioSharing is part of a larger ecosystem of interoperable resources. Here we describe some of the activities under the USA National Institutes of Health (NIH)s Big Data to Knowledge (BD2K) Initiative, illustrating how we track the evolution and use of metadata standards and work to connect them to indexes and annotation tools.

scientific communication and education