bioRxiv ScienceSearch

Biology subjects

Clark, T.

Publications and source records attributed to Clark, T..

6 recordsLinked to original sources

SMRT-Cappable-seq reveals complex operon variants in bacteria

Current methods for genome-wide analysis of gene expression requires shredding original transcripts into small fragments for short-read sequencing. In bacteria, the resulting fragmented information hides operon complexity. Additionally, in-vivo processing of transcripts confounds the accurate identification of the 5 and 3 ends of operons. Here we developed a novel methodology called SMRT-Cappable-seq that combines the isolation of unfragmented primary transcripts with single-molecule long read sequencing. Applied to E. coli, this technology results in an unprecedented definition of the transcriptome with 34% of the known operons being extended by at least one gene. Furthermore, 40% of transcription termination sites have read-through that alters the gene content of the operons. As a result, most of the bacterial genes are present in multiple operon variants reminiscent of eukaryotic splicing. By providing an unprecedented granularity in the operon structure, this study represents an important resource for the study of prokaryotic gene network and regulation.

genomics

Pheno4J: A Gene To Phenotype Graph Database

SummaryEfficient storage and querying of large amounts of genetic and phenotypic data is crucial to contemporary clinical genetic research. This introduces computational challenges for classical relational databases, due to the sparsity and sheer volume of the data. Our Java based solution loads annotated genetic variants and well phenotyped patients into a graph database to allow fast efficient storage and querying of large volumes of structured genetic and phenotypic data. This abstracts technical problems away and lets researchers focus on the science rather than the implementation. We have also developed an accompanying webserver with end-points to facilitate querying of the database.\n\nAvailability and ImplementationThe Java code and python code is available at https://github.com/phenopolis/pheno4i\n\nContactn.pontikos@ucl.ac.uk

bioinformatics

A Data Citation Roadmap for Scientific Publishers

This article presents a practical roadmap for scholarly publishers to implement data citation in accordance with the Joint Declaration of Data Citation Principles (JDDCP), a synopsis and harmonization of the recommendations of major science policy bodies. It was developed by the Publishers Early Adopters Expert Group as part of the Data Citation Implementation Pilot (DCIP) project, an initiative of FORCE11.org and the NIH BioCADDIE program. The structure of the roadmap presented here follows the \"life of a paper\" workflow and includes the categories Pre-submission, Submission, Production, and Publication. The roadmap is intended to be publisher-agnostic so that all publishers can use this as a starting point when implementing JDDCP-compliant data citation. Authors reading this roadmap will also better know what to expect from publishers and how to enable their own data citations to gain maximum impact, as well as complying with what will become increasingly common funder mandates on data transparency.

scientific communication and education

Uniform Resolution of Compact Identifiers for Biomedical Data

Most biomedical data repositories issue locally-unique accessions numbers, but do not provide globally unique, machine-resolvable, persistent identifiers for their datasets, as required by publishers wishing to implement data citation in accordance with widely accepted principles. Local accessions may however be prefixed with a namespace identifier, providing global uniqueness. Such \"compact identifiers\" have been widely used in biomedical informatics to support global resource identification with local identifier assignment.\n\nWe report here on our project to provide robust support for machine-resolvable, persistent compact identifiers in biomedical data citation, by harmonizing the Identifiers.org and N2T.net (Name-To-Thing) meta-resolvers and extending their capabilities. Identifiers.org services hosted at the European Molecular Biology Laboratory - European Bioinformatics Institute (EMBL-EBI), and N2T.net services hosted at the California Digital Library (CDL), can now resolve any given identifier from over 600 source databases to its original source on the Web, using a common registry of prefix-based redirection rules.\n\nWe believe these services will be of significant help to publishers and others implementing persistent, machine-resolvable citation of research data.

bioinformatics

A Data Citation Roadmap for Scholarly Data Repositories

This article presents a practical roadmap for scholarly data repositories to implement data citation in accordance with the Joint Declaration of Data Citation Principles, a synopsis and harmonization of the recommendations of major science policy bodies. The roadmap was developed by the Repositories Expert Group, as part of the Data Citation Implementation Pilot (DCIP) project, an initiative of FORCE11.org and the NIH BioCADDIE (https://biocaddie.org) program. The roadmap makes 11 specific recommendations, grouped into three phases of implementation: a) required steps needed to support the Joint Declaration of Data Citation Principles, b) recommended steps that facilitate article/data publication workflows, and c) optional steps that further improve data citation support provided by data repositories.

scientific communication and education

Phenopolis: an open platform for harmonization and analysis of sequencing and phenotype data

SummaryPhenopolis is an open-source web server which provides an intuitive interface to genetic and phenotypic databases. It integrates analysis tools which include variant filtering and gene prioritisation based on phenotype. The Phenopolis platform will accelerate clinical diagnosis, gene discovery and encourage wider adoption of the Human Phenotype Ontology in the study of rare disease.\n\nAvailability and ImplementationA demo of the website is available at http://phenopolis.github.io (username: demo, password: demo123). If you wish to install a local copy, souce code and installation instruction are available at https://github.com/pontikos/phenopolis. The software is implemented using Python, MongoDB, HTML/Javascript and various bash shell scripts.\n\nContactn.pontikos@ucl.ac.uk\n\nSupplementary informationhttp://phenopolis.github.io

bioinformatics