bioRxiv ScienceSearch

Biology subjects

Andrew I Su

Publications and source records attributed to Andrew I Su.

7 recordsLinked to original sources

CIViC: A knowledgebase for expert-crowdsourcing the clinical interpretation of variants in cancer.

CIViC is an expert crowdsourced knowledgebase for Clinical Interpretation of Variants in Cancer (www.civicdb.org) describing the therapeutic, prognostic, and diagnostic relevance of inherited and somatic variants of all types. CIViC is committed to open source code, open access content, public application programming interfaces (APIs), and provenance of supporting evidence to allow for the transparent creation of current and accurate variant interpretations for use in cancer precision medicine.

Cancer Biology

Knowledge.Bio: A Web application for exploring, building and sharing webs of biomedical relationships mined from PubMed

Knowledge.Bio is a web platform that enhances access and interpretation of knowledge networks extracted from biomedical research literature. The interaction is mediated through a collaborative graphical user interface for building and evaluating maps of concepts and their relationships, alongside associated evidence. In the first release of this platform, conceptual relations are drawn from the Semantic Medline Database and the Implicitome, two compleme ntary resources derived from text mining of PubMed abstracts.\n\nAvailability-- Knowledge.Bio is hosted at http://knowledge.bio/ and the open source code is available at http://bitbucket.org/sulab/kb1/.\n\nContact-- asu@scripps.edu; bgood@scripps.edu

Bioinformatics

A comprehensive and scalable database search system for metaproteomics

BackgroundMass spectrometry-based shotgun proteomics experiments rely on accurate matching of experimental spectra against a database of protein sequences. Existing computational analysis methods are limited in the size of their sequence databases, which severely restricts the proteomic sequencing depth and functional analysis of highly complex samples. The growing amount of public high-throughput sequencing data will only exacerbate this problem. We designed a broadly applicable metaproteomic analysis method (ComPIL) that addresses protein database size limitations.\n\nResultsOur approach to overcome this significant limitation in metaproteomics was to design a scalable set of sequence databases assembled for optimal library querying speeds. ComPIL was integrated with a modified version of the search engine ProLuCID (termed \"Blazmass\") to permit rapid matching of experimental spectra. Proof-of-principle analysis of human HEK293 lysate with a ComPIL database derived from high-quality genomic libraries was able to detect nearly all of the same peptides as a search with a human database (~500x fewer peptides in the database), with a small reduction in sensitivity. We were also able to detect proteins from the adenovirus used to immortalize these cells. We applied our method to a set of healthy human gut microbiome proteomic samples and showed a substantial increase in the number of identified peptides and proteins compared to previous metaproteomic analyses, while retaining a high degree of protein identification accuracy, and allowing for a more in-depth characterization of the functional landscape of the samples.\n\nConclusionsThe combination of ComPIL with Blazmass allows proteomic searches to be performed with database sizes much larger than previously possible. These large database searches can be applied to complex meta-samples with unknown composition or proteomic samples where unexpected proteins may be identified. The protein database, proteomics search engine, and the proteomic data files for the 5 microbiome samples characterized and discussed herein are open source and available for use and additional analysis.

Bioinformatics

Citizen Science for Mining the Biomedical Literature

I.Biomedical literature represents one of the largest and fastest growing collections of unstructured biomedical knowledge. Finding critical information buried in the literature can be challenging. In order to extract information from freeflowing text, researchers need to: 1. identify the entities in the text (named entity recognition), 2. apply a standardized vocabulary to these entities (normalization), and 3. identify how entities in the text are related to one another (relationship extraction). Researchers have primarily approached these information extraction tasks through manual expert curation, and computational methods. We have previously demonstrated that named entity recognition (NER) tasks can be crowdsourced to a group of nonexperts via the paid microtask platform, Amazon Mechanical Turk (AMT); and can dramatically reduce the cost and increase the throughput of biocuration efforts. However, given the size of the biomedical literature even information extraction via paid microtask platforms is not scalable. With our web-based application Mark2Cure (http://mark2cure.org), we demonstrate that NER tasks can also be performed by volunteer citizen scientists with high accuracy. We apply metrics from the Zooniverse Matrices of Citizen Science Success and provide the results here to serve as a basis of comparison for other citizen science projects. Further, we discuss design considerations, issues, and the application of analytics for successfully moving a crowdsourcing workflow from a paid microtask platform to a citizen science platform. To our knowledge, this study is the first application of citizen science to a natural language processing task.

Bioinformatics

MyGene.info and MyVariant.info: Gene and Variant Annotation Query Services

MyGene.info and MyVariant.info provide high-performance data APIs for querying gene and variant annotation information. They demonstrate a new model for organizing biological annotation information by utilizing a cloud-based scalable infrastructure. MyGene.info and MyVariant.info can be accessed at http://mygene.info and http://myvariant.info.

Bioinformatics

Wikidata: A platform for data integration and dissemination for the life sciences and beyond

Wikidata is an open, Semantic Web-compatible database that anyone can edit. This data commons provides structured data for Wikipedia articles and other applications. Every article on Wikipedia has a hyperlink to an editable item in this database. This unique connection to the worlds largest community of volunteer knowledge editors could help make Wikidata a key hub within the greater Semantic Web. The life sciences, as ever, faces crucial challenges in disseminating and integrating knowledge. Our group is addressing these issues by populating Wikidata with the seeds of a foundational semantic network linking genes, drugs and diseases. Using this content, we are enhancing Wikipedia articles to both increase their quality and recruit human editors to expand and improve the underlying data. We encourage the community to join us as we collaboratively create what can become the most used and most central semantic data resource for the life sciences and beyond.

Bioinformatics

Centralizing content and distributing labor: a community model for curating the very long tail of microbial genomes.

The last 20 years of advancement in DNA sequencing technologies have led to the sequencing of thousands of microbial genomes, creating mountains of genetic data. While our efficiency in generating the data improves almost daily, applying meaningful relationships between the taxonomic and genetic entities requires a new approach. Currently, the knowledge is distributed across a fragmented landscape of resources from government-funded institutions such as NCBI and Uniprot to topic-focused databases like the ODB3 database of prokaryotic operons, to the supplemental table of a primary publication. A major drawback to large scale, expert curated databases is the expense of maintaining and extending them over time. No entity apart from a major institution with stable long-term funding can consider this, and their scope is limited considering the magnitude of microbial data being generated daily. Wikidata is an, openly editable, semantic web compatible framework for knowledge representation. Its a project of the Wikimedia Foundation and offers knowledge integration capabilities ideally suited to the challenge of representing the exploding body of information about microbial genomics. We are developing a microbial specific data model, based on Wikidatas semantic web compatibility, that represents bacterial species, strains and the gene and gene products that define them. Currently, we have loaded 1736 gene items and 1741 protein items for two strains of the human pathogenic bacteria Chlamydia trachomatis and used this subset of data as an example of the empowering utility of this model. In our next phase of development we will expand by adding another 118 bacterial genomes and their gene and gene products, totaling over ~900,000 additional entities. This aggregation of knowledge will be a platform for community-driven collaboration, allowing the networking of microbial genetic data through the sharing of knowledge by both the data and domain expert.

Genomics