bioRxiv ScienceSearch

Biology subjects

Friedberg, I.

Publications and source records attributed to Friedberg, I..

7 recordsLinked to original sources

New Drosophila long-term memory genes revealed by assessing computational function prediction methods.

A major bottleneck to our understanding of the genetic and molecular foundation of life lies in the ability to assign function to a gene and, subsequently, a protein. Traditional molecular and genetic experiments can provide the most reliable forms of identification, but are generally low-throughput, making such discovery and assignment a daunting task. The bottleneck has led to an increasing role for computational approaches. The Critical Assessment of Functional Annotation (CAFA) effort seeks to measure the performance of computational methods. In CAFA3 we performed selected screens, including an effort focused on long-term memory. We used homology and previous CAFA predictions to identify 29 key Drosophila genes, which we tested via a long-term memory screen. We identify 11 novel genes that are involved in long-term memory formation and show a high level of connectivity with previously identified learning and memory genes. Our study provides first higher-order behavioral assay and organism screen used for CAFA assessments and revealed previously uncharacterized roles of multiple genes as possible regulators of neuronal plasticity at the boundary of information acquisition and memory formation.

neuroscience

Novel antimicrobial peptide discovery using machine learning and biophysical selection of minimal bacteriocin domains

Bacteriocins are ribosomally produced antimicrobial peptides that represent an untapped source of promising antibiotic alternatives. However, inherent challenges in isolation and identification of natural bacteriocins in substantial yield have limited their potential use as viable antimicrobial compounds. In this study, we have developed an overall pipeline for bacteriocin-derived compound design and testing that combines sequence-free prediction of bacteriocins using a machine-learning algorithm and a simple biophysical trait filter to generate minimal 20 amino acid peptide candidates that can be readily synthesized and evaluated for activity. We generated 28,895 total 20-mer peptides and scored them for charge, -helicity, and hydrophobic moment, allowing us to identify putative peptide sequences with the highest potential for interaction and activity against bacterial membranes. Of those, we selected sixteen sequences for synthesis and further study, and evaluated their antimicrobial, cytotoxicity, and hemolytic activities. We show that bacteriocin-based peptides with the overall highest scores for our biophysical parameters exhibited significant antimicrobial activity against E. coli and P. aeruginosa. Our combined method incorporates machine learning and biophysical-based minimal region determination, to create an original approach to rapidly discover novel bacteriocin candidates amenable to rapid synthesis and evaluation for therapeutic use.

microbiology

Crowdsourcing Image Analysis for Plant Phenomics to Generate Ground Truth Data for Machine Learning

The accuracy of machine learning tasks critically depends on high quality ground truth data. Therefore, in many cases, producing good ground truth data typically involves trained professionals; however, this can be costly in time, effort, and money. Here we explore the use of crowdsourcing to generate a large number of training data of good quality. We explore an image analysis task involving the segmentation of corn tassels from images taken in a field setting. We investigate the accuracy, speed and other quality metrics when this task is performed by students for academic credit, Amazon MTurk workers, and Master Amazon MTurk workers. We conclude that the Amazon MTurk and Master Mturk workers perform significantly better than the for-credit students, but with no significant difference between the two MTurk worker types. Furthermore, the quality of the segmentation produced by Amazon MTurk workers rivals that of an expert worker. We provide best practices to assess the quality of ground truth data, and to compare data quality produced by different sources. We conclude that properly managed crowdsourcing can be used to establish large volumes of viable ground truth data at a low cost and high quality, especially in the context of high throughput plant phenotyping. We also provide several metrics for assessing the quality of the generated datasets.\n\nAuthor SummaryFood security is a growing global concern. Farmers, plant breeders, and geneticists are hastening to address the challenges presented to agriculture by climate change, dwindling arable land, and population growth. Scientists in the field of plant phenomics are using satellite and drone images to understand how crops respond to a changing environment and to combine genetics and environmental measures to maximize crop growth efficiency. However, the terabytes of image data require new computational methods to extract useful information. Machine learning algorithms are effective in recognizing select parts of images, but they require high quality data curated by people to train them, a process that can be laborious and costly. We examined how well crowdsourcing works in providing training data for plant phenomics, specifically, segmenting a corn tassel - the male flower of the corn plant - from the often-cluttered images of a cornfield. We provided images to students, and to Amazon MTurkers, the latter being an on-demand workforce brokered by Amazon.com and paid on a task-by-task basis. We report on best practices in crowdsourcing image labeling for phenomics, and compare the different groups on measures such as fatigue and accuracy over time. We find that crowdsourcing is a good way of generating quality labeled data, rivaling that of experts.

plant biology

Identifying Antimicrobial Peptides using Word Embedding with Deep Recurrent Neural Networks

Antibiotic resistance constitutes a major public health crisis, and finding new sources of antimicrobial drugs is crucial to solving it. Bacteriocins, which are bacterially-produced antimicrobial peptide products, are candidates for broadening the available choices of an-timicrobials. However, the discovery of new bacteriocins by genomic mining is hampered by their sequences low complexity and high variance, which frustrates sequence similarity-based searches. Here we use word embeddings of protein sequences to represent bacteriocins, and apply a word embedding method that accounts for amino acid order in protein sequences,to predict novel bacteriocins from protein sequences without using sequence similarity. Our method predicts, with a high probability, six yet unknown putative bacteriocins in Lactobacil-lus. Generalized, the representation of sequences with word embeddings preserving sequence order information can be applied to protein classification problems for which sequence simi-larity cannot be used.

bioinformatics

MetaRiPPquest: A Peptidogenomics Approach for the Discovery of Ribosomally Synthesized and Post-translationally Modified Peptides

Ribosomally synthesized and post-translationally modified peptides (RiPPs) are an important class of natural products that include many antibiotics and a variety of other bioactive compounds. While recent breakthroughs in RiPP discovery raised the challenge of developing new algorithms for their analysis, peptidogenomic-based identification of RiPPs by combining genome/metagenome mining with analysis of tandem mass spectra remains an open problem. We present here MetaRiPPquest, a software tool for addressing this challenge that is compatible with large-scale screening platforms for natural product discovery. After searching millions of spectra in the Global Natural Products Social (GNPS) molecular networking infrastructure against just six genomic and metagenomic datasets, MetaRiPPquest identified 27 known and discovered 5 novel RiPP natural products.

bioinformatics

Maize GO Annotation - Methods, Evaluation, and Review (maize-GAMER)

1We created a new high-coverage, robust, and reproducible functional annotation of maize protein coding genes based on Gene Ontology (GO) term assignments. Whereas the existing Phytozome and Gramene maize GO annotation sets only cover 41% and 56% of maize protein coding genes, respectively, this study provides annotations for 100% of the genes. We also compared the quality of our newly-derived annotations with the existing Gramene and Phytozome functional annotation sets by comparing all three to a manually annotated gold standard set of 1,619 genes where annotations were primarily inferred from direct assay or mutant phenotype. Evaluations based on the gold standard indicate that our new annotation set is measurably more accurate than those from Phytozome and Gramene. To derive this new high-coverage, high-confidence annotation set we used sequence-similarity and protein-domain-presence methods as well as mixed-method pipelines that developed for the Critical Assessment of Function Annotation (CAFA) challenge. Our project to improve maize annotations is called maize-GAMER (GO Annotation Method, Evaluation, and Review) and the newly-derived annotations are accessible via MaizeGDB (http://download.maizegdb.org/maize-GAMER) and CyVerse (B73 RefGen_v3 5b+ at doi: doi.org/10.7946/P2S62P and B73 RefGen_v4 Zm00001d.2 at doi: doi.org/10.7946/P2M925).

bioinformatics

Ancestral Reconstruction of Gene Blocks using an Event-Based Method

Complexity is a fundamental attribute of life. Complex systems are made of parts that together perform functions that a single component, or subsets containing individual components, cannot. Examples of complex molecular systems include protein structures such as the F1Fo-ATPase, the ribosome, or the flagellar motor: each one of these structures requires most or all of its components to function properly. Given the ubiquity of complex systems in the biosphere, understanding the evolution of complexity is central to biology. At the molecular level, operons are a classic example of a complex system. An operons genes are co-transcribed under the control of a single promoter to a polycistronic mRNA molecule, and the operons gene products often form molecular complexes or metabolic pathways. With the large number of complete bacterial genomes available, we now have the opportunity to explore the evolution of these complex entities, by identifying possible intermediate states of operons. In this work, we developed a maximum parsimony algorithm to reconstruct ancestral operon states, and show a simple vertical evolution model of how operons may evolve from the individual component genes. We describe several ancestral states that are plausible functional intermediate forms leading to the full operon. We also offer Reconstruction of Ancestral Gene blocks Using Events or ROAGUE as a software tool for those interested in exploring gene block and operon evolution.\n\nThe software accompanying this paper is available under GPLv3 license in: https://github.com/nguyenngochuy91/Ancestral-Blocks-Reconstruction.\n\nAll figures in this paper are available in enlarged downloadable form from:https://github.com/nguyenngochuy91/Ancestral-Blocks-Reconstruction/tree/master/images

evolutionary biology