bioRxiv ScienceSearch

Biology subjects

Goodman, A.

Publications and source records attributed to Goodman, A..

7 recordsLinked to original sources

Evaluation of Deep Learning Strategies for Nucleus Segmentation in Fluorescence Images

Identifying nuclei is often a critical first step in analyzing microscopy images of cells, and classical image processing algorithms are most commonly used for this task. Recent developments in deep learning can yield superior accuracy, but typical evaluation metrics for nucleus segmentation do not satisfactorily capture error modes that are relevant in cellular images. Besides, large image data sets with ground truth for evaluation have been limiting. We present an evaluation framework to measure accuracy, types of errors, and computational efficiency; and use it to compare two deep learning strategies (U-Net and DeepCell) alongside a classical approach implemented in CellProfiler. We publicly release a set of 23,165 manually annotated nuclei and source code to reproduce experiments. Our results show that U-Net outperforms both pixel-wise classification networks and classical algorithms. Also, our evaluation framework shows that deep learning improves accuracy and reduces the number of biologically relevant errors by half.

bioinformatics

Weakly supervised learning of single-cell feature embeddings

We study the problem of learning representations for single cells in microscopy images to discover biological relationships between their experimental conditions. Many new applications in drug discovery and functional genomics require capturing the morphology of individual cells as comprehensively as possible. Deep convolutional neural networks (CNNs) can learn powerful visual representations, but require ground truth for training; this is rarely available in biomedical profiling experiments. While we do not know which experimental treatments produce cells that look alike, we do know that cells exposed to the same experimental treatment should generally look similar. Thus, we explore training CNNs using a weakly supervised approach that uses this information for feature learning. In addition, the training stage is regularized to control for unwanted variations using mixup or RNNs. We conduct experiments on two different datasets; the proposed approach yields single-cell embeddings that are more accurate than the widely adopted classical features, and are competitive with previously proposed transfer learning approaches.

bioinformatics

Label-free assessment of red blood cell storage lesions by deep learning

Blood transfusion is a life-saving clinical procedure. With millions of units needed globally each year, it is a growing concern to improve product quality and recipient outcomes.\n\nStored red blood cells (RBCs) undergo continuous degradation, leading to structural and biochemical changes. To analyze RBC storage lesions, complex biochemical and biophysical assays are often employed.\n\nWe demonstrate that label-free imaging flow cytometry and deep learning can characterize RBC morphologies during 42-day storage, replacing the current practice of manually quantifying a blood smear from stored blood units. Based only on bright field and dark field images, our model achieved 90% accuracy in classifying six different RBC morphologies associated with storage lesions versus human-curated manual examination. A model fitted to the deep learning-extracted features revealed a pattern of morphological changes within the aging blood unit that allowed predicting the expiration date of stored blood using solely morphological assessment.\n\nDeep learning and label-free imaging flow cytometry could therefore be applied to reduce complex laboratory procedures and facilitate robust and objective characterization of blood samples.

bioinformatics

High accuracy haplotype-derived allele frequencies from ultra-low coverage pool-seq samples

Evolve-and-resequence (E+R) experiments leverage next-generation sequencing technology to track the allele frequency dynamics of populations as they evolve. While previous work has shown that adaptive alleles can be detected by comparing frequency trajectories from many replicate populations, this power comes at the expense of high-coverage (>100x) sequencing of many pooled samples, which can be cost-prohibitive. Here, we show that accurate estimates of allele frequencies can be achieved with very shallow sequencing depths (<5x) via inference of known founder haplotypes in small genomic windows. This technique can be used to efficiently estimate frequencies for any number of bi-allelic SNPs in populations of any model organism founded with sequenced homozygous strains. Using both experimentally-pooled and simulated samples of Drosophila melanogaster, we show that haplotype inference can improve allele frequency accuracy by orders of magnitude for up to 50 generations of recombination, and is robust to moderate levels of missing data, as well as different selection regimes. Finally, we show that a simple linear model generated from these simulations can predict the accuracy of haplotype-derived allele frequencies in other model organisms and experimental designs. To make these results broadly accessible for use in E+R experiments, we introduce HAF-pipe, an open-source software tool for calculating haplotype-derived allele frequencies from raw sequencing data. Ultimately, by reducing sequencing costs without sacrificing accuracy, our method facilitates E+R designs with higher replication and resolution, and thereby, increased power to detect adaptive alleles.

evolutionary biology

Barriers to Integration of Bioinformatics into Undergraduate Life Sciences Education

Bioinformatics, a discipline that combines aspects of biology, statistics, and computer science, is increasingly important for biological research. However, bioinformatics instruction is rarely integrated into life sciences curricula at the undergraduate level. To understand why, the Network for Integrating Bioinformatics into Life Sciences Education (NIBLSE, \"nibbles\") recently undertook an extensive survey of life sciences faculty in the United States. The survey responses to open-ended questions about barriers to integration were subjected to keyword analysis. The barrier most frequently reported by the ~1,260 respondents was lack of faculty training. Faculty at associates-granting institutions report the least training in bioinformatics and the least integration of bioinformatics into their teaching. Faculty from underrepresented minority groups (URMs) in STEM reported training barriers at a higher rate than others, although the number of URM respondents was small. Interestingly, the cohort of faculty with the most recently awarded PhD degrees reported the most training but were teaching bioinformatics at a lower rate than faculty who earned their degrees in previous decades. Other barriers reported included lack of student interest in bioinformatics; lack of student preparation in mathematics, statistics, and computer science; already overly full curricula; and limited access to resources, including hardware, software, and vetted teaching materials. The results of the survey, the largest to date on bioinformatics education, will guide efforts to further integrate bioinformatics instruction into undergraduate life sciences education.

scientific communication and education

Evolution Of Hierarchy In Bacterial Metabolic Networks

BackgroundIn self-organized systems, the concept of flow hierarchy is a useful way to characterize the movement of information throughout a network. Hierarchical network organizations are shown to arise when there is a cost of maintaining links in the network. A similar constraint exists in metabolic networks, where costs come from reduced efficiency of nonspecific enzymes or from producing unnecessary enzymes. Previous analyses of bacterial metabolic networks have been used to predict the minimal nutrients that a bacterium needs to grow, its mutualistic relationships with other bacteria, and its major ecological niche. Using flow hierarchy, we can also infer the tradeoffs between growth rate and metabolic efficiency that bacteria make given their environmental constraints.\n\nResultsUsing a comparative approach on 2,935 bacterial metabolic networks, we show that flow hierarchy in bacterial metabolic networks tracks a fundamental tradeoff between growth rate and biomass production, and reflects a bacteriums realized ecological strategy. Additionally, by inferring the ancestral metabolic networks, we find that hierarchy decreases with distance from the root of the tree, suggesting the important pressure of increased growth rate relative to efficiency in the face of competition.\n\nConclusionsJust as hierarchical character is an important structural property in efficiently engineered systems, it also evolves in self-organized bacterial metabolic networks, reflects the life-history strategies of those bacteria, and plays an important role in network organization and efficiency.

evolutionary biology

The Landscape Of Type VI Secretion Across Human Gut Microbiomes Reveals Its Role In Community Composition

While the composition of the human gut microbiome has been well defined, the forces governing its assembly are poorly understood. Recently, prominent members of this community from the order Bacteroidales were shown to possess the type VI secretion system (T6SS), which mediates contact-dependent antagonism between Gram-negative bacteria. However, the distribution of the T6SS in human gut microbiomes and its role have not yet been characterized. To address this challenge, we construct an extensive catalog of T6SS effector/immunity (E-I) genes from three genetic architectures (GA1-3) found in Bacteroidales genomes. We then use metagenomic analysis to assess the abundances of these genes across a large set of gut microbiome samples. We find that despite E-I diversity across reference strains, each individual microbiome harbors a limited set of E-I genes representing a single E-I genotype. Importantly, for GA1-2, these genotypes are not associated with a specific species, suggesting selection for compatibility. GA3, in contrast, is restricted to B. fragilis, and its low diversity reflects a single B. fragilis strain per sample. We further show that in infant microbiomes GA3 is enriched and B. fragilis strains are replaced over time, suggesting competition for dominance in developing microbiomes. Finally, we find a strong association between the presence of GA3 and increased abundance of Bacteroides, indicating that this system confers a selective advantage in vivo in Bacteroides rich ecosystems. Combined, our findings provide the first comprehensive characterization of the T6SS landscape in the human microbiome, implicating it in both intra- and inter-species interactions.

microbiology