bioRxiv Science⌕ Search

Biology subjects

Tulpan, D.

Publications and source records attributed to Tulpan, D..

3 recordsLinked to original sources

Predicting environmental stressor levels with machine learning: a comparison between amplicon sequencing, metagenomics, and total RNA sequencing based on taxonomically assigned data

BackgroundMicrobes are increasingly (re)considered for environmental assessments because they are powerful indicators for the health of ecosystems. The complexity of microbial communities necessitates powerful novel tools to derive conclusions for environmental decision-makers, and machine learning is a promising option in that context. While amplicon sequencing is typically applied to assess microbial communities, metagenomics and total RNA sequencing (herein summarized as omics-based methods) can provide a more holistic picture of microbial biodiversity at sufficient sequencing depths. Despite this advantage, amplicon sequencing and omics-based methods have not yet been compared for taxonomy-based environmental assessments with machine learning. In this study, we applied 16S and ITS-2 sequencing, metagenomics, and total RNA sequencing to samples from a stream mesocosm experiment that investigated the impacts of two aquatic stressors, insecticide and increased fine sediment deposition, on stream biodiversity. We processed the data using similarity clustering and denoising (only applicable to amplicon sequencing) as well as multiple taxonomic levels, data types, feature selection, and machine learning algorithms and evaluated the stressor prediction performance of each generated model for a total of 1,536 evaluated combinations of taxonomic datasets and data-processing methods. ResultsSequencing and data-processing methods had a substantial impact on stressor prediction. While omics-based methods detected much more taxa than amplicon sequencing, 16S sequencing outperformed all other sequencing methods in terms of stressor prediction based on the Matthews Correlation Coefficient. However, even the highest observed performance for 16S sequencing was still only moderate. Omics-based methods performed poorly overall, but this was likely due to insufficient sequencing depth. Data types had no impact on performance while feature selection significantly improved performance for omics-based methods but not for amplicon sequencing. ConclusionAmplicon sequencing might be a better candidate for machine-learning-based environmental stressor prediction than omics-based methods, but the latter require further research at higher sequencing depths to confirm this conclusion. More sampling could improve stressor prediction performance, and while this was not possible in the context of our study, thousands of sampling sites are monitored for routine environmental assessments, providing an ideal framework to further refine the approach for possible implementation in environmental diagnostics.

genomics↗

Machine Learning based Genome-Wide Association Studies for Uncovering QTL Underlying Soybean Yield and its Components

Genome-wide association study (GWAS) is currently one of the important approaches for discovering quantitative trait loci (QTL) associated with traits of interest. However, insufficient statistical power is the limiting factor in current conventional GWAS methods for characterizing quantitative traits, especially in narrow genetic bases plants such as soybean. In this study, we evaluated the potential use of machine learning (ML) algorithms such as support vector machine (SVR) and random forest (RF) in GWAS, compared with two conventional methods of mixed linear models (MLM) and fixed and random model circulating probability unification (FarmCPU), for identifying QTL associated with soybean yield components. In this study, important soybean yield component traits, including the number of reproductive nodes (RNP), non-reproductive nodes (NRNP), total nodes (NP), and total pods (PP) per plant along with yield and maturity were assessed using 227 soybean genotypes evaluated across four environments. Our results indicated SVR-mediated GWAS outperformed RF, MLM and FarmCPU in discovering the most relevant QTL associated with the traits, supported by the functional annotation of candidate gene analyses. This study for the first time demonstrated the potential benefit of using sophisticated mathematical approaches such as ML algorithms in GWAS for identifying QTL suitable for genomic-based breeding programs.

plant biology↗

In Pursuit of a Better Broiler: Growth, Efficiency and Mortality of 16 Strains of Broiler Chickens

To meet the growing consumer demand for chicken meat, the poultry industry has selected broiler chickens for increasing efficiency and breast yield. While this high productivity means affordable and consistent product, it has come at a cost to broiler welfare. There has been increasing advocacy and consumer pressure on primary breeders, producers, processors and retailers to improve the welfare of the billions of chickens processed annually. Several small-scale studies have reported better welfare outcomes for slower growing strains compared to fast growing, conventional strains. However, these studies often housed birds with range access or used strains with vastly different growth rates. Additionally, there may be traits other than growth, such as body conformation, that influence welfare. As the global poultry industries consider the implications of using slower growing strains, there was a need for a comprehensive, multidisciplinary examination of broiler chickens with a wide range of genotypes differing in growth rate and other phenotypic traits. To meet this need, our team designed a study to benchmark data on conventional and slower growing strains of broiler chickens reared in standardized laboratory conditions. Over a two-year period, we studied 7,528 broilers from 16 different genetic strains. In this paper, we compare the growth, efficiency and mortality of broilers to one of two target weights (TW): 2.1 kg (TW1) and 3.2 kg (TW2). We categorized strains by their growth rate to TW2 as conventional (CONV), fastest slow strains (FAST), moderate slow strains (MOD) and slowest slow strains (SLOW). When incubated, hatched, housed, managed and fed the same, the categories of strains differed in body weights, growth rates, feed intake and feed efficiency. At 48 days of age, strains in the CONV category were 835-1264 g heavier than strains in the other categories. By TW2, differences in body weights and feed intake resulted in a 22 to 43-point difference in feed conversion ratios. Categories of strains did not differ in their overall mortality rates.

physiology↗