bioRxiv ScienceSearch

Biology subjects

Williams, E.

Publications and source records attributed to Williams, E..

4 recordsLinked to original sources

Machine Learning approach to Predicting Stem-Cell Donor Availability

0. AbstractThe success of Unrelated Donor stem-cell transplants depends not only on finding genetically matched donors but also on donor availability. On average 50% of potential donors in the NMDP database are unavailable for a variety of reasons, after initially matching a patient, with significant variations in availability among subgroups (e.g., by race or age). Several studies have established univariate donor characteristics associated with availability. Individual consideration of each applicable characteristic is laborious. Extrapolating group averages to individual donor level tends to be highly inaccurate. In the current environment with enhanced donor data collection, we can make better estimates of individual donor availability. In this study, we propose a Machine Learning based approach to predict availability of every registered donor, to be used during donor selection and reduce the time taken to complete a transplant.

bioinformatics

Unrelated Donor Selection for Stem Cell Transplants using Predictive Modelling

Unrelated Donor selection for a Hematopoietic Stem Cell Transplant is a complex multi-stage process. Choosing the most suitable donor from a list of Human Leukocyte Antigen (HLA) matched donors can be challenging to even the most experienced physicians and search coordinators. The process involves experts sifting through potentially thousands of genetically compatible donors based on multiple factors. We propose a Machine Learning approach to donor selection based on historical searches performed and selections made for these searches. We describe the process of building a computational model to mimic the donor selection decision process and show benefits of using the proposed model in this study.

bioinformatics

Prospects for genomic selection in cassava breeding

Cassava (Manihot esculenta Crantz) is a clonally propagated staple food crop in the tropics. Genomic selection (GS) reduces selection cycle times by the prediction of breeding value for selection of unevaluated lines based on genome-wide marker data. GS has been implemented at three breeding programs in sub-Saharan Africa. Initial studies provided promising estimates of predictive abilities in single populations using standard prediction models and scenarios. In the present study we expand on previous analyses by assessing the accuracy of seven prediction models for seven traits in three prediction scenarios: (1) cross-validation within each population, (2) cross-population prediction and (3) cross-generation prediction. We also evaluated the impact of increasing training population size by phenotyping progenies selected either at random or using a genetic algorithm. Cross-validation results were mostly consistent across breeding programs, with non-additive models like RKHS predicting an average of 10% more accurately. Accuracy was generally associated with heritability. Cross-population prediction accuracy was generally low (mean 0.18 across traits and models) but prediction of cassava mosaic disease severity increased up to 57% in one Nigerian population, when combining data from another related population. Accuracy across-generation was poorer than within (cross-validation) as expected, but indicated that accuracy should be sufficient for rapid-cycling GS on several traits. Selection of prediction model made some difference across generations, but increasing training population (TP) size was more important. In some cases, using a genetic algorithm, selecting one third of progeny could achieve accuracy equivalent to phenotyping all progeny. Based on the datasets analyzed in this study, it was apparent that the size of a training population (TP) has a significant impact on prediction accuracy for most traits. We are still in the early stages of GS in this crop, but results are promising, at least for some traits. The TPs need to continue to grow and quality phenotyping is more critical than ever. General guidelines for successful GS are emerging. Phenotyping can be done on fewer individuals, cleverly selected, making for trials that are more focused on the quality of the data collected.\n\nAbbreviations

genetics

The Image Data Resource: A Scalable Platform for Biological Image Data Access, Integration, and Dissemination

Access to primary research data is vital for the advancement of science. To extend the data types supported by community repositories, we built a prototype Image Data Resource (IDR) that collects and integrates imaging data acquired across many different imaging modalities. IDR links high-content screening, super-resolution microscopy, time-lapse and digital pathology imaging experiments to public genetic or chemical databases, and to cell and tissue phenotypes expressed using controlled ontologies. Using this integration, IDR facilitates the analysis of gene networks and reveals functional interactions that are inaccessible to individual studies. To enable re-analysis, we also established a computational resource based on IPython notebooks that allows remote access to the entire IDR. IDR is also an open source platform that others can use to publish their own image data. Thus IDR provides both a novel on-line resource and a software infrastructure that promotes and extends publication and re-analysis of scientific image data.

bioinformatics