bioRxiv ScienceSearch

Biology subjects

Biegert, G.

Publications and source records attributed to Biegert, G..

1 recordsLinked to original sources

CancerDiscover: A configurable pipeline for cancer prediction and biomarker identification using machine learning framework

MotivationUse of various high-throughput screening techniques has resulted in an abundance of data, whose complete utility is limited by the tools available for processing and analysis. Machine learning holds great potential for deciphering these data in the context of cancer classification and biomarker identification. However, current machine learning tools require manual processing of raw data from various sequencing platforms, which is both tedious and time-consuming. The current classification tools lack flexibility in choosing the best feature selection algorithms from a range of algorithms and most importantly inability to compare various learning algorithms.\n\nResultsWe developed CancerDiscover, an open-source software pipeline that allows users to efficiently and automatically integrate large high-throughput datasets, preprocess, normalize, and selects best performing features from multiple feature selection algorithms. The pipeline lets users apply various learning algorithms and generates multiple classification models and evaluation reports that distinguish cancer from normal samples, as well as different types and subtypes of cancer.\n\nAvailability and ImplementationThe open source pipeline is freely available for download at https://github.com/HelikarLab/CancerDiscover.\n\nContactelikar2@unl.edu\n\nSupplementary InformationPlease refer to the CancerDiscover README (Supplementary File 1) for detailed instructions on installation and operation of the pipeline. For a list of available feature selection methods, see Supplementary File 2.

bioinformatics