bioRxiv Science⌕ Search

Biology subjects

Webber, J. W.

Publications and source records attributed to Webber, J. W..

3 recordsLinked to original sources

Multi-cancer classification; an analysis of neural network complexity

AO_SCPLOWBSTRACTC_SCPLOWO_ST_ABSBackgroundC_ST_ABSCancer identification is generally framed as binary classification, normally discrimination of a control group from a single cancer group. However, such models lack any cancer-specific information, as they are only trained on one cancer type. The models fail to account for competing cancer risks. For example, an ostensibly healthy individual may have any number of different cancer types, and a tumor may originate from one of several primary sites. Pan-cancer evaluation requires a model trained on multiple cancer types, and controls, simultaneously, so that a physician can be directed to the correct area of the body for further testing. MethodsWe introduce novel neural network models to address multi-cancer classification problems across several data types commonly applied in cancer prediction, including circulating miRNA expression, protein, and mRNA. In particular, we present an analysis of neural network depth and complexity, and investigate how this relates to classification performance. Comparisons of our models with state-of-the-art neural networks from the literature are also presented. ResultsOur analysis evidences that shallow, feed-forward neural net architectures offer greater performance when compared to more complex deep feed-forward, Convolutional Neural Network (CNN), and Graph CNN (GCNN) architectures considered in the literature. ConclusionThe results show that multiple cancers and controls can be classified accurately using the proposed models, across a range of expression technologies in cancer prediction. ImpactThis study addresses the important problem of pan-cancer classification, which is often overlooked in the literature. The promising results highlight the urgency for further research.

bioinformatics↗

Dimensionality reduction by sparse orthogonal projection with applications to miRNA expression analysis and cancer prediction

BackgroundHigh dimensionality, i.e. p > n, is an inherent feature of machine learning. Fitting a classification model directly to p-dimensional data risks overfitting and a reduction in accuracy. Thus, dimensionality reduction is necessary to address overfitting and high dimensionality. ResultsWe present a novel dimensionality reduction method which uses sparse, orthogonal projections to discover linear separations in reduced dimension space. The technique is applied to miRNA expression analysis and cancer prediction. We use least squares fitting and orthogonality constraints to find a set of orthogonal directions which are highly correlated to the class labels. We also enforce L1 norm sparsity penalties, to prevent overfitting and remove the uninformative features from the model. Our method is shown to offer a highly competitive classification performance on synthetic examples and real miRNA expression data when compared to similar methods from the literature which use sparsity ideas and orthogonal projections. DiscussionA novel technique is introduced here, which uses sparse, orthogonal projections for dimensionality reduction. The approach is shown to be highly effective in reducing the dimension of miRNA expression data. The application of focus in this article is miRNA expression analysis and cancer predction. The technique may be generalizable, however, to other high dimensionality datasets.

bioinformatics↗

Fast and robust imputation for miRNA expression data using constrained least squares

High dimensional transcriptome profiling, whether through next generation sequencing techniques or high-throughput arrays, may result in scattered variables with missing data. Data imputation is a common strategy to maximize the inclusion of samples by using statistical techniques to fill in missing values. However, many data imputation methods are cumbersome and risk introduction of systematic bias. Here we present a new data imputation method using constrained least squares and algorithms from the inverse problems literature and present applications for this technique in miRNA expression analysis. The proposed technique is shown to offer an imputation orders of magnitude faster, with greater than or equal accuracy when compared to similar methods from the literature.

bioinformatics↗