bioRxiv ScienceSearch

Biology subjects

Hossein Sharifi Noghabi

Publications and source records attributed to Hossein Sharifi Noghabi.

4 recordsLinked to original sources

Robust Group Fused Lasso for Multisample CNV Detection under Uncertainty

One of the most important needs in the post-genome era is providing the researchers with reliable and efficient computational tools to extract and analyze this huge amount of biological data, in which DNA copy number variation (CNV) is a vitally important one. Array-based comparative genomic hybridization (aCGH) is a common approach in order to detect CNVs. Most of methods for this purpose were proposed for one-dimensional profile. However, slightly this focus has moved from one- to multi-dimensional signals. In addition, since contamination of these profiles with noise is always an issue, it is highly important to have a robust method for analyzing multi-sample aCGH data. In this paper, we propose Robust Grouped Fused Lasso (RGFL) which utilizes the Robust Group Total Variations (RGTV). Instead of l2,1 norm, the l1-l2 M-estimator is used which is more robust in dealing with non-Gaussian noise and high corruption. More importantly, Correntropy (Welsch M-estimator) is also applied for fitting error. Extensive experiments indicate that the proposed method outperforms the state-of-the art algorithms and techniques under a wide range of scenarios with diverse noises.

Bioinformatics

Robust and Stable Gene Selection via Maximum-Minimum Correntropy Criterion

One of the central challenges in cancer research is identifying significant genes among thousands of others on a microarray. Since preventing outbreak and progression of cancer is the ultimate goal in bioinformatics and computational biology, detection of genes that are most involved is vital and crucial. In this article, we propose a Maximum-Minimum Correntropy Criterion (MMCC) approach for selection of biologically meaningful genes from microarray data sets which is stable, fast and robust against diverse noise and outliers and competitively accurate in comparison with other algorithms. Moreover, via an evolutionary optimization process, the optimal number of features for each data set is determined. Through broad experimental evaluation, MMCC is proved to be significantly better compared to other well-known gene selection algorithms for 25 commonly used microarray data sets. Surprisingly, high accuracy in classification by Support Vector Machine (SVM) is achieved by less than10 genes selected by MMCC in all of the cases.

Bioinformatics

Mat-aCGH: a Matlab toolbox for simultaneous multisample aCGH data analysis and visualization

Mat-aCGH is an application toolbox for analysis and visualization of microarray-comparative genomic hybridization (array-CGH or aCGH) data which is based on Matlab. Full process of aCGH analysis, from denoising of the raw data to the visualization of the desired results, can be obtained via Mat-aCGH straightforwardly. The main advantage of this toolbox is that it is collection of recent well-known statistical and information theoretic methods and algorithms for analyzing aCGH data. More importantly, the proposed toolbox is developed for multisample analysis which is one of the current challenges in this area. Mat-aCGH is convenient to apply for any format of data, robust against diverse noise and provides the users with valuable information in the form of diagrams and metrics. Therefore, it eliminates the needs of another software or package for multisample aCGH analysis. aCGH Matlab source codes and datasets are freely available and can be downloaded at: hsharifi.student.um.ac.ir/imagesm/14407/Mat-aCGH.rar.

Bioinformatics

Cancer Classification by Correntropy-Based Sparse Compact Incremental Learning Machine

Cancer prediction is of great importance and significance and it is crucial to provide researchers and scientists with novel, accurate and robust computational tools for this issue. Recent technologies such as Microarray and Next Generation Sequencing have paved the way for computational methods and techniques to play critical roles in this regard. Many important problems in cell biology require the dense nonlinear interactions between functional modules to be considered. The importance of computer simulation in understanding cellular processes is now widely accepted, and a variety of simulation algorithms useful for studying certain subsystems have been designed. In this article, a Sparse Compact Incremental Learning Machine (SCILM) is proposed for cancer classification problem on microarray gene expression data which take advantage of Correntropy cost that makes it robust against diverse noises and outliers. Moreover, since SCILM uses l1-norm of the weights, it has sparseness which can be applied for gene selection purposes as well. Finally, due to compact structure, the proposed method is capable of performing classification tasks in all of the cases with only one neuron in its hidden layer. The experimental analysis is performed on 26 well known microarray datasets regarding diverse kinds of cancers and the results show that the proposed method not only achieved significantly high accuracy but also because of its sparseness, final connectivity weights determined the value and effectivity of each gene regarding the corresponding cancer.

Bioinformatics