bioRxiv ScienceSearch

Biology subjects

Lee, E. K.

Publications and source records attributed to Lee, E. K..

2 recordsLinked to original sources

optSelect: using agent-based modeling and binary PSO techniques for ensemble feature selection and stability assessment

MotivationRecent studies have shown that the ensemble feature selection approaches are essential for generating robust classifiers. Existing methods for aggregating feature lists from different methods require use of arbitrary thresholds for selecting the top ranked features and do not account for classification accuracy while selecting the optimal set. Here we present a two-stage ensemble feature selection framework for finding the optimal set of features without compromising on classification accuracy.\n\nMethods and ResultsWe present herein optSelect, a multi agent-based stochastic optimization approach for nested ensemble feature selection. Stage one involves function perturbation, where ranked list of features are generated using different methods and stage two involves data perturbation, where feature selection is performed within randomly selected subsets of the training data and the optimal set of features is selected within each set using the optSelect. The agents are assigned to different behavior states and move according to a binary PSO algorithm. A multi-objective fitness function is used to evaluate the classification accuracy of the agents. We evaluate the system performance using the random probe method and using five publicly available microarray datasets. The performance of optSelect is compared with single feature selection techniques and existing aggregation methods. The results show that the optSelect algorithm improves the classification accuracy compared to both individual and existing rank aggregation methods. The algorithm is incorporated into an R package, optSelect.\n\nContactkuppal2@emory.edu

bioinformatics

SEACOIN2.0: an interactive mining and visualization tool for information retrieval, summarization, and knowledge discovery

MotivationThe rapidly increasing size of biomedical databases such as MEDLINE requires the use of intelligent data mining methods for information extraction and summarization. Existing biomedical text-mining tools have limited capabilities for inferring topological and network relationships between biomedical terms. Very often too much is returned during summarization leading to information overload.\n\nResultsWe present herein SEACOIN 2.0, an interactive knowledge discovery and hypothesis generation tool for biomedical literature.SEACOIN generates k-ary relational networks of biomedical terms using a novel term ranking scheme to facilitate efficient information retrieval, summarization, and visual data exploration. Summarization is presented via multiple dynamic visualization panels. We evaluate the system performance in information retrieval and features extraction using the BioCreative 2013 Track 3 learning corpus. An average F-measure of 94% was achieved for document retrieval and an average precision of 88% was achieved for identification of top co-occurrence terms. The system allows interactive mining of complex implicit and explicit relationships among biomedical entities (genes, chemicals, diseases/disorders, mutations, etc.) and provides a framework for hypothesis generation. It also improves our understanding of various biological processes and disease mechanisms.\n\nContacteva.lee@gatech.edu

bioinformatics