bioRxiv ScienceSearch

Biology subjects

dos Santos, F. R. C.

Publications and source records attributed to dos Santos, F. R. C..

2 recordsLinked to original sources

Reboot: a straightforward approach to identify genes and splicing isoforms associated with cancer patient prognosis

Nowadays, the massive amount of data generated by modern sequencing technologies provides an unprecedented opportunity to find genes associated with cancer patient prognosis, connecting basic and translational research. However, treating high dimensionality of gene expression data and integrating it with clinical variables are major challenges to carry out these analyses. Here, we present Reboot, an original and efficient algorithm to find genes and splicing isoforms associated with cancer patient survival, disease progression, or other clinical endpoints. Reboot innovates by using a multivariate strategy with penalized Cox regression (LASSO method) combined with a bootstrap approach, in addition to statistical tests for supporting the findings, which are automatically plotted. Applying Reboot on data from 154 glioblastoma patients, we identified a three-gene signature (IKBIP, OSMR, PODNL1) whose increased derived risk score was significantly associated with worse patients prognosis, even in conjunction with other well-established clinical parameters. Similarly, Reboot was able to find a seven-splicing isoforms signature (CENPF-201; MLKL-202; NUP54-201; MCF2L-201; TFDP1-207; BBS1-206; HTT-202) related to worse overall survival in 177 pancreatic adenocarcinoma patients with elevated risk scores after uni- and multivariate analyses. In summary, Reboot is an efficient, intuitive, and straightforward way for finding genes or splicing isoforms (transcripts) signatures relevant to patient prognosis, which can democratize this kind of analysis and shed light on still under-investigated sets of cancer-related genes. Reboot effectively runs on either servers or personal computers and it is freely available at github.com/galantelab/reboot.

bioinformatics

Machine learning approaches reveal genomic regions associated with sugarcane brown rust resistance

Sugarcane is an economically important crop, but its genomic complexity has hindered advances in molecular approaches for genetic breeding. New cultivars are released based on the identification of interesting traits, and for sugarcane, brown rust resistance is a desirable characteristic due to the large economic impact of the disease. Although marker-assisted selection for rust resistance has been successful, the genes involved are still unknown, and the associated regions vary among cultivars, thus restricting methodological generalization. We used genotyping by sequencing of full-sib progeny to relate genomic regions with brown rust phenotypes. We established a pipeline to identify reliable SNPs in complex polyploid data, which were used for phenotypic prediction via machine learning. We identified 14,540 SNPs, which led to a mean prediction accuracy of 50% by using different models. We also tested feature selection algorithms to increase predictive accuracy, resulting in a reduced dataset with more explanatory power for rust phenotypes. Using different feature selection techniques, we achieved accuracy of up to 95% with a dataset of 131 SNPs related to brown rust QTL regions and auxiliary genes. Therefore, our novel strategy has the potential to assist studies of the genomic organization of brown rust resistance in sugarcane.

bioinformatics