bioRxiv Science⌕ Search

Biology subjects

Biggs, P. J.

Publications and source records attributed to Biggs, P. J..

3 recordsLinked to original sources

Visual integration of GWAS and differential expression results with the hidecan R package

SummaryWe present hidecan, an R package for generating visualisations that summarise the results of one or more genome-wide association studies and differential expression analyses, as well as manually curated candidate genes, e.g. extracted from the literature. Availability and ImplementationThe hidecan package is implemented in R and is publicly available on the CRAN repository (https://CRAN.R-project.org/package=hidecan) and on GitHub (https://github.com/PlantandFoodResearch/hidecan). A description of the package, as well as a detailed tutorial are available at https://plantandfoodresearch.github.io/hidecan/. Contactolivia.angelin-bonnet@plantandfood.co.nz. Supplementary informationSupplementary data are available.

genetics↗

Lost In The Forest

Levels of a predictor variable that are absent when a classification tree is grown can not be subject to an explicit splitting rule. This is an issue if these absent levels then present in a new observation for prediction. To date, there remains no satisfactory solution for absent levels in random forest models. Unlike missing data, absent levels are fully observed and known. Ordinal encoding of predictors allows absent levels to be integrated and used for prediction. Using a case study on source attribution of Campylobacter species using whole genome sequencing (WGS) data as predictors, we examine how target-agnostic versus target-based encoding of predictor variables with absent levels affects the accuracy of random forest models. We show that a target-based encoding approach using class probabilities, with absent levels designated the highest rank, is systematically biased, and that this bias is resolved by encoding absent levels according to the a priori hypothesis of equal class probability. We present a novel method of ordinal encoding predictors via principal coordinates analysis (PCO) which capitalizes on the similarity between pairs of predictor levels. Absent levels are encoded according to their similarity to each of the other levels in the training data. We show that the PCO-encoding method performs at least as well as the target-based approach and is not biased.

bioinformatics↗

Distinct gut microbiome patterns associate with consensus molecular subtypes of colorectal cancer

Colorectal cancer (CRC) is a heterogeneous disease and recent advances in subtype classification have successfully stratified the disease using molecular profiling. The contribution of bacterial species to CRC development is increasingly acknowledged, and here, we sought to analyse CRC microbiomes and relate them to tumour consensus molecular subtypes (CMS), in order to better understand the relationship between bacterial species and the molecular mechanisms associated with CRC subtypes. We classified 34 tumours into CRC subtypes using RNA-sequencing derived gene expression and determined relative abundances of bacterial taxonomic groups using 16S rRNA amplicon metabarcoding. 16S rRNA analysis showed enrichment of Fusobacteria and Bacteroidetes, and decreased levels of Firmicutes and Proteobacteria in CMS1. A more detailed analysis of bacterial taxa using non-human RNA-sequencing reads uncovered distinct bacterial communities associated with each molecular subtype. The most highly enriched species associated with CMS1 included Fusobacterium hwasookii and Porphyromonas gingivalis. CMS2 was enriched for Selenomas and Prevotella species, while CMS3 had few significant associations. Targeted quantitative PCR validated these findings and also showed an enrichment of Fusobacterium nucleatum, Parvimonas micra and Peptostreptococcus stomatis in CMS1. In this study, we have successfully associated individual bacterial species to CRC subtypes for the first time.

cancer biology↗