bioRxiv ScienceSearch

Biology subjects

Bongo, L. A.

Publications and source records attributed to Bongo, L. A..

2 recordsLinked to original sources

A Standard Operating Procedure For Outlier Removal In Large-Sample Epidemiological Transcriptomics Datasets

Transcriptome measurements and other -omics type data are increasingly more used in epidemiological studies. Most of omics studies to date are small with samples sizes in the tens, or sometimes low hundreds, but this is changing. Our Norwegian Woman and Cancer (NOWAC) datasets are to date one or two orders of magnitude larger. The NOWAC biobank contains about 50000 blood samples from a prospective study. Around 125 breast cancer cases occur in this cohort each year. The large biological variation in gene expression means that many observations are needed to draw scientific conclusions. This is true for both microarray and RNA-seq type data. Hence, larger datasets are likely to become more common soon.\n\nTechnical outliers are observations that somehow were distorted at the lab or during sampling. If not removed these observations add bias and variance in later statistical analyses, and may skew the results. Hence, quality assessment and data cleaning are important. We find common quality assessment libraries difficult to work with for large datasets for two reasons: slow execution speed and unsuitable visualizations.\n\nIn this paper, we present our standard operating procedure (SOP) for large-sample transcriptomics datasets. Our SOP combines automatic outlier detection with manual evaluation to avoid removing valuable observations. We use laboratory quality measures and statistical measures of deviation to aid the analyst. These are available in the nowaclean R package, currently available on GitHub (https://github.com/3inar/nowaclean). Finally, we evaluate our SOP on one of our larger datasets with 832 observations.

epidemiology

Curve Selection For Predicting Breast Cancer Metastasis From Prospective Gene Expression In Blood

We investigate whether there is information in gene expression levels in blood that predicts breast cancer metastasis. Our data comes from the NOWAC epidemiological cohort study where blood samples were provided at enrollment. This could be anywhere from years to weeks before any cancer diagnosis. When and if a cancer is diagnosed, it could be so in different ways: at a screening, between screenings, or in the clinic, outside of the screening program. To build predictive models we propose that variable selection should include followup time and stratify by detection method. We show by simulations that this improves the probability of selecting relevant predictor genes. We also demonstrate that it leads to improved predictions and more stable gene signatures in our data. There is some indication that blood gene expression levels hold predictive information about metastasis. With further development such information could be used for early detection of metastatic potential and as such aid in cancer treatment.

bioinformatics