bioRxiv ScienceSearch

Biology subjects

Carter, H.

Publications and source records attributed to Carter, H..

4 recordsLinked to original sources

Creating a scalable deep learning based Named Entity Recognition Model for biomedical textual data by repurposing BioSample free-text annotations

High quality metadata annotations for data hosted in large public repositories are essential for research reproducibility, and for conducting fast, powerful and scalable meta-analyses. Currently, a majority of sequencing samples in the National Center for Biotechnology Informations (NCBIs) Sequence Read Archive (SRA) are missing metadata across several categories. In an effort to improve the metadata coverage of these samples, we leveraged almost 44 million attribute-value pairs from SRA BioSample to train a scalable, recurrent neural network that predicts missing metadata via Named Entity Recognition (NER). The network was first trained to classify short text phrases according to 11 metadata categories and achieved an overall accuracy and area under the receiver operating characteristic (AUROC) curve of 85.2% and 0.977 respectively. We then applied our classifier to predict 11 metadata categories from the longer TITLE attribute of samples, evaluating performance on a set of samples withheld from model training. Prediction accuracies were high when extracting sample Species (94.85%), Condition/Disease (95.65%) and Strain (82.03%) from TITLEs, with lower accuracies and lack of predictions for other categories highlighting multiple issues with the current metadata annotations in BioSample. These results indicate the utility of recurrent neural networks for NER-based metadata prediction and the potential for models such as the one presented here to increase metadata coverage in BioSample while minimizing the need for manual curation. AvailabilityAll the analyses, environments, and Jupyter notebooks pertaining to this manuscript are available on Github: https://github.com/cartercompbio/PredictMEE.

bioinformatics

Extracting allelic read counts from 250,000 human sequencing runs in Sequence Read Archive

The Sequence Read Archive (SRA) contains over one million publicly available sequencing runs from various studies using a variety of sequencing library strategies. These data inherently contain information about underlying genomic sequence variants which we exploit to extract allelic read counts on an unprecedented scale. We reprocessed over 250,000 human sequencing runs (>1000 TB data worth of raw sequence data) into a single unified dataset of allelic read counts for nearly 300,000 variants of biomedical relevance curated by NCBI dbSNP, where germline variants were detected in a median of 912 sequencing runs, and somatic variants were detected in a median of 4,876 sequencing runs, suggesting that this dataset facilitates identification of sequencing runs that harbor variants of interest. Allelic read counts obtained using a targeted alignment were very similar to read counts obtained from whole genome alignment. Analyzing allelic read count data for matched DNA and RNA samples from tumors, we find that RNA-seq can also recover variants identified by WXS, suggesting that reprocessed allelic read counts can support variant detection across different library strategies in SRA. This study provides a rich database of known human variants across SRA samples that can support future meta-analyses of human sequence variation.

bioinformatics

Exposure to unpredictable trips and slips while walking can improve balance recovery responses with minimum predictive gait alterations

INTRODUCTIONThis study aimed to determine if repeated exposure to unpredictable trips and slips while walking can improve balance recovery responses when predictive gait alterations (e.g. slowing down) are minimised.\n\nMETHODSTen young adults walked on a 10-m walkway that induced slips and trips in fixed and random locations. Participants were exposed to a total of 12 slips, 12 trips and 6 non-perturbed walks in three conditions: 1) right leg fixed location, 2) left leg fixed location and 3) random leg and location. Kinematics during non-perturbed walks and previous and recovery steps were analysed.\n\nRESULTSThroughout the three conditions, participants walked with similar gait speed, step length and cadence(p>0.05). Participants extrapolated centre of mass (XCoM) was anteriorly shifted immediately before slips at the fixed location (p<0.01), but this predictive gait alteration did not transfer to random perturbation locations. Improved balance recovery from trips in the random location was indicated by increased margin of stability and step length during recovery steps (p<0.05). Changes in balance recovery from slips in the random location was shown by reduced backward XCoM displacement and reduced slip speed during recovery steps (p<0.05).\n\nCONCLUSIONSEven in the absence of most predictive gait alterations, balance recovery responses to trips and slips were improved through exposure to repeated unpredictable perturbations. A common predictive gait alteration to lean forward immediately before a slip was not useful when the perturbation location was unpredictable. Training balance recovery with unpredictable perturbations may be beneficial to fall avoidance in everyday life.

physiology

Pan-Cancer Analysis Reveals Technical Artifacts in The Cancer Genome Atlas (TCGA) Germline Variant Calls

The degree to which germline variation drives cancer development and shapes tumor phenotypes remains largely unexplored, possibly due to a lack of large scale publicly available germline data for a cancer cohort. Here we called germline variants on 9,618 cases from The Cancer Genome Atlas (TCGA) database representing 31 cancer types. We identified batch effects affecting loss of function (LOF) variant calls that can be traced back to differences in the way the sequence data were generated both within and across cancer types. Overall, LOF indel calls were more sensitive to technical artifacts than LOF Single Nucleotide Variant (SNV) calls. In particular, whole genome amplification of DNA prior to sequencing led to an artificially increased burden of LOF indel calls, which confounded association analyses relating germline variants to tumor type despite stringent indel filtering strategies. Due to the inherent noise we chose to remove all 614 amplified DNA samples, including all acute myeloid leukemia and virtually all ovarian cancer samples, from the final dataset. This study demonstrates how insufficient quality control can lead to false positive germlinetumor type associations and draws attention to the need to be sensitive to problems associated with a lack of uniformity in data generation in TCGA data.\n\nAuthor SummaryCancer research to date has largely focused on genetic aberrations specific to tumor tissue. In contrast, the degree to which germline, or inherited, variation contributes to tumorigenesis remains unclear, possibly due to a lack of accessible germline variant data. In this study we identify germline variants in 9,618 samples using raw germline exome data from The Cancer Genome Atlas (TCGA). There are substantial differences in the way exome sequence data was generated both across and within cancer types in TCGA. We observe that differences in sequence data generation introduced batch effects, or variation that is due to technical factors not true biological variation, in our variant data. Most notably, we observe that amplification of DNA prior to sequencing resulted in an excess of predicted damaging indel variants. We show how these batch effects can confound germline association analyses if not properly addressed. Our study highlights the difficulties of working with large public genomic datasets like TCGA where samples are collected over time and across data centers, and particularly cautions the use of amplified DNA samples for genetic association analyses.

genomics