bioRxiv Science⌕ Search

Biology subjects

Wu, C. H.

Publications and source records attributed to Wu, C. H..

3 recordsLinked to original sources

eMIND: Enabling automatic collection of protein variation impacts in Alzheimer's disease from the literature

Alzheimers disease and related dementias (AD/ADRDs) are among the most common forms of dementia, and yet no effective treatments have been developed. To gain insight into the disease mechanism, capturing the connection of genetic variations to their impacts, at the disease and molecular levels, is essential. The scientific literature continues to be a main source for reporting experimental information about the impact of variants. Thus, development of automatic methods to identify publications and extract the information from the unstructured text would facilitate collecting and organizing information for reuse. We developed eMIND, a deep learning-based text mining system that supports the automatic extraction of annotations of variants and their impacts in AD/ADRDs. In particular, we use this method to capture the impacts of protein-coding variants affecting a selected set of protein properties, such as protein activity/function, structure and post-translational modifications. We conducted an evaluation on the efficacy of eMIND to extract variant impact relations and obtained a recall of 0.84 and a precision of 0.94. The publications and extracted information are integrated into the UniProtKB computationally mapped bibliography to expand annotations on protein entries. eMINDs text-mined output are presented using controlled vocabularies and ontologies for variant, disease and impact along with the evidence sentences. A sample of annotated abstracts can be accessed at URL: https://research.bioinformatics.udel.edu/itextmine/emind.

bioinformatics↗

Melanoma clonal subline analysis uncovers heterogeneity-driven immunotherapy resistance mechanisms

Intratumoral heterogeneity (ITH) can promote cancer progression and treatment failure, but the complexity of the regulatory programs and contextual factors involved complicates its study. To understand the specific contribution of ITH to immune checkpoint blockade (ICB) response, we generated single cell-derived clonal sublines from an ICB-sensitive and genetically and phenotypically heterogeneous mouse melanoma model, M4. Genomic and single cell transcriptomic analyses uncovered the diversity of the sublines and evidenced their plasticity. Moreover, a wide range of tumor growth kinetics were observed in vivo, in part associated with mutational profiles and dependent on T cell-response. Further inquiry into melanoma differentiation states and tumor microenvironment (TME) subtypes of untreated tumors from the clonal sublines demonstrated correlations between highly inflamed and differentiated phenotypes with the response to anti-CTLA-4 treatment. Our results demonstrate that M4 sublines generate intratumoral heterogeneity at both levels of intrinsic differentiation status and extrinsic TME profiles, thereby impacting tumor evolution during therapeutic treatment. These clonal sublines proved to be a valuable resource to study the complex determinants of response to ICB, and specifically the role of melanoma plasticity in immune evasion mechanisms.

cancer biology↗

Text mining of CHO bioprocess bibliome: Topic modeling and document classification

Chinese hamster ovary (CHO) cells are widely used for mass production of therapeutic proteins in the pharmaceutical industry. With the growing need in optimizing the performance of producer CHO cell lines, research on CHO cell line development and bioprocess continues to increase in recent decades. Bibliographic mapping and classification of relevant research studies will be essential for identifying research gaps and trends in literature. To qualitatively and quantitatively understand the CHO literature, we have conducted topic modeling using a CHO bioprocess bibliome manually compiled in 2016, and compared the topics uncovered by the Latent Dirichlet Allocation (LDA) models with the human labels of the CHO bibliome. The results show a significant overlap between the manually selected categories and computationally generated topics, and reveal the machine-generated topic-specific characteristics. To identify relevant CHO bioprocessing papers from new scientific literature, we have developed a supervised learning model, Logistic Regression, to identify specific article topics and evaluated the results using three CHO bibliome datasets, Bioprocessing set, Glycosylation set, and Phenotype set. The use of top terms as features supports the explainability of document classification results to yield insights on new CHO bioprocessing papers.

bioinformatics↗