bioRxiv Science⌕ Search

Biology subjects

Ziebarth, J. D.

Publications and source records attributed to Ziebarth, J. D..

2 recordsLinked to original sources

A Machine Learning-Based Investigation of Integrin Expression Patterns in Cancer and Metastasis

BackgroundIntegrins, a family of transmembrane receptor proteins, play complex roles in cancer development and metastasis. These roles could be better delineated through machine learning of transcriptomic data to reveal relationships between integrin expression patterns and cancer. MethodsWe collected publicly available RNA-Seq integrin expression from 8 healthy tissues and their corresponding tumors, along with data from metastatic breast cancer. We then used machine learning methods, including t-SNE visualization and Random Forest classification, to investigate changes in integrin expression patterns. ResultsIntegrin expression varied across tissues and cancers, and between healthy and cancer samples from the same tissue, enabling the creation of models that classify samples by tissue or disease status. The integrins whose expression was important to these classifiers were identified. For example, ITGA7 was key to classification of breast samples by disease status. Analysis in breast tissue revealed that cancer rewires co-expression for most integrins, but the co-expression relationships of some integrins remain unchanged in healthy and cancer samples. Integrin expression in primary breast tumors differed from their metastases, with liver metastasis notably having reduced expression. ConclusionsIntegrin expression patterns vary widely across tissues and are greatly impacted by cancer. Machine learning of these patterns can effectively distinguish samples by tissue or disease status.

cancer biology↗

Predictive Models and Impact of Interfacial Contacts and Amino Acids on Protein-protein Binding Affinity

Protein-protein interactions (PPIs) play a central role in nearly all cellular processes and that require proteins interact with sufficient binding affinity (BA) to form stable or transient complexes. Despite advancements in our understanding of protein-protein binding, much remains unknown about the interfacial region and its association with BA. Here we investigate the correlation of residue and atomic contacts of different types with BA and reveal the impact of the specific amino acids at the binding interface on BA. We create a series of linear regression (LR) models incorporating different contact features at both residue and atomic levels and examine how different methods of identifying and characterizing these properties impact the performance of these models. Particularly, we introduce a new and simple approach to predict BA based on the quantities of specific amino acids in contacts at the protein-protein interface. We show that the interfacial numbers of amino acids can be used to produce models with consistently good performance across different datasets, indicating the importance of the identities of interfacial amino acids in underlying the strength of BA. When trained on a diverse set of 141 complexes from two benchmark datasets, the best performing BA model (Pearson correlation coefficient R=0.68) was generated with an explicit linear equation involving six amino acids (tyrosine, glycine, serine, arginine, valine, and isoleucine). Tyrosine, in particular, was identified as the key amino acid in the quantitative link between specific amino acids and BA, as it had the strongest correlation with BA and was consistently identified as the most important amino acid in feature importance studies. Glycine, serine, and arginine were identified as the next three most important amino acids in predicting BA. The results from this study further our understanding of the importance of specific amino acids in PPIs and can be used to make improved predictions of BA, giving them implications for drug design and screening in the pharmaceutical industry.

biophysics↗