bioRxiv Science⌕ Search

Biology subjects

Banaei-Kashani, F.

Publications and source records attributed to Banaei-Kashani, F..

3 recordsLinked to original sources

EC-Bench: A Benchmark for Enzyme Commission NumberPrediction

Enzymes are proteins that catalyze specific biochemical reactions in cells. Enzyme Commission (EC) numbers are used to annotate enzymes in a four-level hierarchy that classifies enzymes based on the specific chemical reactions they catalyze. Accurate EC number prediction is essential for understanding enzyme functions. Despite the availability of numerous methods for predicting EC numbers from protein sequences, there is no unified framework for evaluating and studying such methods systematically. This gap limits the ability of the community to identify the most effective approaches for enzyme annotation. We introduce EC-Bench, a benchmark for EC number prediction, consisting of 1) an initial representative set of existing methods (including homology-based, deep learning, contrastive learning, and language model methods), 2) existing and novel accuracy and efficiency performance metrics, and 3) selected datasets to allow for comprehensive comparative study. EC-Bench is open-source and provides a framework for researchers to not only compare among existing methods objectively under uniform conditions, but also to introduce and effectively evaluate performance of new methods in a comparative frame-work. To demonstrate the utility of EC-Bench, we perform extensive experimentation to compare the existing EC number prediction methods and establish their advantages and disadvantages in a variety of prediction tasks, namely "exact EC number prediction", "EC number completion" and (partial or additional) "EC number recommendation". We find wide variation in the performance of different methods, but also subtle but potentially useful differences in the performance of different methods across tasks and for different parts of the EC hierarchy.

bioinformatics↗

A Generalized Higher-order Correlation Analysis Framework for Multi-Omics Network Inference

Multiple -omics (genomics, proteomics, etc.) profiles are commonly generated to gain insight into a disease or physiological system. Constructing multi-omics networks with respect to the trait(s) of interest provides an opportunity to understand relationships between molecular features but integration is challenging due to multiple data sets with high dimensionality. One approach is to use canonical correlation to integrate one or two omics types and a single trait of interest. However, these types of methods may be limited due to (1) not accounting for higher-order correlations existing among features, (2) computational inefficiency when extending to more than two omics data when using a penalty term-based sparsity method, and (3) lack of flexibility for focusing on specific correlations (e.g., omics-to-phenotype correlation versus omics-to-omics correlations). In this work, we have developed a novel multi-omics network analysis pipeline called Sparse Generalized Tensor Canonical Correlation Analysis Network Inference (SGTCCA-Net) that can effectively overcome these limitations. We also introduce an implementation to improve the summarization of networks for downstream analyses. Simulation and real-data experiments demonstrate the effectiveness of our novel method for inferring omics networks and features of interest. Author summaryMulti-omics network inference is crucial for identifying disease-specific molecular interactions across various molecular profiles, which helps understand the biological processes related to disease etiology. Traditional multi-omics integration methods focus mainly on pairwise interactions by only considering two molecular profiles at a time. This approach overlooks the complex, higher-order correlations often present in multi-omics data, especially when analyzing more than two types of -omics data and phenotypes. Higher-order correlation, by definition, refers to the simultaneous relationships among more than two types of -omics data and phenotype, providing a more complex and complete understanding of the interactions in biological systems. Our research introduces Sparse Generalized Tensor Canonical Correlation Network Analysis (SGTCCA-Net), a novel framework that effectively utilizes both higher-order and lower-order correlations for multi-omics network inference. SGTCCA-Net is adaptable for exploring diverse correlation structures within multi-omics data and is able to construct complex multi-omics networks in a two-dimensional space. This method offers a comprehensive view of molecular feature interactions with respect to complex diseases. Our simulation studies and real data experiments validate SGTCCA-Net as a potent tool for biomarker identification and uncovering biological mechanisms associated with targeted diseases.

bioinformatics↗

Heterogeneity analysis of acute exacerbations of chronic obstructive pulmonary disease and a deep learning framework with weak supervision and privacy protection

11.1 BackgroundChronic obstructive pulmonary disease (COPD) affects 5-10% of the adult US population and is a major cause of mortality. Acute exacerbations of COPD (AECOPDs) are a major driver of COPD morbidity and mortality, but there are no cost-effective methods to identify early AECOPDs when treatment is most likely to reduce the severity and duration of AECOPDs. 1.2 MethodsWe conducted the first long-term (> 12 months), real-time monitoring studies of AECOPD with wearable sensors and self-reporting. We applied a deep learning-based autoencoder for feature extractions, then applied K-means clustering to detect heterogeneity. Accordingly, we proposed a weakly supervised active learning framework to develop anomaly detection models for robust identification of early AECOPD, and a clustered federated learning approach to personalize the anomaly detection models for early detection of heterogeneous subtypes of AECOPD. We evaluated this model by comparing it with other unsupervised learning models and federated learning models. 1.3 FindingsWe identified two clusters based on the Silhouette score and SHAP analysis.One cluster shows high heart rate, low calories, and low steps; and the other has opposite characteristics. We also found out that a single subject could have exacerbation events from both clusters, indicating that there is not only subject-level heterogeneity but also event-level heterogeneity. Our weakly supervised framework outperformed unsupervised methods by 0.06 in average precision with 25 human annotation labels per subject. Our federated learning framework outperformed standard federated learning methods by 0.14 in F1 score and 0.17 in average precision. 1.4 InterpretationWe showed subject-level and event-level heterogeneity in AECOPD using mobile and wearable device data and developed a practical AECOPD detection framework with limited human annotated labels and keeping data private in each device.

bioinformatics↗