bioRxiv Science⌕ Search

Biology subjects

Hayat, S.

Publications and source records attributed to Hayat, S..

5 recordsLinked to original sources

Fast model-free standardization and integration of single-cell transcriptomics data

Single-cell transcriptomics datasets from the same anatomical sites generated by different research labs are becoming mainstream. However, fast, and computationally inexpensive tools for standardization of cell-type annotation and data integration are still needed to increase research inclusivity. To standardize cell-type annotation and integrate single-cell transcriptomics datasets, we have built a fast, model-free integration method called MASI (Marker-Assisted Standardization and Integration). MASI can run integrative annotation on a personal laptop for approximately one million cells, providing a cheap computational alternative for the single-cell data analysis community. MASI has an average macro F1/overall accuracy of 0.79/0.89 over the 4 benchmark datasets. We demonstrate that MASI outperforms other methods based on speed, and its performance for the tasks of data integration and cell-type annotation is comparable or even superior to other existing methods. We apply MASI for integrative lineage analysis and show that it preserves the underlying biological signal in datasets tested. Finally, to harness knowledge from single-cell atlases, we demonstrate three case studies that cover integration across research groups, biological conditions, and surveyed participants, respectively.

bioinformatics↗

SciViewer- An interactive browser for visualizing single cell datasets

Single-cell sequencing improves our ability to understand biological systems at single-cell resolution and can be used to identify novel drug targets and optimal cell-types for target validation. However, tools that can interactively visualize and provide target-centric views of these large datasets are limited. We present SciViewer (Single-cell Interactive Viewer), a novel tool to interactively visualize, annotate and share single-cell datasets. SciViewer allows visualization of cluster, gene and pathway level information such as clustering annotation, differential expression, pathway enrichment, cell-type specificity, cellular composition, normalized gene expression and comparison across datasets. Further, we provide APIs for SciViewer to interact with publicly available pharmacogenomics databases for systematic evaluation of potential novel drug targets. We provide a module for non-programmatic upload of single-cell datasets. SciViewer will be a useful tool for data exploration and target discovery from single-cell datasets. It is available on GitHub (https://github.com/Dhawal-Jain/SciViewer).

bioinformatics↗

MACA: Marker-based automatic cell-type annotation for single cell expression data

SummaryAccurately identifying cell-types is a critical step in single-cell sequencing analyses. Here, we present marker-based automatic cell-type annotation (MACA), a new tool for annotating single-cell transcriptomics datasets. We developed MACA by testing 4 cell-type scoring methods with 2 public cell-marker databases as reference in 6 single-cell studies. MACA compares favorably to 4 existing marker-based cell-type annotation methods in terms of accuracy and speed. We show that MACA can annotate a large single-nuclei RNA-seq study in minutes on human hearts with ~290k cells. MACA scales easily to large datasets and can broadly help experts to annotate cell types in single-cell transcriptomics datasets, and we envision MACA provides a new opportunity for integration and standardization of cell-type annotation across multiple datasets. Availability and implementationMACA is written in python and released under GNU General Public License v3.0. The source code is available at https://github.com/ImXman/MACA. ContactYang Xu (yxu71@vols.utk.edu), Sikander Hayat (hayat221@gmail.com)

bioinformatics↗

Multiscale interactome analysis coupled with off-target drug predictions reveals drug repurposing candidates for human coronavirus disease

The COVID-19 pandemic has led to an urgent need for the identification of new antiviral drug therapies that can be rapidly deployed to treat patients with this disease. COVID-19 is caused by infection with the human coronavirus SARS-CoV-2. We developed a computational approach to identify new antiviral drug targets and repurpose clinically-relevant drug compounds for the treatment of COVID-19. Our approach is based on graph convolutional networks (GCN) and involves multiscale host-virus interactome analysis coupled to off-target drug predictions. Cellbased experimental assessment reveals several clinically-relevant repurposing drug candidates predicted by the in silico analyses to have antiviral activity against human coronavirus infection. In particular, we identify the MET inhibitor capmatinib as having potent and broad antiviral activity against several coronaviruses in a MET-independent manner, as well as novel roles for host cell proteins such as IRAK1/4 in supporting human coronavirus infection, which can inform further drug discovery studies.

cell biology↗

Evaluation of colorectal cancer subtypes and cell lines using deep learning

Colorectal cancer (CRC) is a common cancer with a high mortality rate and a rising incidence rate in the developed world. The disease shows variable drug response and outcome. Molecular profiling techniques have been used to better understand the variability between tumours as well as cancer models such as cell lines. Drug discovery programs use cell lines as a proxy for human cancers to characterize their molecular makeup and drug response, identify relevant indications and discover biomarkers. In order to maximize the translatability and the clinical relevance of in vitro studies, selection of optimal cancer models is imperative. We have developed a deep learning based method to measure the similarity between CRC tumors and other tumors or disease models such as cancer cell lines. Our method efficiently leverages multi-omics data sets containing copy number alterations, gene expression and point mutations, and learns latent factors that describe the data in lower dimension. These latent factors represent the patterns across gene expression, copy number, and mutational profiles which are clinically relevant and explain the variability of molecular profiles across tumours and cell lines. Using these, we propose a refined colorectal cancer sample classification and provide best-matching cell lines in terms of multi-omics for the different subtypes. These findings are relevant for patient stratification and selection of cell lines for early stage drug discovery pipelines, biomarker discovery, and target identification.

bioinformatics↗